Tag
Qwen3.8-27B is a new AI model that emphasizes wall-clock time over token speed, offering superior intelligence for local deployment on consumer hardware with 32 GB VRAM.
Recent versions of llama.cpp now automatically load MTP tensors for draft-mtp architectures, even if speculative decoding is not enabled, potentially increasing VRAM usage for users with bundled MTP blocks in their GGUF files.