community-ggufs

Tag

Cards List
#community-ggufs

PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

Reddit r/LocalLLaMA ↗ · 2026-07-29

Recent versions of llama.cpp now automatically load MTP tensors for draft-mtp architectures, even if speculative decoding is not enabled, potentially increasing VRAM usage for users with bundled MTP blocks in their GGUF files.

0 favorites 0 likes
← Back to home

Submit Feedback