PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled

Reddit r/LocalLLaMA Tools

Summary

Recent versions of llama.cpp now automatically load MTP tensors for draft-mtp architectures, even if speculative decoding is not enabled, potentially increasing VRAM usage for users with bundled MTP blocks in their GGUF files.

If your GGUF has MTP/NextN tensors baked in (GLM-5.2, hy_v3, qwen35moe, step35, etc.), recent llama.cpp builds load them by default — even if you never pass --spec-type draft-mtp. Before, they were skipped unless you actually enabled speculative decoding. Most community GGUFs bundle the MTP block by default, so this means extra VRAM/RAM use (~1 extra MoE layer) on every load, whether you use MTP or not. See https://github.com/ggml-org/llama.cpp/pull/25980
Original Article

Similar Articles

MTP support merged into llama.cpp

Reddit r/LocalLLaMA

The pull request adding MTP (Multi-Token Prediction) support to llama.cpp has been merged into the master branch.