PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
Summary
Recent versions of llama.cpp now automatically load MTP tensors for draft-mtp architectures, even if speculative decoding is not enabled, potentially increasing VRAM usage for users with bundled MTP blocks in their GGUF files.
Similar Articles
PSA: If you haven’t updated Llama.cpp for a couple of days and find MTP to not be performing well, update llamacpp.
Update Llama.cpp for a significant token generation speed boost, up to 1.5-1.8x, and improved prompt processing.
MTP support merged into llama.cpp
The pull request adding MTP (Multi-Token Prediction) support to llama.cpp has been merged into the master branch.
Remove padding and multiple D2D copies for MTP by gaugarg-nv · Pull Request #24086 · ggml-org/llama.cpp
A pull request for llama.cpp that removes padding and multiple device-to-device copies for Multi-Token Prediction (MTP), improving performance on GPU.
Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp
This pull request enables CUDA graph for MTP draft in llama.cpp to improve performance of LLM inference on GPUs.
llama.cpp MTP speculative simplified for July 2026 big wins on dense models, underwhelming on MoE
An analysis of native MTP speculative decoding in llama.cpp shows significant speedups (1.4x-2.2x) for dense models like Qwen3.6-27B, but underwhelming results on MoE architectures, where gains are minimal due to already low per-step overhead.