NVFP4 + MTP - voilà on llama.cpp

Reddit r/LocalLLaMA Tools

Summary

NVFP4 quantization and Multi-Token Prediction support have been added to llama.cpp in release b9297.

As in title - NVFP4 + MTP at once on llama.cpp [https://github.com/ggml-org/llama.cpp/releases/tag/b9297](https://github.com/ggml-org/llama.cpp/releases/tag/b9297)
Original Article

Similar Articles

Here is my llama.cpp NVFP4/MXFP6 GGUF quantizer tool

Reddit r/LocalLLaMA

The author introduces an open-source GGUF quantizer tool for llama.cpp that creates NVFP4 and MXFP6 quantized models with advanced techniques like RSF, tensor promotion, and dynamic quantization, achieving better quality than existing methods like ModelOpt.

MTP support merged into llama.cpp

Reddit r/LocalLLaMA

The pull request adding MTP (Multi-Token Prediction) support to llama.cpp has been merged into the master branch.

Testing llama.cpp MTP support on Qwen3.6 - RTX 5090

Reddit r/LocalLLaMA

A technical test of llama.cpp's new Multi-Token Prediction (MTP) support using Qwen3.6 models on an RTX 5090, comparing performance with and without MTP across different prompts and GGUF quantizations.

b9200 released - potential mtp pp increase

Reddit r/LocalLLaMA

llama.cpp release b9200 improves prompt processing speed for Multi-Token Prediction by avoiding unnecessary logits copying, reducing memory traffic.