10% faster decode with Q4_K MTP draft model with Gemma 4 31b

Reddit r/LocalLLaMA Models

Summary

A user reports that quantising the f16 MTP draft model to Q4_K for Gemma 4 31b gives roughly 10% faster decode (65 to 72 TPS) on dual 3090s compared to the default Q4_0, while Q2_K performs worse.

(Disclaimer: I am a noob and don’t know what I am doing) Gemma 4 31b unsloth/gemma-4-31B-it-qat-GGUF I took the f16 MTP draft model and quantised it to Q4_K (instead of Q4_0 of unsloth) and gained around 10% in decode: from 65TPs to 72TPs. Dual 3090, split mode layer. Draft KV to Q4_0 Anyone has the same experience or can confirm? Just for fun I tried Q2_K but got worst result
Original Article

Similar Articles

What's your experience with Gemma4 QAT?

Reddit r/LocalLLaMA

User shares positive experience with Gemma4 QAT model, noting quality improvements and speed gains with MTP, and asks others for their experiences.

Gemma4 26b a4b Apex quant is quite good

Reddit r/LocalLLaMA

User benchmarks the APEX quantized version of Gemma4 26B A4B model on AMD RX 9060 XT, achieving 38 tps at 90k context with no quality degradation, finding it better than previous quantizations.