10% faster decode with Q4_K MTP draft model with Gemma 4 31b
Summary
A user reports that quantising the f16 MTP draft model to Q4_K for Gemma 4 31b gives roughly 10% faster decode (65 to 72 TPS) on dual 3090s compared to the default Q4_0, while Q2_K performs worse.
Similar Articles
[3090] Gemma4 QAT + MTP quick TPS numbers [TLDR 1.2-1.8x better]
Benchmark results showing 1.2-1.8x token-per-second speedups on Gemma 4 models (12B and 26B) using QAT and MTP on a 24GB RTX 3090 GPU.
Gemma 4 QAT 31B responds better to KV cache quantization too
The Gemma 4 QAT 31B model demonstrates improved behavior with KV cache quantization, suggesting enhanced inference efficiency.
What's your experience with Gemma4 QAT?
User shares positive experience with Gemma4 QAT model, noting quality improvements and speed gains with MTP, and asks others for their experiences.
Gemma4-12B-QAT Uncensored Balanced is out with MTP (~60% speed boost)!
Release of Gemma4-12B-QAT Uncensored Balanced, a fine-tuned uncensored model with a multi-token-prediction draft head for ~60% faster speculative decoding, optimized for llama.cpp and offering vision support.
Gemma4 26b a4b Apex quant is quite good
User benchmarks the APEX quantized version of Gemma4 26B A4B model on AMD RX 9060 XT, achieving 38 tps at 90k context with no quality degradation, finding it better than previous quantizations.