Unsloth Gemma 4 QAT MTP assistant models now available
Summary
Unsloth released Gemma 4 QAT MTP assistant models as GGUF files on Hugging Face, available in q8_0 and larger quantization formats.
Similar Articles
unsloth/gemma-4-12B-it-qat-GGUF
Unsloth releases GGUF quantized versions of Google DeepMind's Gemma 4 models, optimized with Quantization-Aware Training (QAT) to reduce memory requirements while preserving quality, supporting multiple formats and sizes for diverse deployment.
Unsloth just dropped MTP GGUF weights for Gemma 4!
Unsloth has released Multi-Token Prediction (MTP) GGUF weights for Gemma 4 models (31B, 26B-A4B, 12B) in Q8, F16, and BF16 precisions, available on Hugging Face.
google/gemma-4-12B-it-qat-q4_0-gguf
Google DeepMind releases Gemma 4 models optimized with Quantization-Aware Training (QAT) in multiple formats including GGUF, enabling high quality with reduced memory requirements.
Gemma 4 Quadruple Release, 12B, 12B QAT, 26B-A4B QAT and 31B QAT Uncensored Heretics!
llmfan46 released a quadruple set of uncensored, fine-tuned and quantized Gemma-4 models on Hugging Face, including 12B, 26B-A4B, and 31B variants with QAT and GGUF formats.
moar QAT stuff and hairy ticks
The author releases improved GGUF quantized versions of Gemma 4 models (12B and 31B) using a more accurate quantization-aware training process that achieves lower KLD and higher same-top percentage than stock quantizations.