Unsloth Gemma 4 QAT MTP 辅助模型现已可用

Reddit r/LocalLLaMA 模型

摘要

Unsloth 发布了 Gemma 4 QAT MTP 辅助模型,以 GGUF 文件形式在 Hugging Face 上提供,支持 q8_0 及更大量化格式。

这些模型均以 q8_0 模型形式提供,文件名为 `mtp-gemma-4-*.gguf`,位于目录根目录;同时在 `MTP` 文件夹中以 q8_0 及更大量化规格提供。 - https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-26B-A4B-it-qat-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-E2B-it-qat-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-E2B-it-qat-mobile-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-E4B-it-qat-GGUF/tree/main - https://huggingface.co/unsloth/gemma-4-E4B-it-qat-mobile-GGUF/tree/main
查看原文

相似文章

unsloth/gemma-4-12B-it-qat-GGUF

Hugging Face Models Trending

Unsloth 发布了Google DeepMind的Gemma 4模型的GGUF量化版本,通过量化感知训练(QAT)优化,在保持质量的同时降低内存需求,支持多种格式和大小,适用于不同的部署场景。

google/gemma-4-12B-it-qat-q4_0-gguf

Hugging Face Models Trending

Google DeepMind 发布了 Gemma 4 模型,这些模型通过量化感知训练(QAT)进行了优化,并提供包括 GGUF 在内的多种格式,在降低内存需求的同时实现了高质量。

更多QAT内容以及毛茸茸的tick

Reddit r/LocalLLaMA

作者发布了Gemma 4模型(12B和31B)改进后的GGUF量化版本,采用了更精确的量化感知训练过程,相比原版量化实现了更低的KLD和更高的同top百分比。