@9hills: Qwen3.8-27B 本地部署指南 1. Q4 基本没有质量损失,甚至可以用Q3 2. 用Unsloth的GGUF。 3. 思考开low就行了。 4. 开 dflash2

X AI KOLs Timeline 新闻

摘要

本文提供了 Qwen3.8-27B 模型的本地部署指南,推荐使用 Q4 量化和 Unsloth GGUF 工具,并分享了与 FP8 基准相比较的性能测试结果。

Qwen3.8-27B 本地部署指南 1. Q4 基本没有质量损失,甚至可以用Q3 2. 用Unsloth的GGUF。 3. 思考开low就行了。 4. 开 dflash2
查看原文
查看缓存全文

缓存时间: 2026/08/22 01:19

Qwen3.8-27B 本地部署指南

  1. Q4 基本没有质量损失,甚至可以用Q3
  2. 用Unsloth的GGUF。
  3. 思考开low就行了。
  4. 开 dflash2

Alexey Fateev (@superalesha): Main result. At xhigh every quant landed between 88.0 and 90.0% pass@1 on the full suite:

AWQ INT4: 90.0% NVFP4: 89.3% GGUF Q4_K_M: 89.3% FP8: 88.7% NInfer: 88.0%

Yes, the 4 bit quants scored above the FP8 baseline. McNemar says its a statistical tie, first and last place

相似文章

unsloth/Qwen3.6-27B-MTP-GGUF

Hugging Face Models Trending

Unsloth 发布了 Qwen3.6-27B 模型的 GGUF 权重,该模型支持多令牌预测(MTP),可实现更快的生成速度并增强了智能体(Agentic)编码能力。