2.5x faster Qwen3.6 NVFP4 Unsloth quants

Reddit r/LocalLLaMA Tools

Summary

Unsloth releases quantized Qwen3.6 models using NVFP4 format, achieving 2.5x faster inference speeds.

No content available
Original Article

Similar Articles

unsloth/Qwen3.6-27B-NVFP4

Hugging Face Models Trending

Unsloth releases an NVFP4 quantized checkpoint of Qwen3.6-27B, claiming 2.5x faster throughput and accuracy comparable to FP8 and BF16, with instructions for running on a 24GB GPU via vLLM.

unsloth/Qwen3.8-27B-NVFP4

Hugging Face Models Trending

Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.