2.5x faster Qwen3.6 NVFP4 Unsloth quants
Summary
Unsloth releases quantized Qwen3.6 models using NVFP4 format, achieving 2.5x faster inference speeds.
Similar Articles
unsloth/Qwen3.6-27B-NVFP4
Unsloth releases an NVFP4 quantized checkpoint of Qwen3.6-27B, claiming 2.5x faster throughput and accuracy comparable to FP8 and BF16, with instructions for running on a 24GB GPU via vLLM.
@MiaAI_lab: Nvidia did it again! @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on D…
Nvidia's optimized Qwen 3.6 27B NVFP4 model achieves 41% faster single-session inference and 23-25% faster concurrent inference on DGX Spark compared to Unsloth's version.
unsloth/Qwen3.8-27B-NVFP4
Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.
@MiaAI_lab: FYI the best Qwen 3.6 35b nvfp4 to run is the @NVIDIAAI nvfp4. Do not use unsloth nvfp4, it performs worse. https://hug…
NVIDIA's nvfp4 quantized version of Qwen 3.6 35B is recommended over the Unsloth variant, offering better performance. The model is available on HuggingFace for use in AI applications.
Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD
A fully quantized 4-bit NVFP4 version of the Qwen3.8-27B AI model, trained with the QUASAR method to maintain high quality while reducing model size.