@Ex0byt: Days of model activations, slicing, splicing, fine-tuning + 15 hours of nail-biting NVFP4 calibration/propagation passe…
Summary
A community member released Qwen3.6-35B-A3B-PRISM-NVFP4, a multi-pass, dataset-calibrated zero-loss NVFP4 quantized variant of the Qwen model.
View Cached Full Text
Cached at: 04/23/26, 12:06 PM
Days of model activations, slicing, splicing, fine-tuning + 15 hours of nail-biting NVFP4 calibration/propagation passes. I freely give you Qwen3.6-35B-A3B-PRISM-NVFP4 - Highest-quality multi-pass, 1024 custom dataset-calibrated, zero-loss NVFP4 (support for
Similar Articles
Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD
A fully quantized 4-bit NVFP4 version of the Qwen3.8-27B AI model, trained with the QUASAR method to maintain high quality while reducing model size.
RedHatAI/Qwen3.6-35B-A3B-NVFP4
Red Hat AI released an NVFP4-quantized 35B MoE version of Qwen3.6 that retains 96.28% GSM8K accuracy while enabling 4-bit inference via vLLM.
unsloth/Qwen3.8-27B-NVFP4
Unsloth has released an NVFP4 quantized version of the Qwen3.8-27B AI model, which offers enhanced capabilities in coding, professional work, agentic tasks, and native vision-language understanding.
nvidia/Qwen3.6-35B-A3B-NVFP4 · Hugging Face
NVIDIA releases Qwen3.6-35B-A3B-NVFP4, a quantized version of Alibaba's mixture-of-experts multimodal language model, optimized for deployment on NVIDIA GPUs using Model Optimizer.
2.5x faster Qwen3.6 NVFP4 Unsloth quants
Unsloth releases quantized Qwen3.6 models using NVFP4 format, achieving 2.5x faster inference speeds.