Qwen3.8 Flash AP Quants

Reddit r/LocalLLaMA Models

Summary

The post announces new quantized versions of the Qwen3.8 Flash model that outperform other high-quality quants through modified KLD measurement and optimized prefill performance.

Quite surprised to be beating other high quality quants. It took a lot of benchmarking to get here and we are quite pleased with these, hope they are useful to the community. It required a modified way of measuring KLD with a new dataset, since the NGRAM got in the way by remembering basically all of wikipedia. We tried to not only go for high precision, but also keep prefill performance in mind. Full model card here https://huggingface.co/agentionai/Qwen3.8-Flash-Next-AP-GGUF Let us know if there are any issues. https://preview.redd.it/nq48xxtxj6nh1.png?width=1800&format=png&auto=webp&s=5059384113868ed2717e04b24603a5fbb7d9908c
Original Article

Similar Articles

Qwen3.8 Flash Quants

Reddit r/LocalLLaMA

The author released mainline-compatible imatrix quants for the Qwen3.8-Flash-Next model, offering 20–30GB smaller sizes with competitive PPL, and includes a separate ROCm build for AMD hardware.

Qwen/Qwen3.8-Flash-Next-FP8

Hugging Face Models Trending

The article releases FP8-quantized weights for the Qwen3.8-Flash-Next model, introducing architectural innovations like Hybrid Attention with QSA, Gated Residual, and N-gram Embedding to improve efficiency in large language models.

Qwen3.6-27B Quantization Benchmark

Reddit r/LocalLLaMA

This article benchmarks various Qwen3.6-27B quantizations (Q8 to Q2) using KLD and Same Top P metrics, comparing providers like Unsloth and mradermacher, and offers recommendations for quality-size trade-offs.