Qwen3.8 Flash Quants

Reddit r/LocalLLaMA Models

Summary

The author released mainline-compatible imatrix quants for the Qwen3.8-Flash-Next model, offering 20–30GB smaller sizes with competitive PPL, and includes a separate ROCm build for AMD hardware.

~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly 20–30GB depending on the file you compare against. PPL is in the model card and is competitive with both. Repo: https://huggingface.co/agentionai/Qwen3.8-Flash-Next-AP-GGUF Q4 quants are the ones I would start with. Q3 and Q5 are coming. Recipe is per-layer / tailored, not a blanket lower bpw. AMD / Strix Halo: separate ROCmFP4 build that is a bit better and faster than the Q4_XS on that hardware. (https://huggingface.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF) If you try it, post your quant, RAM/VRAM, tok/s, and whether quality felt on par with Unsloth IQ4_XS / Q4_K. That is the comparison I care about.
Original Article

Similar Articles

Qwen3.8 Flash AP Quants

Reddit r/LocalLLaMA

The post announces new quantized versions of the Qwen3.8 Flash model that outperform other high-quality quants through modified KLD measurement and optimized prefill performance.

Qwen 3.5 122B Heretic ROCmFP4 iMatrix

Reddit r/LocalLLaMA

A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.

Qwen/Qwen3.8-Flash-Next

Hugging Face Models Trending

Release of Qwen3.8-Flash-Next, an open-weight AI model introducing architectural innovations like Hybrid Attention with QSA and Gated Residual for improved efficiency and scalability, previewing the future Qwen4 architecture.