Deceptive model quantization from AtomicChat?

Reddit r/LocalLLaMA Models

Summary

An analysis suggests that AtomicChat's Qwen3.8-Flash-Next quant may be deceptive by misrepresenting IQ2_S tensors as Q4_K_M, highlighting potential transparency issues in model quantization.

I kept seeing guys in this sub saying how AtomicChat's Qwen3.8-Flash-Next quant is so good, fits in their machine when unsloth's can't, runs faster than other quants etc, so I went check out what's happening there. First thing I noticed was that AtomicChat's Q4_K_M quant is suspiciously small when the ngram table is removed (only ~56GB), it seems like most of the tensors in this quant are IQ2_S instead of the usual Q4_K, Q5_K and Q6_K that you usually find in Q4_K_M quants, the GGUF filetype metadata also says IQ2_S instead of Q4_K_M. In their model card, their Q4_K_M also has suspiciously high KLD (0.084). It seems pretty obvious to me that they're pretending a IQ2_S quant as a Q4_K_M, but at the same time I'm genuinely not sure because it can't be only me who found this right? How can nobody be pointing this out? Am I missing something or what may they be doing? Their HF repo ID: AtomicChat/Qwen3.8-Flash-Next-GGUF
Original Article

Similar Articles

Qwen3.8 Flash AP Quants

Reddit r/LocalLLaMA

The post announces new quantized versions of the Qwen3.8 Flash model that outperform other high-quality quants through modified KLD measurement and optimized prefill performance.

Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang

Reddit r/LocalLLaMA

This is a quantized version of the Qwen3.8-27B AI model using IQ4_XS quantization, optimized for 16GB RAM systems, with instructions for local deployment using various tools like llama.cpp and Ollama.