spectralquant

Tag

Cards List
#spectralquant

We built a calibration-aware Q4_K_M quant of Qwen3.5 0.8B that recovers 96.5% of the BF16 gap vs pure llama.cpp Q4_K_M (SpectralQuant)

Reddit r/LocalLLaMA · 2026-06-27

A calibration-aware Q4_K_M quantization of Qwen3.5 0.8B using SpectralQuant recovers 96.5% of the BF16 performance gap compared to the standard llama.cpp Q4_K_M quant.

0 favorites 0 likes
#spectralquant

@anirudhbv_ce: Introducing SpectralQuant.. here to save your KV cache :)

X AI KOLs Timeline · 2026-05-18 Cached

SpectralQuant is a new KV cache quantization technique achieving 5.95× compression on Mistral 7B with only 7.5% perplexity overhead, significantly outperforming TurboQuant while requiring only 15 seconds of calibration per model.

0 favorites 0 likes
← Back to home

Submit Feedback