@ClementDelangue: is this useful?

X AI KOLs Following Tools

Summary

Buun introduces VBR (Variable Bit Rate), a new KV cache format that dynamically quantizes layers to optimize quality within VRAM constraints, available now on master.

is this useful?
Original Article
View Cached Full Text

Cached at: 07/15/26, 09:49 AM

is this useful?

buun (@spiritbuun): Introducing a new KV format: VBR (Variable Bit Rate) VBR dynamically quantizes your KV cache layer-by-layer as your session grows, giving you the highest quality possible within your VRAM constraints. This is my dream format. Pushed to master. Available now. 1/15 🧵

Similar Articles

KVarN: Native vLLM backend for KV-cache quantization by Huawei

Hacker News Top

Huawei CSL releases KVarN, a native vLLM attention backend for KV-cache quantization that delivers 3-5x more KV-cache capacity and up to ~1.3x the throughput of FP16, with no calibration required. It claims up to ~2.4x the throughput of TurboQuant while maintaining FP16-level accuracy on models like Qwen3-32B.