transform-coding

Tag

Cards List
#transform-coding

KV Cache Compression Through the Lens of Transform Coding

arXiv cs.LG ↗ · 2026-08-17 Cached

The paper proposes Attention-Aware Transform Coding (AATC) for compressing KV caches in large language models, achieving near-lossless accuracy at around 5.8x compression by minimizing distortion through attention mechanisms.

0 favorites 0 likes
#transform-coding

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper proposes NOVA-KV, a transform-coding approach to KV cache quantization that uses attention-preserving transforms to allocate bits where queries actually attend, improving long-context retrieval accuracy at low bit rates compared to prior methods.

0 favorites 0 likes
← Back to home

Submit Feedback