Tag
Minima-KV presents a retention-preserving KV cache compression method using mixed-format paged attention to reduce memory footprint in long-context LLM serving, with evaluated performance on benchmarks.