la-approx

Tag

Cards List
#la-approx

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference

arXiv cs.CL ↗ · 2026-05-11 Cached

This paper introduces LaProx, a novel KV Cache eviction strategy for long-context LLM inference that reformulates the problem as an output-aware matrix multiplication approximation, achieving high performance with only 5% cache usage.

0 favorites 0 likes
← Back to home

Submit Feedback