kv-cache-eviction

Tag

Cards List
#kv-cache-eviction

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference

arXiv cs.CL · 2026-05-11 Cached

This paper introduces LaProx, a novel KV Cache eviction strategy for long-context LLM inference that reformulates the problem as an output-aware matrix multiplication approximation, achieving high performance with only 5% cache usage.

0 favorites 0 likes
← Back to home

Submit Feedback