cache-compaction

Tag

Cards List
#cache-compaction

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

arXiv cs.CL · 2026-08-04 Cached

This empirical study examines practical online KV cache compaction for LLM agents, comparing token eviction and attention matching methods under different proxy query sources. It finds that delaying compaction to use future agent queries recovers performance, and token eviction preserves accuracy while reducing KV cache by 80%.

0 favorites 0 likes
← Back to home

Submit Feedback