cache-compaction

标签

Cards List
#cache-compaction

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

arXiv cs.CL · 2026-08-04 缓存

This empirical study examines practical online KV cache compaction for LLM agents, comparing token eviction and attention matching methods under different proxy query sources. It finds that delaying compaction to use future agent queries recovers performance, and token eviction preserves accuracy while reducing KV cache by 80%.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈