Tag
This paper audits six KV-cache compression methods under query-agnostic protocols, finding that rankings change dramatically compared to query-aware evaluations, with implications for cache reuse in long-context inference.