cache-compression

Tag

Cards List
#cache-compression

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

arXiv cs.AI · 4d ago Cached

TaskPress introduces a query-agnostic KV cache compression framework that uses a task guide as a meta-query and quantization scale factors to prune irrelevant tokens, enabling reusable caches across diverse queries with negligible overhead.

0 favorites 0 likes
#cache-compression

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Hugging Face Daily Papers · 2026-07-06 Cached

KVpop introduces a learned KV cache eviction policy supervised by future-attention targets, achieving high compression rates (e.g., 98% performance at 75% compression) on Qwen3 models while maintaining quality.

0 favorites 0 likes
← Back to home

Submit Feedback