eviction

Tag

Cards List
#eviction

PAGE: Partition-Aware Gated KV-Cache Eviction

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper introduces PAGE, a partition-aware gated KV-cache eviction method that uses a scalar metric to predict input classes and apply eviction only when safe, reducing accuracy degradation in large language models.

0 favorites 0 likes
#eviction

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper introduces QEvict, a KV-cache management scheme for LLMs that uses recoverable quantized eviction to handle attention drift during long-context decoding, improving memory efficiency while preserving important historical context.

0 favorites 0 likes
#eviction

Back from the Future: Key-Value Cache Management by Counter-Causal Surprise

arXiv cs.LG ↗ · 2026-07-31 Cached

This paper proposes a KV cache eviction strategy that scores tokens by counter-causal surprise, removing past tokens that are well-predicted from future context. The method is training-free, in-distribution, and achieves competitive performance with a fast single-layer approximation.

0 favorites 0 likes
#eviction

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

Hugging Face Daily Papers ↗ · 2026-07-26 Cached

This paper investigates the memory contract of sparse event-KV serving, showing that evicted source events can still influence answers via cached rows, and that deliberately phrased events can enable donor-aligned recovery without naming the value.

0 favorites 0 likes
#eviction

Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets

arXiv cs.LG ↗ · 2026-06-24 Cached

Introduces Nexus Sampling, a training-free KV-cache eviction method using weighted reservoir sampling instead of deterministic top-k, improving long-context LLM inference under fixed memory budgets, matching dense attention performance at 80% eviction.

0 favorites 0 likes
#eviction

Value-Aware Stochastic KV Cache Eviction for Reasoning Models

Hugging Face Daily Papers ↗ · 2026-06-02 Cached

VaSE is a training-free method for KV cache eviction that protects large-magnitude value states and introduces stochasticity to improve reasoning model accuracy under compression, outperforming existing methods.

0 favorites 0 likes
#eviction

@no_stp_on_snek: first receipts: triattention v3 evicts safely with longctx. ✓HIT every rung 32k → 256k on qwen3.5-2b-4bit (hybrid mamba…

X AI KOLs Following ↗ · 2026-05-08

Introduces triattention v3, a new attention mechanism that enables safe eviction without recall loss for long-context inference, demonstrated on a hybrid mamba+attention model up to 256k tokens.

0 favorites 0 likes
← Back to home

Submit Feedback