key-value-cache

Tag

Cards List
#key-value-cache

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

arXiv cs.AI ↗ · 2026-08-25 Cached

KVBoost is a chunk-level key-value cache reuse system for efficient large language model inference that achieves high cache hit rates and significant speedup in time-to-first-token without quality loss, using dual-hash keying and deviation-guided recomputation.

0 favorites 0 likes
#key-value-cache

MaSRead: Content-Addressed Reading of Replicated Latent Stores

arXiv cs.AI ↗ · 2026-08-13 Cached

Introduces MaSRead, a content-addressed reading mechanism for replicated latent stores where agents share KV cache fragments, enabling later queries to reliably retrieve cached reasoning via opaque keyed tag sets and hard attention masks.

0 favorites 0 likes
#key-value-cache

Addressable Memory for Video World Models

Hugging Face Daily Papers ↗ · 2026-08-07 Cached

This paper introduces WorldTrace, a training-free memory framework for long-horizon video world models that keeps compressed cache addressable, plus LoopBench, a benchmark for episodic recall after long detours. It improves temporal consistency by +15.5% and episodic recall by +19.5% on LoopBench.

0 favorites 0 likes
#key-value-cache

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Hugging Face Daily Papers ↗ · 2026-07-06 Cached

KVpop introduces a learned KV cache eviction policy supervised by future-attention targets, achieving high compression rates (e.g., 98% performance at 75% compression) on Qwen3 models while maintaining quality.

0 favorites 0 likes
#key-value-cache

Information-Aware KV Cache Compression for Long Reasoning

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes InfoKV, an entropy-aware KV cache compression framework that combines token-level predictive uncertainty with attention scores to improve long-context reasoning efficiency. Experiments show it outperforms existing attention-based methods on Llama-3.1, Llama-3.2, and DeepSeek-R1.

0 favorites 0 likes
#key-value-cache

I'm still surprised on how good the kv quantization has become

Reddit r/LocalLLaMA ↗ · 2026-06-15

The author expresses surprise at how effective key-value cache quantization (q4_0) remains even with large context windows, citing accurate retrieval from a 100k context.

0 favorites 0 likes
#key-value-cache

@jiqizhixin: What if your AI’s memory didn’t have to balloon with every extra sentence? University of Oxford, Technion, AITHYRA, and…

X AI KOLs Timeline ↗ · 2026-06-14 Cached

Introduces KV-Compression Aware Training (KV-CAT), a method that encourages transformers to learn compressible key-value caches during training, improving memory efficiency for long-context tasks without sacrificing performance.

0 favorites 0 likes
#key-value-cache

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

arXiv cs.LG ↗ · 2026-05-15 Cached

Introduces Self-Pruned Key-Value Attention (SP-KV), a mechanism that learns to predict future utility of key-value pairs to dynamically prune the KV cache, reducing memory usage and decoding speed by 3-10x with minimal performance degradation. The model and utility predictor are trained end-to-end using next-token prediction.

0 favorites 0 likes
← Back to home

Submit Feedback