long-context-llm

Tag

Cards List
#long-context-llm

Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention

arXiv cs.AI · 2026-08-26 Cached

Minima-KV presents a retention-preserving KV cache compression method using mixed-format paged attention to reduce memory footprint in long-context LLM serving, with evaluated performance on benchmarks.

0 favorites 0 likes
#long-context-llm

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv cs.AI · 2026-07-08 Cached

Introduces FreqDepthKV, a frequency-guided depth sharing method for KV cache compression in long-context LLM inference, which factorizes adjacent-layer KV states into shared low-frequency components and sparse high-frequency residuals, improving memory efficiency and throughput while preserving accuracy on benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback