attention-compensation

Tag

Cards List
#attention-compensation

SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation

arXiv cs.AI · 2026-07-21 Cached

SelKV is a training-free framework for KV cache compression that uses a soft cosine gate for selective merging and an attention-ratio compensation mechanism to correct softmax imbalance, achieving near-lossless generation at 25% cache size and 3.3x decoding speedup on LongBench.

0 favorites 0 likes
← Back to home

Submit Feedback