超越平均注意力:基于多样性的分层KV缓存驱逐评分
摘要
本文提出了一种多样性感知的分层评分方法,用于大型语言模型中的KV缓存驱逐,通过整合注意力分散和冗余性来提升LongBench数据集上的性能。
arXiv:2609.30738v1 Announce Type: new
Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $\lambda_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary with depth, we compare fixed global coefficients with three-segment and quadratic profiles. Only these depth profiles are searched on a development split under a $\sinh$ reparameterization. On all 16 English LongBench datasets with Mistral-7B at a budget of 64 entries per layer, a single global diversification constant improves 13 of 16 datasets (macro +1.1); the gain holds at budget 32 and narrows at 128. Per-dataset search finds no detectable layer structure on most datasets; on passage retrieval it finds a large one: a mid-layer sign flip that rewards similarity and is worth +9.6 over the baseline at budget 64 and, without re-tuning, +13.2 over the global constant at budget 128. Ablations attribute the gain to the redundancy term; replaying every accepted search state on the held-out test set separates genuine structure from tuning noise.
查看缓存全文
缓存时间: 2026/09/28 09:41
# Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction Source: [https://arxiv.org/html/2609.30738](https://arxiv.org/html/2609.30738) Tianfang XieAffiliation:Georgia Institute of TechnologyAffiliation:Atlanta, GA 30332, USAEmail:[tianfang\.xie@gatech\.edu](mailto:[email protected]) ###### Abstract KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by*mean*attention over a small observation window\. We study a unified score,μi\+λ1σi\+λ2corr\(i,S\)\\mu\_\{i\}\+\\lambda\_\{1\}\\sigma\_\{i\}\+\\lambda\_\{2\}\\,\\mathrm\{corr\}\(i,S\), adding attention*dispersion*across window queries and*redundancy*relative to selected tokens\. Forλ2<0\\lambda\_\{2\}<0, the score penalizes similarity to selected tokens as in maximal marginal relevance \(MMR\), without extra forward passes\. To test whether this relevance–diversity balance should vary with depth, we compare fixed global coefficients with three\-segment and quadratic profiles\. Only these depth profiles are searched on a development split under asinh\\sinhreparameterization\. On all 16 English LongBench datasets with Mistral\-7B at a budget of 64 entries per layer, a single global diversification constant improves 13 of 16 datasets \(macro\+1\.1\+1\.1\); the gain holds at budget 32 and narrows at 128\. Per\-dataset search finds no detectable layer structure on most datasets; on passage retrieval it finds a large one: a mid\-layer sign flip that rewards similarity and is worth\+9\.6\+9\.6over the baseline at budget 64 and, without re\-tuning,\+13\.2\+13\.2over the global constant at budget 128\. Ablations attribute the gain to the redundancy term; replaying every accepted search state on the held\-out test set separates genuine structure from tuning noise\. ###### Index Terms: KV cache compression, large language models, long\-context inference, diversified selection ## 1Introduction Serving long\-context large language models \(LLMs\) is dominated by the memory footprint of the key–value \(KV\) cache, which grows linearly with context length and often exceeds the model weights\. Token eviction is a popular remedy: after prefill, each layer keeps only a small budget ofBBcached entries\. State\-of\-the\-art eviction methods such as SnapKV\[[1](https://arxiv.org/html/2609.30738#bib.bib1)\]and PyramidKV\[[2](https://arxiv.org/html/2609.30738#bib.bib2)\]score each historical token by the*mean*attention it receives from an observation window of the last few queries, and keep the top\-BB\. The mean, however, ignores two signals\. First, window queries may agree or disagree about a token; the mean discards this*dispersion*\. Second, and more importantly, top\-mean tokens are often near\-duplicates of one another\. At tight budgets, the cache then fills with*redundant*copies of one salient region while complementary evidence is evicted, the problem that maximal marginal relevance \(MMR\) addresses in information retrieval\[[3](https://arxiv.org/html/2609.30738#bib.bib3)\]\. Inspired by mean–variance portfolio selection\[[4](https://arxiv.org/html/2609.30738#bib.bib4)\], we treat cache eviction as building a portfolio of tokens and study the unified scoreμi\+λ1σi\+λ2corr\(i,S\)\\mu\_\{i\}\+\\lambda\_\{1\}\\sigma\_\{i\}\+\\lambda\_\{2\}\\,\\mathrm\{corr\}\(i,S\)with layer\-wise coefficients\. Our contributions are threefold\. \(i\) A diversified, layer\-wise scoring framework for KV eviction that generalizes the attention\-based token scoring of SnapKV/PyramidKV under a fixed budget allocation and costs no extra forward passes \(Sec\.[3](https://arxiv.org/html/2609.30738#S3)\)\. \(ii\) A comparison of a fixed global constant with two searched layer\-wise parameterizations, Segments and Curve, under a random\-split protocol in which the test set never influences any accepted search move \(Sec\.[3\.3](https://arxiv.org/html/2609.30738#S3.SS3)\)\. \(iii\) A study on all 16 English LongBench datasets that shows a fixed global setting improving 13 of 16 datasets atB=64B\{=\}64without per\-dataset search, and one large, transferable task\-specific structure, together with ablations of the score terms, results for the global constant on a second model, and a*trajectory\-replay*diagnostic that separates genuine structure from tuning noise \(Sec\.[4](https://arxiv.org/html/2609.30738#S4)\)\. ## 2Related Work KV cache compression\.StreamingLLM\[[5](https://arxiv.org/html/2609.30738#bib.bib5)\]keeps attention sinks plus a sliding window; H2O\[[6](https://arxiv.org/html/2609.30738#bib.bib6)\]and Scissorhands\[[7](https://arxiv.org/html/2609.30738#bib.bib7)\]evict by accumulated attention during decoding; FastGen\[[8](https://arxiv.org/html/2609.30738#bib.bib8)\]selects a per\-head compression policy by profiling\. Closest to us, SnapKV\[[1](https://arxiv.org/html/2609.30738#bib.bib1)\]scores prompt tokens by pooled observation\-window attention at the end of prefill, and PyramidKV\[[2](https://arxiv.org/html/2609.30738#bib.bib2)\]adds a depth\-decreasing budget schedule\. These rank tokens*independently*by \(a variant of\) mean attention; our score contains theirs as itsλ1=λ2=0\\lambda\_\{1\}\{=\}\\lambda\_\{2\}\{=\}0point\. Redundancy\-aware compression is a recent and active line: R\-KV\[[9](https://arxiv.org/html/2609.30738#bib.bib9)\]combines attention importance with key\-similarity redundancy for decoding\-time compression; MixKV\[[10](https://arxiv.org/html/2609.30738#bib.bib10)\]balances importance and diversity per head, relative to the global key distribution, for vision\-language caches; and the concurrent TwinKV\[[11](https://arxiv.org/html/2609.30738#bib.bib11)\]adds a repair pass that swaps retained near\-duplicate keys for evicted orphans\. We differ on three axes: redundancy is*conditioned on the set already retained*through a greedy MMR rule, its weight is a searched*function of depth*, and the resulting profiles are tested for transfer across budgets\. Diversified subset selection\.Selecting items that are individually relevant yet mutually complementary is classical: MMR\[[3](https://arxiv.org/html/2609.30738#bib.bib3)\]greedily penalizes similarity to the selected set, and determinantal point processes\[[12](https://arxiv.org/html/2609.30738#bib.bib12)\]model diversity probabilistically; mean–variance portfolio theory\[[4](https://arxiv.org/html/2609.30738#bib.bib4)\]makes the same trade\-off for assets\. We bring this view to KV eviction with a depth\-dependent trade\-off\. ## 3Method ### 3\.1Preliminaries At the end of prefill, SnapKV/PyramidKV use an observation windowWWcomprising the last\|W\|=8\|W\|\{=\}8queries\. For a layer with cache budgetBB, tokeniiis ranked by its mean attention, with the baselines’ max\-pooling retained in the scoring pipeline: score\(i\)=μi=1\|W\|∑q∈WAq,i\.\\mathrm\{score\}\(i\)=\\mu\_\{i\}=\\frac\{1\}\{\|W\|\}\\sum\_\{q\\in W\}A\_\{q,i\}\.\(1\)Here,Aq,iA\_\{q,i\}denotes the attention received by tokeniifrom queryqq\. The top\-BBtokens are retained\. PyramidKV decreases budgets with depth while preserving the total budget of uniform allocation\. ### 3\.2Diversity\-Aware Scoring Our*diversity\-aware, layer\-wise scoring for KV cache eviction*extends mean\-attention ranking to capture query disagreement and redundancy among retained tokens\. Inspired by mean–variance portfolio selection\[[4](https://arxiv.org/html/2609.30738#bib.bib4)\], we combine individual relevance with dispersion and set\-dependent similarity\. This analogy motivates the scoring components rather than a direct portfolio objective: attention statistics characterize each candidate, whereas key similarity relates it to the selected set\. Unified score\.For a candidateiiand the selected token setSS, we define score\(i\)=μi\+λ1σi\+λ2corr\(i,S\),\\mathrm\{score\}\(i\)=\\mu\_\{i\}\+\\lambda\_\{1\}\\sigma\_\{i\}\+\\lambda\_\{2\}\\,\\mathrm\{corr\}\(i,S\),\(2\)where attention dispersion is σi=\(1\|W\|−1∑q∈W\(Aq,i−μi\)2\)1/2,\\sigma\_\{i\}=\\left\(\\frac\{1\}\{\|W\|\-1\}\\sum\_\{q\\in W\}\(A\_\{q,i\}\-\\mu\_\{i\}\)^\{2\}\\right\)^\{1/2\},\(3\)and redundancy is measured by corr\(i,S\)=max\(0,maxj∈Scos\(ki,kj\)\),corr\(i,∅\)=0\.\\mathrm\{corr\}\(i,S\)=\\max\\Bigl\(0,\\;\\max\_\{j\\in S\}\\cos\(k\_\{i\},k\_\{j\}\)\\Bigr\),\\qquad\\mathrm\{corr\}\(i,\\emptyset\)=0\.\(4\)Here,kik\_\{i\}is the cached key vector for tokenii\. The sample standard deviationσi\\sigma\_\{i\}measures disagreement across window queries;corr\(i,S\)\\mathrm\{corr\}\(i,S\)measures the candidate’s greatest similarity to a retained token, with negative similarities clamped to zero\. Thus,λ2<0\\lambda\_\{2\}<0gives an MMR\-style penalty\[[3](https://arxiv.org/html/2609.30738#bib.bib3)\], favoring complementary evidence, whereasλ2\>0\\lambda\_\{2\}\>0rewards similarity\. Because attention magnitudes vary substantially across heads but cosine similarity lies in\[−1,1\]\[\-1,1\], we normalize the base scoresi=μi\+λ1σis\_\{i\}=\\mu\_\{i\}\+\\lambda\_\{1\}\\sigma\_\{i\}per head before greedy selection: s~i\\displaystyle\\tilde\{s\}\_\{i\}=simaxjsj,\\displaystyle=\\frac\{s\_\{i\}\}\{\\max\_\{j\}s\_\{j\}\},\(5\)S\\displaystyle S←S∪\{argmaxi∉S\(s~i\+λ2corr\(i,S\)\)\}\.\\displaystyle\\leftarrow S\\cup\\left\\\{\\arg\\max\_\{i\\notin S\}\\bigl\(\\tilde\{s\}\_\{i\}\+\\lambda\_\{2\}\\,\\mathrm\{corr\}\(i,S\)\\bigr\)\\right\\\}\.The denominator is clamped to a positive value, preserving the base\-score ranking;s~i∈\[0,1\]\\tilde\{s\}\_\{i\}\\in\[0,1\]whenλ1≥0\\lambda\_\{1\}\\geq 0\. The first token maximizess~i\\tilde\{s\}\_\{i\}, and subsequent admissions follow Eq\. \([5](https://arxiv.org/html/2609.30738#S3.E5)\) until\|S\|=B\|S\|=B\. Similarity uses keys as cached, after rotary position embedding \(RoPE\), and therefore reflects both content and relative position\. Settingλ1=λ2=0\\lambda\_\{1\}\{=\}\\lambda\_\{2\}\{=\}0recovers the baseline ranking\. The procedure uses existing attention and keys, requiring no additional forward passes; with candidates restricted to the top4B4Btokens by base score, theBBadmissions costO\(B2d\)O\(B^\{2\}d\)per head\. ### 3\.3Layer\-Wise Parameterization The coefficientsλ1\\lambda\_\{1\}andλ2\\lambda\_\{2\}control complementary aspects of selection and need not be shared across depth\. Layer\-dependent representations and PyramidKV’s nonuniform budget allocation motivate depth\-varying coefficients, but limited development data favor compact profiles over independent tuning at every layer\. Thus, we represent the coefficients with compact depth profiles\. Depth profiles\.ForL=32L\{=\}32layers, independently tuning both coefficients requires2L2Lparameters and risks overfitting 100 development questions\. We instead compare three families for the depth profilea\(ℓ\)a\(\\ell\):*Global*, a shared constant;*Segments*, separate constants for shallow, middle, and deep thirds; and*Curve*, a quadratic: a\(ℓ\)\\displaystyle a\(\\ell\)=c0\+c1t\+c2t2,t=ℓL−1,\\displaystyle=c\_\{0\}\+c\_\{1\}t\+c\_\{2\}t^\{2\},\\qquad t=\\tfrac\{\\ell\}\{L\-1\},\(6\)λ\(ℓ\)\\displaystyle\\lambda^\{\(\\ell\)\}=sinh\(a\(ℓ\)\)\.\\displaystyle=\\sinh\\\!\\bigl\(a\(\\ell\)\\bigr\)\.Layers are indexed byℓ=0,…,L−1\\ell=0,\\ldots,L\-1, and the mapping toλ\(ℓ\)\\lambda^\{\(\\ell\)\}applies separately toλ1\\lambda\_\{1\}andλ2\\lambda\_\{2\}in every family\. Segments and Curve each use three parameters per coefficient and include Global as a special case\. For Curve,c0c\_\{0\},c1c\_\{1\}, andc2c\_\{2\}control the latent profile’s height, tilt, and curvature, allowing at most one turning point\. Thesinh\\sinhmapping is unbounded and sign\-symmetric, approximately identity near zero, and satisfiescosha≥1\\cosh a\\geq 1\. It also spansλ∈\[−3\.6,3\.6\]\\lambda\\in\[\-3\.6,3\.6\]approximately overa∈\[−2,2\]a\\in\[\-2,2\]\. Search protocol\.For each dataset, a fixed\-seed random split yields a development set of 100 questions \(250 for lcc and repobench\-p\) and a held\-out test set of the remaining questions\. Only development scores determine accepted search moves; test scores serve reporting and trajectory replay \(Sec\.[4\.3](https://arxiv.org/html/2609.30738#S4.SS3)\)\. Becauseλ1\\lambda\_\{1\}reshapes base scores andλ2\\lambda\_\{2\}changes the resulting greedy selections, we optimize them jointly by first\-improvement coordinate search over the six parameters of Segments or Curve; Global keeps the fixed settingλ1=0\\lambda\_\{1\}\{=\}0,λ2=−0\.15\\lambda\_\{2\}\{=\}\-0\.15and is not searched\. This setting is also the warm start of every search\. A pilot comparingλ2∈\{−0\.15,−0\.3,−0\.6\}\\lambda\_\{2\}\\in\\\{\-0\.15,\-0\.3,\-0\.6\\\}on qasper, hotpotqa, 2wikimqa, and passage retrieval atB=64B\{=\}64fixed this constant, which is reused across all 16 datasets, budgets, and both models\. Each search allows 14 evaluations, perturbing coordinates by±0\.3\\pm 0\.3inaa\-space\. A move is accepted only when its development\-score improvement exceedsτ\\tau\. Near zero, a0\.30\.3perturbation inaacorresponds approximately to a0\.30\.3change inλ\\lambda; for the redundancy coefficient, this can shift a candidate’s normalized score by about0\.30\.3becausecos∈\[−1,1\]\\cos\\in\[\-1,1\]\. Larger steps help expose effects within the limited evaluation budget, although their effective magnitude varies withaa\. With per\-question score standard deviation≈35\{\\approx\}35, the development standard error is≈3\.5\{\\approx\}3\.5points\. We use an acceptance tolerance ofτ=0\.1\\tau\{=\}0\.1points; this is not a statistical significance threshold\. ## 4Experiments ### 4\.1Experimental Setup DatasetsWe use LongBench\[[13](https://arxiv.org/html/2609.30738#bib.bib13)\]to assess the performance of our method on tasks involving long\-context inputs\. Its 16 English datasets cover single\- and multi\-document question answering, summarization, few\-shot learning, synthetic retrieval and counting, and code completion\. Experiment settingsWe use Mistral\-7B\-Instruct\-v0\.2\[[14](https://arxiv.org/html/2609.30738#bib.bib14)\]\(fp16, SDPA\) in the PyramidKV implementation with a budget ofB=64B\{=\}64entries per layer \(≈1%\{\\approx\}1\\%of a typical prompt\) unless stated otherwise, and all 16 English LongBench\[[13](https://arxiv.org/html/2609.30738#bib.bib13)\]datasets\. Test scores use the official LongBench scorer, which truncates predictions to their first line on trec, triviaqa and samsum; search\-time dev scores use raw outputs\. All comparisons share the same total cache budget\. The full study took≈2500\{\\approx\}2500search trials and≈260\{\\approx\}260GPU\-hours on RTX 4090s\. Table 1:Test scores \(Mistral\-7B, budget 64, official LongBench scorer\)\. Bold: row best\. Wins: number of datasets on which the column is the row best \(ties credited to Seg\.\)\. ### 4\.2Main results Table[1](https://arxiv.org/html/2609.30738#S4.T1)demonstrates the effectiveness of our diversity\-aware, layer\-wise scoring for KV cache eviction atB=64B\{=\}64\. With fixed coefficientsλ1=0\\lambda\_\{1\}\{=\}0andλ2=−0\.15\\lambda\_\{2\}\{=\}\-0\.15, Global improves over PyramidKV on 13 of 16 datasets, increasing the macro score from 33\.31 to 34\.44 \(\+1\.13\+1\.13points\)\. Dataset\-specific search with Segments reaches 35\.15, whereas Curve obtains 34\.33, indicating that greater profile flexibility does not consistently improve performance\. The strongest task\-specific benefit occurs on passage retrieval: Segments achieves 70\.17, exceeding PyramidKV by 9\.57 points and Global by 3\.36 points\. These results support fixed diversification as an effective default, with additional benefits from depth\-dependent scoring on selected tasks\. Table 2:Budget sweep on test sets \(Mistral\-7B\)\. We report the macro average performance; Seg\. profiles are deployed atB=64B\{=\}64and transferred unchanged\.Table[2](https://arxiv.org/html/2609.30738#S4.T2)varies the budget with the profiles deployed atB=64B\{=\}64transferred unchanged\. Global improves the macro score over PyramidKV by 0\.73, 1\.13, and 0\.20 points atB=32B\{=\}32, 64, and 128, and Seg\. by 1\.01, 1\.84, and 0\.60 points, respectively\. The passage\-retrieval profile transfers unchanged \(Fig\.[1](https://arxiv.org/html/2609.30738#S4.F1)c\): it scores 63\.03 atB=32B\{=\}32and 81\.70 atB=128B\{=\}128, 13\.40 above PyramidKV and 13\.16 above Global at the latter budget\. Figure 1:Passage retrieval, Mistral\-7B\. \(a\) Layer\-wise coefficients of the Seg\. profile searched atB=64B\{=\}64; dashed line: Globalλ2=−0\.15\\lambda\_\{2\}\{=\}\-0\.15; vertical lines: segment boundaries\. \(b\) Development best\-so\-far and test replay of every accepted state of the search atB=64B\{=\}64\(14 evaluations,τ=0\.1\\tau\{=\}0\.1\), as change from the initialization; only development scores determine acceptance\. \(c\) Test score vs\. budget: the Seg\. profile tuned atB=64B\{=\}64is applied unchanged atB=32B\{=\}32and128128\. ### 4\.3Ablation studies and further analysis Visualization of passage retrieval results\.Fig\.[1](https://arxiv.org/html/2609.30738#S4.F1)shows the passage\-retrieval profile \(a\), its search trajectory \(b\), and its budget transfer \(c\)\. The profile keepsλ2<0\\lambda\_\{2\}<0in shallow layers but turns it positive in the middle and deep layers\. In the replay, the two acceptedλ1\\lambda\_\{1\}updates do not improve the test score, whereas theλ2\\lambda\_\{2\}sign flip yields a large gain, and the profile tuned atB=64B\{=\}64exceeds both PyramidKV and Global at every budget\. Scoring ruleMacroWinsMean \(baseline\)33\.31—Mean & dispersion33\.508/16Mean & Corr34\.5713/16dispersion & Corr \(fixedλ\\lambda\)33\.169/16All three terms \(ours, 3 segments\)35\.1514/16All three terms, 2 segments34\.7514/16All three terms, 8 segments35\.0314/16 Table 3:Ablation studies of our framework \(Mistral\-7B, budget 64\), with the Seg\. strategy unless noted\. We report the macro average scores \(Macro\) and the number of datasets improved over the baseline \(Wins\)\.Ablation study of our framework\.Table[3](https://arxiv.org/html/2609.30738#S4.T3)removes one score term at a time and varies the depth granularity\. Mean attention is the backbone: a fixed rule without it \(dispersion & Corr,λ1=0\.3\\lambda\_\{1\}\{=\}0\.3,λ2=−0\.15\\lambda\_\{2\}\{=\}\-0\.15, no search\) scores 33\.16, below the 33\.31 baseline, and drops passage retrieval to 57\.41\. The redundancy term carries most of the gain: searchingλ2\\lambda\_\{2\}alone \(Mean & Corr\) reaches 34\.57 and improves 13 datasets, whereas searchingλ1\\lambda\_\{1\}alone \(Mean & dispersion\) reaches 33\.50 and improves 8\. Searching both reaches 35\.15; on passage retrieval the joint search scores 70\.17 against 66\.62 forλ2\\lambda\_\{2\}alone in this round\. Three segments score highest \(35\.15\); eight segments \(35\.03; 16 parameters, 34 evaluations\) and two \(34\.75\) stay within the noise floor of it, so the segment count is not a sensitive choice\. Evaluation on other LLM backbones\.To check that the fixed rule transfers across models, we apply Global \(λ1=0\\lambda\_\{1\}\{=\}0,λ2=−0\.15\\lambda\_\{2\}\{=\}\-0\.15, no re\-tuning\) to Llama\-3\.1\-8B\-Instruct atB=64B\{=\}64\. qasper improves by\+5\.8\+5\.8\(24\.4→30\.124\.4\\to 30\.1\) and hotpotqa by\+2\.0\+2\.0\(51\.1→53\.151\.1\\to 53\.1\), while the other datasets stay within the noise floor; passage retrieval is already at99\.099\.0on this model, leaving limited headroom for a depth profile\. Per\-dataset search on this model is left for future work\. Discussion on efficiency and search cost\.The greedy admission runs once per prompt at the end of prefill and adds no forward passes\. On 12 HotpotQA prompts \(mean 15\.6k tokens, RTX 4090\) selection takes 222 ms per prompt versus 32 ms for top\-kk, increasing end\-to\-end latency for prefill plus 32\-token generation by about 14% \(2\.52 s→\\to2\.87 s\); the overhead comes from performing theBBadmissions sequentially in each layer\. Global needs no per\-dataset search\. A three\-segment profile costs 14 development evaluations per dataset, about 1–3 GPU\-hours on one RTX 4090, and transfers across budgets without re\-tuning \(Table[2](https://arxiv.org/html/2609.30738#S4.T2)\)\. Limitations\.Scores are single\-run on one benchmark suite with test sets of 50–250 questions per dataset, so differences of a few points on a single dataset should be interpreted with caution; greedy search is path\-dependent, and per\-dataset profiles need labeled development data\. ## 5Conclusion Diversified, layer\-wise scoring generalizes mean\-attention token scoring for KV eviction without additional forward passes\. On Mistral\-7B a single global diversification constant is a robust default at small cache budgets, and one task, passage retrieval, rewards a layer\-wise profile with a large gain that survives changes of budget\. Ablations place the benefit in the redundancy term computed on cached keys, and a trajectory\-replay diagnostic makes such findings verifiable\. Next steps are instance\-adaptive coefficients\. ## 6Acknowledgments This work was self\-funded, with no external funding\. The authors declare no conflicts of interest\. GPT \(OpenAI\) was used to polish text and review code\. The authors checked all AI\-assisted content and take full responsibility for it\. ## 7Compliance with Ethical Standards This computational study used the publicly available LongBench benchmark\[[13](https://arxiv.org/html/2609.30738#bib.bib13)\]\. No new data were collected, and no human or animal subjects were involved; therefore, ethical approval was not required\. ## References - \[1\]Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen,“Snapkv: LLM knows what you are looking for before generation,”arXiv preprint arXiv:2404\.14469, 2024\. - \[2\]Zefan Cai, Yefan Zhang, Bofei Gao, Yude Liu, et al\.,“Pyramidkv: Dynamic KV cache compression based on pyramidal information funneling,”arXiv preprint arXiv:2406\.02069, 2024\. - \[3\]Jaime Carbonell and Jade Goldstein,“The use of MMR, diversity\-based reranking for reordering documents and producing summaries,”inProceedings of SIGIR, 1998\. - \[4\]Harry Markowitz,“Portfolio selection,”The Journal of Finance, vol\. 7, no\. 1, pp\. 77–91, 1952\. - \[5\]Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis,“Efficient streaming language models with attention sinks,”arXiv preprint arXiv:2309\.17453, 2023\. - \[6\]Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen,“H2O: Heavy\-hitter oracle for efficient generative inference of large language models,”Advances in Neural Information Processing Systems, 2023\. - \[7\]Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava,“Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,”Advances in Neural Information Processing Systems, 2023\. - \[8\]Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao,“Model tells you what to discard: Adaptive KV cache compression for LLMs,”arXiv preprint arXiv:2310\.01801, 2023\. - \[9\]Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhang, Yuxin Bai, Chao Han, Zhao Zhang, Yun Meng, et al\.,“R\-KV: Redundancy\-aware KV cache compression for reasoning models,”inAdvances in Neural Information Processing Systems, 2025\. - \[10\]Xuyang Liu et al\.,“Mixing importance with diversity: Joint optimization for KV cache compression in large vision\-language models,”inInternational Conference on Learning Representations, 2026\. - \[11\]Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng, Junyan Zhang, and Xuming Hu,“TwinKV: A composable repair pass for KV cache eviction via pairwise key redundancy,”arXiv preprint arXiv:2608\.27128, 2026\. - \[12\]Alex Kulesza and Ben Taskar,“Determinantal point processes for machine learning,”Foundations and Trends in Machine Learning, vol\. 5, no\. 2–3, 2012\. - \[13\]Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al\.,“Longbench: A bilingual, multitask benchmark for long context understanding,”inProceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\), 2024, pp\. 3119–3137\. - \[14\]Albert Q\. Jiang, Alexandre Sablayrolles, Arthur Mensch, et al\.,“Mistral 7B,”arXiv preprint arXiv:2310\.06825, 2023\.
相似文章
Random Attention:重新思考高效推理的KV缓存驱逐策略
本文提出了一种名为Random Attention的KV缓存驱逐方法,它使用随机选择而非评分,匹配选择性方法的同时提高推理任务中的吞吐量。
基于顿悟感知的KV缓存淘汰方法(无需注意力矩阵)
本文介绍了EpiKV,一种基于内部表征变化(顿悟分数)而非注意力权重来评估token重要性的KV缓存淘汰方法,无需具体化注意力矩阵。该方法在推理基准测试中取得了具有竞争力的性能,同时支持长达16倍的上下文长度。
LKV:通过端到端学习多头预算与 Token 选择优化大模型 KV 缓存淘汰机制
本文提出了 LKV,这是一种通过端到端学习基于 Attention Head 的预算分配与 Token 选择策略来优化大语言模型 KV 缓存淘汰的方法,在实现高压缩率的同时取得了最先进的性能表现。
ReST-KV:基于逐层输出重构与时空平滑的鲁棒 KV Cache 驱逐方法
本文介绍了 ReST-KV,一种用于大型语言模型的新型鲁棒 KV Cache 驱逐方法。该方法利用逐层输出重构与时空平滑技术来提升效率,显著降低了解码延迟,并在 LongBench 和 RULER 等长上下文基准测试中超越了现有的最先进基线模型。
信任质量:KV-Cache淘汰中的强制权重
本文分析了稀疏注意力模型中的KV-Cache淘汰策略,表明选择最大权重近乎最优,且已发表的差距源于内存和查询信息,其中ContourKV取得了强劲性能。