Reformulating KV Cache Eviction Problem for Long-Context LLM Inference

arXiv cs.CL Papers

Summary

This paper introduces LaProx, a novel KV Cache eviction strategy for long-context LLM inference that reformulates the problem as an output-aware matrix multiplication approximation, achieving high performance with only 5% cache usage.

arXiv:2605.07234v1 Announce Type: new Abstract: Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth. Existing KV Cache eviction methods primarily rely on local attention weights, neglecting the influence of value representations, output projection, and inter-head interactions. In this work, we reformulate KV Cache eviction from a conventional head-wise, weight-averaging approach into an output-aware, layer-wise matrix multiplication approximation problem. We introduce LaProx, a novel eviction strategy that explicitly models the multiplicative interaction between attention maps and projected value states to accurately quantify token contributions while accounting for inter-head dependencies. Building on this metric, we propose the first unified eviction strategy that assigns globally comparable importance scores to tokens, enabling model-wide selection instead of local, head-wise decisions. Experimental results across 19 datasets on long-context benchmarks LongBench and Needle-In-A-Haystack demonstrate that our approach maintains model performance with only 5\% of the KV cache and consistently outperforms prior works across all configurations. Notably, our method achieves up to 2$\times$ accuracy loss reduction under extreme compression scenarios compared to existing state-of-the-art baselines with minimal overhead.
Original Article
View Cached Full Text

Cached at: 05/11/26, 06:55 AM

# Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
Source: [https://arxiv.org/html/2605.07234](https://arxiv.org/html/2605.07234)
Tho Mai KAIST Daejeon, South Korea thomh1511@kaist\.ac\.kr &Joo\-Young Kim KAIST Daejeon, South Korea jooyoung1203@kaist\.ac\.kr

###### Abstract

Large language models \(LLMs\) support long\-context inference but suffer from substantial memory and runtime overhead due to Key\-Value \(KV\) Cache growth\. Existing KV Cache eviction methods primarily rely on local attention weights, neglecting the influence of value representations, output projection, and inter\-head interactions\. In this work, we reformulate KV Cache eviction from a conventional head\-wise, weight\-averaging approach into an output\-aware, layer\-wise matrix multiplication approximation problem\. We introduce LaProx, a novel eviction strategy that explicitly models the multiplicative interaction between attention maps and projected value states to accurately quantify token contributions while accounting for inter\-head dependencies\. Building on this metric, we propose the first unified eviction strategy that assigns globally comparable importance scores to tokens, enabling model\-wide selection instead of local, head\-wise decisions\. Experimental results across 19 datasets on long\-context benchmarks LongBench and Needle\-In\-A\-Haystack demonstrate that our approach maintains model performance with only 5% of the KV cache and consistently outperforms prior works across all configurations\. Notably, our method achieves up to 2×\\timesaccuracy loss reduction under extreme compression scenarios compared to existing state\-of\-the\-art baselines with minimal overhead\.

## 1Introduction

Recent advances in large language models \(LLMs\) have significantly extended their ability to process long contexts, enabling strong performance in applications such as multi\-turn dialogue[undefam](https://arxiv.org/html/2605.07234#bib.bib40), question answering[undefx](https://arxiv.org/html/2605.07234#bib.bib25), code generation[undefq](https://arxiv.org/html/2605.07234#bib.bib18), and document understanding[undefaw](https://arxiv.org/html/2605.07234#bib.bib50)\. To accelerate autoregressive inference, transformers cache key and value states from previous tokens, avoiding repeated attention computation\. While this Key\-Value \(KV\) cache is essential for efficient decoding, its size grows linearly with context length, quickly becomes a major bottleneck for memory usage and decoding latency in long\-context settings\. While techniques such as head merging or architectural modifications[undefa](https://arxiv.org/html/2605.07234#bib.bib2)can partially alleviate these costs during training, they are often incompatible with fixed, pretrained models commonly used in deployment\. Consequently, managing KV cache efficiently at inference time—without retraining or altering model parameters—becomes a critical challenge for scalable and cost\-effective long\-context LLM deployment under realistic memory and hardware constraints[undefaz](https://arxiv.org/html/2605.07234#bib.bib53)\.

To operate large language models under constrained memory budgets, a common strategy is to dynamically reduce the size of the key\-value \(KV\) cache by evicting entries deemed less influential during inference\. Prior work has shown that, in practice, only a small subset of cached tokens meaningfully contributes to the attention output[undefay](https://arxiv.org/html/2605.07234#bib.bib52);[undefap](https://arxiv.org/html/2605.07234#bib.bib43), motivating a class of eviction\-based methods that selectively retain critical entries while discarding the rest\. Early approaches exploit empirical observations that attention weights are highly concentrated, whereby a minority of tokens consistently receives the majority of attention mass\. Building on this phenomenon, several methods identify important cache entries by averaging attention scores over time, with later refinements introducing observation windows, pooling mechanisms[undefac](https://arxiv.org/html/2605.07234#bib.bib30)or adaptive budget allocation[undefag](https://arxiv.org/html/2605.07234#bib.bib34)to better preserve salient information\. However, such strategies are often heuristics and lack a principled formulation of what constitutes cache entry criticality\. Consequently, the precise relationship between attention behavior, value representations, and their joint impact on the final model output remains insufficiently characterized\.

In this paper, we reformulate KV cache eviction as an optimization problem that preserves layers’ attention outputs under a fixed budget\. By explicitly modeling the output as a product of attention, value, and output matrix, we move beyond conventional attention\-only heuristics, allowing us to rank cache entries by their actual contribution to the multiplicative interactions that form the final layer output\. Critically, this formulation reveals that token importance is fundamentally coupled to the aggregate representation formed within each layer and, consequently, to the model’s eventual output\. This observation suggests that eviction is most effectively managed at the model level rather than through isolated, head\-wise decisions\. Based on this insight, we propose a novel eviction strategy enabling more effective global cache selection\. Our contributions are summarized as follows:

1\. We demonstrate that attention weights alone provide an incomplete measure of token importance, and accurate selection must account for output information and attention layer’s structure itself\.

2\. We reveal that existing independent head\-wise eviction is suboptimal because it neglects inter\-head and inter\-layer interactions, and show that eviction should be done at the model\-level\.

3\. We introduceLayerApproximated Cache \(LaProx\), a new eviction strategy that approximates layer’s output by evaluating tokens across heads and layers simultaneously without any calibration\.

4\. Extensive evaluations on long\-context benchmarks demonstrate that the proposed method consistently outperforms attention\-based eviction strategies, confirming the effectiveness of our proposal\.

## 2Background and Related Works

### 2\.1Basic of Attention and KV Cache Operations

For clarity, we describe the mechanism using Multi\-Head Attention \(MHA\) and omit the layer index, noting that the formulation applies identically to all transformer attention layers\. Let𝐗∈ℝS×D\\mathbf\{X\}\\in\\mathbb\{R\}^\{S\\times D\}be the token embeddings of a sequence of lengthSS, whereDDis the model hidden dimension\. Each attention head operates on a subspace of dimensiondhd\_\{h\}, withD=H⋅dhD=H\\cdot d\_\{h\}forHHheads\. The projection matrices𝐖Q\(h\),𝐖K\(h\),𝐖V\(h\)∈ℝD×dh\\mathbf\{W\}\_\{Q\}^\{\(h\)\},\\mathbf\{W\}\_\{K\}^\{\(h\)\},\\mathbf\{W\}\_\{V\}^\{\(h\)\}\\in\\mathbb\{R\}^\{D\\times d\_\{h\}\}map the shared hidden representations into head\-specific query, key, and value states\. During prompt processing, each head computes

𝐐\(𝐡\)=𝐗𝐖Q\(h\),𝐊\(𝐡\)=𝐗𝐖K\(h\),𝐕\(𝐡\)=𝐗𝐖V\(h\)\\mathbf\{Q^\{\(h\)\}\}=\{\\mathbf\{X\}\\mathbf\{W\}\_\{Q\}^\{\(h\)\}\},\\mathbf\{K^\{\(h\)\}\}=\{\\mathbf\{X\}\\mathbf\{W\}\_\{K\}^\{\(h\)\}\},\\mathbf\{V^\{\(h\)\}\}=\{\\mathbf\{X\}\\mathbf\{W\}\_\{V\}^\{\(h\)\}\}\(1\)with attention weights

𝐀\(𝐡\)=Softmax⁡\(Q\(h\)​K\(h\)⊤dh\)\\mathbf\{A^\{\(h\)\}\}=\\operatorname\{Softmax\}\\left\(\\frac\{Q^\{\(h\)\}\{K^\{\(h\)\}\}^\{\\top\}\}\{\\sqrt\{d\_\{h\}\}\}\\right\)\(2\)The per\-head attention outputs are then concatenated,

𝐀𝐕=Concat⁡\(𝐀\(1\)​𝐕\(1\),…,𝐀\(H\)​𝐕\(H\)\)\\mathbf\{AV\}=\\operatorname\{Concat\}\(\\mathbf\{A\}^\{\(1\)\}\\mathbf\{V\}^\{\(1\)\},\\dots,\\mathbf\{A\}^\{\(H\)\}\\mathbf\{V\}^\{\(H\)\}\)\(3\)and projected to produce the final attention output,

𝐎=𝐀𝐕𝐖O\\mathbf\{O\}=\\mathbf\{AV\}\\mathbf\{W\}\_\{O\}\(4\)Following the projection𝐖O\\mathbf\{W\}\_\{O\}, the final layer output is integrated via a residual connection:

𝐘=Norm⁡\(𝐎\+𝐗\)\\mathbf\{Y\}=\\operatorname\{Norm\}\(\\mathbf\{O\}\+\\mathbf\{X\}\)\(5\)where𝐗\\mathbf\{X\}is the input identity andNorm\\operatorname\{Norm\}denotes a normalization function\.

During autoregressive decoding, at each decoding stepii, only the newly generated token embedding𝐱𝐢∈ℝ1×D\\mathbf\{x\_\{i\}\}\\in\\mathbb\{R\}^\{1\\times D\}is projected to obtain its head\-wise query, key, and value states\. To avoid recomputation of past tokens, the new key\-value pairs are appended to the cache

𝐊\(h\)←Concat⁡\(𝐊\(h\),𝐱i​𝐖K\(h\)\),𝐕\(h\)←Concat⁡\(𝐕\(h\),𝐱i​𝐖V\(h\)\)\\mathbf\{K\}^\{\(h\)\}\\leftarrow\\operatorname\{Concat\}\(\\mathbf\{K\}^\{\(h\)\},\\mathbf\{x\}\_\{i\}\\mathbf\{W\}\_\{K\}^\{\(h\)\}\),\\qquad\\mathbf\{V\}^\{\(h\)\}\\leftarrow\\operatorname\{Concat\}\(\\mathbf\{V\}^\{\(h\)\},\\mathbf\{x\}\_\{i\}\\mathbf\{W\}\_\{V\}^\{\(h\)\}\)\(6\)and the query𝐪𝐢\(𝐡\)=𝐱𝐢​𝐖Q\(h\)\\mathbf\{q\_\{i\}^\{\(h\)\}\}=\{\\mathbf\{x\_\{i\}\}\\mathbf\{W\}\_\{Q\}^\{\(h\)\}\}attends over the cached keys using equation[2](https://arxiv.org/html/2605.07234#S2.E2)\.

While KV caching significantly reduces computation during decoding, the cache grows linearly with sequence length, leading to substantial memory and attention overhead in long\-context inference\.

### 2\.2KV Cache Eviction

KV cache eviction during inference reduces memory and computational overhead without modifying the attention mechanism\. Its objective is to retain important tokens while removing low\-impact ones\. Early methods, such as StreamingLLM[undefar](https://arxiv.org/html/2605.07234#bib.bib45), adopt window\-based strategies that preserve attention sinks and recent tokens while LongFormer[undefc](https://arxiv.org/html/2605.07234#bib.bib4)uses two types of sliding windows cooperating with some pre\-selected input locations\. While efficient, these approaches may discard informative tokens in the middle of long sequences, degrading long\-context performance\. Other works, including H2O[undefay](https://arxiv.org/html/2605.07234#bib.bib52)and Scissorhands[undefae](https://arxiv.org/html/2605.07234#bib.bib32), rank KV entries using accumulated attention scores to better capture token importance\. Building on this line of work, SnapKV[undefac](https://arxiv.org/html/2605.07234#bib.bib30)and CAKE[undefag](https://arxiv.org/html/2605.07234#bib.bib34)further improve performance by averaging attention within an observation window and applying a pooling operation, achieving state\-of\-the\-art \(SOTA\) results\.

Beyond token selection, several studies explore non\-uniform cache budget allocation\. Layer\-wise approaches such as PyramidInfer[undefat](https://arxiv.org/html/2605.07234#bib.bib47)and PyramidKV[undefd](https://arxiv.org/html/2605.07234#bib.bib5)assign budgets based on network depth, while D2O[undefao](https://arxiv.org/html/2605.07234#bib.bib42)and CAKE[undefag](https://arxiv.org/html/2605.07234#bib.bib34)adjust cache sizes using layer\-specific attention variance\. At the head level, AdaKV[undefj](https://arxiv.org/html/2605.07234#bib.bib11)applies top\-k selection across head\-scores with an empirical safeguard, whereas HeadKV[undefl](https://arxiv.org/html/2605.07234#bib.bib13)uses calibration procedures to determine fixed per\-head budgets prior to inference\.

A few works go beyond attention scores\. For example, LAVa[undefai](https://arxiv.org/html/2605.07234#bib.bib36)and CAOTE[undefn](https://arxiv.org/html/2605.07234#bib.bib15)leverage value representations in their eviction indicators but omit the output projection; meanwhile, CriticalKV[undefk](https://arxiv.org/html/2605.07234#bib.bib12)relies on two empirical safeguards to rescale the mean attention scores with output information, disregarding the actual formulation of the attention layer\.

Despite their competitive results, existing methods rely primarily on attention weights for both eviction and budget allocation, or heuristically leverage output informationwithout considering the actual layer’s formulation\.Furthermore, these approaches are limited to performing eviction on a per\-head basis,neglecting cross\-head and cross\-layer interactions\.In contrast, this work proposes a principled eviction criterion that incorporates both attention probabilities andV​WOVW\_\{O\}contributions, and the cross\-heads interaction, providing a more accurate measure of token importance\.

## 3Motivation

![Refer to caption](https://arxiv.org/html/2605.07234v1/x1.png)\(a\)AAandV​WOVW\_\{O\}patterns\.
![Refer to caption](https://arxiv.org/html/2605.07234v1/x2.png)\(b\)Average strength\.

Figure 1:Pattern and magnitude ofAAandV​WOVW\_\{O\}\.In this section, we investigate the relationship between attention weight \(AA\) and the value–output projection \(V​WOVW\_\{O\}\)\. Specifically, we examine whether the average ofAAalone can serve as a faithful proxy for the attention layer output, i\.e\., whether attention weightsAAare sufficient to characterize the whole productA​V​WOAVW\_\{O\}\. This approach assumes two key conditions are satisfied: \(1\) the patterns ofAAandV​WOVW\_\{O\}are well aligned, and \(2\) the magnitude ofAAis not dominated byV​WOVW\_\{O\}\.

Experiment setup\.Our analysis is conducted using the Mistral\-7B\-Instruct\-v0\.3 model\. For visualization clarity, we display only a contiguous subset of tokens\.

Observation\.Figure[1\(a\)](https://arxiv.org/html/2605.07234#S3.F1.sf1)reports the normalized per\-token magnitudes of\|A\|\|A\|and\|V​WO\|\|VW\_\{O\}\|\. While the two quantities share some high\-score tokens \(such as \#25 or \#49\-50\), their overall patterns differ significantly\. Many steps even show opposite peaks; for instance, tokens \#37 and \#39 have highV​WOVW\_\{O\}values but lowAAvalues\. This indicates thatAAandV​WOVW\_\{O\}assess token importance differently, and one cannot be used in place of the other\.

Furthermore, Figure[1\(b\)](https://arxiv.org/html/2605.07234#S3.F1.sf2)reveals that the value range ofAAis much smaller thanV​WOVW\_\{O\}\. As attention weights are normalized probabilities, their values are confined to a narrow range, whereasV​WOVW\_\{O\}has a much wider range of values, which expands in deeper layers\.

These observations demonstrate that attention weights alone are insufficient to represent the attention layer output, motivating the incorporation of value and output projection in cache eviction decisions\.

## 4Methodology

### 4\.1Eviction Indicator

Equations[3](https://arxiv.org/html/2605.07234#S2.E3)and[4](https://arxiv.org/html/2605.07234#S2.E4)show that the standard MHA is defined as the concatenation of all head outputs followed by a linear projection\. Although the output projection mixes attention information from all heads, the computation can be exactly decomposed into a sum of independent head\-wise contributions\.

Algorithm 1Eviction Score ComputationInput:Query

𝐐\\mathbf\{Q\}, KV Cache

\(𝐊,𝐕\)\(\\mathbf\{K\},\\mathbf\{V\}\), Projection

WOW\_\{O\}, Budget

Bt​o​t​a​lB\_\{total\}, Observation Window

ww
Output:Compressed KV cache

\(𝐊~,𝐕~\)\(\\tilde\{\\mathbf\{K\}\},\\tilde\{\\mathbf\{V\}\}\)
// Compute attention weight and projected values

𝐀←Softmax⁡\(𝐐\[−𝐰:,\]𝐊⊤dk\)\\mathbf\{A\}\\leftarrow\\operatorname\{Softmax\}\\\!\\left\(\\frac\{\\mathbf\{Q\[\-w:,\]\}\\mathbf\{K\}^\{\\top\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\)

𝐇←V​WO\\mathbf\{H\}\\leftarrow VW\_\{O\}

// Score tokens

T←T\\leftarrownumber of cached tokens

for

i=0i=0to

TTdo

if

i<T−wi<T\-wthen

𝐩​\[𝐢\]←‖𝐀​\[:,𝐢\]‖2⋅‖𝐇​\[𝐢,:\]‖2\\mathbf\{p\[i\]\}\\leftarrow\\left\\\|\\mathbf\{A\[:,i\]\}\\right\\\|\_\{2\}\\cdot\\left\\\|\\mathbf\{H\[i,:\]\}\\right\\\|\_\{2\}

else

𝐩​\[𝐢\]←∞\\mathbf\{p\[i\]\}\\leftarrow\\infty

endif

endfor

// Evict tokens

𝒮←TopK⁡\(𝐩,Bt​o​t​a​l\)\\mathcal\{S\}\\leftarrow\\operatorname\{TopK\}\(\\mathbf\{p\},B\_\{total\}\)

\(𝐊~,𝐕~\)←\(𝐊​\[𝒮\],𝐕​\[𝒮\]\)\(\\tilde\{\\mathbf\{K\}\},\\tilde\{\\mathbf\{V\}\}\)\\leftarrow\(\\mathbf\{K\}\[\\mathcal\{S\}\],\\mathbf\{V\}\[\\mathcal\{S\}\]\)

return

\(𝐊~,𝐕~\)\(\\tilde\{\\mathbf\{K\}\},\\tilde\{\\mathbf\{V\}\}\)

Remark[4\.1](https://arxiv.org/html/2605.07234#S4.Thmtheorem1)shows thatby integratingV​WO\\boldsymbol\{VW\_\{O\}\}at the head level, we can evaluate token importance across the entire layer\. To quantify token importance, we leverage matrix multiplication associativity and computeV​WOVW\_\{O\}first to preserve the key\-value alignment between the attention matrixAAand the projected valuesV​WOVW\_\{O\}\. A naive but limited approach of independentlyscaling the average attention scores by the magnitude ofV​WOVW\_\{O\}will overlook the fundamental nature of the attention output,which is formed through a matrix multiplication\. Since our objective is to preserve this layer attention output under a constrained KV cache budget, we instead treat cache eviction as a matrix multiplication approximation problem: selecting a subset of tokens that best approximates the full productA×\(V​WO\)A\\times\(VW\_\{O\}\)\. From this perspective, we draw on Monte Carlo analyses of matrix multiplication to provide a rigorous mathematical basis for our eviction criteria\. This theory states that the approximation error of a matrix product is minimized when the selection follows the product of the Euclidean norms of the corresponding column\-row pairs[undefh](https://arxiv.org/html/2605.07234#bib.bib9);[undef](https://arxiv.org/html/2605.07234#bib.bib1)\.

pi=‖𝑨​\[:,i\]‖2​‖𝑽​𝑾𝑶​\[i,:\]‖2∑jS‖𝑨​\[:,j\]‖2​‖𝑽​𝑾𝑶​\[j,:\]‖2p\_\{i\}=\\frac\{\\\|\\boldsymbol\{A\}\[:,i\]\\\|\_\{2\}\\ \\\|\\boldsymbol\{VW\_\{O\}\}\[i,:\]\\\|\_\{2\}\}\{\\sum\_\{j\}^\{S\}\\\|\\boldsymbol\{A\}\[:,j\]\\\|\_\{2\}\\ \\\|\\boldsymbol\{VW\_\{O\}\}\[j,:\]\\\|\_\{2\}\}\(8\)More generally, the score can be represented by:

pi∝‖𝑨​\[:,i\]‖2​‖𝑽​𝑾​𝒐​\[i,:\]‖2p\_\{i\}\\propto\\\|\\boldsymbol\{A\}\[:,i\]\\\|\_\{2\}\\ \\\|\\boldsymbol\{VWo\}\[i,:\]\\\|\_\{2\}\(9\)Equation[9](https://arxiv.org/html/2605.07234#S4.E9)shows the eviction score𝐩𝐢\\mathbf\{p\_\{i\}\}per tokeniiper head\. Intuitively, this criterion favors indices that simultaneously carry significant mass in both matrices, which are the terms that dominate the output\. In this work, we select top indices with the highest scores given by equation[9](https://arxiv.org/html/2605.07234#S4.E9)to ensure that the most influential tokens are preserved\. By grounding our eviction indicator in this matrix approximation principle, we obtain a token importance measure that directlyaligns with the structure of the attention computationand more faithfully preserves the layer output compared to heuristics based solely on attention weights\. The detailed algorithm is shown in Algorithm[1](https://arxiv.org/html/2605.07234#alg1)\.

### 4\.2Eviction Action

The additive structure in equation[7](https://arxiv.org/html/2605.07234#S4.E7)implies that, although we estimate token impact independently within each head, the resulting contribution of a token is not confined to that head alone; rather, it directly participates in the formation of the shared layer output after the output projection\. Consequently, token importance is fundamentally a layer\-level concept: tokens with large estimated contributions impact the layer output equivalently, regardless of which head produces them\.

In contrast, head\-wise selection relies on head local rankings, which can be misleading\. For example, a token may rank highly in its head, but its projected contribution afterV​WOVW\_\{O\}can benumerically insignificant compared to tokens from other heads and provides minimal difference in the aggregation step\.Retaining such “local winners” may waste memory on signals that have little impact on the final layer output\. Conversely, some tokens may not be top\-ranked in their own heads, yet contribute highly across multiple heads\. These tokens are naturally captured by a layer\-level criterion but missed by head\-wise selection\.

Algorithm 2Eviction ActionInput:Eviction scores

\{pl,h,j\}\\\{p\_\{l,h,j\}\\\}, global budget

KK
Output:Selected token set

𝒮\\mathcal\{S\}
// Flatten head\-wise scores

foreach layer

lldo

foreach head

hhdo

foreach token

jjdo

pl,k←pl,h,jp\_\{l,k\}\\leftarrow p\_\{l,h,j\}

endfor

endfor

endfor

// Layer\-wise normalization

foreach layer

lldo

foreach token

jjdo

sl,j←pl,j/∑kpl,ks\_\{l,j\}\\leftarrow p\_\{l,j\}/\\sum\_\{k\}p\_\{l,k\}

endfor

endfor

// Global selection

𝒮←TopK⁡\(\{sl,j\}l,j,K\)\\mathcal\{S\}\\leftarrow\\operatorname\{TopK\}\\\!\\left\(\\\{s\_\{l,j\}\\\}\_\{l,j\},K\\right\)

With the presence ofV​WO\\boldsymbol\{VW\_\{O\}\}, remark[4\.2](https://arxiv.org/html/2605.07234#S4.Thmtheorem2)motivates formulating KV cache eviction as a unified selection problem across an entire layer, rather than a series of independent head\-wise decisions\. Unlike existing weight\-based methods[undefj](https://arxiv.org/html/2605.07234#bib.bib11)which rely on eviction scores that are strictly local to each head, our metric \(Section[4\.1](https://arxiv.org/html/2605.07234#S4.SS1)\) integratesWOW\_\{O\}that maps head contributions onto a unified scale, making token importance comparable across the entire layer, and providing a theoretically grounded basis for global selection\. This allows us to safely flatten eviction scores and perform a joint selection under a fixed layer budget, and naturally enables an adaptive allocation of cache capacity: heads containing influential tokens retain more entries, while those with lower\-impact tokens are pruned more aggressively\. In Table[6](https://arxiv.org/html/2605.07234#A4.T6), we also prove that a greedy selection without our global score can even hurt the performance\.

Similar to how a layer output is formed by accumulating contributions from tokens across all heads,the final model output is also an accumulation across layers\.As shown in Equation[5](https://arxiv.org/html/2605.07234#S2.E5), a Transformer consists of stacked blocks connected through a shared residual path, where each layer’s output is added to this path\. As a result, the final output reflects the aggregated contributions of all layers\. Therefore, if a token has a strong impact on its layer’s output, this impact is directly propagated through the residual path and influences the final output\. In this sense, a token’s importance can be understood by how much it affects/changes this accumulated representation\.

However, while ourV​WOVW\_\{O\}\-based metric accurately captures a token’s contribution to the layer output, these raw scores are not directly comparable across different layers due to the inherent magnitude variations shown in Figure[1\(b\)](https://arxiv.org/html/2605.07234#S3.F1.sf2)\. To resolve this scale disparity, we employ a simple layer\-wise score normalization scheme that transforms flattened scores into a relative importance distribution:

sl,j=pl,j∑kpl,k,s\_\{l,j\}=\\frac\{p\_\{l,j\}\}\{\\sum\_\{k\}p\_\{l,k\}\},\(12\)wherepl,jp\_\{l,j\}denotes the flattened eviction score of tokenjjin layerll\. This normalization neutralizes inter\-layer scale differences for inter\-layer comparison while strictly preserving the relative token rankings within each layer\. The effectiveness of this simple normalization and its necessity is further demonstrated empirically in our evaluation and ablation study in Section[5](https://arxiv.org/html/2605.07234#S5)and Appendix[D\.1](https://arxiv.org/html/2605.07234#A4.SS1)\.

Finally, tokens across the model are jointly selected usingsl,j\{s\_\{l,j\}\}\(Algorithm[2](https://arxiv.org/html/2605.07234#alg2)\)\.This unified rule allows the model to focus on the most critical tokens model\-wide\.Notably, this approach is hyperparameter\-free and requires no calibration or training/finetuning, offering a straightforward implementation without complexity\.

## 5Experiments

Models\.We conduct experiments using three widely\-used open\-sourced LLMs: Llama\-3\.1\-8B\-Instruct[undefo](https://arxiv.org/html/2605.07234#bib.bib16), Mistral\-7B\-Instruct\-v0\.3[undefv](https://arxiv.org/html/2605.07234#bib.bib23), and Qwen3\-8B[undefas](https://arxiv.org/html/2605.07234#bib.bib46)\. These models provide maximum context lengths of 128K, 32K, and 32K tokens, respectively\.

Baselines\.Our method is compared against the full\-cache configuration \(FullKV\) and four SOTA baselines: StreamingLLM \(SLLM\)[undefar](https://arxiv.org/html/2605.07234#bib.bib45), SnapKV[undefac](https://arxiv.org/html/2605.07234#bib.bib30), AdaKV[undefj](https://arxiv.org/html/2605.07234#bib.bib11), CriticalKV[undefk](https://arxiv.org/html/2605.07234#bib.bib12), and CAKE[undefag](https://arxiv.org/html/2605.07234#bib.bib34)\. A summary of these baselines is provided in Table[2](https://arxiv.org/html/2605.07234#A1.T2)\.

Evaluation Scenarios\.Our approach is compared across various memory budgets \(average capacity per head\)\. We use fixed absolute cache sizes rather than ratios relative to the full KV cache to prevent cache size increase with context length, better reflecting real\-world hardware constraints\. For implementation details, see Appendix[A](https://arxiv.org/html/2605.07234#A1)\.

Evaluation Benchmarks\.Following the standard practice of the prior works, our main experiments include two long\-context benchmarks: LongBench[undefb](https://arxiv.org/html/2605.07234#bib.bib3)and RULER[undeft](https://arxiv.org/html/2605.07234#bib.bib21)\. Additionally, we also evaluate on InfiniteBench[undefax](https://arxiv.org/html/2605.07234#bib.bib51), a very\-long\-context benchmark in[G\.1](https://arxiv.org/html/2605.07234#A7.SS1)\.

### 5\.1Evaluations on LongBench Dataset

![Refer to caption](https://arxiv.org/html/2605.07234v1/x3.png)

Figure 2:Average scores among 16 datasets of LongBench under different cache budgets\.Table 1:Comparison across 16 LongBench datasets, withbestandsecond bestresults highlighted\.We evaluate LaProx against SOTA KV cache eviction techniques across the 16 datasets in the LongBench benchmark, using cache budgets ranging from 128 to 1024 tokens\. Table[1](https://arxiv.org/html/2605.07234#S5.T1)details the performance across three models at a budget of 128 tokens, while Figure[2](https://arxiv.org/html/2605.07234#S5.F2)illustrates the average performance across the full range of budget constraints\.

As shown in Table[1](https://arxiv.org/html/2605.07234#S5.T1), LaProx consistently outperforms previous works in nearly every LongBench’s dataset, leading to significant improvements in total performance\. Figure[2](https://arxiv.org/html/2605.07234#S5.F2)further demonstrates our superior results across all budget sizes and models\. Notably, the performance gap between LaProx and the baselines widens as the memory budget becomes more constrained, highlighting the robustness of LaProx under extreme hardware limitations\.

We can also see that SLLM consistently exhibits the weakest performance, which is expected given its aggressive removal of intermediate tokens\. By contrast, other approaches improve the performance by actively selecting important tokens\. However, these baselines introduce only heuristic refinements to SnapKV, lacking a principled foundation or consideration of the underlying model structure, thereby limiting their gains over the vanilla SnapKV\. Nevertheless, they remain strong baselines due to their improvements across many settings\. In some cases, however,such heuristic strategies can even degrade performance,as observed for Qwen3 at budgets of 256\-1024 by AdaKV and CriticalKV\. Meanwhile, LaProx achieves superior performance across the majority of datasets and memory configurations\. These results underscore the advantages of our unified global eviction strategy\.

### 5\.2Evaluations on Needle\-in\-A\-Haystack

![Refer to caption](https://arxiv.org/html/2605.07234v1/x4.png)

Figure 3:Comparison across 3 NIAH variants at 32K context length\.To evaluate retrieval performance, we employ Needle\-in\-A\-Haystack \(NIAH\) test, where a target sentence is embedded within a long\-context distractor\. Following RULER[undeft](https://arxiv.org/html/2605.07234#bib.bib21), we examine three representative configurations: \(1\) Single\-Needle with 1 needle and 1 target \(1N\-1T\): There is a single needle that the model must retrieve from the context; \(2\) Multi\-Needle with 4 needles and 1 target \(4N\-1T\): 4 needles \(1 target and 3 distractors\) are inserted to the context and the model must isolate a specific target from three distracting needles; and \(3\) Multi\-Needle with 4 needles and 4 targets \(4N\-4T\): 4 needles are inserted and the model must retrieve all of them\. To match the Mistral and Qwen models’ context windows and to balance the evaluation cost, each configuration is evaluated at a 32K context length across 100 samples per task, using cache budgets ranging from 256 to 2048 tokens\. Further evaluations on RULER tasks are provided in Appendix[G\.2](https://arxiv.org/html/2605.07234#A7.SS2)and[G\.3](https://arxiv.org/html/2605.07234#A7.SS3)\.

Consistent with LongBench results, LaProx demonstrates significant performance gains on NIAH tests across all evaluated models and cache budgets\. Notably, on Mistral\-7B\-Instruct\-v0\.3 with a highly constrained budget of 256 tokens, LaProx achieves a 1\.5 to 3×\\timesimprovement in retrieval accuracy over CriticalKV and other prior methods, respectively, across all three NIAH variants\.

Meanwhile, although the baselines perform more competitively on Llama\-3\.1\-8B\-Instruct and Qwen3\-8B, LaProx consistently maintains the highest scores and reaches FullKV performance sooner than other approaches\. The performance gap becomes particularly pronounced in resource\-constrained settings \(256 tokens\) or complex tasks \(4 needles\)\.

Furthermore, similar to the LongBench observations,heuristic refinement methods continue to exhibit unstable behaviorand, in many cases, even degrade performance relative to the vanilla SnapKV baseline \(such as CAKE and AdaKV\)\.

### 5\.3Analysis of Matrix\-Approximation\-based Eviction Criterion

![Refer to caption](https://arxiv.org/html/2605.07234v1/x5.png)

Figure 4:Similarity score between the full and approximated attention layer outputsBeyond benchmark accuracy, we further investigate whether our matrix\-based eviction criterion improves the similarity between the full and approximated attention outputs\. Specifically, for each layer, we measure the cosine similarity between the full and compressed attention outputs for the first decoding token in Mistral\-7B\-Instruct\-v0\.3\. The evaluation is conducted on the TREC dataset, with the KV cache of 128 tokens\. As shown in Figure[4](https://arxiv.org/html/2605.07234#S5.F4), our method consistently achieves higher cosine similarity across all layers compared to the vanilla approach based solely on local attention weights, indicating a more faithful approximation of the attention output\.

### 5\.4Evaluations on Efficiency

![Refer to caption](https://arxiv.org/html/2605.07234v1/x6.png)

Figure 5:Efficiency Analysis\.To evaluate efficiency, we report both prefill latency, which includes initial prompt processing and eviction overhead, per\-token decoding latency, and peak memory usage\. All measurements are conducted on a single NVIDIA H100 \(80GB\) GPU using the Meta\-Llama\-3\.1\-8B\-Instruct model with a fixed cache budget of 128 tokens\. In this section, we compare our approach against AdaKV and CriticalKV, representing budget\-allocation\-based and output\-aware eviction baselines, respectively\.

As shown in Figure[5](https://arxiv.org/html/2605.07234#S5.F5), our method introduces only marginal overhead during prefilling across all context lengths\. During decoding, all eviction methods achieve comparable efficiency and consistently outperform the FullKV setting\. In contrast to FullKV, whose decoding latency grows rapidly with sequence length, our approach maintains stable per\-token latency by enforcing a strict cache budget\. Consequently, at 128K context length, our method achieves a 2\.3×\\timesdecoding speedup over FullKV\.

In addition, all eviction methods substantially reduce peak memory usage compared to FullKV\. For instance, at 128K context length, eviction\-based approaches require only around 47\.5GB of memory, whereas FullKV consumes 63\.3GB\.

### 5\.5Analysis on the Number of Retained Tokens

Implicitly, our model\-wide eviction strategy \(Algorithm[2](https://arxiv.org/html/2605.07234#alg2)\) retains different numbers of tokens across layers and heads according to their estimated importance\. To analyze this behavior, we conduct experiments on the NIAH dataset using two representative models, Llama and Mistral, and record the number of retained tokens for each head\. The total cache size is fixed to 128 tokens, corresponding to an average of 96 retained entries per head \(with an additional 32\-entry history window\)\.

![Refer to caption](https://arxiv.org/html/2605.07234v1/x7.png)Figure 6:Retained tokens and variation per headAs shown in Figure[6](https://arxiv.org/html/2605.07234#S5.F6), the number of retained tokens varies substantially across layers and heads\. For example, in Llama, layers 13\-14 retain significantly more tokens than layers 3 and 24\. Similarly, within layer 27, a single head retains the majority of tokens\.

We also observe distinct patterns across models\. Llama concentrates retained tokens in the middle layers, whereas Mistral progressively shifts importance toward later layers\. Moreover, these patterns are not stable and can vary by up to 500 tokens per head across inputs\.

These results show that token contributions are highly non\-uniform across layers and heads, vary across models, and are strongly input\-dependent\.This underscores the need of our model\-wide eviction strategy,which dynamically selects tokens without offline calibration, and highlights a key limitation of calibration\-based prior methods[undefag](https://arxiv.org/html/2605.07234#bib.bib34);[undefl](https://arxiv.org/html/2605.07234#bib.bib13);[undefd](https://arxiv.org/html/2605.07234#bib.bib5), which assume stable optimal budgets across diverse inputs\.

## 6Conclusion

This paper revisits KV cache eviction for long\-context LLM inference and addresses two fundamental oversights in current methods\. First, we identify that current heuristic approaches neglect the structure of the attention layers; by considering the attention layer formulation, we provide a more complete measure of token importance\. Second, we demonstrate that traditional head\-level eviction is inherently suboptimal, as token importance is more accurately captured at the model\-level\. Based on these insights, we introduce LaProx that reformulates cache eviction as a global matrix approximation problem rather than a head\-wise weight\-averaging task\.Our work is the first to explore a unified model\-level cache eviction strategy, exposing the weakness of the long\-standing traditional head\-based heuristic schemes\.Extensive evaluations across 19 datasets from LongBench and NIAH benchmarks confirm that this approach consistently outperforms prior strategies\. Furthermore, our analysis validates that our method achieves higher fidelity to full\-cache attention outputs than standard heuristics\. Ultimately, this work introduces a different, global point of view to the study of KV cache, offering a principled foundation for efficient long\-context inference\.

## References

- \(1\)Menachem Adelman, Kfir Levy, Ido Hakimi and Mark Silberstein“Faster neural network training with approximate tensor operations”In*Advances in Neural Information Processing Systems*34, 2021, pp\. 27877–27889
- \(2\)Joshua Ainslie et al\.“Gqa: Training generalized multi\-query transformer models from multi\-head checkpoints”In*arXiv preprint arXiv:2305\.13245*, 2023
- \(3\)Yushi Bai et al\.“Longbench: A bilingual, multitask benchmark for long context understanding”In*Proceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\)*, 2024, pp\. 3119–3137
- \(4\)Iz Beltagy, Matthew E Peters and Arman Cohan“Longformer: The long\-document transformer”In*arXiv preprint arXiv:2004\.05150*, 2020
- \(5\)Zefan Cai et al\.“Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling”In*arXiv preprint arXiv:2406\.02069*, 2024
- \(6\)Tri Dao“Flashattention\-2: Faster attention with better parallelism and work partitioning”In*arXiv preprint arXiv:2307\.08691*, 2023
- \(7\)Pradeep Dasigi et al\.“A dataset of information\-seeking questions and answers anchored in research papers”In*arXiv preprint arXiv:2105\.03011*, 2021
- \(8\)Alessio Devoto, Yu Zhao, Simone Scardapane and Pasquale Minervini“A Simple and EffectiveL​\_​2L\\\_2Norm\-Based Strategy for KV Cache Compression”In*arXiv preprint arXiv:2406\.11430*, 2024
- \(9\)Petros Drineas, Ravi Kannan and Michael W Mahoney“Fast Monte Carlo algorithms for matrices I: Approximating matrix multiplication”In*SIAM Journal on Computing*36\.1SIAM, 2006, pp\. 132–157
- \(10\)Alexander Richard Fabbri et al\.“Multi\-news: A large\-scale multi\-document summarization dataset and abstractive hierarchical model”In*Proceedings of the 57th annual meeting of the association for computational linguistics*, 2019, pp\. 1074–1084
- \(11\)Yuan Feng et al\.“Ada\-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference”In*arXiv preprint arXiv:2407\.11550*, 2024
- \(12\)Yuan Feng et al\.“Identify critical kv cache in llm inference from an output perturbation perspective”In*arXiv preprint arXiv:2502\.03805*, 2025
- \(13\)Yu Fu et al\.“Not all heads matter: A head\-level kv cache compression method with integrated retrieval and reasoning”In*arXiv preprint arXiv:2410\.19258*, 2024
- \(14\)Bogdan Gliwa, Iwona Mochol, Maciej Biesek and Aleksander Wawer“SAMSum corpus: A human\-annotated dialogue dataset for abstractive summarization”In*arXiv preprint arXiv:1911\.12237*, 2019
- \(15\)Raghavv Goel et al\.“CAOTE: KV Cache Selection for LLMs via Attention Output Error\-Based Token Eviction”In*arXiv preprint arXiv:2504\.14051*, 2025
- \(16\)Aaron Grattafiori et al\.“The llama 3 herd of models”In*arXiv preprint arXiv:2407\.21783*, 2024
- \(17\)Yifeng Gu et al\.“AhaKV: Adaptive Holistic Attention\-Driven KV Cache Eviction for Efficient Inference of Large Language Models”In*arXiv preprint arXiv:2506\.03762*, 2025
- \(18\)Daya Guo et al\.“Longcoder: A long\-range pre\-trained language model for code completion”In*International Conference on Machine Learning*, 2023, pp\. 12098–12107PMLR
- \(19\)Xanh Ho, Anh\-Khoa Duong Nguyen, Saku Sugawara and Akiko Aizawa“Constructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps”In*arXiv preprint arXiv:2011\.01060*, 2020
- \(20\)Coleman Hooper et al\.“Kvquant: Towards 10 million context length llm inference with kv cache quantization”In*Advances in Neural Information Processing Systems*37, 2024, pp\. 1270–1303
- \(21\)Cheng\-Ping Hsieh et al\.“RULER: What’s the Real Context Size of Your Long\-Context Language Models?”In*arXiv preprint arXiv:2404\.06654*, 2024
- \(22\)Luyang Huang et al\.“Efficient attentions for long document summarization”In*arXiv preprint arXiv:2104\.02112*, 2021
- \(23\)Albert Q\. Jiang et al\.“Mistral 7B”, 2023arXiv:[https://arxiv\.org/abs/2310\.06825](https://arxiv.org/abs/2310.06825)
- \(24\)Mandar Joshi, Eunsol Choi, Daniel S Weld and Luke Zettlemoyer“Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension”In*arXiv preprint arXiv:1705\.03551*, 2017
- \(25\)Ehsan Kamalloo, Nouha Dziri, Charles Clarke and Davood Rafiei“Evaluating open\-domain question answering in the era of large language models”In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, 2023, pp\. 5591–5606
- \(26\)Gregory Kamradt“Needle In A Haystack \- pressure testing LLMs”, 2023URL:[https://github\.com/gkamradt/LLMTest\_NeedleInAHaystack](https://github.com/gkamradt/LLMTest_NeedleInAHaystack)
- \(27\)Tomáš Kočiskỳ et al\.“The narrativeqa reading comprehension challenge”In*Transactions of the Association for Computational Linguistics*6MIT Press One Rogers Street, Cambridge, MA 02142\-1209, USA journals\-info …, 2018, pp\. 317–328
- \(28\)Wonbeom Lee, Jungi Lee, Junghwan Seo and Jaewoong Sim“\{\\\{InfiniGen\}\\\}: Efficient generative inference of large language models with dynamic\{\\\{KV\}\\\}cache management”In*18th USENIX Symposium on Operating Systems Design and Implementation \(OSDI 24\)*, 2024, pp\. 155–172
- \(29\)Xin Li and Dan Roth“Learning question classifiers”In*COLING 2002: The 19th International Conference on Computational Linguistics*, 2002
- \(30\)Yuhong Li et al\.“Snapkv: Llm knows what you are looking for before generation”In*Advances in Neural Information Processing Systems*37, 2024, pp\. 22947–22970
- \(31\)Tianyang Liu, Canwen Xu and Julian McAuley“Repobench: Benchmarking repository\-level code auto\-completion systems”In*arXiv preprint arXiv:2306\.03091*, 2023
- \(32\)Zichang Liu et al\.“Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time”In*Advances in Neural Information Processing Systems*36, 2023, pp\. 52342–52364
- \(33\)Zirui Liu et al\.“Kivi: A tuning\-free asymmetric 2bit quantization for kv cache”In*arXiv preprint arXiv:2402\.02750*, 2024
- \(34\)Ziran Qin et al\.“Cake: Cascading and adaptive kv cache eviction with layer preferences”In*arXiv preprint arXiv:2503\.12491*, 2025
- \(35\)Colin Raffel et al\.“Exploring the limits of transfer learning with a unified text\-to\-text transformer”In*Journal of machine learning research*21\.140, 2020, pp\. 1–67
- \(36\)Yiqun Shen et al\.“LAVa: Layer\-wise KV Cache Eviction with Dynamic Budget Allocation”, 2025arXiv:[https://arxiv\.org/abs/2509\.09754](https://arxiv.org/abs/2509.09754)
- \(37\)Ying Sheng et al\.“Flexgen: High\-throughput generative inference of large language models with a single gpu”In*International Conference on Machine Learning*, 2023, pp\. 31094–31116PMLR
- \(38\)Hanlin Tang et al\.“Razorattention: Efficient kv cache compression through retrieval heads”In*arXiv preprint arXiv:2407\.15891*, 2024
- \(39\)Jiaming Tang et al\.“Quest: Query\-aware sparsity for efficient long\-context llm inference”In*arXiv preprint arXiv:2406\.10774*, 2024
- \(40\)Vicuna Team“Vicuna: An open\-source chatbot impressing gpt\-4 with 90% chatgpt quality”In*Vicuna: An open\-source chatbot impressing gpt\-4 with*90, 2023
- \(41\)Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot and Ashish Sabharwal“MuSiQue: Multihop Questions via Single\-hop Question Composition”In*Transactions of the Association for Computational Linguistics*10MIT Press One Broadway, 12th Floor, Cambridge, Massachusetts 02142, USA …, 2022, pp\. 539–554
- \(42\)Zhongwei Wan et al\.“D2o: Dynamic discriminative operations for efficient long\-context inference of large language models”In*arXiv preprint arXiv:2406\.13035*, 2024
- \(43\)Hanrui Wang, Zhekai Zhang and Song Han“Spatten: Efficient sparse attention architecture with cascade token and head pruning”In*2021 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\)*, 2021, pp\. 97–110IEEE
- \(44\)Guangxuan Xiao et al\.“Duoattention: Efficient long\-context llm inference with retrieval and streaming heads”In*arXiv preprint arXiv:2410\.10819*, 2024
- \(45\)Guangxuan Xiao et al\.“Efficient streaming language models with attention sinks, 2024”In*URL https://arxiv\. org/abs/2309\.17453*1, 2024
- \(46\)An Yang et al\.“Qwen3 technical report”In*arXiv preprint arXiv:2505\.09388*, 2025
- \(47\)Dongjie Yang et al\.“Pyramidinfer: Pyramid kv cache compression for high\-throughput llm inference”In*arXiv preprint arXiv:2405\.12532*, 2024
- \(48\)Zhilin Yang et al\.“HotpotQA: A Dataset for Diverse, Explainable Multi\-hop Question Answering”, 2018arXiv:[https://arxiv\.org/abs/1809\.09600](https://arxiv.org/abs/1809.09600)
- \(49\)Hailin Zhang et al\.“Pqcache: Product quantization\-based kvcache for long context llm inference”In*Proceedings of the ACM on Management of Data*3\.3ACM New York, NY, USA, 2025, pp\. 1–30
- \(50\)Tianyi Zhang et al\.“Benchmarking large language models for news summarization”In*Transactions of the Association for Computational Linguistics*12MIT Press One Broadway, 12th Floor, Cambridge, Massachusetts 02142, USA …, 2024, pp\. 39–57
- \(51\)Xinrong Zhang et al\.“∞\\inftyBench: Extending Long Context Evaluation Beyond 100K Tokens”, 2024arXiv:[https://arxiv\.org/abs/2402\.13718](https://arxiv.org/abs/2402.13718)
- \(52\)Zhenyu Zhang et al\.“H2o: Heavy\-hitter oracle for efficient generative inference of large language models”In*Advances in Neural Information Processing Systems*36, 2023, pp\. 34661–34710
- \(53\)Youpeng Zhao, Di Wu and Jun Wang“Alisa: Accelerating large language model inference via sparsity\-aware kv caching”In*2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture \(ISCA\)*, 2024, pp\. 1005–1017IEEE
- \(54\)Ming Zhong et al\.“QMSum: A new benchmark for query\-based multi\-domain meeting summarization”In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, 2021, pp\. 5905–5921

## Appendix AImplementation details

Methods with uniform allocation assign equal cache capacity to each attention head, whereas non\-uniform strategies, including ours, adaptively distribute cache capacity across heads or layers while keeping the total memory budget fixed\. All baseline hyperparameters follow their default official implementations\. For SLLM[undefar](https://arxiv.org/html/2605.07234#bib.bib45), we retain four sink tokens and allocate the remaining budget to a sliding recent window\. All other methods employ an observation window of 32 tokens to limit overhead and apply average pooling with a kernel size of 7 to mitigate information fragmentation\. The summary of the baselines is shown in Table[2](https://arxiv.org/html/2605.07234#A1.T2)\.

To ensure compatibility with Grouped Query Attention \(GQA\)[undefa](https://arxiv.org/html/2605.07234#bib.bib2), we follow standard practice in prior works by using the mean attention weight within each query group as the selection criterion\. All experiments are accelerated using FlashAttention\-2[undefe](https://arxiv.org/html/2605.07234#bib.bib6)\. And consistent with prior works[undefac](https://arxiv.org/html/2605.07234#bib.bib30);[undefj](https://arxiv.org/html/2605.07234#bib.bib11), cache eviction is applied once after the prefilling phase per layer\.

Table 2:Summary of evaluated methods on eviction score\.
## Appendix BAdditional Related Works on KV Cache Management

#### KV Cache Eviction Metrics\.

Beyond standard attention weights, several works explore alternative definitions of token importance\. The work in[undefg](https://arxiv.org/html/2605.07234#bib.bib8)utilizes theL​2L2norm of key states to reduce computational overhead at the expense of accuracy\. Other approaches, such as[undefp](https://arxiv.org/html/2605.07234#bib.bib17), enhance attention\-based metrics by incorporating value representations\. However, these methods rely solely on value states, which only partially capture a token’s contribution to the final representation\. In practice, the output is further modulated by the output projection matrix \(WOW\_\{O\}\)—a factor our ablation study \(Section[D\.1](https://arxiv.org/html/2605.07234#A4.SS1)\) identifies as a critical\.

CriticalKV[undefk](https://arxiv.org/html/2605.07234#bib.bib12)is the most closely related to our work, as it also incorporates both value states and output projections\. However, a fundamental conceptual gap remains: CriticalKV acts as a helper for other works and it treats output information merely as a scaling factor applied to the average\-based eviction score\. In contrast,we reformulate KV cache eviction as a principled approximation of the attention matrix productandprovide a complete solution\. By explicitly modeling the interaction between attention weights and projected value states as a unified operation, our method provides a more accurate estimation of token influence, as validated in Section[5](https://arxiv.org/html/2605.07234#S5)\.

Furthermore, because CriticalKV lacks this rigorous matrix\-product formulation,it must rely on 2 hyper\-parameters to maintain model performance\(an empirical safeguard and anϵ\\epsilonto mitigate information loss\)\. Its scoring also remains strictly head\-centric, failing to address eviction as a layer\-level optimization problem\. Our approach moves beyond simple "scaling helpers" by providing a unified, theoretical\-bound layer\-wise solution that naturally handles budget allocation across heads without requiring empirical tuning\.

#### KV Cache Selection\.

While KV cache eviction methods reduce memory usage by retaining only a small subset of critical key–value entries, sparse attention approaches such as Quest[undefal](https://arxiv.org/html/2605.07234#bib.bib39)and InfiniGen[undefaa](https://arxiv.org/html/2605.07234#bib.bib28)preserve the entire KV cache during inference but restrict computation to a selected subset of entries at each attention step\. Or, some works[undefaq](https://arxiv.org/html/2605.07234#bib.bib44);[undefak](https://arxiv.org/html/2605.07234#bib.bib38)combine the Dynamic Selection with Eviction approach by classifying attention heads into Streaming Heads with a limited Cache size and Retrieval Heads with a full Cache size\. By limiting attention computation rather than storage, sparse attention methods can significantly accelerate inference and often achieve high output quality\. However, because all KV entries are still stored, these methodsdo not reduce the memory footprint of the KV cache\.

#### KV Cache Quantization\.

KV cache quantization methods reduce memory and computation costs by representing cached values using lower\-precision formats\. These approaches can be broadly divided into fixed\-precision quantization, where all tokens share the same bit\-width[undefav](https://arxiv.org/html/2605.07234#bib.bib49);[undefaj](https://arxiv.org/html/2605.07234#bib.bib37), and mixed\-precision schemes, which allocate different bit\-widths to different tokens[undefs](https://arxiv.org/html/2605.07234#bib.bib20);[undefaf](https://arxiv.org/html/2605.07234#bib.bib33)\. However, while quantization effectively lowers per\-token storage cost, the totalKV cache size continues to grow linearly with context length, limiting its ability to address memory bottlenecks in very long\-context settings\.

## Appendix CDetails of Evaluation Benchmarks

#### LongBench Benchmark\.

LongBench[undefb](https://arxiv.org/html/2605.07234#bib.bib3)is a widely used long\-context benchmark, serving as a standardized evaluation protocol commonly adopted by prior works and baselines[undefac](https://arxiv.org/html/2605.07234#bib.bib30);[undefag](https://arxiv.org/html/2605.07234#bib.bib34);[undefj](https://arxiv.org/html/2605.07234#bib.bib11)\. It consists of 16 datasets across 6 task domains: Single\-Doc QA[undefz](https://arxiv.org/html/2605.07234#bib.bib27);[undeff](https://arxiv.org/html/2605.07234#bib.bib7), Multi\-Doc QA[undefau](https://arxiv.org/html/2605.07234#bib.bib48);[undefr](https://arxiv.org/html/2605.07234#bib.bib19);[undefan](https://arxiv.org/html/2605.07234#bib.bib41), Summarization[undefu](https://arxiv.org/html/2605.07234#bib.bib22);[undefaaa](https://arxiv.org/html/2605.07234#bib.bib54);[undefi](https://arxiv.org/html/2605.07234#bib.bib10), Few\-shot Learning[undefab](https://arxiv.org/html/2605.07234#bib.bib29);[undefw](https://arxiv.org/html/2605.07234#bib.bib24);[undefm](https://arxiv.org/html/2605.07234#bib.bib14), Synthetic Task[undefah](https://arxiv.org/html/2605.07234#bib.bib35), and Code Completion[undefq](https://arxiv.org/html/2605.07234#bib.bib18);[undefad](https://arxiv.org/html/2605.07234#bib.bib31)\. The average token length across all 16 datasets is 6,711\. More details can be found in Table[3](https://arxiv.org/html/2605.07234#A3.T3)\.

Table 3:Details of each dataset in LongBench\.
#### Ruler Benchmark\.

Ruler[undeft](https://arxiv.org/html/2605.07234#bib.bib21)is a diagnostic benchmark designed to evaluate long\-context capabilities beyond simple retrieval\. It consists of 13 datasets across 4 task domains:

- •Retrieval: An extension of NIAH[undefy](https://arxiv.org/html/2605.07234#bib.bib26)that tests retrieval robustness using diverse needle types and varying quantities of hidden information\.
- •Multi\-hop Tracing: Evaluates the model’s ability to track variable assignments and identify co\-occurrence patterns that require connecting multiple pieces of information across the sequence\.
- •Aggregation: Tests the ability to identify the most frequent or common words distributed throughout the text\.
- •Question Answering: Tests the capability to answer questions where the answer is deeply embedded within extensive distracting or irrelevant content\.

Further evaluations on more tasks from Ruler are provided in Appendix[G\.2](https://arxiv.org/html/2605.07234#A7.SS2)and[G\.3](https://arxiv.org/html/2605.07234#A7.SS3)\.

## Appendix DAblation Study

### D\.1Eviction Criteria Analysis

Table 4:Comparison with different Eviction Indicators\.Eviction IndicatorValue StateOutput WeightAvg\.40\.53✓41\.69✓✓41\.84Table 5:Comparison with different Allocation Strategies\.
Allocation strategyHead\-FlattenLayer\-FlattenNormalizeAvg\.41\.84✓42\.96✓✓41\.14✓✓✓44\.00

In this section, we present a series of ablation studies to evaluate the effectiveness of our proposed eviction strategy\. We use Mistral\-7B\-Instruct\-v0\.3 with cache budgetBt​o​t​a​l=128​LB\_\{total\}=128Lon LongBench as the default settings\.

Effectiveness of Proposed Eviction Indicator\.To validate the necessity of our metric, we isolate the influence ofVVfromV​WOVW\_\{O\}then compare them against the baseline\. As shown in Table[4](https://arxiv.org/html/2605.07234#A4.T4), incorporatingVVimproves performance over attention\-only baselines, but remains sub\-optimal as it still provides an incomplete computation of the attention output\. The highest accuracy is achieved by our fullA​V​WOAVW\_\{O\}indicator, demonstrating that the output projection is essential for quantifying a token’s contribution to the layer’s output\.

Effectiveness of Proposed Allocation Strategy\.We further analyze the necessity of each component in our allocation method, specifically flattening \(head\-wise and layer\-wise\) and normalization \(Table[5](https://arxiv.org/html/2605.07234#A4.T5)\)\. While head\-wise allocation provides immediate performance gains, layer\-wise allocation degrades performance to below the baseline when the scores are used in their raw form\. This decline stems from raw score scale differences across layers; without normalization, the budget focuses only on high\-magnitude layers while starving others\. Despite its simplicity, our normalization scheme resolves this bias, allowing layer\-wise allocation to effectively complement head\-wise settings and further improve performance\.

While the allocation gain shown in Table[5](https://arxiv.org/html/2605.07234#A4.T5)is more noticeable, it is worth noting that our method still outperforms prior approaches without it\. Furthermore,projection and allocation strategy are not independent; rather, the former is the mathematical prerequisite for the latter\.This is because without our projection, attention\-based score is limited in its own head and not comparable globally\.

To validate this, we experiment with 4 different settings: Uniform budget and Global allocation with Attention score \(ALA\_\{L\}andAGA\_\{G\}\), Uniform budget and Global allocation with LaProx score \(LLL\_\{L\}andLGL\_\{G\}\)\. The experiments were conducted using LongBench benchmark under the budget of 128 tokens\. The results are shown in Table[6](https://arxiv.org/html/2605.07234#A4.T6)\.

Table 6:Comparison across 16 LongBench datasets for the cache budget of 128L\. The best result is highlighted inbold\.As seen from the results,withoutV​WOVW\_\{O\}, global allocation offers no improvement compared to the baseline\.Particularly, the performance ofAGA\_\{G\}is similar toALA\_\{L\}, and is lower thanLLL\_\{L\}, which is our method without any allocation\.

### D\.2Sensitivity Analysis

Window size robustness\.Following the configurations of the baselines[undefac](https://arxiv.org/html/2605.07234#bib.bib30);[undefag](https://arxiv.org/html/2605.07234#bib.bib34);[undefj](https://arxiv.org/html/2605.07234#bib.bib11), we utilized a default historical window size of 32 for our primary experiments\. To assess the sensitivity of our method to this hyperparameter, we conducted an ablation study on the Mistral\-7B\-Instruct\-v0\.3 model \(with a cache budget of 128\) across four window sizes: 8, 16, 32, and 64\. The results are summarized in Table[7](https://arxiv.org/html/2605.07234#A4.T7)\. LaProx maintains stable performance across window sizes 8 through 32, with scores ranging between 43\.79 and 44\.10\. However, performance degrades at a window size of 64; this occurs because the large observation window reduces the candidate space for eviction, thereby limiting the flexibility of our eviction strategy\. Overall, these results demonstrate that LaProx is highly robust to variations in the historical window size\.

Table 7:Analysis on different observation window sizes\.MethodSingle\-Document QAMulti\-Document QASummarizationFew\-shot LearningSyntheticCodeAvg\.NrtvQAQasperMF\-enHotpotQA2WikiMQAMusiqueGovRepQMSumMultiNewsTRECTriviaQASAMSumPCountPR\-enLccRB\-PMistral\-7B\-Instruct\-v0\.3,Btotal=F​u​l​lB\_\{\\text\{total\}\}=FullFullKV29\.0741\.5852\.8849\.3739\.0128\.5834\.8125\.6627\.827688\.5947\.45\.59861\.462\.5348\.01Mistral\-7B\-Instruct\-v0\.3,Btotal=128​LB\_\{\\text\{total\}\}=128Lw=827\.8730\.5153\.8548\.235\.9525\.9822\.5823\.222\.247288\.6643\.685\.59456\.9954\.444\.10w=1625\.7531\.4352\.347\.4635\.5725\.5522\.9322\.9822\.2669\.588\.4444\.175\.59457\.4355\.3443\.79w=322830\.2654\.0247\.8637\.5425\.5421\.2623\.3122\.163\.589\.743\.4659658\.8557\.6644\.00w=6423\.2424\.9945\.748\.4335\.2824\.2919\.1321\.219\.95288\.7443\.283\.59255\.5355\.0840\.77

## Appendix ELimitations

Our work is the first to introduce a unified eviction framework that assigns globally comparable importance scores to tokens, enabling model\-wide joint selection of KV cache entries\. To achieve practical efficiency, our current formulation relies on relatively simple approximation and normalization schemes, which we view as an initial foundation for future model\-wide KV cache research\. Future works may explore more sophisticated matrix multiplication approximation schemes or different normalization methods for global tokens comparison to further improve cache eviction performance\.

## Appendix FProofs

### F\.1Proof of Remark[4\.1](https://arxiv.org/html/2605.07234#S4.Thmtheorem1)

Let:

- •hhbe the number of attention heads\.
- •dvd\_\{v\}be the dimension of each head\.
- •dm​o​d​e​ld\_\{model\}be the model dimension\.
- •Hi∈ℝn×dvH^\{i\}\\in\\mathbb\{R\}^\{n\\times d\_\{v\}\}be the output of theii\-th head \(Hi=Attention​\(Qi,Ki,Vi\)H^\{i\}=\\text\{Attention\}\(Q^\{i\},K^\{i\},V^\{i\}\)\)\.
- •WO∈ℝ\(h⋅dv\)×dm​o​d​e​lW\_\{O\}\\in\\mathbb\{R\}^\{\(h\\cdot d\_\{v\}\)\\times d\_\{model\}\}be the output weight matrix\.

Then, the standard MHA is defined as the concatenation of all head outputs followed by a linear projection:

Output=Concat​\(H1,H2,…,Hh\)​WO\\text\{Output\}=\\text\{Concat\}\(H^\{1\},H^\{2\},\\dots,H^\{h\}\)W\_\{O\}\(13\)
We can represent the concatenation as a partitioned row matrix:

H=\[H1H2…Hh\]H=\\begin\{bmatrix\}H^\{1\}&H^\{2\}&\\dots&H^\{h\}\\end\{bmatrix\}\(14\)
We can similarly partition the weight matrixWOW\_\{O\}vertically intohhblocks, where eachWOi∈ℝdv×dm​o​d​e​lW\_\{O\}^\{i\}\\in\\mathbb\{R\}^\{d\_\{v\}\\times d\_\{model\}\}:

WO=\[WO1WO2⋮WOh\]W\_\{O\}=\\begin\{bmatrix\}W\_\{O\}^\{1\}\\\\ W\_\{O\}^\{2\}\\\\ \\vdots\\\\ W\_\{O\}^\{h\}\\end\{bmatrix\}\(15\)
By applying the rules of block matrix multiplication:

Output=H​WO\\displaystyle=HW\_\{O\}\(16\)=\[H1H2…Hh\]​\[WO1WO2⋮WOh\]\\displaystyle=\\begin\{bmatrix\}H^\{1\}&H^\{2\}&\\dots&H^\{h\}\\end\{bmatrix\}\\begin\{bmatrix\}W\_\{O\}^\{1\}\\\\ W\_\{O\}^\{2\}\\\\ \\vdots\\\\ W\_\{O\}^\{h\}\\end\{bmatrix\}=\(H1​WO1\)\+\(H2​WO2\)\+⋯\+\(Hh​WOh\)\\displaystyle=\(H^\{1\}W\_\{O\}^\{1\}\)\+\(H^\{2\}W\_\{O\}^\{2\}\)\+\\dots\+\(H^\{h\}W\_\{O\}^\{h\}\)=∑i=1hHi​WOi\\displaystyle=\\sum\_\{i=1\}^\{h\}H^\{i\}W\_\{O\}^\{i\}□\\square

### F\.2Proof of Remark[4\.2](https://arxiv.org/html/2605.07234#S4.Thmtheorem2)

Consider a transformer layerllwithHHattention heads\. Letiidenote the current query position andjja cached token position\. The output of multi\-head attention at layerllis

𝐨l​\(i\)=∑h=1H𝐨l,h​\(i\)​WOl,h,\\mathbf\{o\}^\{l\}\(i\)=\\sum\_\{h=1\}^\{H\}\\mathbf\{o\}^\{l,h\}\(i\)\\,W\_\{O\}^\{l,h\},\(17\)where the output of headhhis

𝐨l,h​\(i\)=∑jAl,h​\(i,j\)​Vl,h​\(j\)\.\\mathbf\{o\}^\{l,h\}\(i\)=\\sum\_\{j\}A^\{l,h\}\(i,j\)\\,V^\{l,h\}\(j\)\.\(18\)Substituting and reordering terms yields

𝐨l​\(i\)=∑j∑h=1HAl,h​\(i,j\)​Vl,h​\(j\)​WOl,h\.\\mathbf\{o\}^\{l\}\(i\)=\\sum\_\{j\}\\sum\_\{h=1\}^\{H\}A^\{l,h\}\(i,j\)\\,V^\{l,h\}\(j\)\\,W\_\{O\}^\{l,h\}\.\(19\)
Equation \([19](https://arxiv.org/html/2605.07234#A6.E19)\) admits a natural decomposition over cached tokens\. In particular, the contribution of tokenjjto the layer output can be written as

Δ​𝐨l​\(i,j\)=∑h=1HAl,h​\(i,j\)​Vl,h​\(j\)​WOl,h\.\\Delta\\mathbf\{o\}^\{l\}\(i,j\)=\\sum\_\{h=1\}^\{H\}A^\{l,h\}\(i,j\)\\,V^\{l,h\}\(j\)\\,W\_\{O\}^\{l,h\}\.\(20\)This expression shows that a token’s effect on the layer output is not localized to any single head, but instead accumulates additively across all heads without any nonlinear operations\.□\\square

## Appendix GExtended Experimental Results

### G\.1Extended Evaluation on InfiniteBench Benchmark

To further assess the performance of our proposal under extreme long\-context settings, we conducted additional evaluations on the InfiniteBench benchmark[undefax](https://arxiv.org/html/2605.07234#bib.bib51)and compared against 2 representative baselines, AdaKV and CriticalKV\.

Since InfiniteBench involves extremely long contexts \(including some datasets with an average input length of upto 190K\+ tokens\), we evaluate on Meta\-Llama\-3\.1\-8B\-Instruct, which supports 128K context lengths\. To balance evaluation cost, we set the cache budget to 1024 tokens and the dataset limit to 100 samples per dataset\. The results are summarized in Table[8](https://arxiv.org/html/2605.07234#A7.T8)\.

Table 8:Comparison of different methods on InfiniteBench\. Best result is highlighted inbold\.MethodEn\.SumEn\.QAEn\.MCEn\.DiaCode\.DebugCode\.RunMath\.FindRet\.PassRet\.NumRet\.KVAvg\.Meta\-Llama\-3\.1\-8B\-Instruct,Btotal=1024​LB\_\{\\text\{total\}\}=1024LAdaKV23\.328\.667192215310094238\.40CriticalKV23\.349\.317082225310093038\.07LaProx21\.649\.8371102225310096438\.95

The results show that, on this highly challenging long\-context benchmark, LaProx consistently achieves stronger performance than our strongest baseline across most datasets for both models\. These results demonstrate the robustness of our method under extreme long\-context settings, as it maintains strong performance across different models and evaluation tasks\.

### G\.2Evaluations on Question\-Agnostic Setting

In our main experiments, we follow the standard practice of the prior works[undefac](https://arxiv.org/html/2605.07234#bib.bib30);[undefd](https://arxiv.org/html/2605.07234#bib.bib5);[undefag](https://arxiv.org/html/2605.07234#bib.bib34), where the question is compressed together with the context, to provide a unified evaluation scenario\. In this section, we adopt a question\-agnostic setting, where the context is compressed without any knowledge of the question, providing a more challenging evaluation\. The results are summarized in Tables[9](https://arxiv.org/html/2605.07234#A7.T9)\-[10](https://arxiv.org/html/2605.07234#A7.T10)\.

Evaluating RULER[undeft](https://arxiv.org/html/2605.07234#bib.bib21)is particularly computationally expensive, as it consists of 13 synthetic tasks and requires approximately 17 GPU hours for a single 32K\-context evaluation per configuration\. As a result, conducting a comprehensive sweep across all settings becomes prohibitively costly\. Therefore, in this Appendix, we restrict comparisons to 2 representative baselines, AdaKV and CriticalKV\.

Table 9:Comparison across 16 LongBench datasets\. The best result is highlighted inbold\.MethodSingle\-Document QAMulti\-Document QASummarizationFew\-shot LearningSyntheticCodeAvg\.NrtvQAQasperMF\-enHotpotQA2WikiMQAMusiqueGovReportQMSumMultiNewsTRECTriviaQASAMSumPCountPR\-enLccRB\-PMeta\-Llama\-3\.1\-8B\-Instruct, 20% CacheFullKV29\.2144\.655\.7358\.1449\.2632\.6134\.6525\.2426\.97392\.4543\.427\.6299\.563\.5852\.7349\.29AdaKV27\.0229\.0133\.1249\.1929\.4521\.2426\.8621\.8422\.835491\.3343\.846\.128066\.5955\.3841\.11CriticalKV29\.5630\.0233\.2551\.3733\.9825\.528\.8422\.5322\.925491\.7543\.857\.6195\.566\.7354\.6843\.25LaProx31\.3441\.2651\.9955\.0439\.2526\.3629\.9524\.0424\.056892\.243\.836\.9199\.568\.0256\.0947\.36

Table 10:Comparison across 13 Ruler datasets\. The best result is highlighted inbold\.MethodCWEFWENIAH\_MK1NIAH\_MK2NIAH\_MK3NIAH\_MQNIAH\_MVNIAH\_S1NIAH\_S2NIAH\_S3QA1QA2VTAvg\.Meta\-Llama\-3\.1\-8B\-Instruct, 20% CacheFullKV5194\.6710010010097\.599\.5100100100845410090\.82AdaKV1085\.338221257462\.5989421344792\.257\.38CriticalKV12\.675\.338814586\.2577\.25991001844489058\.26LaProx1290\.33100999797\.596\.7510010097655499\.285\.21

The results in Table[9](https://arxiv.org/html/2605.07234#A7.T9)and[10](https://arxiv.org/html/2605.07234#A7.T10)reveal that, under this challenging setting, methods that do not leverageV​WOVW\_\{O\}information suffer significant performance degradation\. In contrast, LaProx consistently maintains performance close to the full KV cache and outperforms all competing approaches\. Furthermore, on the more challenging NIAH variants \(multi\-needle\), most prior methods fail to reliably retrieve the needles, while LaProx maintains performance comparable to the FullKV setting\.

### G\.3Evaluations on Mixture\-of\-Experts

To demonstrate that our approach easily adapts across architectures, we evaluate it on Mixture\-of\-Experts \(MoE\) models\. Although MoE introduces sparsity in the Feed\-Forward Networks \(FFN\), the attention layers—where the KV cache resides—remain unchanged\. Since cache eviction operates within attention, our method transfers directly to MoE without any architectural modifications\. In this experiment, we use Qwen3\-30B\-A3B\-Instruct\-2507 and perform the same tasks and settings with Appendix[G\.2](https://arxiv.org/html/2605.07234#A7.SS2)\. The results are presented in Tables[11](https://arxiv.org/html/2605.07234#A7.T11)and[12](https://arxiv.org/html/2605.07234#A7.T12)\.

Table 11:Comparison across 16 LongBench datasets\. The best result is highlighted inbold\.MethodSingle\-Document QAMulti\-Document QASummarizationFew\-shot LearningSyntheticCodeAvg\.NrtvQAQasperMF\-enHotpotQA2WikiMQAMusiqueGovReportQMSumMultiNewsTRECTriviaQASAMSumPCountPR\-enLccRB\-PQwen3\-30B\-A3B\-Instruct\-2507, 20% CacheFullKV27\.7140\.0953\.8366\.9962\.132\.6930\.5622\.2424\.257688\.8248\.171410074\.6270\.852\.05AdaKV27\.5431\.1930\.394839\.6718\.3825\.8118\.5819\.24575\.5831\.1111\.095336\.8141\.7834\.57CriticalKV32\.0630\.7933\.1954\.0647\.4926\.1329\.0819\.8821\.047088\.2447\.021298\.575\.3469\.7947\.16LaProx32\.9437\.4146\.159\.4947\.4325\.3828\.8219\.9722\.077887\.2948\.541210075\.7871\.7449\.56

Table 12:Comparison across 13 Ruler datasets\. The best result is highlighted inbold\.MethodCWEFWENIAH\_MK1NIAH\_MK2NIAH\_MK3NIAH\_MQNIAH\_MVNIAH\_S1NIAH\_S2NIAH\_S3QA1QA2VTAvg\.Qwen3\-30B\-A3B\-Instruct\-2507, 20% CacheFullKV83\.299\.3310010010010098\.5100100100866410094\.69AdaKV57\.89281021601604324628\.424\.02CriticalKV60\.696100142100981001004344897\.265\.67LaProx64\.5971001009910092\.75100100100716399\.891\.31

The results in Tables[11](https://arxiv.org/html/2605.07234#A7.T11)and[12](https://arxiv.org/html/2605.07234#A7.T12)further demonstrate the effectiveness of our method in the MoE setting, where it limits performance degradation to only around 3

### G\.4Detailed Performance Analysis Across Cache Budgets on LongBench Benchmark

Tables[13](https://arxiv.org/html/2605.07234#A7.T13),[14](https://arxiv.org/html/2605.07234#A7.T14), and[15](https://arxiv.org/html/2605.07234#A7.T15)report detailed LongBench results shown in Figure[2](https://arxiv.org/html/2605.07234#S5.F2)for LaProx and competing baselines under cache budgets ranging from 128 to 1024 tokens on Meta\-Llama\-3\.1\-8B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, and Qwen3\-8B, respectively\.

Overall, the results show that LaProx consistently outperforms prior methods across all LongBench task categories on all evaluated models\.

Table 13:Comparison across 16 LongBench datasets on Meta\-Llama\-3\.1\-8B\-Instruct for cache budgets from 128L to 1024L\. The best result is highlighted inboldand the second best inunderline\.Table 14:Comparison across 16 LongBench datasets on Mistral\-7B\-Instruct\-v0\.3 for cache budgets from 128L to 1024L\. The best result is highlighted inboldand the second best inunderline\.MethodSingle\-Document QAMulti\-Document QASummarizationFew\-shot LearningSyntheticCodeAvg\.NrtvQAQasperMF\-enHotpotQA2WikiMQAMusiqueGovReportQMSumMultiNewsTRECTriviaQASAMSumPCountPR\-enLccRB\-PMistral\-7B\-Instruct\-v0\.3,Btotal=F​u​l​lB\_\{\\text\{total\}\}=FullFullKV29\.0741\.5852\.8849\.3739\.0128\.5834\.8125\.6627\.827688\.5947\.45\.59861\.462\.5348\.01Mistral\-7B\-Instruct\-v0\.3,Btotal=128​LB\_\{\\text\{total\}\}=128LSLLM21\.4222\.2826\.8237\.2833\.3117\.6116\.7319\.7717\.8645\.585\.6440\.475\.58055\.1352\.0736\.08SnapKV23\.8125\.0546\.9244\.9835\.0623\.1120\.4921\.5519\.3243\.589\.7842\.87693\.555\.9554\.640\.41AdaKV24\.0825\.5247\.0847\.1635\.1623\.8320\.2322\.2820\.74588\.9243\.776\.592\.556\.6355\.8140\.95CAKE25\.0427\.3248\.4947\.4835\.2424\.6321\.3222\.5420\.7845\.589\.4143\.3779556\.4754\.9941\.54CriticalKV25\.4225\.9546\.4346\.4435\.9223\.9321\.1222\.3519\.264789\.2343\.626\.593\.555\.9954\.5141\.07LaProx2830\.2654\.0247\.8637\.5425\.5421\.2623\.3122\.163\.589\.743\.4659658\.8557\.6644\.00Mistral\-7B\-Instruct\-v0\.3,Btotal=256​LB\_\{\\text\{total\}\}=256LSLLM22\.4923\.2129\.7539\.6232\.5116\.9619\.1419\.2920\.1254\.585\.1242\.985\.58057\.7155\.0537\.75SnapKV27\.3331\.0251\.5448\.9536\.5427\.0422\.1622\.9422\.0655\.589\.444\.52596\.558\.8758\.0343\.58AdaKV27\.0131\.1152\.3247\.1437\.0628\.0122\.4523\.0522\.6562\.589\.4444\.31696\.558\.6559\.3244\.22CAKE27\.131\.4253\.349\.3837\.5627\.6322\.8723\.1723\.085789\.2444\.574\.596\.559\.559\.4744\.14CriticalKV26\.831\.3252\.3548\.8536\.8527\.3322\.5523\.122\.9558\.589\.2343\.964\.596\.559\.1859\.3843\.96LaProx28\.5334\.7154\.8748\.8837\.6228\.9923\.5324\.0723\.6462\.589\.0344\.295\.59659\.1559\.9545\.08Mistral\-7B\-Instruct\-v0\.3,Btotal=512​LB\_\{\\text\{total\}\}=512LSLLM24\.1925\.8930\.4540\.632\.3617\.3522\.0220\.223\.2865\.586\.9543\.7568159\.2956\.3439\.70SnapKV28\.0833\.9353\.9148\.9637\.6327\.392423\.7724\.076789\.2345\.26596\.559\.9560\.2245\.31AdaKV29\.0434\.6453\.3548\.9936\.7227\.6624\.4123\.9924\.477189\.0345\.6959760\.0360\.9645\.74CAKE28\.6835\.5453\.6848\.6438\.628\.3225\.0723\.7625\.0569\.589\.4445\.645\.597\.559\.9860\.8745\.98CriticalKV28\.0937\.0753\.7149\.5237\.3428\.0824\.4824\.2524\.847189\.3344\.283\.597\.560\.3361\.3345\.91LaProx29\.737\.2952\.949\.5738\.7328\.4725\.4125\.0924\.7774\.589\.6645\.9469760\.8361\.8646\.74Mistral\-7B\-Instruct\-v0\.3,Btotal=1024​LB\_\{\\text\{total\}\}=1024LSLLM24\.7927\.8830\.9942\.9132\.6518\.0324\.6420\.725\.4568\.588\.7145\.375\.582\.561\.159\.2441\.19SnapKV28\.3537\.4753\.4149\.1239\.4628\.3826\.2924\.6725\.947188\.8946\.84597\.561\.0861\.7746\.57AdaKV28\.1437\.3552\.8648\.9738\.5928\.2426\.624\.8225\.8172\.589\.1946\.12598\.560\.8262\.3646\.61CAKE29\.9337\.653\.1250\.338\.6328\.3627\.4324\.4926\.737389\.1945\.95697\.561\.0162\.1346\.96CriticalKV29\.0637\.2853\.949\.9238\.2428\.2126\.6224\.1426\.5673\.589\.1945\.27597\.561\.6962\.6746\.79LaProx28\.9939\.1652\.150\.3938\.8728\.0527\.7725\.1226\.367689\.61475\.59961\.4762\.3147\.36

Table 15:Comparison across 16 LongBench datasets on Qwen3\-8B for cache budgets from 128L to 1024L\. The best result is highlighted inboldand the second best inunderline\.MethodSingle\-Document QAMulti\-Document QASummarizationFew\-shot LearningSyntheticCodeAvg\.NrtvQAQasperMF\-enHotpotQA2WikiMQAMusiqueGovReportQMSumMultiNewsTRECTriviaQASAMSumPCountPR\-enLccRB\-PQwen3\-8B,Btotal=F​u​l​lB\_\{\\text\{total\}\}=FullFullKV32\.2646\.5952\.7359\.3551\.2133\.232\.4424\.1325\.687165\.8240110053\.0650\.1746\.17Qwen3\-8B,Btotal=128​LB\_\{\\text\{total\}\}=128LSLLM12\.1828\.0521\.7936\.3439\.076\.8815\.7818\.5415\.954362\.5633\.337345\.1645\.4331\.25SnapKV16\.3431\.3542\.3752\.6742\.3318\.2616\.0319\.6816\.495375\.6934\.4419946\.8647\.8638\.33AdaKV21\.6632\.5843\.9650\.8142\.9520\.916\.4519\.3316\.025173\.9234\.2419845\.7746\.9238\.47CAKE19\.6932\.6745\.0757\.548\.0220\.517\.8220\.8716\.734769\.5135\.3959945\.5545\.9439\.14CriticalKV18\.4735\.8345\.7848\.5942\.8419\.0916\.6719\.9417\.615675\.3135\.309947\.7948\.2839\.15LaProx17\.5136\.9151\.2258\.6446\.3123\.8518\.6821\.2117\.326176\.4935\.8619948\.8651\.2141\.57Qwen3\-8B,Btotal=256​LB\_\{\\text\{total\}\}=256LSLLM14\.0128\.8322\.0536\.9938\.977\.8318\.618\.718\.735270\.434\.7956749\.8646\.8833\.16SnapKV18\.237\.5449\.4463\.1348\.9324\.4520\.421\.719\.25970\.6237\.04010050\.6453\.1742\.09AdaKV22\.3136\.0547\.7557\.0744\.6222\.3320\.4720\.4618\.836474\.0737\.82110050\.1249\.9141\.67CAKE20\.3738\.7850\.3164\.0148\.2725\.4722\.0221\.5619\.816173\.1738\.12910049\.1449\.6643\.16CriticalKV22\.1540\.4148\.8958\.0146\.8327\.8921\.6221\.7220\.16870\.5237\.64010051\.3552\.1442\.97LaProx23\.541\.2650\.9662\.9846\.5530\.4422\.5522\.2921\.637072\.9738\.17010052\.515344\.30Qwen3\-8B,Btotal=512​LB\_\{\\text\{total\}\}=512LSLLM15\.1331\.5624\.9239\.0741\.298\.1622\.2118\.8722\.256274\.2135\.375152\.2250\.9234\.76SnapKV22\.1142\.5953\.0561\.2849\.6830\.6523\.9423\.2621\.886769\.2739\.44010054\.1253\.1644\.46AdaKV22\.5640\.5548\.4660\.7647\.5730\.5723\.9321\.8821\.476768\.9738\.13010052\.2852\.5243\.54CAKE23\.0643\.7751\.5360\.5349\.8431\.8425\.2523\.1122\.766475\.4338\.42610053\.0253\.4545\.12CriticalKV25\.4743\.9851\.3760\.748\.8128\.4425\.522\.7522\.676971\.4738\.47010054\.8953\.7544\.82LaProx24\.7645\.8452\.6861\.9649\.9828\.8725\.3723\.324\.337174\.7738\.3010053\.853\.7645\.55Qwen3\-8B,Btotal=1024​LB\_\{\\text\{total\}\}=1024LSLLM19\.1833\.3427\.0344\.740\.999\.5925\.620\.1521\.026380\.0237\.2373953\.2950\.5335\.73SnapKV26\.3843\.4953\.262\.2548\.9229\.7727\.3123\.4920\.876971\.9738\.09110054\.5151\.9545\.13AdaKV24\.8842\.5951\.6860\.1846\.7830\.5626\.6821\.9620\.887074\.7738\.43010053\.2851\.4244\.63CAKE24\.0944\.9752\.0260\.5850\.6231\.5427\.8323\.2524\.337073\.9937\.71110054\.6753\.5645\.63CriticalKV24\.9743\.6151\.9560\.0248\.7328\.828\.232321\.287073\.737\.59010054\.7153\.845\.02LaProx27\.7146\.1253\.0461\.8350\.4830\.7227\.5822\.921\.077076\.2741\.35110054\.1852\.3746\.04

Similar Articles

KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs

Hugging Face Daily Papers

KV Packet proposes a recomputation-free cache reuse framework for LLMs that uses trainable soft-token adapters to bridge context discontinuities, eliminating overhead while maintaining performance comparable to full recomputation baselines on Llama-3.1 and Qwen2.5.