为LLM推理压缩长上下文至答案对齐的记忆嵌入
摘要
本文提出一种上下文与答案对齐的记忆压缩(CMC)框架,通过将长输入上下文压缩为紧凑的记忆嵌入,在不修改解码器权重的情况下降低LLM推理成本,显著提升性能和效率。
arXiv:2609.25537v1 Announce Type: new
Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing soft-compression methods either lack query-guided memory selection at inference time, train without answer-targeted supervision, or couple compression tightly to a specific decoder architecture. We propose a Context-to-Answer-Aligned Memory Compression (CMC) framework, which compresses long input contexts into compact Context Memory Embeddings (CMEs) aligned to any frozen decoder's embedding space, reducing inference costs without modifying decoder weights. CMC introduces a two-tier KV cache that combines question-guided CME selection with a local context window, and trains the compressor with answer-targeted distillation from a frozen LLM. Experiments across nine encoder-decoder combinations and four QA benchmarks show that CMC consistently outperforms the baseline, achieving up to 7.3 EM and 4.0 F1 point gains on SQuAD, while reducing inference time and energy consumption by up to 20% and peak reserved GPU memory by up to 50% at 3,000 generation tokens. Ablation studies confirm that each architectural component and training objective contributes to the performance.
查看缓存全文
缓存时间: 2026/09/23 09:16
# Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference
Source: [https://arxiv.org/html/2609.25537](https://arxiv.org/html/2609.25537)
Md Mostafizer RahmanMd Faizul Ibne AminAffiliation:The University of Aizu, Aizuwakamatsu, JapanMd Shahajada MiaAffiliation:The University of Aizu, Aizuwakamatsu, JapanYutaka WatanobeAffiliation:The University of Aizu, Aizuwakamatsu, JapanFang Liu††thanks:Corresponding authorAffiliation:Lucy Family Institute for Data & Society, University of Notre Dame, IN, USAAffiliation:Department of Applied and Computational Mathematics and StatisticsUniversity of Notre Dame, IN, USA\{mrahman3, fliu2\}@nd\.edu,\{fiamin, d8262103, yutaka\}@u\-aizu\.ac\.jp
###### Abstract
Large language model \(LLM\) inference is constrained by the quadratic scaling of self\-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales\. Existing soft\-compression methods either lack query\-guided memory selection at inference time, train without answer\-targeted supervision, or couple compression tightly to a specific decoder architecture\. We propose a Context\-to\-Answer\-Aligned Memory Compression \(CMC\) framework, which compresses long input contexts into compact Context Memory Embeddings \(CMEs\) aligned to any frozen decoder’s embedding space, reducing inference costs without modifying decoder weights\. CMC introduces a two\-tier KV cache that combines question\-guided CME selection with a local context window, and trains the compressor with answer\-targeted distillation from a frozen LLM\. Experiments across nine encoder\-decoder combinations and four QA benchmarks show that CMC consistently outperforms the baseline, achieving up to 7\.3 EM and 4\.0 F1 point gains on SQuAD, while reducing inference time and energy consumption by up to 20% and peak reserved GPU memory by up to 50% at3,0003,000generation tokens\. Ablation studies confirm that each architectural component and training objective contributes to the performance\.
## 1Introduction
Figure 1:Overview of the CMC framework: \(ii\) context denoising and chunked compression via ContextEncoder,iiii\) cross\-architecture projection via MemoryBridge, and \(iiiiii\) two\-tier KV cache inference via DecoderLLM with question\-guided Tier\-1 CME selection and Tier\-2 local window\.Large language models \(LLMs\) have demonstrated exceptional performance across a wide range of natural language understanding and generation tasks[Brown et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib3);[Touvron et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib4);[Jiang et al\. \(2023a\)](https://arxiv.org/html/2609.25537#bib.bib5);[Gemma Team et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib6)\. Central to their success is the Transformer architecture[Vaswani et al\. \(2017\)](https://arxiv.org/html/2609.25537#bib.bib1), whose self\-attention mechanism enables rich contextual reasoning over input sequences\. However, as the demand for processing longer documents, multi\-turn dialogues, and multi\-document reasoning grows, a fundamental bottleneck emerges: the computational cost of self\-attention scales quadratically with sequence length, and the KV cache grows linearly, consuming substantial GPU memory, increasing inference latency, and driving up energy consumption with every additional token[Pope et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib27);[Sheng et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib21)\. These constraints become particularly acute in real\-world deployments where contexts routinely span thousands of tokens, and throughput and energy efficiency are critical requirements alongside model quality\.
Architectural methods address this bottleneck by redesigning the attention mechanism itself, either through sparse or local attention patterns[Beltagy et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib7);[Zaheer et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib8)or by extending positional encodings to accommodate longer sequences[Chen et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib9);[Dong et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib23)\. While effective, these approaches require substantial architectural modification or continued pretraining, making them difficult to apply to already\-deployed LLMs\. An alternative is hard\-prompt compression, which discards tokens deemed redundant by self\-information scoring or perplexity\-based ranking[Li et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib10);[Pan et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib22)\. Although model\-agnostic, hard\-prompt methods incur irreversible information loss and degrade on tasks requiring evidence synthesis across distant passages, such as multi\-hop reasoning and long\-document QA[Li et al\. \(2025a\)](https://arxiv.org/html/2609.25537#bib.bib15)\. In contrast, soft\-compression methods encode the full context into a compact set of dense embeddings that a frozen decoder can attend to in place of the original tokens[Chevalier et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib12);[Ge et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib11);[Mu et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib13);[Kim et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib14);[Dai et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib16), achieving higher compression rates while preserving semantic content\.[Dai et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib16)propose PCC, a compressor\-LLM framework that pretrains a lightweight encoder and converter to compress context into dense memory slots consumed by a frozen decoder\. However, PCC supplies all compressed memory slots uniformly at inference time with no mechanism for query\-guided selection, and trains solely on text reconstruction without answer\-targeted supervision\. More broadly, no existing soft\-compression method simultaneously addresses cross\-architecture deployment, answer\-targeted training supervision, and query\-guided memory selection at inference stage\.
To address these gaps, we propose aContext\-to\-Answer\-Aligned Memory Compression \(CMC\)framework, illustrated in Figure[1](https://arxiv.org/html/2609.25537#S1.F1)\. CMC compresses long input contexts into compact Context Memory Embeddings \(CMEs\) projected into any frozen decoder’s own embedding space, enabling efficient inference without modifying decoder weights\. CMC comprises three key components: \(ii\) a trainableContextEncoderthat applies graph\-based context denoising and chunked compression to produce CMEs via fixed placeholder tokens; \(iiii\) aMemoryBridge, a norm\-calibrated two\-layer MLP that projects CMEs into the frozen decoder’s embedding space, enabling cross\-architecture pairing of any encoder with any decoder; and \(iiiiii\) a frozenDecoderLLMthat generates responses conditioned on a two\-tier KV cache combining top\-KKCMEs selected by cosine similarity to the input question \(Tier\-1\) and a local context window at full token resolution \(Tier\-2\)\. The ContextEncoder and MemoryBridge are trained jointly in two phases: Phase\-1 on autoencoding and autoregressive objectives to align CMEs with the decoder, and Phase\-2 on knowledge distillation and contrastive memory\-answer alignment using a frozen DecoderLLM to incorporate answer\-targeted supervision\. The main contributions of this work are as follows\.
- •We propose CMC, a soft\-compression framework that incorporates graph\-based context denoising to improve CME quality and a two\-tier KV cache inference policy that combines question\-guided CME selection with a local context window, providing*both global*semantic coverage*and local*span\-level precision without full\-context attention\.
- •We propose a two\-phase training strategy\. Phase\-1 aligns CMEs with the frozen decoder via reconstruction objectives, and Phase\-2 distills answer\-targeted knowledge from a frozen DecoderLLM \(educator\) via cross\-entropy, KL divergence, and contrastive memory\-answer alignment losses, introducing answer\-targeted training supervision for cross\-architecture soft\-compression\.
- •We conduct comprehensive experiments across nine encoder\-decoder combinations and four QA benchmarks \(SQuAD, AdversarialQA, HotpotQA, CovidQA\) and show that CMC consistently outperforms the baseline in EM and F1 with notable reductions in inference time, energy consumption, multiply–accumulate operation \(MAC\), and peak GPU memory\.
## 2Methods
We propose the CMC framework to compress a long input contextCCinto compact CMEs that a frozen DecoderLLM attends to in place of the original tokens\. Given contextCCand questionqq, CMC learns a ContextEncoderfθcf\_\{\\theta\_\{c\}\}and MemoryBridgegϕg\_\{\\phi\}such that:
a^\\displaystyle\\hat\{a\}=fθd\(⋅∣M∗,Clocal,q\),where\\displaystyle=f\_\{\\theta\_\{d\}\}\(\\cdot\\mid M^\{\*\},C\_\{\\text\{local\}\},q\),\\mbox\{ where\}\(1\)M∗\\displaystyle M^\{\*\}=TopK\(gϕ\(fθc\(C\)\),q,K\)\.\\displaystyle=\\text\{TopK\}\(g\_\{\\phi\}\(f\_\{\\theta\_\{c\}\}\(C\)\),\\,q,\\,K\)\.\(2\)a^\\hat\{a\}is the answer generated by the DecoderLLMfθdf\_\{\\theta\_\{d\}\},M∗M^\{\*\}is the set of query\-selected decoder\-ready CMEs \(formulated in Sections[2\.2](https://arxiv.org/html/2609.25537#S2.SS2)and[2\.3](https://arxiv.org/html/2609.25537#S2.SS3)\), andClocalC\_\{\\text\{local\}\}is the local window \(formulated in Section[2\.5](https://arxiv.org/html/2609.25537#S2.SS5)\)\. The ContextEncoderfθcf\_\{\\theta\_\{c\}\}and MemoryBridgegϕg\_\{\\phi\}are the only trainable components, whereasfθdf\_\{\\theta\_\{d\}\}remains frozen throughout\. The following subsections detail each component as illustrated in Figure[1](https://arxiv.org/html/2609.25537#S1.F1)\.
### 2\.1Context Denoising, Segmentation, and Memory Tokens
We apply graph\-based context denoising to remove low\-salience tokens – tokens that are semantically isolated from their immediate neighbors – from the full contextCCbefore compression\. Given contextC=\(x1,…,xn\)C=\(x\_\{1\},\\ldots,x\_\{n\}\), we compute the cosine similarity between each token and itsggimmediately downstream neighbors in embedding space\. A tokenxix\_\{i\}is retained if its cosine similarity to at least one of itsggsubsequent neighbors meets or exceeds the tunable thresholdτ\\tau:
Ci′=\{xi∈C\|maxj=i\+1min\(i\+g,n\)𝐞i⋅𝐞j‖𝐞i‖‖𝐞j‖≥τ\}\.C^\{\\prime\}\_\{i\}=\\left\\\{x\_\{i\}\\in C\\;\\middle\|\\;\\max\_\{j=i\+1\}^\{\\min\(i\+g,\\,n\)\}\\frac\{\\mathbf\{e\}\_\{i\}\\cdot\\mathbf\{e\}\_\{j\}\}\{\\\|\\mathbf\{e\}\_\{i\}\\\|\\\|\\mathbf\{e\}\_\{j\}\\\|\}\\geq\\tau\\right\\\}\.\(3\)xix\_\{i\}is discarded ifCi′C\_\{i\}^\{\\prime\}is an empty set\. For retainedxix\_\{i\}, the denoised contextC′C^\{\\prime\}ofk∈\(1,g\]k\\in\(1,g\]tokens is divided intonsn\_\{s\}non\-overlapping segments, each of fixed lengthtt, yielding𝒮=\{s1,s2,…,sns\}\\mathcal\{S\}=\\\{s\_\{1\},s\_\{2\},\\ldots,s\_\{n\_\{s\}\}\\\}, wherens=⌈k/t⌉n\_\{s\}=\\lceil k/t\\rceil\. To each segmentsis\_\{i\}, we appendm\(c\)m^\{\(c\)\}fixed placeholder tokens bounded by special<MEM\>and</MEM\>tokens, forming the augmented sequences~i\\tilde\{s\}\_\{i\}:
s~i=\[si∥<MEM\>∥p1,…,pm\(c\)⏟m\(c\)placeholders∥</MEM\>\],\\tilde\{s\}\_\{i\}\\\!=\\\!\\left\[s\_\{i\}\\Big\\\|\\texttt\{<MEM\>\}\\Big\\\|\\underbrace\{\\texttt\{p\}\_\{1\},\\ldots,\\texttt\{p\}\_\{m^\{\(c\)\}\}\}\_\{m^\{\(c\)\}\\text\{ placeholders\}\}\\Big\\\|\\texttt\{</MEM\>\}\\right\],\(4\)The number of placeholder tokens per segment is controlled by the compression rater=t/m\(c\)r=t/m^\{\(c\)\}; a higherrryields a more compact representation at the cost of reduced context coverage\. We evaluater∈\{2,4,8\}r\\in\\\{2,4,8\\\}in our experiments\.
### 2\.2CME Extraction with ContextEncoder
Each augmented segments~i\\tilde\{s\}\_\{i\}in Eq\. \([4](https://arxiv.org/html/2609.25537#S2.E4)\) is processed independently by the ContextEncoderfθcf\_\{\\theta\_\{c\}\}, a trainable decoder\-only causal language model\. In particular, via causal self\-attention, each of them\(c\)m^\{\(c\)\}placeholder tokens attends to all preceding chunk tokens insis\_\{i\}, condensing the segment content intom\(c\)m^\{\(c\)\}hidden state vectorsHi\(c\)∈ℝm\(c\)×HcH^\{\(c\)\}\_\{i\}\\in\\mathbb\{R\}^\{m^\{\(c\)\}\\times H\_\{c\}\}extracted as the segment\-level CMEs from the final transformer layer:
Hi\(c\)\\displaystyle H^\{\(c\)\}\_\{i\}=fθc\(s~i\)\[placeholder positions\],\\displaystyle=f\_\{\\theta\_\{c\}\}\(\\tilde\{s\}\_\{i\}\)\[\\text\{placeholder positions\}\],\(5\)whereHcH\_\{c\}is the dimension of each hidden state vector\. Though the placeholder tokens are fixed, their hidden states become informative CMEs through training the ContextEncoder weights\. After processing allnsn\_\{s\}segments, the full CME matrix is formed by concatenation:
𝐇\(c\)=\[H1\(c\);H2\(c\);…;Hns\(c\)\]∈ℝMb×Hc,\\mathbf\{H\}^\{\(c\)\}=\[H^\{\(c\)\}\_\{1\};\\,H^\{\(c\)\}\_\{2\};\\,\\ldots;\\,H^\{\(c\)\}\_\{n\_\{s\}\}\]\\in\\mathbb\{R\}^\{M\_\{b\}\\times H\_\{c\}\},\(6\)whereMb=ns×m\(c\)M\_\{b\}=n\_\{s\}\\times m^\{\(c\)\}is the total memory vectors representing the full context\.
### 2\.3Cross\-Architecture Projection with MemoryBridge
The ContextEncoder and the frozen DecoderLLM may belong to different model families with different hidden dimensionsHcH\_\{c\}andHdH\_\{d\}\. MemoryBridge is a trainable two\-layer MLP that projects each CME vector𝐡∈ℝHc\\mathbf\{h\}\\in\\mathbb\{R\}^\{H\_\{c\}\}from ContextEncoder into the DecoderLLM’s embedding spaceℝHd\\mathbb\{R\}^\{H\_\{d\}\}\. Let𝐦^\\hat\{\\mathbf\{m\}\}denote the output of the two\-layer MLP after LayerNorm111See Appendix[A\.1](https://arxiv.org/html/2609.25537#A1.SS1)for the LayerNorm formulation used\.\. We apply
𝐦~=νtanh\(𝐦^\)⋅min\(1,2ν‖νtanh\(𝐦^\)‖2\),\\tilde\{\\mathbf\{m\}\}\\\!=\\\!\\nu\\tanh\(\\hat\{\\mathbf\{m\}\}\)\\cdot\\min\\\!\\left\(1,\\ \\frac\{2\\nu\}\{\\left\\\|\\nu\\tanh\(\\hat\{\\mathbf\{m\}\}\)\\right\\\|\_\{2\}\}\\right\),\(7\)wheretanh\\tanhis applied coordinatewise andν=s⋅ν¯\\nu=s\\cdot\\bar\{\\nu\}, in whichssis a learnable scalar andν¯\\bar\{\\nu\}denotes the averageℓ2\\ell\_\{2\}norm of the frozen decoder’s input token embedding vectors, computed across all vocabulary tokens once at initialization\. The first factor,νtanh\(𝐦^\)\\nu\\,\\tanh\(\\hat\{\\mathbf\{m\}\}\), bounds each coordinate but not theℓ2\\ell\_\{2\}norm of the vector; the second factor rescales the full vector whenever itsℓ2\\ell\_\{2\}norm exceeds2ν2\\nu, and leaves it unchanged otherwise\. Eq\. \([7](https://arxiv.org/html/2609.25537#S2.E7)\) guarantees‖𝐦~‖2≤2ν\\\|\\tilde\{\\mathbf\{m\}\}\\\|\_\{2\}\\leq 2\\nu, preventing the projected CMEs from exceeding twice the typicalℓ2\\ell\_\{2\}norm of the DecoderLLM’s own token embeddings\.
Applying MemoryBridgegϕg\_\{\\phi\}to each segment’s CME matrixHi\(c\)H^\{\(c\)\}\_\{i\}yields its representation:
M\(i\)=gϕ\(Hi\(c\)\)∈ℝm\(c\)×Hd\.M^\{\(i\)\}=g\_\{\\phi\}\(H^\{\(c\)\}\_\{i\}\)\\in\\mathbb\{R\}^\{m^\{\(c\)\}\\times H\_\{d\}\}\.After allnsn\_\{s\}segments are processed, the decoder\-ready representations are concatenated to
M=\[M\(1\);M\(2\);…;M\(ns\)\]∈ℝMb×Hd\.M=\[M^\{\(1\)\};\\,M^\{\(2\)\};\\,\\ldots;\\,M^\{\(n\_\{s\}\)\}\]\\in\\mathbb\{R\}^\{M\_\{b\}\\times H\_\{d\}\}\.\(8\)to form the full memory matrixMMthat provides the DecoderLLM with a compressed sequence ofHdH\_\{d\}\-dimensional memory vectors in place of the original context tokens\.
### 2\.4Two\-Phase Training
The ContextEncoderfθcf\_\{\\theta\_\{c\}\}and MemoryBridgegϕg\_\{\\phi\}are the same trainable modules optimized continuously across both phases in sequence, with the DecoderLLMfθdf\_\{\\theta\_\{d\}\}frozen throughout\.
##### Phase\-1: Decoder alignment\.
Phase\-1 aligns the CMEs with the decoder’s embedding space via autoencoding \(AE\) and autoregressive \(AR\) objectives\. The AE loss conditions the decoder onMMin Eq\. \([8](https://arxiv.org/html/2609.25537#S2.E8)\) to reconstruct the original context; the AR loss conditions the decoder on the sameMMto reconstruct only the first half of the context, leading to the loss function in Phase 1:
ℒ1=ℒAE\+λARℒAR\.\\mathcal\{L\}\_\{1\}=\\mathcal\{L\}\_\{\\text\{AE\}\}\+\\lambda\_\{\\text\{AR\}\}\\mathcal\{L\}\_\{\\text\{AR\}\}\.\(9\)This loss formulation ensures that CMEs carry sufficient information for the frozen decoder to reconstruct the context, where the AR term, with hyperparameterλAR\>0\\lambda\_\{\\text\{AR\}\}\>0, up\-weights the reconstruction of the earlier portion relative to the full\-context AE objective\.
##### Phase\-2: Answer\-aligned CMEs\.
Phase\-2 fine\-tunes the CMEs from Phase\-1 toward answer\-relevant content using the frozen DecoderLLM \(educator\)fθdf\_\{\\theta\_\{d\}\}, which receives the local windowClocalC\_\{\\text\{local\}\}and questionqq\. The same frozen weightsfθdf\_\{\\theta\_\{d\}\}are also invoked as the DecoderLLM \(learner\), conditioned onMMandClocalC\_\{\\text\{local\}\}\. The two DecoderLLMs share identical parameters and differ only in their input\.
Three objectives act jointly: a cross\-entropy lossℒCE\\mathcal\{L\}\_\{\\text\{CE\}\}supervises the DecoderLLM \(learner\) on gold answer tokens; a KL divergence lossℒKL\\mathcal\{L\}\_\{\\text\{KL\}\}aligns the DecoderLLM \(learner\)’s output distribution with the DecoderLLM \(educator\)’s; and a contrastive lossℒCL\\mathcal\{L\}\_\{\\text\{CL\}\}\([van den Oord et al\., 2018](https://arxiv.org/html/2609.25537#bib.bib19)\)pulls the element\-wise mean of the CMEs toward the DecoderLLM \(educator\)’s hidden state at the answer span\. Taken together, the Phase\-2 loss is:
ℒ2=ℒCE\+λKLℒKL\+λCLℒCL,\\mathcal\{L\}\_\{2\}=\\mathcal\{L\}\_\{\\text\{CE\}\}\+\\lambda\_\{\\text\{KL\}\}\\,\\mathcal\{L\}\_\{\\text\{KL\}\}\+\\lambda\_\{\\text\{CL\}\}\\,\\mathcal\{L\}\_\{\\text\{CL\}\},\(10\)whereλKL\>0\\lambda\_\{\\text\{KL\}\}\>0andλCL\>0\\lambda\_\{\\text\{CL\}\}\>0are hyperparameters\. Our ablation study confirms that the three components in the Phase\-2 loss in Eq\. \([10](https://arxiv.org/html/2609.25537#S2.E10)\) act jointly – removing any component degrades performance to the level of Phase\-1 alone\.
### 2\.5Inference with Two\-Tier KV Cache Policy
At the inference stage, the frozen DecoderLLM generates responses conditioned on a two\-tier KV cache built from the decoder\-ready CME representationMMand questionqq\.
##### Tier\-1: Question\-guided top\-KKselection\.
The question tokens are embedded via the frozen decoder’s embedding layer and mean\-pooled to obtain a query vector𝐪vec∈ℝHd\\mathbf\{q\}\_\{\\text\{vec\}\}\\in\\mathbb\{R\}^\{H\_\{d\}\}\. Each CME𝐦~j∈M\\tilde\{\\mathbf\{m\}\}\_\{j\}\\in Mis scored by cosine similarity:
scorej=𝐦~j⋅𝐪vec‖𝐦~j‖‖𝐪vec‖\.\\text\{score\}\_\{j\}=\\frac\{\\tilde\{\\mathbf\{m\}\}\_\{j\}\\cdot\\mathbf\{q\}\_\{\\text\{vec\}\}\}\{\\\|\\tilde\{\\mathbf\{m\}\}\_\{j\}\\\|\\\|\\mathbf\{q\}\_\{\\text\{vec\}\}\\\|\}\.\(11\)The top\-KKCMEs are selected and re\-ordered by original document position to formM∗M^\{\*\}, whereKKis a tunable hyperparameter controlling the breadth of semantic coverage\.
##### Tier\-2: Local context window\.
AWW\-token windowClocalC\_\{\\text\{local\}\}is extracted fromCCcentered at the starting\-token positionα\\alphaof the gold answer span, subject to context boundary constraints:
Clocal\\displaystyle C\_\{\\text\{local\}\}=C\[start:start\+W\],where\\displaystyle=C\[\\text\{start\}:\\text\{start\}\+W\],\\mbox\{ where\}\(12\)start=max\(0,min\(\|C\|−W,α−W/2\)\)\.\\displaystyle=\\max\\left\(\\\!0,\\min\\left\(\|C\|\\\!\-\\\!W,\\alpha\\\!\-W/2\\right\)\\right\)\.Tier\-2 preserves the information critical region at full token resolution, providing the local textual precision that CMEs alone cannot supply\. In our experiments,α\\alphais taken from benchmark annotations when available and otherwise computed by string matching of the information\-critical region within the contextCC\. Both CMC and the baseline use the sameα\\alphain every comparison\. In deployment settings whereα\\alphais unavailable, a lightweight span localization step such as CME\-guided or BM25 retrieval can be used; see Appendix[A\.8](https://arxiv.org/html/2609.25537#A1.SS8)\.
##### KV cache assembly\.
The prefill input is
input=\[M∗‖Clocal‖q\]\.\\text\{input\}=\[\\,M^\{\*\}\\;\\\|\\;C\_\{\\text\{local\}\}\\;\\\|\\;q\\,\]\.\(13\)During autoregressive generation, the total KV cache length is bounded byK\+W\+\|q\|K\+W\+\|q\|tokens at all generation steps regardless of document length, where\|q\|\|q\|denotes the number of tokens in the questionqq\.
## 3Experiments
### 3\.1Datasets
We evaluate CMC on four extractive QA benchmarks covering diverse reasoning types\. SQuAD[Rajpurkar et al\. \(2016\)](https://arxiv.org/html/2609.25537#bib.bib20)is a single\-hop extractive QA dataset where answers are spans within a single passage\. AdversarialQA[Bartolo et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib25)extends this with adversarially constructed questions designed to challenge model comprehension\. HotpotQA[Yang et al\. \(2018\)](https://arxiv.org/html/2609.25537#bib.bib24)requires multi\-hop reasoning across multiple supporting passages, while CovidQA[Möller et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib26)evaluates domain\-specific QA on biomedical literature\. We conduct experiments under two data settings\. Thecontrolled\-data settinguses 10,000 training and 1,000 validation samples for SQuAD, AdversarialQA, and HotpotQA, and 1,500 training and 200 validation samples for CovidQA\. Thefull\-data settinguses the complete training and validation splits of SQuAD, AdversarialQA, and HotpotQA to enable direct comparison with published baselines; CovidQA is evaluated under the small\-data setting only due to its limited corpus size\. This selection covers single\-hop, adversarial, multi\-hop, and domain\-specific reasoning, providing a broad evaluation of CMC across varied context lengths and answer types\.
SQuADAdversarialQAHotpotQACovidQA†AverageDecoderLLMModelEMF1EMF1EMF1EMF1EMF1Controlled\-data settingLlama\-38B\-InstructLlama \(baseline\)62\.481\.034\.353\.344\.168\.122\.563\.840\.866\.6CMC67\.082\.936\.855\.149\.669\.329\.564\.245\.767\.9Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)61\.075\.929\.647\.953\.071\.119\.057\.640\.763\.1CMC60\.275\.532\.250\.353\.871\.722\.063\.142\.165\.2Gemma\-29B\-ITGemma \(baseline\)55\.178\.132\.956\.440\.167\.010\.562\.134\.765\.9CMC61\.181\.236\.859\.542\.070\.614\.063\.238\.568\.6Full\-data settingLlama\-38B\-InstructLlama \(baseline\)48\.2874\.2732\.4752\.5544\.3167\.28\-\-41\.6964\.70CMC55\.5678\.2935\.4353\.4247\.7767\.83\-\-46\.2566\.51Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)45\.3669\.2626\.3745\.9153\.4171\.13\-\-41\.7162\.10CMC30\.6755\.1629\.7749\.2648\.3365\.49\-\-35\.3755\.38Gemma\-29B\-ITGemma \(baseline\)40\.9171\.6829\.7753\.1737\.8065\.60\-\-36\.1663\.48CMC44\.1373\.2733\.3356\.2441\.9069\.10\-\-39\.7966\.20Table 1:EM and F1 on four QA benchmarks under controlled\-data and full\-data settings\.Bolddenotes the higher score per row\.†CovidQA is evaluated under the controlled\-data setting only due to its limited corpus size\.
### 3\.2Experimental Settings
All experiments are conducted on NVIDIA A100 80 GB GPU\. The DecoderLLM is loaded in 4\-bit NF4 quantisation with double quantisation enabled and kept frozen throughout training and inference\. We evaluate three ContextEncoders – GPT2\-Large, OPT\-1\.3B, and OPT\-2\.7B – paired with three frozen DecoderLLMs – Llama\-3\-8B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, and Gemma\-2\-9B\-IT – yielding nine encoder\-decoder combinations\. The ContextEncoder and MemoryBridge are trained jointly for one epoch using AdamW \(learning rate=1×10−5=1\\times 10^\{\-5\},β=\(0\.9,0\.98\)\\beta=\(0\.9,\\,0\.98\)\)\. Full hyperparameter details are provided in Appendix[A\.1](https://arxiv.org/html/2609.25537#A1.SS1)\.
##### Evaluation Metrics\.
We evaluate CMC on task performance and inference efficiency\. For task performance, we report Exact Match \(EM\) and F1\. For efficiency, we measure inference time decomposed into compression, prefill, and decode stages, energy consumption in kWh via CodeCarboncourty2024codecarbon, MACs, and peak allocated and reserved GPU memory\. Efficiency metrics are evaluated on SQuAD using GPT2\-Large, Mistral, and Llama atr=4r=4, across generation token budgetsT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}\.
### 3\.3Baselines
We evaluate CMC against two groups of baselines\. Theprimary baselinepasses the local context window directly to the frozen DecoderLLM without compression, using the gold answer offset to center the window\. We run this baseline across all three DecoderLLMs using identical training data, evaluation splits, and inference settings as CMC, ensuring a controlled comparison\. The second baseline type is thepublished baselines\.For full\-data experiments on SQuAD, HotpotQA, and AdversarialQA, we compare against one hard\-prompt method – LLMLingua\-2[Pan et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib22)– and four soft\-compression methods – AutoCompressor[Chevalier et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib12), xRAG[Cheng et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib34), ICAE[Ge et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib11), and PCC\-lite and PCC\-large[Dai et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib16)\.
### 3\.4Results
#### 3\.4\.1Task Performance
Table[1](https://arxiv.org/html/2609.25537#S3.T1)reports EM and F1 under both data settings atr=4r=4; full results across all nine encoder\-decoder combinations andr∈\{2,4,8\}r\\in\\\{2,4,8\\\}are in Appendix[A\.2](https://arxiv.org/html/2609.25537#A1.SS2)\.Controlled\-data setting\.CMC outperforms the baselines on the majority of datasets\. With Llama, CMC achieves the largest gains, improving EM by 4\.6, 2\.5, 5\.5, and 7\.0 points on SQuAD, AdversarialQA, HotpotQA, and CovidQA, respectively \(average\+\+4\.9 EM\)\. For Mistral, CMC improves on three of four datasets for a net average gain of 1\.4 EM\. For Gemma, CMC improves consistently across all four datasets, with the largest F1 gains on AdversarialQA \(\+\+3\.1\) and HotpotQA \(\+\+3\.6\)\.Full\-data setting\.CMC continues to benefit Llama and Gemma, with Llama gaining 7\.28 EM on SQuAD \(48\.28→\\rightarrow55\.56\) and Gemma gaining 3\.22 EM \(40\.91→\\rightarrow44\.13\)\. Mistral shows a mixed pattern\. CMC improves over the Mistral baseline on AdversarialQA \(26\.37→\\rightarrow29\.77, \+3\.40 EM\) but underperforms the Mistral baseline by 14\.7 EM on SQuAD \(45\.4→\\rightarrow30\.7\) and regresses on HotpotQA \(53\.4→\\rightarrow48\.3\)\. All three decoders share an identical Phase\-2 configuration, including a fixed KD step budget that does not scale with training set size, while Phase\-1 steps scale directly with it \(∼\\sim8\.8×\\timesmore steps on full\-data vs\. controlled\-data SQuAD\); we attribute Mistral’s SQuAD/HotpotQA regressions to this training imbalance, to which Llama and Gemma appear less sensitive under the identical settings\.
MethodSQuADHotpotQAAdv\.QAEMF1EMF1EMF1AutoCompressor0\.3521\.460\.2916\.292\.0014\.09xRAG3\.4618\.1916\.2927\.513\.4713\.75ICAE21\.6345\.6926\.6835\.1611\.7027\.98LLMLingua\-232\.1851\.2044\.1855\.7224\.8035\.41PCC \(Lite, 4×\\times\)57\.4475\.8342\.2050\.3737\.8350\.36PCC \(Large, 4×\\times\)60\.0477\.7639\.9748\.1939\.3752\.56CMC55\.5678\.2947\.7767\.8335\.4353\.42Table 2:Comparison of CMC with published baselines under the full\-data setting on SQuAD, HotpotQA, and AdversarialQA\. The best EM and F1 scores for each dataset are highlighted in bold\.Table[2](https://arxiv.org/html/2609.25537#S3.T2)compares CMC against published baselines on the full\-data setting\. Although PCC\-lite shares the same GPT2\-Large encoder as CMC, all published baselines use different decoder architectures[Dai et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib16)\. CMC achieves the best F1 on all three datasets, surpassing PCC\-Large by 0\.53 F1 on SQuAD, 19\.64 on HotpotQA, and 0\.86 on AdversarialQA\. On HotpotQA, CMC also leads on EM \(47\.77 vs 39\.97 for PCC\-Large\), a 7\.8\-point gain particularly notable given the multi\-hop reasoning requirement\. CMC underperforms PCC\-Large on EM for SQuAD \(55\.56 vs 60\.04\) and AdversarialQA \(35\.43 vs 39\.37\), likely reflecting PCC’s use of a larger encoder/decoder architecture\. AutoCompressor and xRAG achieve low EM on SQuAD, HotpotQA, and AdversarialQA, while ICAE and LLMLingua\-2 perform substantially below CMC across all three datasets\. A qualitative case study illustrating CMC behavior across four outcome categories is provided in Appendix[A\.9](https://arxiv.org/html/2609.25537#A1.SS9)\.
r=2r=4r=86464656566666767EM\(%\)OPT\-2\.7BOPT\-1\.3BGPT2\-Large\(a\)CMC \(Llama\): EMr=2r=4r=88080818182828383F1 \(%\)\(b\)CMC \(Llama\): F1r=2r=4r=856565757585859596060EM\(%\)\(c\)CMC \(Mistral\): EMr=2r=4r=87272737374747575F1 \(%\)OPT\-2\.7BOPT\-1\.3BGPT2\-Large\(d\)CMC \(Mistral\): F1r=2r=4r=860606161EM\(%\)\(e\)CMC \(Gemma\): EMr=2r=4r=880808181F1 \(%\)\(f\)CMC \(Gemma\): F1
Figure 2:EM and F1 scores on SQuAD \(controlled\-data setting\) across compression ratesr∈\{2,4,8\}r\\in\\\{2,4,8\\\}for three DecoderLLMs paired with three ContextEncoders\.
#### 3\.4\.2Generation and Use of CMEs
The number of CMEs generated per context chunk is controlled by the compression raterr, where a smallerrrproduces more CMEs per chunk and a largerrrproduces fewer\. Figure[2](https://arxiv.org/html/2609.25537#S3.F2)shows the EM and F1 sensitivity tor∈\{2,4,8\}r\\in\\\{2,4,8\\\}across all nine encoder\-decoder combinations on SQuAD\.r=2r=2consistently yields the lowest performance\. The larger CME count introduces noise into the Tier\-1 prefix, diluting answer\-relevant signals and making it harder for the frozen DecoderLLMs to locate the correct span\. In contrast,r=8r=8reduces the CME count to the point where important contextual information is lost, particularly for longer passages\. The intermediate settingr=4r=4achieves the best or tied\-best EM in 6 of 9 combinations\. However, the optimal rate is mildly decoder\-dependent: for example, Llama strongly favorsr=4r=4across all encoders \(EM gains of 1\.1–1\.7 points overr=8r=8\), while Mistral and Gemma show near\-flat performance betweenr=4r=4andr=8r=8\(differences≤0\.1\\leq 0\.1EM\), suggesting decoder\-specific sensitivity to compression granularity\.
#### 3\.4\.3Computational Efficiency
Figure[3](https://arxiv.org/html/2609.25537#S3.F3)shows inference time and energy cost forT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}tokens on SQuAD for three evaluation sample sizes\. CMC consistently reduces both metrics relative to the baseline across all token budgets and sample sizes\. At 1,000 samples, the total time reduction ranges from 4\.6% atT=512T=512\(6,043 s→\\rightarrow5,763 s\) to 20\.0% atT=3,000T=3\{,\}000\(41,963 s→\\rightarrow33,562 s\), with corresponding energy reductions of 4\.6% and 20\.3%, respectively\. This scaling effect arises because CMC’s compression overhead is approximately constant \(≈\\approx25 s regardless ofTT\), while decode savings grow with generation length due to the bounded two\-tier KV cache\. The same trend holds across all sample sizes, confirming the gains reflect the two\-tier KV cache architecture rather than dataset\-specific effects\. Stage\-wise inference time and energy breakdowns for Llama and Mistral are provided in Appendix[A\.3](https://arxiv.org/html/2609.25537#A1.SS3)\.
5121024204830000011223344⋅104\\cdot 10^\{4\}Token GenerationTotal Time \(s\)Base\(1k\)CMC\(1k\)Base\(512\)CMC\(512\)Base\(256\)CMC\(256\)\(a\)Inference Time5121024204830000011223344556677Token GenerationTotal Energy \(kWh\)Base\(1k\)CMC\(1k\)Base\(512\)CMC\(512\)Base\(256\)CMC\(256\)\(b\)Energy Cost
Figure 3:Inference time and energy consumption of CMC vs\. the baseline on SQuAD acrossT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}tokens\. Parenthetical labels denote sample size \(256, 512, 1,000\)\.Figure[4](https://arxiv.org/html/2609.25537#S3.F4)shows peak allocated and peak reserved GPU memory for Llama acrossT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}on SQuAD \(1,000 samples, GPT2\-Large,r=4r=4\)\. Peak allocated memory remains nearly constant for CMC \(≈\\approx8\.85 GB\) while the baseline increases from 10\.0 GB atT=512T=512to 10\.2 GB atT=3,000T=3\{,\}000, yielding a reduction of≈\\approx1\.2 GB\. Peak reserved memory reveals a more substantial difference: the baseline spikes to 19\.7 GB atT=3,000T=3\{,\}000while CMC maintains a near\-constant 9\.8 GB, corresponding to a 50% reduction\. Mistral results in the appendix show an even larger effect \(24\.3 GB→\\rightarrow9\.1 GB, a 62\.5% reduction\)\. These reductions arise because baseline inference accumulates KV cache entries proportionally toTT, whereas CMC’s two\-tier policy caps the cache at a fixed budget\. Detailed memory results for both decoders are provided in Appendix[A\.4](https://arxiv.org/html/2609.25537#A1.SS4)\. Analytical attention MAC analysis in Appendix[A\.5](https://arxiv.org/html/2609.25537#A1.SS5)confirms that these savings are architectural rather than hardware\-specific\.
512102420483000224466881010Token GenerationPeak allocated \(GB\)BaselineCMC\(a\)Peak allocated memory5121024204830002266101014141818Token GenerationPeak reserved \(GB\)BaselineCMC\(b\)Peak reserved memory
Figure 4:Peak allocated and peak reserved GPU memory \(GB\) for CMC vs\. the baseline on 1,000 SQuAD samples acrossT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}tokens\.
## 4Ablation Study
Table[3](https://arxiv.org/html/2609.25537#S4.T3)presents the ablation results on SQuAD\.Architectural ablations \(A1–A3\):Removing graph\-based denoising \(A1\) reduces EM by 26\.6 points relative to CMC, 22\.0 points below the baseline, confirming that denoising is essential for CME quality rather than merely a computational optimization\. Random CME selection \(A2\) reduces EM by 8\.1 points: the same number of CME slots reach the decoder but are decorrelated from the question, reducing Tier\-1 prefix informativeness\. Removing Tier\-2 \(A3\) produces the largest drop \(−\-50\.0 EM\)\. The near\-identical AE/AR losses between A3 and CMC \(1\.40/0\.73 vs 1\.38/0\.70\) confirm the deficit arises at inference rather than from degraded CME quality, establishing Tier\-2 as structurally indispensable\.Phase\-2 training objective ablations \(A4–A6\)\.All converge to EM==0\.589 – identical to Phase\-1 only \(A6\)\. Removing eitherℒKL\\mathcal\{L\}\_\{\\text\{KL\}\}orℒCL\\mathcal\{L\}\_\{\\text\{CL\}\}individually is as damaging as removing Phase\-2 entirely, revealing that all three objectives must act jointly:ℒCE\\mathcal\{L\}\_\{\\text\{CE\}\}provides token\-level supervision,ℒKL\\mathcal\{L\}\_\{\\text\{KL\}\}aligns the learner with the educator’s distribution, andℒCL\\mathcal\{L\}\_\{\\text\{CL\}\}grounds the CME prefix in the answer embedding space\.
ModelEMF1𝚫\\boldsymbol\{\\Delta\}EM𝚫\\boldsymbol\{\\Delta\}F1Llama \(baseline\)0\.6240\.810−\-0\.046−\-0\.019CMC0\.6700\.829——Architecture ablationsw/o graph denoising \(A1\)0\.4040\.564−\-0\.266−\-0\.265w/o query\-guided top\-KK\(A2\)0\.5890\.757−\-0\.081−\-0\.072w/o Tier\-2 local window \(A3\)0\.1700\.247−\-0\.500−\-0\.582Training objective ablationsw/oℒCL\\mathcal\{L\}\_\{\\text\{CL\}\}\(A4\)0\.5890\.757−\-0\.081−\-0\.072w/oℒKL\\mathcal\{L\}\_\{\\text\{KL\}\}\(A5\)0\.5890\.757−\-0\.081−\-0\.072Phase\-1 only, no Phase\-2 \(A6\)0\.5890\.757−\-0\.081−\-0\.072Table 3:Ablation study on SQuAD \(controlled\-data setting, GPT2\-Large,r=4r=4, Llama\)\.Δ\\DeltaEM andΔ\\DeltaF1 are relative to CMC\.Sensitivity analysesof local window size and denoising threshold are provided in Appendices[A\.6](https://arxiv.org/html/2609.25537#A1.SS6)and[A\.7](https://arxiv.org/html/2609.25537#A1.SS7), respectively\. Adeployment analysisof Tier\-2 window localization without gold annotations is provided in Appendix[A\.8](https://arxiv.org/html/2609.25537#A1.SS8)\.
## 5Related Work
Architectural methods\.Sparse and local attention mechanisms[Beltagy et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib7);[Zaheer et al\. \(2020\)](https://arxiv.org/html/2609.25537#bib.bib8)reduce quadratic complexity by restricting each token’s receptive field, while positional interpolation[Chen et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib9), YaRN[Peng et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib30), and LongLoRA[Chen et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib29)extend the context window of pretrained LLMs through rotary embedding modifications\. FlashAttention[Dao et al\. \(2022\)](https://arxiv.org/html/2609.25537#bib.bib28)reduces memory bandwidth costs without approximation, and KV cache management strategies such as H2O[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib17), and StreamingLLM[Xiao et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib18)evict low\-importance entries during decoding\. These methods require architectural modification, continued pretraining, or operate only on already\-computed caches without reducing prefill cost, and none produce a compact input representation consumable by an unmodified frozen LLM\.
Hard\-prompt compression\.Selective Context[Li et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib10)prunes tokens by self\-information, LLMLingua[Jiang et al\. \(2023b\)](https://arxiv.org/html/2609.25537#bib.bib31)and LongLLMLingua[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib32)apply perplexity\-based budget control, LLMLingua\-2[Pan et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib22)distils compression decisions into a lightweight classifier, and QGC[Cao et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib33)weights tokens by query relevance\. These methods are model\-agnostic but incur irreversible information loss, as discarded tokens cannot be recovered by the decoder, making them susceptible to performance degradation on tasks requiring cross\-passage evidence synthesis such as multi\-hop reasoning\.
Soft\-compression\.Soft\-compression methods encode the full context into dense embeddings consumed by a frozen decoder, avoiding hard token removal\. RMT[Bulatov et al\. \(2022\)](https://arxiv.org/html/2609.25537#bib.bib36)propagates memory tokens recurrently across segments\. AutoCompressor[Chevalier et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib12)and ICAE[Ge et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib11)compress long documents into soft summary vectors via autoencoding objectives\. Gisting[Mu et al\. \(2023\)](https://arxiv.org/html/2609.25537#bib.bib13)compresses instructions into gist tokens, and CCM[Kim et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib14)targets online interaction scenarios by compressing accumulating KV pairs via a conditional LoRA adapter\. 500xCompressor[Li et al\. \(2025b\)](https://arxiv.org/html/2609.25537#bib.bib2)compresses the entire context into a small set of learned tokens, and then uses the KV representation of these tokens for the downstream QA task\. In RAG settings, xRAG[Cheng et al\. \(2024\)](https://arxiv.org/html/2609.25537#bib.bib34)and COCOM[Rau et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib35)compress retrieved passages into dense representations for a single specific target decoder\.[Dai et al\. \(2025\)](https://arxiv.org/html/2609.25537#bib.bib16)introduce a decoupled compressor\-LLM framework with cross\-architecture projection but without query\-guided memory selection or answer\-targeted training supervision\. ATACompressor[Li et al\. \(2025a\)](https://arxiv.org/html/2609.25537#bib.bib15)introduces query\-guided selective encoding but does not address cross\-architecture deployment or answer\-targeted memory alignment\. No existing method simultaneously addresses all three gaps, which CMC is designed to fill\.
## 6Conclusion
We presented CMC, a soft\-compression framework that compresses long contexts into compact CMEs for computation\- and energy\- efficient inference with frozen LLMs\. Across nine encoder\-decoder combinations and four extractive QA benchmarks, CMC consistently outperforms the baseline in the majority of configurations under the controlled\-data setting, and achieves the best F1 against published soft\-compression baselines under full\-data conditions\. AtT=3,000T=3\{,\}000tokens, CMC reduces inference time and energy by 20% and peak reserved GPU memory by up to 62\.5%, enabling deployment at generation lengths that would otherwise exhaust GPU memory\. Ablation studies confirm that graph\-based denoising, the Tier\-2 local window, and all three Phase\-2 objectives are each necessary; removing any single component causes substantial performance degradation\.
## Limitations
While CMC demonstrates consistent improvements in EM, F1, and inference efficiency, we acknowledge the following limitations\. First, all experiments are conducted on extractive QA benchmarks; the effectiveness of CMC on abstractive QA, summarisation, and RAG has not been evaluated and may require adjustments to the answer\-targeted distillation objective\. Moreover, CMC is trained and evaluated within each benchmark independently; cross\-dataset generalization \(e\.g\., training on SQuAD and evaluating on HotpotQA\) has not been evaluated\. Second, CMC uses a fixed compression rater∈\{2,4,8\}r\\in\\\{2,4,8\\\}, which may over\-compress short passages or under\-compress long ones; adaptive rate scheduling conditioned on context length is left for future work\. Lastly, CMC is evaluated on English\-language QA benchmarks only; its effectiveness on non\-English or multilingual contexts, where tokenisation and semantic similarity properties differ substantially, has not been investigated\.
## Ethics Statement
No ethical approval was needed for this study\.
## Availability Statement
## References
- Bartoloet al\.\(2020\)M\. Bartolo, A\. Roberts, J\. Welbl, S\. Riedel, and P\. StenetorpBeat the AI: investigating adversarial human annotation for reading comprehension\.Transactions of the Association for Computational Linguistics8,pp\. 662–678\.External Links:[Link](https://aclanthology.org/2020.tacl-1.43/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00338)Cited by:[§3\.1](https://arxiv.org/html/2609.25537#S3.SS1.p1.1)\.
- Beltagyet al\.\(2020\)I\. Beltagy, M\. E\. Peters, and A\. CohanLongformer: the long\-document transformer\.arXiv preprint arXiv:2004\.05150\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2004.05150)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Bulatovet al\.\(2022\)A\. Bulatov, Y\. Kuratov, and M\. BurtsevRecurrent memory transformer\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 11079–11091\.External Links:[Document](https://dx.doi.org/10.52202/068431-0805),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/47e288629a6996a17ce50b90a056a0e1-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Caoet al\.\(2024\)Z\. Cao, Q\. Cao, Y\. Lu, N\. Peng, L\. Huang, S\. Cheng, and J\. SuRetaining key information under high compression ratios: query\-guided compressor for LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12685–12695\.External Links:[Link](https://aclanthology.org/2024.acl-long.685/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.685)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p2.1)\.
- Chenet al\.\(2023\)S\. Chen, S\. Wong, L\. Chen, and Y\. TianExtending context window of large language models via positional interpolation\.arXiv preprint arXiv:2306\.15595\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2306.15595)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Chenet al\.\(2024\)Y\. Chen, S\. Qian, H\. Tang, X\. Lai, Z\. Liu, S\. Han, and J\. JiaLongLoRA: efficient fine\-tuning of long\-context large language models\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 8220–8238\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/211ab571cc9f3802afa6ffff52ae3e5b-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Chenget al\.\(2024\)X\. Cheng, X\. Wang, X\. Zhang, T\. Ge, S\. Chen, F\. Wei, H\. Zhang, and D\. ZhaoxRAG: extreme context compression for retrieval\-augmented generation with one token\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 109487–109516\.External Links:[Document](https://dx.doi.org/10.52202/079017-3476),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/c5cf13bfd3762821ef7607e63ee90075-Paper-Conference.pdf)Cited by:[§3\.3](https://arxiv.org/html/2609.25537#S3.SS3.p1.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Chevalieret al\.\(2023\)A\. Chevalier, A\. Wettig, A\. Ajith, and D\. ChenAdapting language models to compress contexts\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 3829–3846\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.232/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.232)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§3\.3](https://arxiv.org/html/2609.25537#S3.SS3.p1.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Daiet al\.\(2025\)Y\. Dai, J\. Lian, Y\. Huang, W\. Zhang, M\. Zhou, M\. Wu, X\. Xie, and H\. LiaoPretraining context compressor for large language models with embedding\-based memory\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 28715–28732\.External Links:[Link](https://aclanthology.org/2025.acl-long.1394/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1394),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§3\.3](https://arxiv.org/html/2609.25537#S3.SS3.p1.1),[§3\.4\.1](https://arxiv.org/html/2609.25537#S3.SS4.SSS1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Daoet al\.\(2022\)T\. Dao, D\. Fu, S\. Ermon, A\. Rudra, and C\. RéFlashAttention: fast and memory\-efficient exact attention with io\-awareness\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 16344–16359\.External Links:[Document](https://dx.doi.org/10.52202/068431-1189),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/67d57c32e20fd0a7a302cb81d36e40d5-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Donget al\.\(2024\)Z\. Dong, J\. Li, X\. Men, W\. X\. Zhao, B\. Wang, Z\. Tian, W\. Chen, and J\. WenExploring context window of large language models via decomposed positional vectors\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 10320–10347\.External Links:[Document](https://dx.doi.org/10.52202/079017-0330),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/1403ab1a427050538ec59c7f570aec8b-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1)\.
- Geet al\.\(2024\)T\. Ge, H\. Jing, L\. Wang, X\. Wang, S\. Chen, and F\. WeiIn\-context autoencoder for context compression in a large language model\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 2591–2607\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/0b276510ec2d3f6613a8b60c41ff0438-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§3\.3](https://arxiv.org/html/2609.25537#S3.SS3.p1.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Gemma Teamet al\.\(2024\)Gemma Team, T\. Mesnard, C\. Hardin, R\. Dadashi, S\. Bhupatiraju, S\. Pathak, L\. Sifre, M\. Rivière, M\. S\. Kale, J\. Love,et al\.Gemma: open models based on Gemini research and technology\.arXiv preprint arXiv:2403\.08295\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2403.08295)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Jianget al\.\(2023a\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Jianget al\.\(2023b\)H\. Jiang, Q\. Wu, C\. Lin, Y\. Yang, and L\. QiuLLMLingua: compressing prompts for accelerated inference of large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 13358–13376\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.825/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.825)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p2.1)\.
- Jianget al\.\(2024\)H\. Jiang, Q\. Wu, X\. Luo, D\. Li, C\. Lin, Y\. Yang, and L\. QiuLongLLMLingua: accelerating and enhancing LLMs in long context scenarios via prompt compression\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 1658–1677\.External Links:[Link](https://aclanthology.org/2024.acl-long.91/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.91)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p2.1)\.
- Kimet al\.\(2024\)J\. Kim, J\. Yeom, S\. Yun, and H\. O\. SongCompressed context memory for online language model interaction\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 35316–35334\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/97dc07f1253ab33ee514f395a82fa7cc-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Liet al\.\(2025a\)X\. Li, H\. Li, Y\. Zhou, Q\. Ai, and Y\. LiuATACompressor: adaptive task\-aware compression for efficient long\-context processing in llms\.InProceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region,SIGIR\-AP 2025,pp\. 343–352\.External Links:ISBN 9798400722189,[Link](https://doi.org/10.1145/3767695.3769499),[Document](https://dx.doi.org/10.1145/3767695.3769499)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Liet al\.\(2023\)Y\. Li, B\. Dong, F\. Guerin, and C\. LinCompressing context to enhance inference efficiency of large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6342–6353\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.391/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.391)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p2.1)\.
- Liet al\.\(2025b\)Z\. Li, Y\. Su, and N\. Collier500xCompressor: generalized prompt compression for large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 25081–25091\.External Links:[Link](https://aclanthology.org/2025.acl-long.1219/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1219),ISBN 979\-8\-89176\-251\-0Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Mölleret al\.\(2020\)T\. Möller, A\. Reina, R\. Jayakumar, and M\. PietschCOVID\-QA: a question answering dataset for COVID\-19\.InProceedings of the 1st Workshop on NLP for COVID\-19 at ACL 2020,K\. Verspoor, K\. B\. Cohen, M\. Dredze, E\. Ferrara, J\. May, R\. Munro, C\. Paris, and B\. Wallace \(Eds\.\),External Links:[Link](https://aclanthology.org/2020.nlpcovid19-acl.18/)Cited by:[§3\.1](https://arxiv.org/html/2609.25537#S3.SS1.p1.1)\.
- Muet al\.\(2023\)J\. Mu, X\. Li, and N\. GoodmanLearning to compress prompts with gist tokens\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 19327–19352\.External Links:[Document](https://dx.doi.org/10.52202/075280-0848),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/3d77c6dcc7f143aa2154e7f4d5e22d68-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Panet al\.\(2024\)Z\. Pan, Q\. Wu, H\. Jiang, M\. Xia, X\. Luo, J\. Zhang, Q\. Lin, V\. Rühle, Y\. Yang, C\. Lin, H\. V\. Zhao, L\. Qiu, and D\. ZhangLLMLingua\-2: data distillation for efficient and faithful task\-agnostic prompt compression\.InFindings of the Association for Computational Linguistics \(ACL\) 2024,Bangkok, Thailand,pp\. 963–981\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.57)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§3\.3](https://arxiv.org/html/2609.25537#S3.SS3.p1.1),[§5](https://arxiv.org/html/2609.25537#S5.p2.1)\.
- Penget al\.\(2024\)B\. Peng, J\. Quesnelle, H\. Fan, and E\. ShippoleYaRN: efficient context window extension of large language models\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 31932–31951\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/874a4d89f2d04b4bcf9a2c19545cf040-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Popeet al\.\(2023\)R\. Pope, S\. Douglas, A\. Chowdhery, J\. Devlin, J\. Bradbury, J\. Heek, K\. Xiao, S\. Agrawal, and J\. DeanEfficiently scaling transformer inference\.InProceedings of Machine Learning and Systems,D\. Song, M\. Carbin, and T\. Chen \(Eds\.\),Vol\.5,pp\. 606–624\.External Links:[Link](https://proceedings.mlsys.org/paper_files/paper/2023/file/c4be71ab8d24cdfb45e3d06dbfca2780-Paper-mlsys2023.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Rajpurkaret al\.\(2016\)P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. LiangSQuAD: 100,000\+ questions for machine comprehension of text\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,pp\. 2383–2392\.External Links:[Link](https://aclanthology.org/D16-1264/),[Document](https://dx.doi.org/10.18653/v1/D16-1264)Cited by:[§3\.1](https://arxiv.org/html/2609.25537#S3.SS1.p1.1)\.
- Rauet al\.\(2025\)D\. Rau, S\. Wang, H\. Déjean, S\. Clinchant, and J\. KampsContext embeddings for efficient answer generation in retrieval\-augmented generation\.InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining,WSDM ’25,pp\. 493–502\.External Links:ISBN 9798400713293,[Link](https://doi.org/10.1145/3701551.3703527),[Document](https://dx.doi.org/10.1145/3701551.3703527)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p3.1)\.
- Robertson and Zaragoza \(2009\)S\. Robertson and H\. ZaragozaThe probabilistic relevance framework: bm25 and beyond\.Found\. Trends Inf\. Retr\.3\(4\),pp\. 333–389\.External Links:ISSN 1554\-0669,[Link](https://doi.org/10.1561/1500000019),[Document](https://dx.doi.org/10.1561/1500000019)Cited by:[§A\.8](https://arxiv.org/html/2609.25537#A1.SS8.p1.1)\.
- Shenget al\.\(2023\)Y\. Sheng, L\. Zheng, B\. Yuan, Z\. Li, M\. Ryabinin, B\. Chen, P\. Liang, C\. Re, I\. Stoica, and C\. ZhangFlexGen: high\-throughput generative inference of large language models with a single GPU\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 31094–31116\.External Links:[Link](https://proceedings.mlr.press/v202/sheng23a.html)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Touvronet al\.\(2023\)H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.2307.09288)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- van den Oordet al\.\(2018\)A\. van den Oord, Y\. Li, and O\. VinyalsRepresentation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.48550/arXiv.1807.03748)Cited by:[§2\.4](https://arxiv.org/html/2609.25537#S2.SS4.SSS0.Px2.p2.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p1.1)\.
- Xiaoet al\.\(2024\)G\. Xiao, Y\. Tian, B\. Chen, S\. Han, and M\. LewisEfficient streaming language models with attention sinks\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 21875–21895\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/5e5fd18f863cbe6d8ae392a93fd271c9-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 2369–2380\.External Links:[Link](https://aclanthology.org/D18-1259/),[Document](https://dx.doi.org/10.18653/v1/D18-1259)Cited by:[§3\.1](https://arxiv.org/html/2609.25537#S3.SS1.p1.1)\.
- Zaheeret al\.\(2020\)M\. Zaheer, G\. Guruganesh, K\. A\. Dubey, J\. Ainslie, C\. Alberti, S\. Ontanon, P\. Pham, A\. Ravula, Q\. Wang, L\. Yang, and A\. AhmedBig bird: transformers for longer sequences\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 17283–17297\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/c8512d142a2d849725f31a9a7a361ab9-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.25537#S1.p2.1),[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
- Zhanget al\.\(2023\)Z\. Zhang, Y\. Sheng, T\. Zhou, T\. Chen, L\. Zheng, R\. Cai, Z\. Song, Y\. Tian, C\. Ré, C\. Barrett, Z\. "\. Wang, and B\. ChenH2O: heavy\-hitter oracle for efficient generative inference of large language models\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 34661–34710\.External Links:[Document](https://dx.doi.org/10.52202/075280-1506),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/6ceefa7b15572587b78ecfcebb2827f8-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.25537#S5.p1.1)\.
## Appendix AAppendix
### A\.1Hyperparameters
Table[4](https://arxiv.org/html/2609.25537#A1.T4)reports the full set of hyperparameters used across all experiments\. The ContextEncoder and MemoryBridge are optimised jointly; the frozen DecoderLLM receives no gradient updates at any stage\. The two\-tier KV cache hyperparametersKKandWWare tuned per dataset and encoder\-decoder configuration within the ranges reported below\.
HyperparameterValueHardwareGPUNVIDIA A100 80 GBDecoderLLMQuantisation4\-bit NF4, double quantisationCompute dtypebfloat16Decoder trainingFrozen \(no gradient updates\)Optimiser \(ContextEncoder \+ MemoryBridge\)OptimiserAdamWLearning rate1×10−51\\times 10^\{\-5\}Weight decay0\.010\.01β\\beta\(0\.9,0\.98\)\(0\.9,\\ 0\.98\)Gradient clip norm0\.60\.6Training epochs11Batch size\{4,8\}\\\{4,8\\\}\(tuned per config\)ContextEncoderArchitecturesGPT2\-Large, OPT\-1\.3B, OPT\-2\.7BMax input length3,072 tokensMemoryBridgeArchitectureTwo\-layer MLPNormalisationLayerNorm \(ϵ=10−5\\epsilon=10^\{\-5\}\)ActivationGELUOutput activationNorm\-calibrated tanhGraph\-based token denoisingCosine similarity thresholdτ\\tau0\.050\.05Neighbourhood windowkk33Minimum keep intervalEvery 16 tokensIG warmup steps200200Chunking and compressionChunk sizett256256tokensCompression ratesrr\{2,4,8\}\\\{2,4,8\\\}Memory token cap per chunk128128Min memory tokens per chunk44Warmup memory tokens88Two\-tier KV cacheTier\-1 top\-KKCMEsK∈\{30,40\}K\\in\\\{30,40\\\}Tier\-2 local windowWWW∈\{130,256\}W\\in\\\{130,256\\\}tokensPhase\-1 AE/ARAR loss weightλAR\\lambda\_\{\\text\{AR\}\}1\.0Phase\-2 QA distillationKD steps per epoch800800KD batch size22CE loss weightλCE\\lambda\_\{\\text\{CE\}\}1\.01\.0KL loss weightλKL\\lambda\_\{\\text\{KL\}\}0\.50\.5Contrastive loss weightλCL\\lambda\_\{\\text\{CL\}\}0\.10\.1Contrastive temperatureδ\\delta0\.070\.07Local window drop probability0\.50\.5Local window chars \(educator\)1,200−2,0001\{,\}200\-2,000Local window tokens \(student\)130130InferenceMax new tokens3232Decoder max length1,5361\{,\}536tokensTable 4:Hyperparameters used in all CMC experiments\.##### LayerNorm details\.
All LayerNorm operations in MemoryBridge use PyTorch’s defaultnn\.LayerNorm:LayerNorm\(𝐱\)=𝜸⊙\(𝐱−μ\)/σ2\+ϵ\+𝜷\\text\{LayerNorm\}\(\\mathbf\{x\}\)=\\boldsymbol\{\\gamma\}\\odot\(\\mathbf\{x\}\-\\mu\)/\\sqrt\{\\sigma^\{2\}\+\\epsilon\}\+\\boldsymbol\{\\beta\}, whereμ\\muandσ2\\sigma^\{2\}are the mean and variance of𝐱\\mathbf\{x\}over the feature dimension,ϵ=10−5\\epsilon=10^\{\-5\}\(Table[4](https://arxiv.org/html/2609.25537#A1.T4)\), and𝜸,𝜷\\boldsymbol\{\\gamma\},\\boldsymbol\{\\beta\}are learnable per\-dimension scale and shift parameters trained jointly with the rest of MemoryBridge\.
DecoderLLMModelContextEncoderrrEMF1Llama\-38B\-InstructLlama \(baseline\)——0\.62400\.8097OPT\-2\.7B2/4/80\.6570/0\.6680/0\.66700\.8113/0\.8283/0\.8274CMCOPT\-1\.3B2/4/80\.6570/0\.6680/0\.66700\.8113/0\.8283/0\.8274GPT2\-Large2/4/80\.6530/0\.6700/0\.66700\.8107/0\.8286/0\.8274Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)——0\.61000\.7586OPT\-2\.7B2/4/80\.5740/0\.5990/0\.60000\.7391/0\.7533/0\.7532CMCOPT\-1\.3B2/4/80\.5740/0\.5990/0\.60000\.7391/0\.7533/0\.7532GPT2\-Large2/4/80\.5740/0\.6020/0\.60000\.7390/0\.7551/0\.7532Gemma\-29B\-ITGemma \(baseline\)——0\.55100\.7805OPT\-2\.7B2/4/80\.6090/0\.6100/0\.61100\.8102/0\.8115/0\.8118CMCOPT\-1\.3B2/4/80\.6090/0\.6100/0\.61000\.8102/0\.8115/0\.8115GPT2\-Large2/4/80\.6090/0\.6100/0\.61000\.8102/0\.8115/0\.8115Table 5:EM and F1 results onSQuAD\(controlled\-data setting\)\.DecoderLLMModelContextEncoderrrEMF1Llama\-38B\-InstructLlama \(baseline\)——0\.34300\.5329OPT\-2\.7B2/4/80\.3740/0\.3660/0\.36700\.5328/0\.5501/0\.5495CMCOPT\-1\.3B2/4/80\.3740/0\.3660/0\.36700\.5328/ 0\.5501/0\.5495GPT2\-Large2/4/80\.3750/0\.3680/0\.36500\.5330/0\.5509/0\.5464Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)——0\.29600\.4786OPT\-2\.7B2/4/80\.3050/0\.3200/0\.32000\.4888/0\.5013/0\.5012CMCOPT\-1\.3B2/4/80\.3050/0\.3200/0\.32000\.4888/0\.5013/0\.5012GPT2\-Large2/4/80\.3060/0\.3220/0\.31800\.4913/0\.5029/0\.5008Gemma\-29B\-ITGemma \(baseline\)——0\.32900\.5641OPT\-2\.7B2/4/80\.3660/0\.3680/0\.36800\.5921/0\.5952/0\.5952CMCOPT\-1\.3B2/4/80\.3660/0\.3680/0\.36800\.5921/0\.5952/0\.5952GPT2\-Large2/4/80\.3650/0\.3680/0\.36800\.5916/0\.5952/0\.5952Table 6:EM and F1 results onAdversarialQA\(controlled\-data setting\)\.DecoderLLMModelContextEncoderrrEMF1Llama\-38B\-InstructLlama \(baseline\)——0\.44100\.6814OPT\-2\.7B2/4/80\.4610/0\.4650/0\.46100\.6554/0\.6634/0\.6619CMCOPT\-1\.3B2/4/80\.4880/0\.4810/0\.47900\.6907/0\.6933/0\.6922GPT2\-Large2/4/80\.4960/0\.4840/0\.46100\.6934/0\.6943/0\.6619Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)——0\.53000\.7106OPT\-2\.7B2/4/80\.4810/0\.4910/0\.49300\.6540/0\.6668/0\.6672CMCOPT\-1\.3B2/4/80\.5260/0\.4910/0\.53800\.7042/0\.6668/0\.7169GPT2\-Large2/4/80\.4810/0\.5360/0\.53800\.6544/0\.7158/0\.7169Gemma\-29B\-ITGemma \(baseline\)——0\.40100\.6700OPT\-2\.7B2/4/80\.4200/0\.4200/0\.41500\.7059/0\.7059/0\.7020CMCOPT\-1\.3B2/4/80\.4200/0\.4200/0\.41500\.7059/0\.7059/0\.7020GPT2\-Large2/4/80\.4200/0\.4200/0\.41500\.7059/0\.6949/0\.6939Table 7:EM and F1 results onHotpotQA\(controlled\-data setting\)\.DecoderLLMModelContextEncoderrrEMF1Llama\-38B\-InstructLlama \(baseline\)——0\.22500\.6383OPT\-2\.7B2/4/80\.2950/0\.2950/0\.29500\.6416/0\.6416/0\.6416CMCOPT\-1\.3B2/4/80\.2900/0\.2900/0\.29000\.6412/0\.6392/0\.6392GPT2\-Large2/4/80\.2900/0\.2900/0\.29000\.6412/0\.6392/0\.6412Mistral\-7BInstruct\-v0\.3Mistral \(baseline\)——0\.19000\.5761OPT\-2\.7B2/4/80\.2150/0\.2200/0\.21500\.6238/0\.6305/0\.6238CMCOPT\-1\.3B2/4/80\.2150/0\.2200/0\.21500\.6238/0\.6305/0\.6238GPT2\-Large2/4/80\.2150/0\.2200/0\.21500\.6238/0\.6305/0\.6238Gemma\-29B\-ITGemma \(baseline\)——0\.10500\.6209OPT\-2\.7B2/4/80\.14/0\.14/0\.140\.6323/0\.6323/0\.6323CMCOPT\-1\.3B2/4/80\.14/0\.14/0\.140\.6323/0\.6323/0\.6323GPT2\-Large2/4/80\.14/0\.14/0\.140\.6323/0\.6323/0\.6323Table 8:EM and F1 results onCovidQA\.r=2r=2r=4r=4r=8r=8DecoderLLMContextEncoderEMF1EMF1EMF1Llama\-38B\-InstructOPT\-2\.7B0\.44680\.66030\.44850\.67090\.44750\.6701OPT\-1\.3B0\.45220\.66900\.45120\.67770\.45070\.6771GPT2\-Large0\.45350\.66960\.45300\.67830\.44570\.6692Mistral\-7BInstruct\-v0\.3OPT\-2\.7B0\.39370\.62640\.40750\.63800\.40700\.6363OPT\-1\.3B0\.40500\.63900\.40750\.63800\.41830\.6488GPT2\-Large0\.39400\.62710\.42000\.65110\.41780\.6487Gemma\-29B\-ITOPT\-2\.7B0\.38370\.68510\.38450\.68620\.38350\.6853OPT\-1\.3B0\.38370\.68510\.38450\.68620\.38320\.6852GPT2\-Large0\.38350\.68500\.38450\.68350\.38320\.6832Overall average0\.41070\.66070\.41570\.66780\.41520\.6671Table 9:Average EM and F1 across four datasets \(SQuAD, AdversarialQA, HotpotQA, CovidQA\) under controlled\-data setting, grouped by DecoderLLM, ContextEncoder, and compression rater∈\{2,4,8\}r\\in\\\{2,4,8\\\}\.Bolddenotes the best average per DecoderLLM–ContextEncoder pair\.
### A\.2Detailed Results per Dataset
Tables[5](https://arxiv.org/html/2609.25537#A1.T5)–[8](https://arxiv.org/html/2609.25537#A1.T8)report EM and F1 for all nine encoder\-decoder combinations acrossr∈\{2,4,8\}r\\in\\\{2,4,8\\\}under the controlled\-data setting\.Bolddenotes the best CMC result per DecoderLLM group\. Three findings across these tables support the design choices made in CMC\. First, OPT\-1\.3B and OPT\-2\.7B produce near\-identical results across all datasets and decoders, while GPT2\-Large — the smallest encoder — matches or outperforms both, confirming that ContextEncoder scale does not drive performance\. Second, Gemma shows near\-zero variation across all encoder and compression rate combinations \(≤\\leq0\.001 EM\), indicating that the two\-tier KV cache policy is effective regardless of decoder sensitivity to compression granularity\. Third, CovidQA shows consistent improvements across all nine configurations despite only 1,500 training samples, demonstrating that the two\-phase training strategy generalises to low\-resource domain\-specific settings\.
Additionally, Table[9](https://arxiv.org/html/2609.25537#A1.T9)reports average EM and F1 across all encoder\-decoder combinations and datasets\. The overall average confirmsr=4r=4as the best compression rate \(EM==0\.4157, F1==0\.6678\), marginally outperformingr=8r=8\(EM==0\.4152\) and more substantially outperformingr=2r=2\(EM==0\.4107\)\.
SamplesModelTTCom\. \(s\)Pre\. \(s\)Dec\. \(s\)Total \(s\)256Baseline512\-7\.611,535\.771,543\.38CMC6\.296\.791,464\.631,477\.71Baseline1024\-7\.713,194\.183,201\.91CMC6\.216\.872,930\.892,943\.98Baseline2048\-7\.405,815\.225,822\.62CMC6\.606\.315,067\.515,080\.43Baseline3000\-7\.6110,717\.0910,724\.70CMC6\.506\.898,576\.128,589\.51512Baseline512\-16\.602,574\.882,591\.49CMC12\.7816\.592,445\.592,474\.97Baseline1024\-17\.515,353\.995,371\.51CMC12\.9612\.614,931\.134,956\.70Baseline2048\-14\.8811,639\.9311,654\.82CMC13\.4912\.7410,496\.7410,522\.99Baseline3000\-17\.6718,233\.1218,250\.79CMC12\.8215\.0714,312\.2914,340\.181000Baseline512\-31\.576,011\.686,043\.28CMC25\.1426\.945,711\.325,763\.45Baseline1024\-31\.4612,498\.0512,529\.52CMC25\.2627\.1111,451\.6211,504\.01Baseline2048\-31\.6226,954\.1626,985\.80CMC25\.5127\.7522,958\.7523,012\.03Baseline3000\-31\.3741,931\.8241,963\.20CMC25\.6927\.1333,508\.7433,561\.58Table 10:Stage\-wise inference time \(seconds\) forLlama\-3\-8B\-Instructon SQuAD across three sample sizes \(GPT2\-Large ContextEncoder,r=4r=4, batch size 8\)\. Com\. = compression; Pre\. = prefill; Dec\. = decode\. Baseline compression time is zero\.
### A\.3Inference Efficiency
Tables[10](https://arxiv.org/html/2609.25537#A1.T10)–[13](https://arxiv.org/html/2609.25537#A1.T13)report stage\-wise inference time and energy for Llama and Mistral on SQuAD across three sample sizes andT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}\. These tables extend the 1,000\-sample Llama results in Section[3\.4\.3](https://arxiv.org/html/2609.25537#S3.SS4.SSS3)to smaller evaluation sets and Mistral\. We present two additional observations\. First, the compression overhead scales linearly with sample size — approximately 6 s at 256 samples, 13 s at 512, and 26 s at 1,000 — and is independent ofTT, confirming that CMC’s compression cost is determined by the number of contexts processed rather than the generation length\. Second, CMC prefill time is marginally lower than the baseline across all conditions, reflecting the shorter local context window relative to the uncompressed prompt\. Moreover, Mistral shows consistent savings matching Llama \(≈\\approx20% atT=3,000T=3\{,\}000, 1,000 samples\), confirming that the efficiency gains are decoder\-independent\.
SamplesModelTTCom\. \(kWh\)Pre\. \(kWh\)Dec\. \(kWh\)Total \(kWh\)256Baseline512\-0\.00120\.23920\.2404CMC0\.00100\.00110\.22780\.2298Baseline1024\-0\.00120\.49660\.4978CMC0\.00100\.00110\.45550\.4576Baseline2048\-0\.00171\.30631\.3079CMC0\.00140\.00131\.06811\.0709Baseline3000\-0\.00121\.71871\.7200CMC0\.00100\.00111\.36961\.3718512Baseline512\-0\.00380\.59000\.5938CMC0\.00290\.00370\.54960\.5562Baseline1024\-0\.00401\.22271\.2267CMC0\.00290\.00281\.10251\.1082Baseline2048\-0\.00332\.60222\.6056CMC0\.00280\.00272\.20012\.2056Baseline3000\-0\.00404\.09964\.1036CMC0\.00280\.00333\.10683\.11281000Baseline512\-0\.00500\.95140\.9564CMC0\.00400\.00430\.90430\.9125Baseline1024\-0\.00501\.97231\.9773CMC0\.00400\.00431\.80091\.8092Baseline2048\-0\.00504\.25754\.2625CMC0\.00400\.00443\.61243\.6208Baseline3000\-0\.00506\.72046\.7254CMC0\.00410\.00435\.35385\.3623Table 11:Stage\-wise energy consumption \(kWh\) forLlama\-3\-8B\-Instructon SQuAD across three sample sizes \(GPT2\-Large ContextEncoder,r=4r=4, batch size 8\)\. Stage\-wise energy is apportioned proportionally to wall\-clock time\.SamplesModelTTCom\. \(s\)Pre\. \(s\)Dec\. \(s\)Total \(s\)256Baseline512\-8\.641,537\.281,545\.92CMC6\.427\.271,461\.701,475\.39Baseline1024\-8\.443,190\.703,199\.14CMC6\.337\.032,919\.232,932\.59Baseline2048\-8\.476,864\.736,873\.19CMC6\.416\.945,846\.335,859\.69Baseline3000\-8\.869,094\.459,103\.32CMC6\.426\.217,429\.817,442\.45512Baseline512\-16\.733,061\.613,078\.34CMC12\.0814\.142,911\.432,937\.66Baseline1024\-16\.755,415\.735,432\.49CMC12\.9412\.425,179\.115,204\.47Baseline2048\-15\.9811,465\.2611,481\.24CMC13\.0312\.3410,250\.0610,275\.45Baseline3000\-15\.5418,180\.7418,196\.29CMC12\.9113\.2514,883\.9714,910\.141000Baseline512\-33\.085,179\.305,212\.39CMC26\.0028\.635,058\.985,113\.63Baseline1024\-31\.3410,657\.5210,688\.87CMC27\.1724\.4510,295\.9210,347\.56Baseline2048\-34\.2426,854\.9426,889\.19CMC24\.8728\.0522,832\.3522,885\.29Baseline3000\-34\.2941,807\.0441,841\.34CMC26\.0927\.7933,354\.4233,408\.33Table 12:Stage\-wise inference time \(seconds\) forMistral\-7B\-Instruct\-v0\.3on SQuAD across three sample sizes \(GPT2\-Large ContextEncoder,r=4r=4, batch size 8\)\. Com\. = compression; Pre\. = prefill; Dec\. = decode\.SamplesModelTTCom\. \(kWh\)Pre\. \(kWh\)Dec\. \(kWh\)Total \(kWh\)256Baseline512\-0\.00140\.24640\.2478CMC0\.00100\.00120\.23350\.2357Baseline1024\-0\.00140\.51990\.5213CMC0\.00100\.00110\.47270\.4749Baseline2048\-0\.00141\.10761\.1089CMC0\.00100\.00110\.93510\.9373Baseline3000\-0\.00202\.08742\.0895CMC0\.00130\.00131\.60391\.6066512Baseline512\-0\.00270\.49660\.4993CMC0\.00200\.00230\.47130\.4756Baseline1024\-0\.00371\.19191\.1956CMC0\.00280\.00271\.13561\.1411Baseline2048\-0\.00372\.61722\.6209CMC0\.00270\.00262\.13152\.1367Baseline3000\-0\.00364\.17074\.1743CMC0\.00280\.00293\.21803\.22371000Baseline512\-0\.00721\.12381\.1310CMC0\.00540\.00601\.05501\.0664Baseline1024\-0\.00682\.32372\.3306CMC0\.00550\.00502\.10092\.1115Baseline2048\-0\.00554\.32934\.3349CMC0\.00400\.00453\.65033\.6587Baseline3000\-0\.00566\.79466\.8001CMC0\.00420\.00445\.33975\.3483Table 13:Stage\-wise energy consumption \(kWh\) forMistral\-7B\-Instruct\-v0\.3on SQuAD across three sample sizes \(GPT2\-Large ContextEncoder,r=4r=4, batch size 8\)\. Stage\-wise energy is apportioned proportionally to wall\-clock time\.
### A\.4Peak GPU Memory
Tables[14](https://arxiv.org/html/2609.25537#A1.T14)and[15](https://arxiv.org/html/2609.25537#A1.T15)report peak allocated and peak reserved GPU memory for Llama and Mistral across three sample sizes andT∈\{512,1024,2048,3000\}T\\in\\\{512,1024,2048,3000\\\}, extending the 1,000\-sample Llama results in Section[3\.4\.3](https://arxiv.org/html/2609.25537#S3.SS4.SSS3)\. Peak allocated memory is stable acrossTTfor CMC \(≈\\approx8\.85 GB for Llama,≈\\approx8\.43 GB for Mistral\) regardless of sample size\. The baseline is also stable for Llama \(≈\\approx10\.0 GB\) but grows withTTfor Mistral \(9\.1 GB atT=512T=512to 11\.1 GB atT=3,000T=3\{,\}000\), reflecting Mistral’s longer tokenised prompts and larger KV cache growth per step\. The key finding is in peak reserved memory\. CMC maintains a near\-constant footprint across all sample sizes and token budgets \(≈\\approx9\.8 GB for Llama,≈\\approx9\.1 GB for Mistral\), while the baseline spikes sharply at largeTT: up to 19\.7 GB for Llama and 24\.3 GB for Mistral atT=3,000T=3\{,\}000\. This confirms that the bounded two\-tier KV cache prevents memory growth regardless of decoder architecture, sample size, or generation length\.
Peak Alloc\. \(GB\)Peak Res\. \(GB\)SamplesModelTTBaseCMCBaseCMC256BaselinevsCMC51210\.0338\.91611\.2349\.74210249\.9688\.84811\.1259\.74220489\.9708\.80611\.1009\.717300010\.1518\.84815\.3329\.742512BaselinevsCMC51210\.0188\.85011\.1729\.807102410\.0278\.85912\.43611\.381204810\.0188\.85011\.1729\.807300010\.1608\.85018\.6179\.8071000BaselinevsCMC51210\.0178\.85311\.1979\.836102410\.0178\.85311\.1979\.836204810\.0178\.85311\.1999\.838300010\.1598\.85319\.6899\.836Table 14:Peak allocated and peak reserved GPU memory \(GB\) forLlama\-3\-8B\-Instructon SQuAD \(GPT2\-Large,r=4r=4, batch size 8\)\. Peak allocated memory is measured viatorch\.cuda\.max\_memory\_allocated\(\)and peak reserved memory viatorch\.cuda\.max\_memory\_reserved\(\)\.Peak Alloc\. \(GB\)Peak Res\. \(GB\)SamplesModelTTBaseCMCBaseCMC256BaselinevsCMC5129\.0528\.43210\.1979\.08410249\.2668\.43410\.3469\.084204810\.2038\.43414\.1629\.084300011\.0758\.39724\.2629\.082512BaselinevsCMC5129\.0738\.44910\.4349\.14610249\.2768\.43110\.3289\.082204810\.1928\.43114\.1259\.082300011\.0858\.43124\.2629\.0821000BaselinevsCMC5129\.0748\.43110\.2219\.08210249\.2768\.43110\.3289\.082204810\.2138\.44914\.1669\.109300011\.0848\.44924\.3039\.109Table 15:Peak allocated and peak reserved GPU memory \(GB\) forMistral\-7B\-Instruct\-v0\.3on SQuAD \(GPT2\-Large,r=4r=4, batch size 8\)\. Peak reserved memory is measured viatorch\.cuda\.max\_memory\_reserved\(\)\.
### A\.5Attention MACs Analysis
Attention MACs are estimated analytically asMACst=2⋅nh⋅dh⋅nl⋅Lt\\text\{MACs\}\_\{t\}=2\\cdot n\_\{h\}\\cdot d\_\{h\}\\cdot n\_\{l\}\\cdot L\_\{t\}per generation step, whereLt≤K\+W\+\|q\|L\_\{t\}\\leq K\+W\+\|q\|for CMC vs\.Lt=P\+tL\_\{t\}=P\+tfor the baseline, withPPdenoting the combined length of the baseline’s context window and question, andttthe number of tokens generated so far\. Absolute values differ across decoders due to architecture \(nl=42n\_\{l\}=42for Gemma vs\. 32 for Llama and Mistral\) and tokeniser\-dependent prompt lengths\. Table[16](https://arxiv.org/html/2609.25537#A1.T16)shows that theK\+WK\{\+\}Wbound reduces attention MACs by 27\.0% on average \(6\.6%–51\.3%\) across all nine combinations\. Reductions are largest on SQuAD \(34\.0% average\) where longer prompt lengths make the fixed cache bound more effective\. Smaller reductions on HotpotQA for Llama and Gemma \(6\.6% and 8\.3%\) reflect shorter baseline prompts in that split\. These reductions are consistent with the empirical efficiency gains in Section[3\.4\.3](https://arxiv.org/html/2609.25537#S3.SS4.SSS3), confirming that CMC’s efficiency advantage is architectural rather than hardware\-specific\.
DatasetDecoderBaselineCMCReductionSQuADLlama1\.611\.1230\.6%Mistral1\.821\.1238\.1%Gemma3\.002\.0133\.0%AdversarialQALlama0\.710\.5522\.0%Mistral0\.770\.5627\.4%Gemma2\.581\.9424\.9%HotpotQALlama0\.600\.566\.6%Mistral0\.720\.3551\.3%Gemma1\.561\.438\.3%Table 16:Estimated attention MACs \(Billion\) for CMC vs\. the baseline \(r=4r=4,W=130W=130,K=40K=40\)\.
### A\.6Local Window Size Sensitivity
Figure[5](https://arxiv.org/html/2609.25537#A1.F5)shows EM and F1 as a function of the Tier\-2 local window sizeWWon SQuAD\. EM increases monotonically from 0\.140 atW=50W=50to 0\.354 atW=200W=200, then plateaus atW=256W=256\(EM==0\.355\), confirming that larger local windows consistently benefit extractive QA by providing more surrounding context for precise span extraction\. The plateau atW≥200W\\geq 200suggests diminishing returns once the window captures the full answer region for most passages\. The defaultW=130W=130achieves EM==0\.304, which is 5\.0 points below theW=200W=200\(EM==0\.354\), indicating thatW=200W=200is a stronger default for deployment when the additional KV budget \(K\+W=230K\+W=230tokens\) is available\. The finding is consistent with the 200\-sample evaluation, which shows the same monotonic pattern and plateau boundary, confirming robustness to sample size\.
50508080100100130130160160200200256256000\.10\.10\.20\.20\.30\.30\.40\.40\.50\.50\.60\.60\.70\.7Local windowWW\(tokens\)ScoreEMF1Figure 5:Local windowWWsensitivity on SQuAD \(full\-data setting, GPT2\-Large,r=4r=4, Llama,K=30K=30, 2000 validation samples\)\. EM and F1 increase monotonically withWW, plateauing atW≥200W\\geq 200\.
### A\.7Denoising Threshold \(τ\\tau\) Sensitivity
Figure[6](https://arxiv.org/html/2609.25537#A1.F6)shows EM and F1 as a function of the graph\-based context denoising thresholdτ\\tauon SQuAD\. The results reveal a binary pattern\. Atτ=0\\tau=0\(no denoising\), EM drops to 0\.307 — a 19\.9\-point degradation relative to any denoising setting \(τ≥0\.02\\tau\\geq 0\.02, EM==0\.507\)\. This confirms that graph\-based context denoising is critical for CME quality: without it, the ContextEncoder compresses all tokens including highly redundant ones, producing a noisier and less informative CME prefix\. Forτ≥0\.02\\tau\\geq 0\.02, EM and F1 are identical across all tested values \(EM==0\.507, F1==0\.739\), with token retention stabilising at 14\.0%\. This plateau arises because thekeep\_min\_every==16 constraint – which forces retention of at least one token every 16 positions – becomes the binding constraint forτ≥0\.02\\tau\\geq 0\.02\. Once this floor dominates, the exact value ofτ\\tauhas no further effect on which tokens are retained\. The defaultτ=0\.05\\tau=0\.05is therefore well within the stable region and robust to any choice in the range\[0\.02,0\.20\]\[0\.02,0\.20\]\.
0\.000\.020\.050\.080\.100\.150\.20000\.10\.10\.20\.20\.30\.30\.40\.40\.50\.50\.60\.60\.70\.70\.80\.8Denoising thresholdτ\\tauScoreEMF1Figure 6:Denoising thresholdτ\\tausensitivity on SQuAD \(full\-data setting, GPT2\-Large,r=4r=4, Llama,K=30K=30,W=140W=140, 2,000 validation samples\)\. No denoising \(τ=0\\tau=0\) reduces EM by 19\.9 points\. Forτ≥0\.02\\tau\\geq 0\.02, EM and F1 are identical as thekeep\_min\_every=16=16floor becomes the binding constraint\. The defaultτ=0\.05\\tau=0\.05\(dashed line\) lies within the stable region\.
### A\.8Tier\-2 Window Localization Analysis
In main experiments, the Tier\-2 local window is centered using the gold answer character offsetα\\alphaprovided by benchmark annotations, applied identically to both CMC and the baseline so thatα\\alphais a controlled variable rather than information available only to CMC\. While this is standard practice in extractive QA evaluation,α\\alphais unavailable at deployment time\. Table[17](https://arxiv.org/html/2609.25537#A1.T17)quantifies the performance gap between the oracle offset and three deployment\-realistic alternatives on 2,000 validation samples from SQuAD and AdversarialQA \(full\-data setting\) and HotpotQA \(controlled\-data setting\) \(GPT2\-Large,r=4r=4, Llama,K=30K=30,W=140W=140\)\.BM25 top\-1 sentence[Robertson and Zaragoza \(2009\)](https://arxiv.org/html/2609.25537#bib.bib37)ranks context sentences by relevance to the question and uses the midpoint of the top\-ranked sentence asα\\alpha\.CME\-guided localizationuses the character offset of the chunk that produced the highest\-norm CME asα\\alpha, requiring no external library\.Midpoint heuristicalways centres atα=\|C\|/2\\alpha=\|C\|/2, using no information about the question or answer at all\.
On SQuAD, oracle\-free localization stays close to the oracle: BM25 costs only 0\.4 EM points \(Δ\\DeltaEM==−\-0\.004, 99% retention\), CME\-guided costs 2\.0 points \(−\-0\.020, 95%\), and even the midpoint heuristic retains 91% of oracle EM\. HotpotQA shows the same pattern, and more strongly: BM25, CME\-guided, and the midpoint heuristic all retain 97–98% of oracle EM, with CME\-guided \(0\.415\) and the midpoint heuristic \(0\.414\) essentially matching BM25 \(0\.412\)\. We attribute this robustness to HotpotQA’s multi\-hop structure, where answer\-supporting evidence is distributed across multiple passages rather than concentrated at a single span\. AdversarialQA shows the largest oracle dependence \(−\-0\.052 to−\-0\.069, 78–84% retention\), consistent with its adversarially constructed questions being designed to resist the shallow lexical and positional cues that BM25 and the midpoint heuristic rely on\. These experimental results show that CMC’s Tier\-2 benefit comes from the presence of a local context window, not from privileged access to the gold answer location: oracle\-free localization retains the large majority of oracle performance \(78−\-99%\) across single\-hop, adversarial, and multi\-hop QA datasets\.
DatasetTier\-2 localisationEMF1𝚫\\boldsymbol\{\\Delta\}EMSQuADOracle\(α\\alphafrom annotations\)0\.4080\.591—BM25 top\-1 sentence0\.4040\.585−\-0\.004CME\-guided \(top\-1 CME chunk\)0\.3880\.556−\-0\.020Midpoint heuristic0\.3730\.533−\-0\.035AdversarialQAOracle0\.3210\.501—BM250\.2690\.428−\-0\.052CME\-guided0\.2640\.405−\-0\.057Midpoint heuristic0\.2520\.389−\-0\.069HotpotQAOracle0\.4230\.634—BM250\.4120\.621−\-0\.011CME\-guided0\.4150\.622−\-0\.008Midpoint heuristic0\.4140\.624−\-0\.009Table 17:Effect of Tier\-2 window localization method on SQuAD \(full\-data\), AdversarialQA \(full\-data\), and HotpotQA \(controlled\-data\) \(GPT2\-Large,r=4r=4, Llama,K=30K=30,W=140W=140, 2,000 validation samples per dataset\)\.Δ\\DeltaEM is relative to the oracle within each dataset\.
### A\.9Case Study
Table[18](https://arxiv.org/html/2609.25537#A1.T18)presents four representative examples from the SQuAD validation set illustrating the behaviour of CMC relative to the baseline across four outcome categories\.Case 1shows a straightforward factual question where both CMC and the baseline correctly extract the gold answer \(oxides\), confirming that CMC preserves answer quality on simple single\-token spans\.Case 2demonstrates a clear CMC advantage\. The context contains two monetary figures — £43 million \(the original gallery budget\) and £76 million \(the revised cost estimate\)\. The baseline anchors to the more prominent earlier mention and returns the incorrect figure\. The CME prefix supplies sufficient global context for CMC to distinguish between the two figures and return the correct answer\.Case 3illustrates a limitation of CMC\. Both CMC and the baseline locate the correct passage, but CMC produces1979rather than the goldJune 1979, losing the month through compression\. This suggests that fine\-grained temporal spans are susceptible to information loss when the relevant token is not well\-represented in the CME prefix\.Case 4shows a partial improvement\. The gold answer is a full sentence \(The city’s residents fled to the north\), which neither model matches exactly\. The baseline retrieves onlythe north\(F1==0\.40\), while CMC retrievesto the north\(F1==0\.55\), a longer span that better overlaps with the gold answer\. This illustrates that CMC can improve partial answer coverage even when an exact match is not achieved\.
Case 1Both Win \(Baseline✓ / CMC✓\)Question:What is the usual form of oxygen bound compounds?Gold answer:oxidesBaseline:oxidesEM==1 / F1==1\.00CMC:oxidesEM==1 / F1==1\.00Context:Due to its electronegativity, oxygen forms chemical bonds with almost all other elements to give corresponding oxides\. The surface of most metals, such as aluminium and titanium, are oxidized in the presence of air and become coated with a thin film of oxide that passivates the metal and slows further corrosion\. Many oxides of the transition metals are non\-stoichiometric compounds, with slightly less metal than the chemical formula would show\. For example, the mineral FeO \(wüstite\) is written as Fe 1 \- xO, where x is usually around 0\.05\.Case 2CMC wins \(Baseline✗ / CMC✓\)Question:What is the estimated cost of the V&A branded gallery?Gold answer:£76 millionBaseline:£43 millionEM==0 / F1==0\.50CMC:£76 millionEM==1 / F1==1\.00Context:The V&A is in discussion with the University of Dundee, University of Abertay, Dundee City Council and the Scottish Government with a view to opening a new £43 million gallery in Dundee that would use the V&A brand although it would be funded through and operated independently\. As of 2015, with costs estimated at £76 million, it is the most expensive gallery project ever undertaken in Scotland\. The V&A Dundee will be on the city’s waterfront and is intended to focus on fashion, architecture, product design, graphic arts and photography\. It is planned that it could open within five years\. Dundee City Council is expected to pay a major part of the running costs\. The V&A is not contributing financially, but will be providing expertise, loans and exhibitions\.Case 3Baseline wins \(Baseline✓ / CMC✗\)Question:When was 7 Lincoln Square completed?Gold answer:June 1979Baseline:June 1979EM==1 / F1==1\.00CMC:1979EM==0 / F1==0\.67Context:Meanwhile, ABC News, which formed as a newly separate division, sought to become a global leader in television news\. In 1977, Roone Arledge was named president of the new ABC News in addition to being president of ABC Sports\. That same year, ABC launched a major expansion of its office facilities in New York City\. The company first constructed a new 10\-story building on land previously occupied by an abandoned warehouse on the corner of Columbus Avenue and West 66th Street; the facility that was built in its place is nicknamed "7 Lincoln Square" \(although it is actually located at 149 Columbus Avenue\)\. Meanwhile, a former parking lot, located at 30 West 67th Street, was transformed into an impressive 15\-story building\. Both buildings were completed in June 1979\. WABC\-TV moved its operations from offices at 77 West 66th Street to 149 Columbus Avenue, freeing up space for the ABC network to house some of its operations\.Case 4CMC partial gain \(Baseline∼\\sim/ CMC∼\\sim\)Question:Where did the residents of Antioch flee to?Gold answer:The city’s residents fled to the northBaseline:the northEM==0 / F1==0\.40CMC:to the northEM==0 / F1==0\.55Context:The plague struck various countries in the Middle East during the pandemic, leading to serious depopulation and permanent change in both economic and social structures\. As it spread to western Europe, the disease entered the region from southern Russia also\. By autumn 1347, the plague reached Alexandria in Egypt, probably through the port’s trade with Constantinople, and ports on the Black Sea\. During 1347, the disease travelled eastward to Gaza, and north along the eastern coast to cities in Lebanon, Syria and Palestine, including Ashkelon, Acre, Jerusalem, Sidon, Damascus, Homs, and Aleppo\. In 1348–49, the disease reached Antioch\. The city’s residents fled to the north, most of them dying during the journey, but the infection had been spread to the people of Asia Minor\.Table 18:Qualitative case study on SQuAD validation set \(full\-data setting, GPT2\-Large,r=4r=4, Llama\)\.Baselinereceives the local context window without compression\. Gold answers are shown inblue, correct predictions ingreen, and incorrect predictions inred\. EM and F1 are per\-example scores\.相似文章
大规模端到端上下文压缩
本文提出隐上下文语言模型(LCLMs),这是一系列编码器-解码器压缩器,通过架构搜索和大规模预训练高效处理长上下文,在准确性、速度和内存使用上优于传统KV缓存方法。
理解在早期完成:大型语言模型中的深度分工及其用于无界上下文记忆
本文介绍了CoMem,一种利用LLM中深度方向分工的方法,它缓存中间残差张量,并且仅重新计算上层以进行检索,从而实现有界的读取计算和内存,与存储的上下文长度无关。在Qwen3-8B上评估,CoMem在长上下文任务上取得了强劲性能,同时显著节省内存并加速预填充。
面向长上下文大语言模型的训练-推理一致性分段执行
本文提出了一种面向长上下文大语言模型的训练-推理一致性分段执行框架,旨在解决全上下文训练与受限推理机制之间的不匹配问题,在显著降低内存占用的同时实现了相当的性能。
更少的上下文,更高的准确性:一种用于LLM代理的双时态记忆引擎,其中精简检索的上下文胜过了完整历史
本文介绍了Engram,一个开源的用于LLM代理的双时态记忆引擎,它通过检索一个紧凑的上下文片段(约9.6k token),在LongMemEval上以混合读取路径融合稠密、词汇、图和时间信号,比完整历史基线(79k token)高出10.4个准确率点。
决策感知记忆卡:面向工具使用LLM代理的反事实启发式上下文选择与压缩
介绍了CICL,一种决策感知上下文层,通过将上下文视为决策时刻的干预,使用反事实启发式评分和类型化记忆卡(受令牌预算限制),为工具使用的LLM代理选择和压缩证据。在SWE-bench和RepoBench上的实验显示,在检索准确性和行动关键性方面取得了实际提升。