Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Hugging Face Daily Papers 05/16/26, 12:00 AM Papers

long-context sparse-attention efficient-inference llm attention-mechanism kv-cache

Summary

RTPurbo leverages intrinsic sparsity in full-attention LLMs to achieve efficient long-context inference with minimal training overhead, enabling significant speedups while maintaining near-lossless accuracy.

Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-p selection more suitable than fixed top-k sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36times prefill speedup at 1M context and about a 2.01times decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.

Original Article

View Cached Full Text

Cached at: 05/22/26, 06:23 AM

Paper page - Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Source: https://huggingface.co/papers/2605.16928

Abstract

Long-context inferencein large language models is bottlenecked by the quadratic cost offull attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset ofattention headstruly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, makingdynamic top-p selectionmore suitable than fixed top-k sparsification. Based on these insights, we proposeRTPurbo, which retains the fullKV cacheonly for retrieval heads and introduces a lightweighttoken indexerforsparse attention. By exploiting the model’sintrinsic sparsity,RTPurboachieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show thatRTPurbopreserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36timesprefill speedupat 1M context and about a 2.01timesdecode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.

View arXiv page View PDF Add to collection

Get this paper in your agent:

hf papers read 2605\.16928

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.16928 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.16928 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.16928 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Paper page - Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Abstract

Models citing this paper0

Datasets citing this paper0

Spaces citing this paper0

Collections including this paper0

Similar Articles

I built a 2.5D visual compiler for AI agents: It separates topology from geometry so LLMs stop generating spaghetti diagrams.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

Epiphany-Aware KV Cache Eviction Without the Attention Matrix

Information-Aware KV Cache Compression for Long Reasoning

Submit Feedback

Similar Articles

I built a 2.5D visual compiler for AI agents: It separates topology from geometry so LLMs stop generating spaghetti diagrams.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

Epiphany-Aware KV Cache Eviction Without the Attention Matrix

Information-Aware KV Cache Compression for Long Reasoning