@VukRosic99: Long-context Transformers hit two walls: quadratic attention compute and a KV cache that reaches hundreds of GB at 1M t…

X AI KOLs Timeline Papers

Summary

MiniCPM-SALA is a 9B-parameter hybrid attention model that interleaves sparse and linear attention to overcome the quadratic compute and large KV cache bottlenecks of long-context Transformers. It achieves 3.5x faster inference than Qwen3-8B at 256K tokens and supports up to 1M tokens on consumer GPUs, with a cost-effective continual training approach that reduces training costs by ~75%.

Long-context Transformers hit two walls: quadratic attention compute and a KV cache that reaches hundreds of GB at 1M tokens. Sparse attention cuts compute but keeps dense storage; linear attention fixes both but compresses context. MiniCPM-SALA (OpenBMB, 9B) interleaves them - 1 sparse layer per 3 linear ones, placement chosen automatically, stabilized by QK-Norm everywhere, RoPE only on the linear layers, and an output gate after every block. The economic headline: it converts an intermediate MiniCPM-4.0 checkpoint into the hybrid in 5 stages for about 2T tokens - roughly a quarter of the 8T scratch cost - dodging the cold-start instability hybrid-from-scratch recipes pay. Results: matches Qwen3-8B on standard benchmarks, pulls clearly ahead on long context (RULER-128K 89.4 vs 71.7). Trained only to 520K, it extrapolates to 2M tokens with no length-extension trick. Inference: 3.5x faster than Qwen3-8B at 256K on one A6000D, and on a 32GB RTX 5090 it scales to 1M tokens where Qwen3-8B runs out of memory at 128K. Made a short visual breakdown - one diagram per trick. Swipe through. --- paper - https://arxiv.org/abs/2602.11761 code - https://github.com/OpenBMB/MiniCPM model - https://huggingface.co/openbmb/MiniCPM-SALA… full summary pdf - https://gist.github.com/vukrosic/0c77bcf28e99843e4f47694b5b866b2b… Every Sunday I run a hands-on live AI research with 1 on 1 help: https://skool.com/become-ai-researcher-2669/about…
Original Article
View Cached Full Text

Cached at: 07/11/26, 03:26 PM

Long-context Transformers hit two walls: quadratic attention compute and a KV cache that reaches hundreds of GB at 1M tokens. Sparse attention cuts compute but keeps dense storage; linear attention fixes both but compresses context.

MiniCPM-SALA (OpenBMB, 9B) interleaves them - 1 sparse layer per 3 linear ones, placement chosen automatically, stabilized by QK-Norm everywhere, RoPE only on the linear layers, and an output gate after every block.

The economic headline: it converts an intermediate MiniCPM-4.0 checkpoint into the hybrid in 5 stages for about 2T tokens - roughly a quarter of the 8T scratch cost - dodging the cold-start instability hybrid-from-scratch recipes pay.

Results: matches Qwen3-8B on standard benchmarks, pulls clearly ahead on long context (RULER-128K 89.4 vs 71.7). Trained only to 520K, it extrapolates to 2M tokens with no length-extension trick. Inference: 3.5x faster than Qwen3-8B at 256K on one A6000D, and on a 32GB RTX 5090 it scales to 1M tokens where Qwen3-8B runs out of memory at 128K.

Made a short visual breakdown - one diagram per trick. Swipe through.


paper - https://arxiv.org/abs/2602.11761 code - https://github.com/OpenBMB/MiniCPM model - https://huggingface.co/openbmb/MiniCPM-SALA… full summary pdf - https://gist.github.com/vukrosic/0c77bcf28e99843e4f47694b5b866b2b…

Every Sunday I run a hands-on live AI research with 1 on 1 help: https://skool.com/become-ai-researcher-2669/about…


MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling

Source: https://arxiv.org/html/2602.11761

Abstract

The evolution of large language models (LLMs) towards applications with ultra-long contexts faces challenges posed by the high computational and memory costs of the Transformer architecture. While existing sparse and linear attention mechanisms attempt to mitigate these issues, they typically involve a trade-off between memory efficiency and model performance. This paper introduces MiniCPM-SALA111SALA stands for Sparse Attention and Linear Attention., a 9B-parameter hybrid architecture that integrates the high-fidelity long-context modeling of sparse attention (InfLLM-V2) with the global efficiency of linear attention (Lightning Attention). By employing a layer selection algorithm to integrate these mechanisms in a 1:3 ratio and utilizing a hybrid positional encoding (HyPE), the model maintains efficiency and performance for long-context tasks. Furthermore, we introduce a cost-effective continual training framework that transforms pre-trained Transformer-based models into hybrid models, which reduces training costs by approximately 75% compared to training from scratch. Extensive experiments show that MiniCPM-SALA maintains general capabilities comparable to full-attention models while offering improved efficiency. On a single NVIDIA A6000D GPU, the model achieves up to 3.5×\timesthe inference speed of the full-attention model at the sequence length of 256K tokens and supports context lengths of up to 1M tokens, a scale where traditional full-attention 8B models fail because of memory constraints.

1Introduction

As large language models (LLMs)(OpenAIet al.,2024; Comaniciet al.,2025; Grattafioriet al.,2024; Yanget al.,2025a; DeepSeek-AIet al.,2025)become increasingly effective, the application scenarios of LLMs are undergoing a profound paradigm shift, transitioning from simple question-answering(Brownet al.,2020)to more advanced applications, such as deep understanding and generation of ultra-long contexts(Baiet al.,2024,2025; Zhouet al.,2025; Shaoet al.,2024), repository-scale code engineering(Guoet al.,2024; Jimenezet al.,2024; Liuet al.,2024), and long-horizon agents for complex tasks(Qianet al.,2024; Mialonet al.,2023; Liet al.,2026). For these advanced applications, models are no longer confined to processing fragmented information. Instead, they must demonstrate the capacity to handle ultra-long contexts, such as grasping entire technical manuals at once, analyzing comprehensive project dependency trees containing tens of thousands of lines of code, and maintaining coherent task states and memory over multi-day human-AI collaborations. This pursuit of holistic contextual information makes the ability to process millions of tokens a critical aspect for advanced LLMs(Kimi Teamet al.,2025; NVIDIAet al.,2025b).

However, the Transformer architecture(Vaswaniet al.,2017), which is the foundation of modern LLMs, encounters severe computational bottlenecks when handling ultra-long contexts due to its core full-attention mechanism. This bottleneck manifests primarily in two dimensions: (1) thecompute bottleneckof computational complexity: for the standard attention mechanism, the computational cost grows quadratically with the sequence lengthNN, i.e., its complexity is𝒪​(N2)\mathcal{O}(N^{2}). When the context scales to the level of millions of tokens, the huge overhead causes the inference latency to increase dramatically; (2) thememory bottleneckof KV-Cache: during the auto-regressive generation process, the model must store the key and value states (KVs) of all historical contextual tokens to avoid redundant computation. For a typical 8B-parameter model, even when utilizing Grouped Query Attention (GQA)(Ainslieet al.,2023), the KV-Cache required for millions of tokens can reach dozens or even hundreds of gigabytes.

To address the aforementioned challenges, existing solutions have developed two primary paradigms: Sparse Attention(Yuanet al.,2025; DeepSeek-AIet al.,2025; Xiaoet al.,2024; Zhaoet al.,2025)and Linear Attention(Yanget al.,2024a; Gu and Dao,2024; Penget al.,2023; Yanget al.,2024b,2025b). Both paradigms present distinct advantages and inherent limitations. Sparse attention methods attempt to break the compute bottleneck by computing only the most salient portions of the attention matrix, such as adopting sliding windows or global anchors. However, these methods are hindered by a “sparse computation, dense storage” limitation. While local computation reduces immediate processing overhead, the model must still retain the full KV-Cache to support contextual information retrieval. Linear attention utilizes recurrent formulations to successfully reduce computational complexity to𝒪​(N)\mathcal{O}(N). Nevertheless, this extreme efficiency is achieved by the lossy compression of contextual information and inevitably results in performance degradation.

MiniCPM-SALA employs a hybrid architecture of sparse and linear attention(Chenet al.,2026), specifically designed to achieve efficient ultra-long sequence modeling. This architecture combines the high-fidelity long-context modeling capabilities of InfLLM-V2(Zhaoet al.,2025)and the global computational efficiency of Lightning Attention(Qinet al.,2024). Through this integrated approach, the model significantly mitigates inference overhead and memory consumption, while simultaneously addressing the precision bottleneck typical of pure linear architectures in long-range information processing. Consequently, MiniCPM-SALA provides a balanced solution that maintains both efficiency and high performance for long-context tasks. Furthermore, we employ the continual training paradigm to transform a pre-trained Transformer model into our hybrid model. By eschewing training from scratch, this approach significantly reduces the computational costs of model development. While several works have begun exploring the integration of sparse and linear attention(Huet al.,2025; Houet al.,2025; He and Garner,2025), to the best of our knowledge, MiniCPM-SALA is the first to demonstrate through large-scale experimentation that these hybrids can match the performance of full-attention baselines. Furthermore, the model exhibits high efficiency and strong performance in long-context processing.

In summary, the main contributions of this study can be outlined as follows:

  • •We introduce a Sparse-Linear hybrid attention mechanism integrating 25% InfLLM-V2 and 75% Lightning Attention to strike a balance between throughput and precision. By leveraging the granular focus of sparse attention for local details and the𝒪​(N)\mathcal{O}(N)efficiency of linear attention for broad context, the architecture maintains high semantic accuracy as the sequence length scales up.
  • •We demonstrate that the Transformer-to-hybrid paradigm is a highly effective strategy for building strong hybrid models. This approach circumvents the inefficiencies of cold-start training by performing an architectural transformation on the pre-trained weights, thereby reducing the total training budget to approximately 25% relative to training a comparable model from scratch.
  • •We adopt HyPE (Hybrid Positional Encoding)(Chenet al.,2026)to effectively harmonize the performance across both short and long contexts. While maintaining general capabilities (e.g., knowledge, mathematics, and coding) comparable to modern full-attention models like Qwen3-8B, MiniCPM-SALA has substantial advantages across multiple long-context benchmarks.
  • •MiniCPM-SALA demonstrates substantial resource savings and speed advantages in long-context scenarios. On the NVIDIA A6000D GPU, MiniCPM-SALA achieves up to 3.5×\timesthe inference speed of Qwen3-8B at a sequence length of 256K tokens. Furthermore, MiniCPM-SALA supports inference at context lengths of up to 1M tokens on both NVIDIA A6000D and 5090 GPUs, whereas Qwen3-8B fails at this length due to out-of-memory (OOM) errors. These results demonstrate the broad prospects of MiniCPM-SALA in edge-side information-intensive applications.

2Model Development

In this section, we introduce the model architecture and training strategies for MiniCPM-SALA. Specifically, we combine the efficient sparse attention for long-context modeling and linear attention for global efficiency in MiniCPM-SALA. Moreover, we also introduce an efficient training method, which can transform a standard Transformer model into sparse-linear hybrid attention.

Refer to captionFigure 1:Architecture of MiniCPM-SALA. The model adopts an efficient hybrid design that combines InfLLM-V2(Zhaoet al.,2025)and Lightning Attention(Qinet al.,2024)modules in a 1:3 ratio. Building on an intermediate MiniCPM-4.0(MiniCPM-Teamet al.,2025)checkpoint, MiniCPM-SALA undergoes a continual training phase to convert a standard Transformer model into a sparse-linear hybrid model.### 2.1Model Architecture

The overall architecture of MiniCPM-SALA is illustrated in Figure1. MiniCPM-SALA adopts a hybrid architecture that interleaves sparse attention layers and linear attention layers. We retain the Feed-Forward Network (FFN) block after each attention block in the Transformer architecture to ensure high-capacity knowledge representation. Inspired by the architectural designs of recent representative studies, such as Qwen3-Next(Qwen Team,2025)and Kimi-Linear(Kimi Teamet al.,2025), as well as our internal small-scale preliminary experiments, we employ a 1:3 mixing ratio: 25% of the layers adopt sparse attention while the remaining 75% employ linear attention.

This hybrid configuration leverages the complementary strengths of both attention mechanisms. Linear attention layers have constant computational and memory complexities with respect to sequence length, facilitating efficient processing of long contexts. On the other hand, sparse attention layers facilitate effective modeling of long-range dependencies. Rather than naively uniformly interleaving the two attention variants, we determine the placement of sparse attention modules using the layer selection mechanism proposed byChenet al.(2026), which results in superior downstream performance.

Training StrategyExisting paradigms for training hybrid models generally fall into two categories: (1) training from scratch(Zuoet al.,2025; Qwen Team,2025; Kimi Teamet al.,2025; NVIDIAet al.,2025b)and (2) converting a pre-trained Transformer model into a hybrid model via cross-architecture distillation(Wanget al.,2024a; Hoshinoet al.,2025; Liet al.,2025; Guet al.,2025). Although training from scratch offers simplicity and maximum architectural flexibility, continual-training conversion is a more resource-efficient alternative that leverages parameter inheritance from established pre-trained models. By recycling pre-trained weights and representations, the continual-training method significantly reduces the immense computational cost typically associated withde novotraining, achieving competitive performance with a fraction of the budget. Accordingly, MiniCPM-SALA leverages a conversion-based framework that uses continual training to adapt a Transformer into an efficient hybrid version while preserving its core capabilities.

Sparse Attention and Linear AttentionFor the sparse attention layers, we incorporate InfLLM-V2(Zhaoet al.,2025), which offers the distinct advantage of introducing no additional parameters to the architecture. Its inherent flexibility and ability to switch seamlessly between dense and sparse modes are highly compatible with our conversion process. This compatibility facilitates a stable training initialization by allowing sparse modules to inherit dense weights without architectural discrepancies, ensuring that the conversion to a hybrid structure does not compromise the model capacity. For the linear attention layers, we utilize Lightning Attention(Qinet al.,2024). Given our Transformer-to-hybrid conversion paradigm, Lightning Attention is selected for its functional proximity to the standard softmax attention. This structural alignment is intended to mitigate the complexities of parameter adaptation, thereby preserving pre-trained knowledge and ensuring robust downstream performance. Lightning Attention also provides better length generalization capabilities according toChenet al.(2026), which may improve data efficiency during long-context continual-training.

Other Architectural ImprovementsFollowing HypeNet(Chenet al.,2026), we also introduce several architectural modifications to enhance the expressivity and training stability of MiniCPM-SALA. These include QK-Normalization(Henryet al.,2020), HyPE(Chenet al.,2026), and the integration of output gates.

  • •QK-Normalization:This is applied to all attention layers (both sparse and linear layers) to prevent the activation spikes that often occur in long-context training and further improve and boost the expressivity of linear attention modules.
  • •HyPE (Hybrid Positional Encoding):To balance rich positional awareness and long-range information retention, we employ a hybrid approach to positional encoding. We apply Rotary Positional Embedding (RoPE)(Suet al.,2023)to the linear attention layers to facilitate position-sensitive memory, allowing the model to preserve the relative order of tokens within the global context. On the other hand, we remove RoPE in the sparse attention layers. This strategic omission prevents the decay of long-distance information often associated with RoPE, thereby enabling more precise recall over extended contexts.
  • •Output gates:Furthermore, we incorporate an output gate after each attention block (both sparse and linear). This architectural choice aligns with recent advances in the gated attention mechanism(Qiuet al.,2025), in which the output gate has been shown to effectively mitigate issues such as attention sink. By regulating the information flow, the output gate prevents excessive focus on specific tokens and ensures a more flexible distribution of attention weights. Empirically, we observe that integrating output gates into both linear and sparse attention significantly improves model stability and performance.

Table 1:Overview of the whole training process to build MiniCPM-SALA.StageTrainableSparseSequence# TokensParametersAttentionLengthArchitecture Conversion (HALO)Linear AttentionDisabled0.5K1.3BContinual Stable-TrainingAll ParametersDisabled4K314.6BShort-Decay TrainingAll ParametersDisabled4K1006.6BLong-Decay TrainingAll ParametersEnabled32K102.2B160K62.9B520K50.6BSupervised Fine-TuningAll ParametersEnabled64K204.5B140K213.3B

2.2Model Training

The training of MiniCPM-SALA is conducted through a multi-stage process that starts from an intermediate checkpoint of MiniCPM-4.0(MiniCPM-Teamet al.,2025), which has already been trained on 7T tokens. This methodology represents an extended implementation of Hybrid Attention via Layer Optimization (HALO)(Chenet al.,2026). In the initial phase, we use the HALO framework to convert softmax attention to linear attention. This conversion serves as the starting point for subsequent pipeline stages, including continual pre-training and post-training. By leveraging this approach, the model can transition from a dense architecture to a hybrid structure while preserving the general capabilities acquired during the backbone’s earlier training phases. The entire conversion process, consisting of five stages, is shown in Table1. It is worth noting that the Transformer-to-hybrid training of MiniCPM-SALA consumes approximately 2T tokens. This corresponds to roughly 25% of the data volume required to train MiniCPM-4.0 from scratch (8T tokens).

Architecture Conversion (HALO)The first stage uses HALO to convert the Transformer model from a full attention architecture to a hybrid architecture. During this phase, the training configuration of MiniCPM-SALA differs from the standard HALO approach in two aspects. First, regarding layer selection, we keep the first and last layers unconverted to improve training stability. For the remaining layers, we utilize the HALO selection algorithm to determine which layers are preserved as softmax attention layers. These preserved softmax attention layers are subsequently trained as sparse attention in later stages. The second difference from standard HALO is that we do not perform the final fine-tuning step of the original HALO process. Instead, we conduct more extensive continual pre-training and post-training, which comprise the subsequent stages of our methodology. The training process at this stage is highly efficient, using only 1.3B tokens with a sequence length of 512 tokens. Furthermore, only the converted linear-attention layers are trainable during this stage, while all other parameters remain frozen.

Continual Stable-TrainingThe second stage is continual stable-training. We use the checkpoint from the previous stage as the starting point for further training on the MiniCPM-4.0 pre-training dataset. The primary objective of this phase is to facilitate better coordination between the converted linear attention layers and other model components, including full attention layers, FFN layers, and embeddings. The sequence length for this process is set to 4K tokens, with a total training volume of 314.6B tokens. Since the sequence length remains relatively short, the sparse attention is disabled at this stage to maintain computational efficiency. For the hyperparameter configuration, the learning rate (LR) is set to7.5×10−37.5\times 10^{-3}and held constant after a 2,000-step LR warmup period. Accounting for the sequence length and the number of GPUs, the global batch size is set to 7.8M tokens.

Short-Decay TrainingThe third stage is short-decay training, during which the LR undergoes exponential decay from7.5×10−37.5\times 10^{-3}to3.75×10−43.75\times 10^{-4}. This process utilizes a sequence length of 4K tokens and a global batch size of 7.8M tokens. This stage involves training on 1T tokens, representing the most extensive data volume in the entire development pipeline. Building on the MiniCPM-4.0 decay strategy, we significantly increase the weight of L2 high-quality selection data(Wanget al.,2026)and introduce a large volume of PDF corpora and L3 synthetic data. This approach aims to enhance general capabilities and logical reasoning using high-information-density training data, achieving the efficient compression and internalization of massive amounts of knowledge.

Long-Decay TrainingThe fourth stage, long-decay, progressively extends the context length from 4K to 32K, 160K, and finally 520K tokens. These processes use data volumes of 102.2B tokens, 62.9B tokens, and 50.6B tokens, respectively. To accommodate the increased sequence lengths, the global batch size is adjusted to 7.8M, 9.8M, and 10.1M tokens, while the LR is systematically decays from3×10−43\times 10^{-4}to2×10−42\times 10^{-4}at 32K, then to1×10−41\times 10^{-4}at 160K, and finally to3.75×10−53.75\times 10^{-5}at 520K to conclude the process. At this stage, we up-sample the proportion of long-context data to better align the model with long-sequence distributions. Given the growing computational advantages of sparse attention at longer sequences, we enable the sparse attention mechanism at this stage and maintain full-parameter training, thereby allowing the model to effectively learn the synergy between sparse attention and linear attention.

Supervised Fine-TuningThe SFT corpus for this stage is composed of high-quality reasoning-intensive data, encompassing code, mathematics, knowledge, function calls, and general dialogue. This selection is designed to fully catalyze the reasoning and task-execution capabilities under complex logic. Furthermore, we specifically synthesize long-context data to enhance the precision of information retrieval and cross-document comprehension within extended sequences. During the SFT stage, the context length is set to 64K and increased to 140K afterwards, utilizing 204.5B and 213.3B tokens, respectively. Sparse attention remains enabled throughout this entire process. By bridging shorter and longer contexts, this strategy allows the model to better balance general capabilities with long-context proficiency. For both phases, the LR follows a schedule with a 1,000-step warmup to a peak of1×10−31\times 10^{-3}before decaying to1×10−41\times 10^{-4}, while the global batch sizes are set to 15.7M for the 64K phase and 17.8M for the 140K phase.

3Experiments

Table 2:Standard evaluation results of MiniCPM-SALA and other open-source LLMs.ModelsQwen3Nemotron-Nano-v2MiniCPM-4.1Ministral-3-RFalcon-H1RMiniCPM-SALA# Param.8B9B8B8B7B9BKnowledgeCMMLU81.6861.5984.7271.7463.5581.55MMLU-Pro73.2671.7972.7068.7570.9867.04CodeHumanEval93.9093.9091.4696.9596.3495.12LCB-v556.8968.2656.8965.8767.6660.48LCB-v648.5760.0051.4353.7157.7152.00MBPP81.3293.3991.0594.1691.0589.11MathAIME2473.3371.6780.8381.4686.6783.75AIME2566.6756.6772.0875.0081.0478.33OtherBBH74.1774.2882.6864.3963.1781.55IFEval84.6686.6977.4570.0686.3276.34Average73.4573.8276.1374.2176.4576.53

Table 3:Long-context evaluation results of MiniCPM-SALA and other open-source LLMs.ModelsQwen3Nemotron-Nano-v2Ministral-3-RFalcon-H1RMiniCPM-SALA# Param.8B9B8B7B9BRULER64K80.5388.7770.6656.5092.65128K71.7468.0145.0936.3389.37MRCR64K-2N29.2020.9144.0213.1829.7764K-4N21.5613.6935.809.0620.5764K-8N17.8213.2417.236.9316.56128K-2N26.5014.6150.309.1728.62128K-4N14.7512.2022.668.2219.62128K-8N12.157.5514.477.5410.12NoLiMa32K43.4019.693.7814.8954.5464K23.3511.822.489.8742.95128K11.255.803.484.7323.86Average32.0225.1228.1816.0438.97

Table 4:Ultra-long context evaluation results of MiniCPM-SALA and other open-source LLMs.∗denotes results cited from the official Qwen3-Next documentation.RULER128K512K1000K2048KQwen3-30B-A3B-Instruct-2507∗89.178.472.8-Qwen3-235B-A22B-Instruct-2507∗93.990.984.5-Qwen3-Next-80B-A3B-Instruct∗96.086.980.3-MiniCPM-SALA (9B)89.487.186.381.6Refer to caption(a)TTFT (s) on A6000D (non-quantized). Refer to caption(b)End-to-end (s) latency on A6000D (non-quantized). Refer to caption(c)TTFT (s) on A6000D (quantized). Refer to caption(d)End-to-end (s) latency on A6000D (quantized).

Figure 2:Inference speed comparison between Qwen3-8B and MiniCPM-SALA. For each tested sequence length, the models process a specified input (prefilling) and generate 1K tokens (decoding). “TTFT” denotes Time To First Token, representing the prefilling latency, while “End-to-end” measures the total latency including both prefilling and decoding phases.Refer to caption(a)TTFT (s) on 5090 (non-quantized). Refer to caption(b)End-to-end (s) latency on 5090 (non-quantized). Refer to caption(c)TTFT (s) on 5090 (quantized). Refer to caption(d)End-to-end (s) latency on 5090 (quantized).

Figure 3:Inference speed comparison between Qwen3-8B and MiniCPM-SALA. For each tested sequence length, the models process a specified input (prefilling) and generate 1K tokens (decoding).### 3.1Model Performance

Benchmarks

To thoroughly assess the general capabilities of the model, we conducted evaluations across a diverse array of benchmarks. These include knowledge-intensive tasks (CMMLU(Liet al.,2023), MMLU-Pro(Wanget al.,2024b)), coding benchmarks (HumanEval(Chenet al.,2021), LCB-v5/v6Jainet al.(2025), MBPP(Austinet al.,2021)), and mathematical reasoning sets (AIME24/25(AIME,2025)), alongside other representative benchmarks such as BBH(Suzgunet al.,2022)and IFEval(Zhouet al.,2023). We further evaluated long-context capabilities using RULER(Hsiehet al.,2024), MRCR222https://huggingface.co/datasets/openai/mrcr, and NoliMa(Modarressiet al.,2025). We utilized the OpenCompass framework(Contributors,2023)to conduct the evaluations.

Baseline ModelsGiven that MiniCPM-SALAis a 9B-parameter model, we selected a series of modern baselines of comparable size, encompassing both hybrid and full-attention architectures. Specifically, the baselines include Qwen3-8B(Yanget al.,2025a), Nemotron-Nano-v2-9B(NVIDIAet al.,2025a), MiniCPM-4.1-8B(MiniCPM-Teamet al.,2025), Ministral-3-Reasoning-8B(Liuet al.,2026), and Falcon-H1R-7B(Teamet al.,2026). We exclude MiniCPM-4.1-8B from the evaluation of long contexts because of its limitation to a context length of 64K.

Results of Standard EvaluationTable2presents the performance of MiniCPM-SALA across a variety of standard benchmarks. The model achieves an average score of 76.53, which represents a competitive level among open-source models of a similar scale. In coding tasks, the model demonstrates high proficiency with scores of 95.12 on HumanEval and 89.11 on MBPP. Mathematical reasoning capabilities also remain robust, as evidenced by the scores of 83.75 on AIME24 and 78.33 on AIME25. These results indicate that the integration of long-context mechanisms does not result in a significant degradation of general capabilities or short-context performance. The model maintains a performance profile that is comparable to, and in some cases exceeds, the performance of models such as Qwen3-8B and Falcon-H1R-7B in standard evaluation settings.

Results of Long-Context EvaluationThe evaluation of long-context capabilities is summarized in Table3, covering benchmarks such as RULER, MRCR, and NoLiMa. MiniCPM-SALA shows a notable proficiency in managing extended input sequences. On the RULER benchmark at a 128K context length, the model maintains a score of 89.37, while many other baselines exhibit a more pronounced decrease in accuracy at the same scale. The advantage of the model is particularly visible in the NoLiMa benchmark, where it achieves a score of 23.86 at the 128K level. This performance is substantially higher than the scores recorded for other models in the comparison. With an overall average long-context score of 38.97, the model demonstrates improved stability and effective information retrieval across large context windows.

Results of Ultra-Long ContextAs demonstrated in Table4, MiniCPM-SALA exhibits surprising length extrapolation capabilities. The results for the Qwen3 models are sourced from the official Qwen3-Next documentation333https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct. Despite being restricted to a 520K training length, the model successfully extrapolates to 2048K tokens without a significant degradation in performance, maintaining a score of 81.6. It is worth noting that this extrapolation requires no auxiliary techniques (e.g., YaRN(Penget al.,2024)). This result highlights the efficacy of our approach in handling context windows far beyond the training stage. Additionally, MiniCPM-SALA shows remarkable parameter efficiency, surpassing the performance of the Qwen3-Next-80B-A3B-Instruct model at the 1000K context length (86.3 vs. 80.3), proving that effective long-context processing does not necessarily require massive parameter counts. The length extrapolation capabilities of MiniCPM-SALA can be attributed to the NoPE configuration within the sparse attention layers. In this design, the stored KV-Cache does not require combination with positional information, which can otherwise hinder the capture of long-range dependencies.

3.2Inference Speed

We assessed the inference speed of MiniCPM-SALA and Qwen3-8B across different hardware and sequence lengths. To verify the long-text processing capabilities of the model in edge computing scenarios, we conducted experiments not only on cloud-grade inference chips, such as the NVIDIA A6000D, but also on consumer-grade edge GPUs, such as the NVIDIA 5090. For each sequence length, we measured both the Time To First Token (TTFT) and the end-to-end latency. The former serves as an indicator of the prefilling speed, while the latter reflects the combined performance of the prefilling and decoding phases. To align the evaluation with practical deployment scenarios, we assessed the inference latency for both non-quantized models and models compressed via GPTQ(Frantaret al.,2023)INT4 quantization.

Figure2presents a comprehensive comparison of inference latency between Qwen3-8B and MiniCPM-SALA on an NVIDIA A6000D GPU (96GB VRAM). We evaluated performance across sequence lengths ranging from 64K to 1024K tokens. As illustrated, MiniCPM-SALA demonstrates a significant performance advantage over the baseline across all tested configurations. In non-quantized settings, MiniCPM-SALA consistently achieves lower latency. Notably, at a sequence length of 256K, MiniCPM-SALA reduces the TTFT from 180.8s (Qwen3) to just 51.6s.

Crucially, the results highlight a distinct advantage in memory efficiency. While Qwen3-8B encounters OOM failures at sequence lengths of 512K and 1024K, MiniCPM-SALA successfully processes these extended contexts. For example, at 1024K tokens, MiniCPM-SALA maintains a TTFT of 250.3s (non-quantized) and 256.9s (quantized), whereas the baseline fails to complete the inference. This trend persists in the end-to-end latency metrics, proving that MiniCPM-SALA is robust enough for ultra-long context generation tasks where full-attention models fail.

Figure3demonstrates the critical advantage of MiniCPM-SALA on memory-constrained hardware. On the RTX 5090 (32GB VRAM), the baseline Qwen3-8B hits a “memory wall” significantly earlier than on the A6000D, triggering OOM errors at just 128K tokens in non-quantized settings and 256K in quantized settings. In stark contrast, MiniCPM-SALA successfully scales to 1024K context lengths without memory failure. This suggests that MiniCPM-SALA effectively democratizes long-context inference, enabling 1M-token processing on consumer-level GPUs where full-attention architectures are unusable.

4Conclusion

In this paper, we presented MiniCPM-SALA, a hybrid architecture that combines sparse and linear attention to overcome the computational and memory bottlenecks of ultra-long context modeling. By utilizing a cost-effective Transformer-to-hybrid training paradigm, we successfully retained the general capabilities of full-attention models while reducing training costs by approximately 75%. Experimental results confirm that MiniCPM-SALA achieves a substantial inference speedup and enables 1M-token context processing on single GPUs (e.g., NVIDIA A6000D), surpassing the limitations of standard 8B models. These results establish MiniCPM-SALA as a scalable and accessible solution for next-generation, information-intensive applications.

5Contributions and Acknowledgments

MiniCPM-SALA is the result of the collective efforts of all members of our team. Please refer toChenet al.(2026)andZhaoet al.(2025)for model architecture details.

Contributors(Ordered by the last name) Wenhao An, Yingfa Chen, Yewei Fang, Jiayi Li, Xin Li, Yaohui Li, Yishan Li, Yuxuan Li, Biyuan Lin, Chuan Liu⋆, Hezi Liu, Siyuan Liu, Hongya Lyu, Yinxu Pan, Shixin Ren, Xingyu Shen, Zhou Su, Haojun Sun, Yangang Sun, Zhen Leng Thai, Xin Tian, Rui Wang⋆, Xiaorong Wang, Yudong Wang, Bo Wu, Xiaoyue Xu, Dong Xu, Shuaikang Xue, Jiawei Yang, Bowen Zhang, Jinqian Zhang, Letian Zhang, Shengnan Zhang, Xinyu Zhang, Xinyuan Zhang⋆, Zhu Zhang, Hengyu Zhao, Jiacheng Zhao⋆, Zhi Zheng, Jie Zhou, Zihan Zhou

Project Design and CoordinationShuo Wang, Chaojun Xiao, Xu Han, Zhiyuan Liu, Maosong Sun

AffiliationsContributors marked with⋆are affiliated with XCORE SIGMA, while the remaining contributors are affiliated with OpenBMB.

References

  • AIME (2025)AIME problems and solutions.External Links:LinkCited by:§3.1.
  • J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai (2023)GQA: training generalized multi-query transformer models from multi-head checkpoints.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,External Links:LinkCited by:§1.
  • J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le,et al.(2021)Program synthesis with large language models.arXiv preprint arXiv:2108.07732.Cited by:§3.1.
  • Y. Bai, X. Lv, J. Zhang, Y. He, J. Qi, L. Hou, J. Tang, Y. Dong, and J. Li (2024)LongAlign: a recipe for long context alignment of large language models.InFindings of the Association for Computational Linguistics: EMNLP 2024,External Links:LinkCited by:§1.
  • Y. Bai, J. Zhang, X. Lv, L. Zheng, S. Zhu, L. Hou, Y. Dong, J. Tang, and J. Li (2025)LongWriter: unleashing 10,000+ word generation from long context LLMs.InThe Thirteenth International Conference on Learning Representations,External Links:LinkCited by:§1.
  • T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell,et al.(2020)Language models are few-shot learners.InAdvances in Neural Information Processing Systems,H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.),Vol.33,pp. 1877–1901.External Links:LinkCited by:§1.
  • M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman,et al.(2021)Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374.Cited by:§3.1.
  • Y. Chen, Z. L. Thai, Z. Zhou, Z. Zhang, X. Shen, S. Wang, C. Xiao, X. Han, and Z. Liu (2026)Hybrid linear attention done right: efficient distillation and effective architectures for extremely long contexts.External Links:2601.22156,LinkCited by:3rd item,§1,§2.1,§2.1,§2.1,§2.2,§5.
  • G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen,et al.(2025)Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.External Links:2507.06261,LinkCited by:§1.
  • O. Contributors (2023)OpenCompass: a universal evaluation platform for foundation models.Note:https://github.com/open-compass/opencompassCited by:§3.1.
  • DeepSeek-AI, A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong,et al.(2025)DeepSeek-v3.2: pushing the frontier of open large language models.External Links:2512.02556,LinkCited by:§1,§1.
  • E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh (2023)GPTQ: accurate post-training quantization for generative pre-trained transformers.External Links:2210.17323,LinkCited by:§3.2.
  • A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan,et al.(2024)The llama 3 herd of models.External Links:2407.21783,LinkCited by:§1.
  • A. Gu and T. Dao (2024)Mamba: linear-time sequence modeling with selective state spaces.InFirst Conference on Language Modeling,External Links:LinkCited by:§1.
  • Y. Gu, Q. Hu, S. Yang, H. Xi, J. Chen, S. Han, and H. Cai (2025)Jet-nemotron: efficient language model with post neural architecture search.External Links:2508.15884,LinkCited by:§2.1.
  • D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang (2024)DeepSeek-coder: when the large language model meets programming – the rise of code intelligence.External Links:2401.14196,LinkCited by:§1.
  • M. He and P. N. Garner (2025)Alleviating forgetfulness of linear attention by hybrid sparse attention and contextualized learnable token eviction.External Links:2510.20787,LinkCited by:§1.
  • A. Henry, P. R. Dachapally, S. S. Pawar, and Y. Chen (2020)Query-key normalization for transformers.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp. 4246–4253.Cited by:§2.1.
  • Y. Hoshino, H. Tachibana, M. Inahara, and H. Takegawa (2025)RAD: redundancy-aware distillation for hybrid models via self-speculative decoding.External Links:2505.22135,LinkCited by:§2.1.
  • H. Hou, Z. Huang, K. Tan, R. Lu, and F. R. Yu (2025)RWKV-x: a linear complexity hybrid language model.External Links:2504.21463,LinkCited by:§1.
  • C. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, and B. Ginsburg (2024)RULER: what’s the real context size of your long-context language models?.InFirst Conference on Language Modeling,External Links:LinkCited by:§3.1.
  • X. Hu, J. Leng, J. Zhao, K. Tu, and W. Wu (2025)Hardware-aligned hierarchical sparse attention for efficient long-term memory access.External Links:2504.16795,LinkCited by:§1.
  • N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica (2025)LiveCodeBench: holistic and contamination free evaluation of large language models for code.InThe Thirteenth International Conference on Learning Representations,External Links:LinkCited by:§3.1.
  • C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan (2024)SWE-bench: can language models resolve real-world github issues?.InThe Twelfth International Conference on Learning Representations,External Links:LinkCited by:§1.
  • Kimi Team, Y. Zhang, Z. Lin, X. Yao, J. Hu, F. Meng, C. Liu, X. Men, S. Yang, Z. Li, W. Li,et al.(2025)Kimi linear: an expressive, efficient attention architecture.External Links:2510.26692,LinkCited by:§1,§2.1,§2.1.
  • H. Li, Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin (2023)CMMLU: measuring massive multitask language understanding in chinese.External Links:2306.09212Cited by:§3.1.
  • K. Li, J. Shi, Y. Xiao, M. Jiang, J. Sun, Y. Wu, S. Xia, X. Cai, T. Xu, W. Si, W. Li, D. Wang, and P. Liu (2026)AgencyBench: benchmarking the frontiers of autonomous agents in 1m-token real-world contexts.External Links:2601.11044,LinkCited by:§1.
  • Y. Li, S. Yang, S. Tan, M. Mishra, R. Panda, J. Zhou, and Y. Kim (2025)Distilling to hybrid attention models via kl-guided layer selection.External Links:2512.20569,LinkCited by:§2.1.
  • A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan,et al.(2026)Ministral 3.External Links:2601.08584,LinkCited by:§3.1.
  • T. Liu, C. Xu, and J. McAuley (2024)RepoBench: benchmarking repository-level code auto-completion systems.InThe Twelfth International Conference on Learning Representations,External Links:LinkCited by:§1.
  • G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom (2023)GAIA: a benchmark for general ai assistants.External Links:2311.12983,LinkCited by:§1.
  • MiniCPM-Team, C. Xiao, Y. Li, X. Han, Y. Bai, J. Cai, H. Chen, W. Chen, X. Cong, G. Cui, N. Ding,et al.(2025)MiniCPM4: ultra-efficient llms on end devices.External Links:2506.07900,LinkCited by:Figure 1,§2.2,§3.1.
  • A. Modarressi, H. Deilamsalehy, F. Dernoncourt, T. Bui, R. A. Rossi, S. Yoon, and H. Schuetze (2025)NoLiMa: long-context evaluation beyond literal matching.InForty-second International Conference on Machine Learning,External Links:LinkCited by:§3.1.
  • NVIDIA, A. Basant, A. Khairnar, A. Paithankar, A. Khattar, A. Renduchintala, A. Malte, A. Bercovich, A. Hazare, A. Rico,et al.(2025a)NVIDIA nemotron nano 2: an accurate and efficient hybrid mamba-transformer reasoning model.External Links:2508.14444,LinkCited by:§3.1.
  • NVIDIA, A. Blakeman, A. Grattafiori, A. Basant, A. Gupta, A. Khattar, A. Renduchintala, A. Vavre, A. Shukla, A. Bercovich, A. Ficek,et al.(2025b)Nemotron 3 nano: open, efficient mixture-of-experts hybrid mamba-transformer model for agentic reasoning.External Links:2512.20848,LinkCited by:§1,§2.1.
  • OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat,et al.(2024)GPT-4 technical report.External Links:2303.08774,LinkCited by:§1.
  • B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski, X. Du, M. Grella, K. Gv, X. He, H. Hou, P. Kazienko, J. Kocon, J. Kong, B. Koptyra, H. Lau, J. Lin, K. S. I. Mantri, F. Mom, A. Saito, G. Song, X. Tang, J. Wind, S. Woźniak, Z. Zhang, Q. Zhou, J. Zhu, and R. Zhu (2023)RWKV: reinventing RNNs for the transformer era.InFindings of the Association for Computational Linguistics: EMNLP 2023,External Links:LinkCited by:§1.
  • B. Peng, J. Quesnelle, H. Fan, and E. Shippole (2024)YaRN: efficient context window extension of large language models.InThe Twelfth International Conference on Learning Representations,External Links:LinkCited by:§3.1.
  • C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun (2024)ChatDev: communicative agents for software development.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),External Links:LinkCited by:§1.
  • Z. Qin, W. Sun, D. Li, X. Shen, W. Sun, and Y. Zhong (2024)Various lengths, constant speed: efficient language modeling with lightning attention.InForty-first International Conference on Machine Learning,External Links:LinkCited by:§1,Figure 1,§2.1.
  • Z. Qiu, Z. Wang, B. Zheng, Z. Huang, K. Wen, S. Yang, R. Men, L. Yu, F. Huang, S. Huang, D. Liu, J. Zhou, and J. Lin (2025)Gated attention for large language models: non-linearity, sparsity, and attention-sink-free.InThe Thirty-ninth Annual Conference on Neural Information Processing Systems,External Links:LinkCited by:3rd item.
  • Qwen Team (2025)Qwen3-Next: Towards Ultimate Training & Inference Efficiency.External Links:LinkCited by:§2.1,§2.1.
  • Y. Shao, Y. Jiang, T. Kanell, P. Xu, O. Khattab, and M. Lam (2024)Assisting in writing Wikipedia-like articles from scratch with large language models.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers),External Links:LinkCited by:§1.
  • J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu (2023)RoFormer: enhanced transformer with rotary position embedding.External Links:2104.09864,LinkCited by:2nd item.
  • M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou,et al.(2022)Challenging big-bench tasks and whether chain-of-thought can solve them.arXiv preprint arXiv:2210.09261.Cited by:§3.1.
  • F. L. Team, I. Chaabane, P. Khanna, S. Mohmad, S. Frikha, S. Hu, A. Abubaker, R. Alami, M. Lubinets, M. E. A. Seddik, and H. Hacid (2026)Falcon-h1r: pushing the reasoning frontiers with a hybrid model for efficient test-time scaling.External Links:2601.02346,LinkCited by:§3.1.
  • A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017)Attention is all you need.InAdvances in Neural Information Processing Systems,I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.),Vol.30,pp..External Links:LinkCited by:§1.
  • J. Wang, D. Paliotta, A. May, A. M. Rush, and T. Dao (2024a)The mamba in the llama: distilling and accelerating hybrid models.InAdvances in Neural Information Processing Systems,A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.),Vol.37,pp. 62432–62457.External Links:Document,LinkCited by:§2.1.
  • Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen (2024b)MMLU-pro: a more robust and challenging multi-task language understanding benchmark.InAdvances in Neural Information Processing Systems,A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.),Vol.37,pp. 95266–95290.External Links:Document,LinkCited by:§3.1.
  • Y. Wang, Z. Fu, H. Zhao, C. Zhao, C. Zhou, X. Lin, H. Lyu, S. Xue, Y. Yi, Y. Wang, Z. Zheng, Y. Zhang, J. Zhou, C. Xiao, X. Han, Z. Liu, and M. Sun (2026)Data science and technology towards agi part i: tiered data management.External Links:2602.09003,LinkCited by:§2.2.
  • C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, and M. Sun (2024)InfLLM: training-free long-context extrapolation for llms with an efficient context memory.InAdvances in Neural Information Processing Systems,A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.),Vol.37,pp. 119638–119661.External Links:Document,LinkCited by:§1.
  • A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv,et al.(2025a)Qwen3 technical report.External Links:2505.09388,LinkCited by:§1,§3.1.
  • S. Yang, J. Kautz, and A. Hatamizadeh (2025b)Gated delta networks: improving mamba2 with delta rule.InThe Thirteenth International Conference on Learning Representations,External Links:LinkCited by:§1.
  • S. Yang, B. Wang, Y. Shen, R. Panda, and Y. Kim (2024a)Gated linear attention transformers with hardware-efficient training.InForty-first International Conference on Machine Learning,External Links:LinkCited by:§1.
  • S. Yang, B. Wang, Y. Zhang, Y. Shen, and Y. Kim (2024b)Parallelizing linear transformers with the delta rule over sequence length.InThe Thirty-eighth Annual Conference on Neural Information Processing Systems,External Links:LinkCited by:§1.
  • J. Yuan, H. Gao, D. Dai, J. Luo, L. Zhao, Z. Zhang, Z. Xie, Y. Wei, L. Wang, Z. Xiao, Y. Wang, C. Ruan, M. Zhang, W. Liang, and W. Zeng (2025)Native sparse attention: hardware-aligned and natively trainable sparse attention.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),External Links:LinkCited by:§1.
  • W. Zhao, Z. Zhou, Z. Su, C. Xiao, Y. Li, Y. Li, Y. Zhang, W. Zhao, Z. Li, Y. Huang, A. Sun, X. Han, and Z. Liu (2025)InfLLM-v2: dense-sparse switchable attention for seamless short-to-long adaptation.External Links:2509.24663,LinkCited by:§1,§1,Figure 1,§2.1,§5.
  • J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou (2023)Instruction-following evaluation for large language models.External Links:2311.07911,LinkCited by:§3.1.
  • Z. Zhou, C. Li, X. Chen, S. Wang, Y. Chao, Z. Li, H. Wang, Q. Shi, Z. Tan, X. Han, X. Shi, Z. Liu, and M. Sun (2025)LLM×\timesMapReduce: simplified long-sequence processing using large language models.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),External Links:LinkCited by:§1.
  • J. Zuo, M. Velikanov, I. Chahed, Y. Belkada, D. E. Rhayem, G. Kunsch, H. Hacid, H. Yous, B. Farhat, I. Khadraoui, M. Farooq, G. Campesan, R. Cojocaru, Y. Djilali, S. Hu, I. Chaabane, P. Khanna, M. E. A. Seddik, N. D. Huynh, P. L. Khac, L. AlQadi, B. Mokeddem, M. Chami, A. Abubaker, M. Lubinets, K. Piskorski, and S. Frikha (2025)Falcon-h1: a family of hybrid-head language models redefining efficiency and performance.External Links:2507.22448,LinkCited by:§2.1.

Similar Articles

Subquadratic AI introduces SubQ-1.1-Small, a new model using Smart Sparse Attention

Reddit r/singularity

Subquadratic AI introduces SubQ-1.1-Small, a model leveraging Smart Sparse Attention to achieve near-perfect long-context retrieval up to 12M tokens with up to 1,000x attention compute reduction. It balances long-context optimization with strong general reasoning, outperforming baselines on benchmarks like NIAH and RULER.

MiniMax Sparse Attention

Hugging Face Daily Papers

MiniMax Sparse Attention introduces a blockwise sparse attention mechanism that achieves significant speedups for ultra-long-context LLMs, reducing per-token attention compute by 28.4x at 1M context with wall-clock speedups of 14.2x for prefill and 7.6x for decoding on H800 GPUs. The method is accompanied by an open-source inference kernel and a publicly released multimodal model.

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

arXiv cs.CL

This paper presents a novel method for fine-tuning transformer language models with sparse attention to enable efficient long-context inference, often outperforming models trained with exact attention, and introduces an efficient implementation and a new open-source library.