机制的转变:位置编码的选择如何塑造上下文内检索

arXiv cs.CL 论文

摘要

本文对位置编码的选择(RoPE 与 SWA NoPE 混合方案)如何影响上下文内检索进行了机制层面的分析,在 22 个开源权重模型上展开研究,表明混合位置编码会将检索机制从基于位置的方式转向基于语义的方式,从而解释了它们在长上下文上的性能提升以及相应的权衡取舍。

arXiv:2609.38530v1 Announce Type: new Abstract: Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. We further show on a controlled pre-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: SWA NoPE improves over RoPE on multiple-target retrieval and QA, but degrades when distinguishing competing keys. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long-context retrieval.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:44

# How Positional Encoding Choice Shapes In-Context Retrieval
Source: [https://arxiv.org/html/2609.38530](https://arxiv.org/html/2609.38530)
###### Abstract

Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding\-window attention and NoPE with global attention \(SWA NoPE\)\. However, how these choices shape in\-context retrieval remains unclear\. To study this question, we take a mechanistic view, tracing how positional encoding \(PE\) choice shapes the internal mechanisms models use for in\-context retrieval\. Across 22 open\-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval\. We further show on a controlled pre\-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information\. Finally, we show that the reported long\-context gains of PE hybrids mask a retrieval trade\-off: SWA NoPE improves over RoPE on multiple\-target retrieval and QA, but degrades when distinguishing competing keys\. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long\-context retrieval\.

## 1Introduction

Language model architectures have increasingly moved away from standard RoPE\([Su et al\., 2021](https://arxiv.org/html/2609.38530#bib.bib11)\)toward hybrid designs that vary attention span and positional encoding across layers, architectural choices we jointly refer to as positional encoding \(PE\) choice\. Prior work has shown that PE hybrids can improve long\-context retrieval performance: SWA NoPE\([Puvvada et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib32);[Yang et al\., 2025b](https://arxiv.org/html/2609.38530#bib.bib8)\), which combines local sliding\-window attention \(SWA\) layers using RoPE with global attention layers without PE \(NoPE\), outperforms RoPE on several long\-context benchmarks\. Yet the underlying explanation for how PE choice shapes long\-context retrieval ability remains unclear\.

Recent interpretability work provides a way to investigate this question, identifying three core mechanisms that language models use for in\-context retrieval: positional, based on an entity’s location in context, and lexical and reflexive, based on an entity’s semantic content\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib2)\)\. Using this mechanistic analysis as an intermediate lens between model architecture and retrieval behavior, we ask whether PE choice changes how models retrieve information and whether these underlying mechanism differences can explain long\-context retrieval performance\.

We first find that PE hybrids have a consistently more semantic mechanism allocation\. Across 22 open\-weight models spanning eight families, RoPE models rely primarily on the positional mechanism, while PE hybrids rely primarily on semantic \(lexical and reflexive\) mechanisms\. To study why this shift occurs, we analyze a controlled set of pre\-training ablations of PE choice \(from RoPE to SWA RoPE to SWA NoPE\), finding that SWA alone does not produce the shift, while additionally removing RoPE from global layers does\. We find confining positional encoding to local layers degrades representations of positional information, matching the observed semantic shift\.

Finally, we show how our mechanistic analysis explains the impact of PE choice on long\-context retrieval behavior\. We find that SWA NoPE outperforms RoPE on multiple\-target needle\-in\-a\-haystack \(NIAH\) and QA tasks, where semantic mechanisms excel at finding relevant content, but degrades when trying to distinguish between semantically similar keys\. These results demonstrate that retrieval strategies produce different strengths and failure modes at long context, suggesting that SWA NoPE’s long\-context gains better reflect a shift in retrieval strategy than a uniform improvement\.

We summarize our contributions below:

- •Positional Encoding Choice→\\toMechanism Allocation\.We show that retrieval strategies are architecture\-dependent and replacing uniform RoPE with SWA NoPE systematically shifts mechanism allocation from positional toward semantic retrieval\.
- •Why does this shift occur?Linear probing suggests that confining positional encoding to local layers degrades representations of positional information, providing a potential explanation for the shift away from positional retrieval\.
- •Mechanism Allocation→\\toLong\-Context Retrieval\.We show that the mechanism shift predicts long\-context retrieval behavior\. The shift toward semantic mechanisms improves retrieval of relevant content but increases sensitivity to semantic confusability, which we validate experimentally across long\-context retrieval tasks\.

![Refer to caption](https://arxiv.org/html/2609.38530v1/Images/ShiftMech_Motivation.png)Figure 1:Overview\.We show that positional encoding \(PE\) choice shapes mechanism allocation: the SWA NoPE architecture shifts toward semantic mechanisms compared to RoPE\. We then show that this mechanism shift predicts long\-context retrieval performance: SWA NoPE improves on multiple\-target retrieval and QA tasks, but struggles to distinguish competing keys\.
## 2Related Work

Positional Encoding and Hybrid Architectures\.RoPE\([Su et al\., 2021](https://arxiv.org/html/2609.38530#bib.bib11)\)is widely used in modern LLMs, but can struggle to extrapolate beyond its training window\([Liu, 2026](https://arxiv.org/html/2609.38530#bib.bib18);[Du et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib17);[Du et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib1)\), motivating substantial work on how positional encoding can scale to longer contexts\. One line of work extends RoPE through interpolation and frequency scaling\([bloc97, 2023b](https://arxiv.org/html/2609.38530#bib.bib19);[bloc97, 2023a](https://arxiv.org/html/2609.38530#bib.bib20);[emozilla, 2023](https://arxiv.org/html/2609.38530#bib.bib21);[Peng et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib10)\), while another modifies or removes explicit positional encoding altogether\([Gelberg et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib7);[Barbero et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib5);[Khan et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib6);[Gopalakrishnan et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib30);[Movahedi et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib37)\)\. Furthermore, transformers without explicit positional encodings can still achieve competitive length generalization in some settings\([Haviv et al\., 2022](https://arxiv.org/html/2609.38530#bib.bib9);[Kazemnejad et al\., 2023](https://arxiv.org/html/2609.38530#bib.bib39)\)\.

Most relevant to our work are the hybrid SWA NoPE architectures proposed by[Yang et al\. \(2025b\)](https://arxiv.org/html/2609.38530#bib.bib8);[Puvvada et al\. \(2025\)](https://arxiv.org/html/2609.38530#bib.bib32)that interleave local RoPE and global NoPE, improving aggregate long\-context performance\. Prior work has sought to explain these gains through properties of the architecture\.[Qiao et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib31), for example, argue that long\-range retrieval is primarily carried by global attention layers, while local attention shapes how these retrieval capabilities emerge during training, motivating the use of NoPE in global layers\. In contrast, we find that SWA NoPE does not uniformly improve long\-context retrieval\. Instead, its mechanism shift from positional toward semantic retrieval predicts distinct strengths and failure modes across retrieval settings\.

Mechanistic Analysis of In\-Context Retrieval\.Early mechanistic work on entity binding identified a positional retrieval mechanism\([Feng and Steinhardt, 2024](https://arxiv.org/html/2609.38530#bib.bib12);[Prakash et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib13);[Prakash et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib14)\)\. More recently,[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)showed that models also use lexical and reflexive mechanisms, which we refer to collectively as semantic mechanisms\. Complementary work characterizes individual attention heads as positional or symbolic, finding theoretically and empirically that symbolic mechanisms generalize more robustly to longer sequences\([Urrutia et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib3);[Urrutia et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib4)\)\. While this work characterizes the mechanisms models use for retrieval, we study how architectural choices in positional encoding shape which of these mechanisms models learn to rely on\.

## 3Mechanistic Lens

In this section we explain the mechanistic lens we use throughout this work\. Prior work\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib2)\)found three core mechanisms that language models use to bind and retrieve entities in\-context\. We first describe the counterfactual approach used to disentangle these mechanisms, and then explain how we use it to measure mechanism allocation\.111Code is available at[https://github\.com/ericenouen/shiftmech](https://github.com/ericenouen/shiftmech)\.

Figure 2:Example of counterfactual activation patching\.The prompt pair is constructed so that each retrieval mechanism corresponds to a different answer after activation patching\. The positional mechanism follows location and predicts ‘B’, the lexical mechanism follows the ‘hat’ key and predicts ‘C’, and the reflexive mechanism follows the self\-referential pointer and predicts ‘A’\. If patching has no effect, the model retains the original prediction \(‘D’\)\.Counterfactual Patching\.To understand how a model uses these mechanisms,[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)use carefully paired original and counterfactual prompts and apply activation patching on the residual stream to identify which mechanism dominates a given prediction\. We illustrate their counterfactual patching framework using the Boxes task in Figure[2](https://arxiv.org/html/2609.38530#S3.F2), where objects are assigned to boxes and the model is queried for which box contains a particular object\.

The prompts are constructed so that the patched prediction uniquely identifies the dominant mechanism\. We define each mechanism below:

- •Positional Mechanism \(P\)\.The model stores the position of the queried item and later retrieves from that same position\. In[Figure2](https://arxiv.org/html/2609.38530#S3.F2), the counterfactual answer ‘A’ appears first, so the mechanism retrieves from the first position in the original prompt, returning ‘B’\.
- •Lexical Mechanism \(L\)\.The model stores the queried item and later retrieves using that item\. In[Figure2](https://arxiv.org/html/2609.38530#S3.F2), the counterfactual answer stores the semantic key ‘hat’, so the mechanism retrieves the box associated with ‘hat’ in the original prompt, returning ‘C’\.
- •Reflexive Mechanism \(R\)\.The model stores a self\-referential pointer to the answer and later retrieves that answer directly\. In[Figure2](https://arxiv.org/html/2609.38530#S3.F2), the self\-referential pointer directly stores ‘A’, so the mechanism returns ‘A’\.

In addition to these three mechanisms, patching can produce two other outcomes\. We classify a sample asno effectwhen the model retains the original prompt’s prediction \(‘D’ in[Figure2](https://arxiv.org/html/2609.38530#S3.F2)\), and asunknownwhen the prediction does not correspond to any of the three mechanisms or the original prediction\. With a larger number of entities \(we study 20 entities throughout this work\), there are many additional predictions that can therefore fall into the unknown category\.

Mechanism Allocation\.A model may use different retrieval mechanisms across different samples\. We therefore repeat the counterfactual patching procedure over many synthetic samples from a task and estimate the rate of each outcome as the fraction of samples classified as positional \(PP\), lexical \(LL\), reflexive \(RR\), unknown, or no effect\. We refer to these estimated rates as the model’s mechanism allocation\. We refer to the mechanism with the highest rate amongPP,LL, andRRas the model’s dominant mechanism\.

Given this mechanism allocation, our primary comparison is between positional and semantic retrieval\. We group the lexical and reflexive mechanisms as semantic mechanisms, since both retrieve based on entity information, and compare them to the positional mechanism which retrieves based on location\. We define the positional ratio asrpos=PL\+Rr\_\{\\mathrm\{pos\}\}=\\frac\{P\}\{L\+R\}\. Higherrposr\_\{\\mathrm\{pos\}\}indicates a more positional mechanism allocation, while lowerrposr\_\{\\mathrm\{pos\}\}indicates a more semantic mechanism allocation\. We use the mechanism allocation and positional ratio throughout this work to understand the internal mechanisms guiding in\-context retrieval\.

## 4Positional Encoding Choice Shapes Mechanism Allocation

We use this mechanistic framework to understand how positional encoding choice shapes the retrieval mechanisms a model learns to use\. We first show across 22 open\-weight instruction\-tuned models that PE hybrids exhibit a semantic shift compared to RoPE \([Section4\.1](https://arxiv.org/html/2609.38530#S4.SS1)\)\. We then study a controlled progression of pre\-training ablations on PE choice \(from RoPE to SWA RoPE to SWA NoPE\), finding that SWA alone does not produce the shift, while additionally removing RoPE from global layers does \([Section4\.2](https://arxiv.org/html/2609.38530#S4.SS2)\)\. Finally, we find that restricting RoPE to local layers reduces the quality of positional information represented by the model, providing a potential explanation for the semantic shift \([Section4\.3](https://arxiv.org/html/2609.38530#S4.SS3)\)\.

Table 1:Model families and descriptions\.We compare RoPE architectures that apply RoPE consistently throughout the model with PE hybrid architectures that vary the application of RoPE across layers\. We report the PE configuration, as well as window size and local:global ratio where applicable\.†Gemma\-4 2B/4B have a window size of 512 and Gemma\-4 2B uses a 4:1 local:global ratio\.### 4\.1Model Battery

In this section, we analyze mechanism allocation across eight model families to examine how PE choice shapes the learned mechanisms for in\-context retrieval\. We divide these models into the following two groups based on their PE configuration \([Table1](https://arxiv.org/html/2609.38530#S4.T1)\):

- •RoPE Architectures\.Models that use RoPE\([Su et al\., 2021](https://arxiv.org/html/2609.38530#bib.bib11)\)at every attention layer to encode positional information\. Prior analyses of in\-context retrieval mechanisms\([Prakash et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib14);[Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib2)\)studied only models of this type\.
- •PE Hybrid Architectures\.We adopt a broad definition of PE hybrid in this work, considering any model that varies how positional encodings are applied across layers\. These designs often combine this variation in PE with local and global attention layers\. For example, Gemma\-3 varies the RoPE base across local and global layers, Gemma\-4 varies the RoPE base and uses p\-RoPE\([Barbero et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib5)\)in global layers, and Command R7B and SmolLM3 both use NoPE layers\.[Table1](https://arxiv.org/html/2609.38530#S4.T1)summarizes both the PE configuration and local/global attention structure of each architecture\. Window size indicates the sliding window size\([Beltagy et al\., 2020](https://arxiv.org/html/2609.38530#bib.bib29)\)where applicable\.

These PE hybrids reflect a broader architectural trend away from applying RoPE uniformly throughout the model and toward differentiating PE across local and global layers\. Very recent architectures, including Llama 4\([Meta, 2025](https://arxiv.org/html/2609.38530#bib.bib16)\)and MAI\-Thinking\-1\([MAI, 2026](https://arxiv.org/html/2609.38530#bib.bib25)\), continue this trend\. Our model battery therefore lets us ask what impact this shift in PE design has on retrieval mechanisms\.

Experimental Details\.We evaluate 22 instruction\-tuned models across eight families: thirteen RoPE models \(Gemma\-2 \{2B, 9B, 27B\}, Llama\-3\.1 \{8B, 70B\}, Qwen\-2\.5 \{3B, 7B, 32B, 72B\}, Qwen\-3 \{1\.7B, 4B, 8B, 14B\}\), and nine PE hybrid models \(Gemma\-3 \{4B, 12B, 27B\}, Gemma\-4 \{2B, 4B, 12B, 31B\}, Command R7B, SmolLM3\-3B\)\. We evaluate each model on all 29 entity binding tasks from[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)\. For each model, we perform the mechanism analysis from[Section3](https://arxiv.org/html/2609.38530#S3)with 1000 samples per task at its selected retrieval layer \([Table6](https://arxiv.org/html/2609.38530#A1.T6), Appendix[B](https://arxiv.org/html/2609.38530#A2)\)\. We classify each prediction as positional, lexical, reflexive, or unknown, pool predictions across tasks, and computerpos=PL\+Rr\_\{\\mathrm\{pos\}\}=\\frac\{P\}\{L\+R\}from the pooled counts\. See Appendix[A\.2](https://arxiv.org/html/2609.38530#A1.SS2)for additional details\.

PE Hybrids Shift Mechanisms\.We reportrposr\_\{\\mathrm\{pos\}\}and mechanism fractions normalized overPP,LL, andRRin[Figure3](https://arxiv.org/html/2609.38530#S4.F3), with the full outcome rates in Appendix[C](https://arxiv.org/html/2609.38530#A3)\. We find that PE hybrids exhibit a systematic shift toward semantic retrieval\. While every RoPE model hasrpos≥0\.59r\_\{\\mathrm\{pos\}\}\\geq 0\.59\(Llama3\.1 8B\), every PE hybrid hasrpos≤0\.53r\_\{\\mathrm\{pos\}\}\\leq 0\.53\(Command R7B\)\. At the extremes, the positional mechanism accounts for65%65\\%of predictions in Qwen2\.5 72B \(rpos=1\.85r\_\{\\mathrm\{pos\}\}=1\.85\), compared to only22%22\\%in Gemma\-4 2B \(rpos=0\.28r\_\{\\mathrm\{pos\}\}=0\.28\), with semantic mechanisms correspondingly accounting for35%35\\%and78%78\\%\. Despite substantial variation in model family, parameter count, context length, and training data, we find a clear semantic shift going from RoPE to PE hybrid architectures\.

Prior mechanistic analyses of in\-context retrieval studied RoPE models\([Prakash et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib14);[Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib2)\), and our extension to PE hybrids reveals a different regime\. The positional mechanism accounts for the largest share amongPP,LL, andRRfor every RoPE model we test, while for every PE hybrid, either the lexical or reflexive mechanism has the largest share \([Table7](https://arxiv.org/html/2609.38530#A3.T7)\)\. In other words, the dominant mechanism consistently changes from positional to semantic across the two architecture classes\. Although PE hybrids use the same retrieval mechanisms identified by prior work, they exhibit a qualitatively different mechanism allocation\.

Our model battery reveals a clear difference in mechanism allocation between RoPE and PE hybrid models despite substantial variation in model family, scale, and architecture\. However, this variation also makes it difficult to determine which architectural differences drive the shift toward semantic retrieval\. We turn to a controlled set of pre\-training ablations to isolate the effect of PE choice\.

Figure 3:Mechanism allocation across the model battery pooled over 29 tasks and sorted byrposr\_\{\\mathrm\{pos\}\}\. \(Top\) Relative allocation to positional, lexical, and reflexive retrieval mechanisms\. \(Bottom\) Positional ratio for each model\. The architecture classes separate completely: every RoPE model has a higherrposr\_\{\\mathrm\{pos\}\}\(≥0\.59\\geq 0\.59\) than every PE hybrid \(≤0\.53\\leq 0\.53\)\.Table 2:Mechanism allocation of three model architectures: RoPE, SWA RoPE, and SWA NoPE at two training checkpoints \(16K and 32K training length\) on the Boxes task\. We report the percentage \(%\) of samples classified as positional, lexical, reflexive, and unknown\. We additionally report the ratio of positional to semantic \(lexical, reflexive\) mechanisms\. At both training lengths, introducing SWA RoPE alone does not produce the semantic shift, but SWA with NoPE global layers substantially shifts the mechanism allocation, reducingrposr\_\{\\mathrm\{pos\}\}by a factor of two\.
### 4\.2Ablating Positional Encoding Choice

To isolate whether positional encoding choice drives the shift in mechanism allocation, we utilize the controlled pre\-training checkpoints from[Qiao et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib31):

- •RoPE\.Global attention for all layers, with RoPE\([Su et al\., 2021](https://arxiv.org/html/2609.38530#bib.bib11)\)applied in every layer\.
- •SWA RoPE\.Interleaving global attention and local sliding\-window attention \(SWA\)\([Beltagy et al\., 2020](https://arxiv.org/html/2609.38530#bib.bib29)\), where tokens only attend to a fixed window size\. RoPE still applied in every layer\.
- •SWA NoPE\.The same alternating attention pattern as SWA RoPE, but RoPE is applied only in local layers, while global layers use NoPE\([Puvvada et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib32);[Yang et al\., 2025b](https://arxiv.org/html/2609.38530#bib.bib8)\)\.

Each architecture is first trained at 16K context length for 100B tokens, and then extended to 32K context length for an additional 5B tokens, totaling six checkpoints\. These models each have 665M parameters and share the same training hyperparameters\([Qiao et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib31), see\)\. Both SWA variants use a 1:1 ratio of local to global layers with a sliding window size of 128 tokens\. We utilize the framework from[Section3](https://arxiv.org/html/2609.38530#S3)to classify each sample as positional \(PP\), lexical \(LL\), reflexive \(RR\), or unknown, based on which token the model predicts under activation patching\. We measure at layer 15 using 10K samples per model\. Finally, since these are small base models, we evaluate on the Boxes task from[Figure2](https://arxiv.org/html/2609.38530#S3.F2)to ensure reliable task performance\. For further details see Appendix[A\.3](https://arxiv.org/html/2609.38530#A1.SS3)\.

We report our results in Table[2](https://arxiv.org/html/2609.38530#S4.T2)\. At both training lengths, SWA NoPE learns a substantially more semantic mechanism allocation than RoPE:rposr\_\{\\mathrm\{pos\}\}falls from0\.300\.30to0\.140\.14for the 16K checkpoints and0\.260\.26to0\.130\.13for the 32K checkpoints\. At both checkpoints, SWA NoPE reduces the positional mechanism and unknown rate, and increases the lexical mechanism\.

In contrast, SWA RoPE does not exhibit this semantic shift and even increasesrposr\_\{\\mathrm\{pos\}\}at the 16K training checkpoint from0\.300\.30to0\.400\.40, and nearly matches RoPE at 32K \(0\.260\.26to0\.270\.27\)\. The positional ratios are relatively low across all three architectures, consistent with the strong lexical bias of the Boxes task \(see[Figure8](https://arxiv.org/html/2609.38530#A3.F8)and[SectionC\.2](https://arxiv.org/html/2609.38530#A3.SS2)for more details\)\. These controlled experiments show that SWA alone does not produce the semantic shift, while additionally removing RoPE from global layers does\.

### 4\.3Confining positional encoding degrades Ordering ID creation

Having established that the SWA NoPE architecture drives the shift toward semantic retrieval, we next investigate why this shift occurs\. We first use our understanding of the three mechanisms guiding in\-context retrieval to identify where PE choice could impact mechanism allocation\. This analysis suggests that PE choice may affect how models form ordering IDs\. We therefore examine whether SWA NoPE weakens the positional representations used to form ordering IDs\.

Where can PE shape mechanisms?To begin, we look at the three in\-context retrieval mechanisms\. The lexical and reflexive mechanisms are both semantic by design, storing either a semantic key into the target entity or a self\-referential pointer to the target entity\. Neither mechanism explicitly relies on positional information for retrieval, making the positional mechanism the natural candidate through which PE choice could directly affect mechanism allocation\. We therefore dig deeper into how the positional mechanism operates\.

Recent work studying in\-context retrieval found that the positional mechanism occurs in two distinct stages via a lookback mechanism\([Prakash et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib14)\)\. During the first stage, the model associates each answer with an ordering ID \(OI\) that represents its ordinal position in the list\. For example, the answer token appearing third in the list stores a representation of its ordinal position \(“third”\)\.

During retrieval, the model identifies the position specified by the query and matches it to the corresponding OI to recover the answer\. Once created, the OI acts as a positional key for retrieval, analogous to the semantic key used by the lexical mechanism\. This decomposition suggests that PE choice may affect positional retrieval by altering how these positional keys are formed\. We therefore investigate whether OI creation differs across the same controlled pre\-training ablations used above\.

SWA NoPE Confines PE\.Our controlled ablation showed that SWA NoPE shifts toward semantic retrieval, while SWA RoPE does not\. Since ordering IDs encode position, their formation may depend on how positional information is provided by these architectures\. SWA RoPE applies RoPE in both local and global layers, while SWA NoPE confines RoPE to local layers and uses NoPE in global layers\. Next, we test whether confining PE to local layers makes OI creation more difficult\.

Table 3:Linear probe accuracy \(%\) on the residual stream at layer 15 across items in the list, predicting either the index in the list or the attached semantic key, across the six model checkpoints\. Standard deviation reported in parentheses over eight runs\. SWA NoPE has the worst ordering ID probe accuracy, while semantic key probe accuracy remains high\.Experimental Details\.To test whether ordering ID creation is impacted by PE choice, we train linear probes on the residual stream of the model\. We use the same Boxes task from Section[4](https://arxiv.org/html/2609.38530#S4), extracting the residual stream at each of the 20 answer tokens, and train a linear probe to predict the ordering ID \(the entity’s position in the list1−201\-20\)\. As a control, we train a linear probe to predict the semantic key \(the object paired with the box\), which checks that semantic information is still represented\. We train probes at layer 15 since we found in[Section4\.2](https://arxiv.org/html/2609.38530#S4.SS2)that this is where retrieval occurs, and extract the 1280\-dimensional activations at this layer\.

We fit anL2L\_\{2\}\-regularized logistic regression for each linear probe\. We train on10001000prompts and test on a disjoint set of10001000prompts, forming training and test sets of20×1000=20,00020\\times 1000=20\{,\}000samples each\. We report the mean and standard deviation across eight independently sampled runs\. We analyze the six checkpoints: RoPE, SWA RoPE, SWA NoPE, each trained at 16K context length and then extended to 32K\. Additional details about the experimental setup can be found in Appendix[A\.4](https://arxiv.org/html/2609.38530#A1.SS4)\.

SWA NoPE degrades OI creation\.Our results are shown in Table[3](https://arxiv.org/html/2609.38530#S4.T3)\. Ordering IDs are substantially less linearly decodable under SWA NoPE, with probe accuracy decreasing by12\.112\.1points at 16K and11\.011\.0points at 32K compared to RoPE\. SWA RoPE exhibits a much smaller decrease of3\.23\.2and2\.22\.2points, respectively\. In contrast, semantic key accuracy remains nearly unchanged across all six checkpoints\. As a result, the gap between semantic key and ordering ID accuracy grows from roughly44points under RoPE to1616points under SWA NoPE\. Overall, SWA NoPE selectively degrades the representation of OIs relative to semantic keys, consistent with its semantic shift\.

Mechanistic Explanation\.We hypothesize that confining PE to local layers impairs OI formation during training, shifting the model toward semantic retrieval\. Under SWA NoPE, layers with explicit positional information have only a local view of the sequence, while layers with global context must infer position without explicit positional encoding, making globally consistent OIs more difficult to learn\. Weaker OI formation may make positional retrieval less reliable, leading the model to rely more on semantic mechanisms whose underlying semantic representations remain intact\.

This explanation predicts that as the sliding window shrinks relative to context length, OI formation should become more difficult and the mechanism shift should strengthen\. Conversely, as the window approaches the full context length, the receptive field of the RoPE layers approaches the full context and this constraint should weaken\. We leave testing this prediction with controlled checkpoints ablating window size for future work\.

## 5Mechanism Allocation Shapes Long\-Context Retrieval

Having established in[Section4](https://arxiv.org/html/2609.38530#S4)that PE choice shifts the mechanisms models learn to use for retrieval, we next ask whether these differences predict long\-context retrieval behavior\. Prior work reports that SWA NoPE improves performance over RoPE\([Yang et al\., 2025b](https://arxiv.org/html/2609.38530#bib.bib8);[Puvvada et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib32);[Qiao et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib31)\)on long\-context benchmarks such as RULER\([Hsieh et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib15)\)\. Given that SWA NoPE also shifts retrieval toward semantic mechanisms, these gains raise a natural question: do semantic retrieval mechanisms simply scale better to long contexts? We investigate this question by examining how performance varies across the RULER subtasks using the same pre\-training checkpoints studied in[Section4](https://arxiv.org/html/2609.38530#S4)\.

We re\-evaluate the released checkpoints from[Qiao et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib31)on RULER\([Hsieh et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib15)\)\. We evaluate both the 16K and 32K checkpoints at their respective training context lengths, using 500 samples per task\. We evaluate on all thirteen tasks: eight needle\-in\-a\-haystack \(NIAH\) tasks \(three single key, three multi key, one multi value, and one multi query task\), two QA tasks, a variable tracking task, and two aggregation tasks \(common word extraction, frequent word extraction\)\. We group the subtasks by retrieval setting to examine which tasks drive the aggregate RULER improvement\. Additional experimental details can be found in Appendix[A\.5](https://arxiv.org/html/2609.38530#A1.SS5)\.

Table 4:RULER scores across the three model architectures: RoPE, SWA RoPE, and SWA NoPE\. The 16K checkpoints are trained for 100B tokens, and the 32K checkpoints extend them for a further 5B tokens, and each is evaluated at their training context length\. SWA NoPE substantially improves the overall RULER score, with gains concentrated in most NIAH and QA tasks\. However, it does not improve over RoPE on the first NIAH multi\-key task or consistently improve on variable tracking and aggregation\.Table 5:We test each checkpoint on our confusable key task and compare to the RULER multi\-value and multi\-query subtasks\. SWA NoPE struggles to select amongst competing keys, especially as they become more semantically similar, flipping the model ranking compared to multi\-item NIAH where all keys are returned\.SWA NoPE is not uniformly better\.We report the RULER results in Table[4](https://arxiv.org/html/2609.38530#S5.T4)\. Consistent with prior work\([Qiao et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib31)\), SWA NoPE substantially improves the overall RULER score over RoPE and SWA RoPE\. However, this aggregate gain is not uniform across subtasks\. SWA NoPE substantially improves most needle\-in\-a\-haystack \(NIAH\) and QA tasks, but does not improve over RoPE on the first NIAH multi\-key task or the variable tracking and aggregation tasks\. We report the full task breakdown in Appendix[D](https://arxiv.org/html/2609.38530#A4)\. These differences suggest that SWA NoPE changes which retrieval settings the model handles well rather than uniformly improving long\-context retrieval\.

Mechanism Shift Predicts a Retrieval Trade\-off\.The semantic shift identified in[Section4](https://arxiv.org/html/2609.38530#S4)suggests that SWA NoPE may excel when retrieving relevant content without discrimination among competing keys, but struggle when such discrimination is required\. This distinction is reflected in the RULER results: SWA NoPE performs particularly well on multi\-value and multi\-query NIAH, where the requested values should all be returned rather than selecting one target while rejecting competing keys\. To test the other side of this trade\-off directly, we construct a retrieval task in which each context contains four needles, one queried key and three distractor keys, with each key consisting of four hyphenated words \(e\.g\.,daughter\-others\-prop\-sudo\)\. Importantly, even when the keys share no words \(0/40/4\), retrieval requires distinguishing the queried key from the three distractors\. We increase semantic confusability by varying the number of words shared across keys from zero to three\.

We report the results in Table[5](https://arxiv.org/html/2609.38530#S5.T5)\. Requiring the model to distinguish among competing keys reverses the model ordering even when the keys share no words \(0/40/4\): RoPE outperforms SWA NoPE, whereas SWA NoPE substantially outperforms RoPE on multi\-value and multi\-query retrieval\. Increasing the semantic similarity among these competing keys further amplifies this difference, as predicted by the shift toward semantic retrieval\. At 16K, the gap between RoPE and SWA NoPE grows from 8\.8 points at0/40/4to 27\.2 points at2/42/4, with the same pattern at 32K, before narrowing at3/43/4as performance degrades across all architectures\. Together, these results show that SWA NoPE’s long\-context gains reflect a shift in retrieval strategy with distinct strengths and failure modes rather than a uniform improvement in retrieval ability\.

## 6Conclusion

In this work we analyzed how positional encoding \(PE\) choice shapes the in\-context retrieval mechanisms a language model learns during training\. Across controlled pre\-training checkpoints and a battery of open\-weight instruction\-tuned models, we find that PE choice substantially shifts whether models rely on positional or semantic retrieval\. Our controlled analysis further links this shift to degraded ordering ID representations when positional encoding is confined to local layers\. These differences in mechanism allocation translate into different long\-context retrieval profiles, rather than a uniform improvement or degradation in retrieval ability\. As recent open\-weight models increasingly adopt architectures that vary PE across layers, understanding these mechanistic consequences becomes increasingly important\.

Intentional Design\.Through our mechanistic analysis we showed where and why SWA NoPE improves or degrades long\-context retrieval, highlighting a broader role for interpretability in architectural design\([Orgad et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib38)\)\. By allowing model developers to better understand how architectural choices shape the internal mechanisms learned by the model, mechanistic analysis can help developers shape for their desired downstream behavior\. More broadly, our results suggest that mechanistic interpretability can serve not only to explain existing models, but also as a tool for intentional design\.

Avoiding Long\-Range RoPE\.Recent work has highlighted theoretical limitations of RoPE at long context, leading to either token or positional confusion\([Liu, 2026](https://arxiv.org/html/2609.38530#bib.bib18);[Du et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib1)\)\. Long\-context extensions based on RoPE scaling, such as YaRN\([Peng et al\., 2024](https://arxiv.org/html/2609.38530#bib.bib10)\), additionally modify the positional encoding to extrapolate beyond the context lengths seen during training\. Instead, the SWA NoPE architecture restricts RoPE to fixed\-size sliding\-window layers, and so the maximum positional distance remains fixed as the global context length increases\. At the same time, our results show that confining positional encoding in this way substantially changes the model’s learned retrieval mechanisms\. We leave exploring the benefits and drawbacks of these choices to future work\.

## Limitations and Future Work

Scaling\.One limitation of this work is the cost associated with training models from scratch\. While we are able to test small\-scale checkpoints ablating the positional encoding choice\([Qiao et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib31)\), we are unable to train larger models, exhaustively ablate combinations of design choices, or analyze how the mechanisms develop during training\. We leave these directions for future work\.

Data\.Mechanism allocation depends on which retrieval mechanisms are reinforced by the training objective, and therefore on the training data distribution\. It would be interesting to study how different data distributions shape the in\-context retrieval mechanisms learned by the model, or whether particular mechanisms could be targeted directly through the choice of long\-context training data\.

Linear hybrids\.Many recent open\-weight models have also adopted linear hybrids instead of PE hybrids\([Kimi et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib34);[Blakeman et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib35);[Qwen, 2026](https://arxiv.org/html/2609.38530#bib.bib36)\)\. We provide a preliminary analysis of these models in Appendix[E](https://arxiv.org/html/2609.38530#A5), but leave a more comprehensive study for future work\. In particular, it would be interesting to understand whether alternative efficient attention strategies shift the allocation of the same retrieval mechanisms studied here, or lead models to learn fundamentally different retrieval mechanisms\.

## AI Statement

In this work, we used generative AI tools for implementing methods in code and feedback on research methodology\. Generating data sets, developing theoretical models, formulating/proving mathematical claims are not applicable to this work\. Additionally, we used generative AI tools for help with plot and figure creation, as well as polishing the writing of the paper\. We have reviewed all AI\-assisted work, and stand by the research contributions made in this work\. We take full responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

## Acknowledgements

We thank Linxi Zhao and Sofian Zalouk for insightful discussions on the ideas and manuscript\. This research was supported by a gift to the LinkedIn–Cornell Bowers Strategic Partnership, and ARO grant W911NF\-25\-1\-0254\. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect those of the sponsors\.

## References

- Bakouchet al\.\(2025\)E\. Bakouch, L\. Ben Allal, A\. Lozhkov, N\. Tazi, L\. Tunstall, C\. M\. Patiño, E\. Beeching, A\. Roucher, A\. J\. Reedi, Q\. Gallouédec, K\. Rasul, N\. Habib, C\. Fourrier, H\. Kydlicek, G\. Penedo, H\. Larcher, M\. Morlon, V\. Srivastav, J\. Lochner, X\. Nguyen, C\. Raffel, L\. von Werra, and T\. WolfSmolLM3: smol, multilingual, long\-context reasoner\.Note:[https://huggingface\.co/blog/smollm3](https://huggingface.co/blog/smollm3)Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.10.1.1)\.
- Barberoet al\.\(2024\)F\. Barbero, A\. Vitvitskyi, C\. Perivolaropoulos, R\. Pascanu, and P\. VeličkovićRound and round we go\! what makes rotary positional encodings useful?\.arXiv preprint arXiv:2410\.06205\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1),[2nd item](https://arxiv.org/html/2609.38530#S4.I1.i2.p1.1)\.
- Beltagyet al\.\(2020\)I\. Beltagy, M\. E\. Peters, and A\. CohanLongformer: the long\-document transformer\.arXiv preprint arXiv:2004\.05150\.Cited by:[2nd item](https://arxiv.org/html/2609.38530#S4.I1.i2.p1.1),[2nd item](https://arxiv.org/html/2609.38530#S4.I2.i2.p1.1)\.
- Blakemanet al\.\(2025\)A\. Blakeman, A\. Grattafiori, A\. Basant, A\. Gupta, A\. Khattar, A\. Renduchintala, A\. Vavre, A\. Shukla, A\. Bercovich, A\. Ficek,et al\.Nvidia nemotron 3: efficient and open intelligence\.arXiv preprint arXiv:2512\.20856\.Cited by:[Appendix E](https://arxiv.org/html/2609.38530#A5.p1.1),[Limitations and Future Work](https://arxiv.org/html/2609.38530#Sx1.p3.1)\.
- bloc97 \(2023a\)bloc97Add NTK\-Aware interpolation ”by parts” correction\.External Links:[Link](https://github.com/jquesnelle/scaled-rope/pull/1)Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- bloc97 \(2023b\)bloc97NTK\-Aware Scaled RoPE allows LLaMA models to have extended \(8k\+\) context size without any fine\-tuning and minimal perplexity degradation\.\.External Links:[Link](https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/)Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Duet al\.\(2026\)Y\. Du, P\. Harris, M\. Tian, E\. A\. Huerta, S\. Ronanki, S\. Rongali, A\. Galstyan, and H\. PengRoPE distinguishes neither positions nor tokens in long contexts, provably\.arXiv preprint arXiv:2605\.15514\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1),[§6](https://arxiv.org/html/2609.38530#S6.p3.1)\.
- Duet al\.\(2025\)Y\. Du, M\. Tian, S\. Ronanki, S\. Rongali, S\. Bodapati, A\. Galstyan, A\. Wells, R\. Schwartz, E\. A\. Huerta, and H\. PengContext length alone hurts llm performance despite perfect retrieval\.arXiv preprint arXiv:2510\.05381\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- emozilla \(2023\)emozillaDynamically Scaled RoPE further increases performance of long context LLaMA with zero fine\-tuning\.External Links:[Link](https://www.reddit.com/r/LocalLLaMA/comments/14mrgpr/dynamically_scaled_rope_further_increases/)Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Feng and Steinhardt \(2024\)J\. Feng and J\. SteinhardtHow do language models bind entities in context?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 36391–36413\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p3.1)\.
- Gelberget al\.\(2025\)Y\. Gelberg, K\. Eguchi, T\. Akiba, and E\. CetinExtending the context of pretrained llms by dropping their positional embeddings\.arXiv preprint arXiv:2512\.12167\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Gemma \(2026\)GemmaGemma 4 technical report\.arXiv preprint arXiv:2607\.02770\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.9.1.1)\.
- Gopalakrishnanet al\.\(2025\)A\. Gopalakrishnan, R\. Csordás, J\. Schmidhuber, and M\. C\. MozerDecoupling the” what” and” where” with polar coordinate positional embeddings\.arXiv preprint arXiv:2509\.10534\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.4.1.1)\.
- Gur\-Ariehet al\.\(2026\)Y\. Gur\-Arieh, M\. Geva, and A\. GeigerMixing mechanisms: how language models retrieve bound entities in\-context\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 27757–27786\.Cited by:[§A\.1](https://arxiv.org/html/2609.38530#A1.SS1.p1.1),[§A\.1](https://arxiv.org/html/2609.38530#A1.SS1.p2.1),[§A\.3](https://arxiv.org/html/2609.38530#A1.SS3.p3.1),[Appendix B](https://arxiv.org/html/2609.38530#A2.p1.1),[§C\.2](https://arxiv.org/html/2609.38530#A3.SS2.p1.1),[§1](https://arxiv.org/html/2609.38530#S1.p2.1),[§2](https://arxiv.org/html/2609.38530#S2.p3.1),[§3](https://arxiv.org/html/2609.38530#S3.p1.1),[§3](https://arxiv.org/html/2609.38530#S3.p2.1),[1st item](https://arxiv.org/html/2609.38530#S4.I1.i1.p1.1),[§4\.1](https://arxiv.org/html/2609.38530#S4.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.38530#S4.SS1.p5.1)\.
- Havivet al\.\(2022\)A\. Haviv, O\. Ram, O\. Press, P\. Izsak, and O\. LevyTransformer language models without positional encodings still learn positional information\.InFindings of the Association for Computational Linguistics: EMNLP 2022,pp\. 1382–1390\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Hsiehet al\.\(2024\)C\. Hsieh, S\. Sun, S\. Kriman, S\. Acharya, D\. Rekesh, F\. Jia, Y\. Zhang, and B\. GinsburgRULER: what’s the real context size of your long\-context language models?\.arXiv preprint arXiv:2404\.06654\.Cited by:[§A\.5](https://arxiv.org/html/2609.38530#A1.SS5.p1.1),[§5](https://arxiv.org/html/2609.38530#S5.p1.1),[§5](https://arxiv.org/html/2609.38530#S5.p2.1)\.
- Kamathet al\.\(2025\)A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard,et al\.Gemma 3 technical report\.arXiv preprint arXiv:2503\.197864\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.8.1.1)\.
- Kazemnejadet al\.\(2023\)A\. Kazemnejad, I\. Padhi, K\. Natesan Ramamurthy, P\. Das, and S\. ReddyThe impact of positional encoding on length generalization in transformers\.Advances in Neural Information Processing Systems36,pp\. 24892–24928\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Khanet al\.\(2026\)M\. A\. Khan, K\. P\. Gummadi, M\. Gupta, and A\. RavichanderFractional rotation, full potential? investigating performance and convergence of partial rope\.arXiv preprint arXiv:2603\.11611\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Kimiet al\.\(2025\)Kimi, Y\. Zhang, Z\. Lin, X\. Yao, J\. Hu, F\. Meng, C\. Liu, X\. Men, S\. Yang, Z\. Li,et al\.Kimi linear: an expressive, efficient attention architecture\.arXiv preprint arXiv:2510\.26692\.Cited by:[Appendix E](https://arxiv.org/html/2609.38530#A5.p1.1),[Limitations and Future Work](https://arxiv.org/html/2609.38530#Sx1.p3.1)\.
- Liu \(2026\)F\. LiuRotary positional embeddings as phase modulation: theoretical bounds on the rope base for long\-context transformers\.arXiv preprint arXiv:2602\.10959\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1),[§6](https://arxiv.org/html/2609.38530#S6.p3.1)\.
- MAI \(2026\)MAIMAI\-thinking\-1: building a hill\-climbing machine\.Technical reportMicrosoft AI\.External Links:[Link](https://microsoft.ai/pdf/mai-thinking-1.pdf)Cited by:[§4\.1](https://arxiv.org/html/2609.38530#S4.SS1.p2.1)\.
- Meta \(2025\)MetaThe llama 4 herd\.External Links:[Link](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)Cited by:[§4\.1](https://arxiv.org/html/2609.38530#S4.SS1.p2.1)\.
- Movahediet al\.\(2026\)S\. Movahedi, T\. Carstensen, A\. Afzal, F\. Hutter, A\. Orvieto, and V\. CevherSelective rotary position embedding\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 10219–10247\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1)\.
- Orgadet al\.\(2026\)H\. Orgad, F\. Barez, T\. Haklay, I\. Lee, M\. Mosbach, A\. Reusch, N\. Saphra, B\. Wallace, S\. Wiegreffe, E\. Wong,et al\.Interpretability can be actionable\.arXiv preprint arXiv:2605\.11161\.Cited by:[§6](https://arxiv.org/html/2609.38530#S6.p2.1)\.
- Penget al\.\(2024\)B\. Peng, J\. Quesnelle, H\. Fan, and E\. ShippoleYarn: efficient context window extension of large language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 31932–31951\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p1.1),[§6](https://arxiv.org/html/2609.38530#S6.p3.1)\.
- Prakashet al\.\(2024\)N\. Prakash, T\. Shaham, T\. Haklay, Y\. Belinkov, and D\. BauFine\-tuning enhances existing mechanisms: a case study on entity tracking\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 8057–8082\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p3.1)\.
- Prakashet al\.\(2025\)N\. Prakash, N\. Shapira, A\. S\. Sharma, C\. Riedl, Y\. Belinkov, T\. R\. Shaham, D\. Bau, and A\. GeigerLanguage models use lookbacks to track beliefs\.arXiv preprint arXiv:2505\.14685\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p3.1),[1st item](https://arxiv.org/html/2609.38530#S4.I1.i1.p1.1),[§4\.1](https://arxiv.org/html/2609.38530#S4.SS1.p5.1),[§4\.3](https://arxiv.org/html/2609.38530#S4.SS3.p3.1)\.
- Puvvadaet al\.\(2025\)K\. C\. Puvvada, F\. Ladhak, S\. A\. Serrano, C\. Hsieh, S\. Acharya, S\. Majumdar, F\. Jia, S\. Kriman, S\. Sun, D\. Rekesh,et al\.Swan\-gpt: an efficient and scalable approach for long\-context language modeling\.arXiv preprint arXiv:2504\.08719\.Cited by:[§1](https://arxiv.org/html/2609.38530#S1.p1.1),[§2](https://arxiv.org/html/2609.38530#S2.p2.1),[3rd item](https://arxiv.org/html/2609.38530#S4.I2.i3.p1.1),[§5](https://arxiv.org/html/2609.38530#S5.p1.1)\.
- Qiaoet al\.\(2026\)Z\. Qiao, Y\. Xu, C\. Xiao, Z\. Su, Z\. Zhou, Y\. Chen, X\. Xu, X\. Han, and Z\. LiuRethinking the role of efficient attention in hybrid architectures\.arXiv preprint arXiv:2606\.15378\.Cited by:[§A\.3](https://arxiv.org/html/2609.38530#A1.SS3.p1.1),[§A\.3](https://arxiv.org/html/2609.38530#A1.SS3.p2.1),[§2](https://arxiv.org/html/2609.38530#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.38530#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.38530#S4.SS2.p2.1),[§5](https://arxiv.org/html/2609.38530#S5.p1.1),[§5](https://arxiv.org/html/2609.38530#S5.p2.1),[§5](https://arxiv.org/html/2609.38530#S5.p3.1),[Limitations and Future Work](https://arxiv.org/html/2609.38530#Sx1.p1.1)\.
- Qwen \(2026\)QwenQwen3\.5\-omni technical report\.arXiv preprint arXiv:2604\.15804\.Cited by:[Appendix E](https://arxiv.org/html/2609.38530#A5.p1.1),[Limitations and Future Work](https://arxiv.org/html/2609.38530#Sx1.p3.1)\.
- Riviereet al\.\(2024\)M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé,et al\.Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.3.1.1)\.
- Suet al\.\(2021\)J\. Su, Y\. Lu, S\. Pan, B\. Wen, and Y\. LiuRoFormer: enhanced transformer with rotary position embedding\.CoRRabs/2104\.09864\.External Links:[Link](https://arxiv.org/abs/2104.09864),2104\.09864Cited by:[§1](https://arxiv.org/html/2609.38530#S1.p1.1),[§2](https://arxiv.org/html/2609.38530#S2.p1.1),[1st item](https://arxiv.org/html/2609.38530#S4.I1.i1.p1.1),[1st item](https://arxiv.org/html/2609.38530#S4.I2.i1.p1.1)\.
- Urrutiaet al\.\(2026\)F\. Urrutia, J\. J\. Alegría, C\. S\. Macias, J\. Salas, C\. B\. Calderon, and C\. RojasPositional versus symbolic attention heads: learning dynamics, rope geometry, and length generalization\.arXiv preprint arXiv:2605\.31558\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p3.1)\.
- Urrutiaet al\.\(2025\)F\. Urrutia, J\. Salas, A\. Kozachinskiy, C\. B\. Calderon, H\. Pasten, and C\. RojasDecoupling positional and symbolic attention behavior in transformers\.arXiv preprint arXiv:2511\.11579\.Cited by:[§2](https://arxiv.org/html/2609.38530#S2.p3.1)\.
- Yanget al\.\(2025a\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.6.1.1)\.
- Yanget al\.\(2024\)A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.5.1.1)\.
- Yanget al\.\(2025b\)B\. Yang, B\. Venkitesh, D\. G\. Talupuru, H\. Lin, D\. Cairuz, P\. Blunsom, and A\. LocatelliRope to nope and back again: a new hybrid attention strategy\.Advances in Neural Information Processing Systems38,pp\. 64133–64157\.Cited by:[§1](https://arxiv.org/html/2609.38530#S1.p1.1),[§2](https://arxiv.org/html/2609.38530#S2.p2.1),[3rd item](https://arxiv.org/html/2609.38530#S4.I2.i3.p1.1),[Table 1](https://arxiv.org/html/2609.38530#S4.T1.6.1.11.1.1),[§5](https://arxiv.org/html/2609.38530#S5.p1.1)\.

## Appendix AExperimental Details

In this section, we describe the experimental details for our main experiments\. We first describe the mechanistic lens used throughout this work \([SectionA\.1](https://arxiv.org/html/2609.38530#A1.SS1)\), then the model battery \([SectionA\.2](https://arxiv.org/html/2609.38530#A1.SS2)\), the pre\-training ablation \([SectionA\.3](https://arxiv.org/html/2609.38530#A1.SS3)\), linear probing \([SectionA\.4](https://arxiv.org/html/2609.38530#A1.SS4)\), and long\-context behavior \([SectionA\.5](https://arxiv.org/html/2609.38530#A1.SS5)\)\. All experiments are run in bf16 on a single NVIDIA RTX 6000 Ada \(48GB\), except models larger than 25B parameters, which exceed its memory and are run on an NVIDIA B200\.

### A\.1Mechanistic Lens

In this section we describe general details about the framework we use for our main mechanistic analysis\. We utilize the datasets and framework from[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)that disentangles three mechanisms that language models use to perform in\-context retrieval: positional, lexical, and reflexive\. In total they propose 29 datasets: ten synthetic groupings \(Filling Liquids, People and Objects, Programming Dictionary, Music, Biology Experiment, Chemistry Experiment, Transportation, Sports Events, Space Observations, Boxes\)\. While the Boxes task only has two entities \(Object, Box\), each other task queries up to three different entities \(first, second, third\), totaling 29 task/query pairs\.

The given LM is run on two prompts: the counterfactual prompt and original prompt\. Then, at the given layerℓ\\ell, the activations from the LM run on the counterfactual prompt are patched to the LM run on the original prompt\. Through this patching procedure, each of the three mechanisms predicts a different token in the original prompt \(see[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)for additional details on how the mechanisms are disentangled\)\. Across 1000 generated pairs, the mechanism allocation can be measured by counting the number of samples that follow each of the mechanisms\.

We prompt at multiple tokens for all our results, and patch at a different layer for each model\. The full breakdown for our battery is shown in[Table6](https://arxiv.org/html/2609.38530#A1.T6)highlighting our choices for each model\.

Table 6:Patched layer and token positions per model\. Positions are matched by token role: the ‘?’ and ‘:’ of the query, plus the final two chat\-scaffold positions\. Offsets differ with template length\.
### A\.2Model Battery

In this section we describe details about the larger model battery\. We evaluate 22 different post\-trained models\. RoPE: Qwen2\.5 \(3B, 7B, 32B, 72B\), Qwen3 \(1\.7B, 4B, 8B, 14B\), Llama\-3\.1 \(8B, 70B\), and Gemma\-2 \(2B, 9B, 27B\)\. PE hybrid: Gemma\-3 \(4B, 12B, 27B\), Gemma\-4 \(2B, 4B, 12B, 31B\), Command R7B, and SmolLM3 3B\. We also average across the full task battery using 20 entities, and evaluate with 1000 samples for each task\. Two important decisions for each model are the choice of positions to patch, as well as the layer selection\. Since these are all instruction\-tuned models with their own chat templates, we match the patched tokens by role \(final token, newline, ‘:’, ‘?’\) and perform layer selection in[AppendixB](https://arxiv.org/html/2609.38530#A2)to select these for each model\. Our results are shown in[Table6](https://arxiv.org/html/2609.38530#A1.T6)\. Gemma\-2, Qwen2\.5, and Llama3\.1 follow prior work, while we compute the selections for the new models ourselves\.

### A\.3Pre\-training Ablation

In this section we describe details about the model checkpoints from the pre\-training ablation\([Qiao et al\., 2026](https://arxiv.org/html/2609.38530#bib.bib31)\), and our mechanistic analysis on these models\.

We evaluate these models on Boxes because this task is relatively simple and they are 665M parameter base models that struggle on most tasks\. Additionally, we use a cloze format \(appending ‘The object is in Box’ to the end of the prompt\) to ensure these models can retain high accuracy on this task\. We run 10K samples per model to get a robust measure of the mechanism allocation, use only the last token position for patching, and patch at layer 15\. The 16K checkpoints were trained for 100B tokens at 16K context length, and the 32K checkpoints were trained for an additional 5B tokens at 32K context length\.[Qiao et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib31)trains three different architectures: RoPE, SWA RoPE, and SWA NoPE\. Both SWA architectures use a sliding window size of 128 tokens, with a ratio of 1:1 local:global layers\. Further, SWA NoPE removes RoPE from the global layers but retains RoPE in the sliding window layers\. These models also use an attention sink\.

To validate our choice of layer 15, we present layer scans for these models\. Our results are shown in[Figure4](https://arxiv.org/html/2609.38530#A1.F4)\. We show that layer 15 is the last layer before the answer is retrieved for each model in our pre\-training checkpoints\. The messiness 4 setting differentiates the reflexive self\-referential pointer from the counterfactual prompt’s answer by ensuring the object does not exist in the original prompt, so it can only be reflexive in messiness 4 if the answer has already been retrieved\. See[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)for additional details on the layer\-selection procedure\.

Figure 4:For every checkpoint in our pre\-training ablation, layer 15 is the last layer before the answer is retrieved\.
### A\.4Linear Probing

In this section we provide additional details for the linear probing experiments\.

We probe the three pre\-training checkpoints \(RoPE, SWA RoPE, SWA NoPE\), each with 30 layers anddmodel=1280d\_\{\\text\{model\}\}=1280\. We evaluate both the 16K checkpoints trained for 100B tokens and the corresponding 32K checkpoints extended for an additional 5B tokens\. For all models, we probe at layer 15, the same layer used for our mechanism analysis\. We extract the residual stream entering block 15 using a forward pre\-hook, with no KV cache\.

We use the Boxes task withN=20N=20entities \(“the \{Object\} is in Box \{Box\}”\)\. For each entity, we extract the residual at the Box\-letter token\. For the*ordering ID*probe, the label is the entity’s position in the list \(2020classes, chance5%5\\%\)\. For the*Semantic Key*control, we predict the object bound to each box \(8282classes, chance≈1\.2%\\approx 1\.2\\%\)\. Objects are randomized across positions between prompts, making the semantic\-key label independent of position\.

We fit multinomial logistic regression with L2 regularization \(C=1C=1\) on standardized features using scikit\-learn \(StandardScalerfollowed byLogisticRegression\)\. We use all 1280 dimensions without dimensionality reduction\. Each repeat uses 1000 training prompts \(20,000 examples\) and a separate set of 1000 held\-out test prompts\. We run eight repeats, redrawing both sets and refitting the probe each time, and report mean held\-out accuracy±\\pmSD across repeats\.

### A\.5Long\-Context Behavior

In this section we provide additional details for our RULER evaluation and confusable\-key task\. We evaluate our pre\-training checkpoints on RULER following[Hsieh et al\. \(2024\)](https://arxiv.org/html/2609.38530#bib.bib15), using 500 samples per task\. We evaluate the 16K checkpoints at 16K context length and the 32K checkpoints at 32K context length\.

Our confusable\-key task builds on NIAH\-mk1, where we replace each key with four hyphenated words and vary the number of words shared across competing keys\. Each prompt contains four key\-value pairs inserted into a haystack of Paul Graham essays\.

For ak/4k/4example, we samplekkwords shared across all four keys and independently sample the remaining4−k4\-kwords for each key, with no overlap between the remaining words\. We then independently shuffle the four words within each key before joining them with hyphens\. Thus,0/40/4contains four disjoint keys, while at3/43/4all four keys share three words and differ by only one\. Values are uniformly sampled 7\-digit numbers, and key\-value pairs use the standard RULER needle template\. We uniformly sample one of the four keys to query, giving a chance level of25%25\\%\.

We score a prediction as correct when the target value appears in the model generation\. We generate 500 samples at each similarity level \(0/40/4,1/41/4,2/42/4,3/43/4\) for each checkpoint at both 16K and 32K context lengths\. The0/40/4condition uses the same haystack, needle and query structure as NIAH\-mk1, but replaces its single\-word keys with disjoint four\-word keys\. It therefore controls for the change in key format while varying only key similarity across the0/40/4–3/43/4ladder\.

## Appendix BLayer Selection Validation

In this section we describe the protocol we use to select the patching layer for our main results\. We mostly follow the setup of[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2)and select the last layer at which patching does not transfer an already\-retrieved answer\. To test this, we use their messiness\-4 counterfactual, in which the counterfactual answer is a new item that does not appear in the original prompt\. If the model outputs this answer after patching, the answer can only have come from the patched activation, so retrieval has already happened by that layer\. Scans patch the same positions as the main experiments and use the boxes task with query category 1\. We show layer scans for three models to illustrate the process\.

- •Gemma\-3 4B \([Figure5](https://arxiv.org/html/2609.38530#A2.F5)\) has a clear shift from retrieval in layer 23 to almost entirely reflexive in layer 24 and the messiness\-4 condition returns entirely the counterfactual answer, so we selectℓ=23\\ell=23\.
- •Gemma\-4 4B \([Figure6](https://arxiv.org/html/2609.38530#A2.F6)\) is also clean: after layer 29 the messiness\-4 scan shows a substantial share of samples returning the counterfactual answer, so we selectℓ=29\\ell=29\.
- •Gemma\-4 2B \([Figure7](https://arxiv.org/html/2609.38530#A2.F7)\) is less clean: after layer 24 there is a small increase in the counterfactual answer\. We therefore selectℓ=24\\ell=24, erring on the side of preventing an already\-retrieved counterfactual answer from entering the mechanism allocation analysis\.

Figure 5:Layer scan for Gemma\-3 4B with the messiness\-0 \(top\) and messiness\-4 \(bottom\) counterfactuals\. Bars show the share of each mechanism when patching at each layer\.Figure 6:Layer scan for Gemma\-4 4B with the messiness\-0 \(top\) and messiness\-4 \(bottom\) counterfactuals\. Bars show the share of each mechanism when patching at each layer\.Our selected retrieval layer for all PE hybrids also ended up being the global layer when there was a mix of local and global layers\. Gemma\-4 2B, which has a 4:1 ratio of local to global layers \(global layers 4, 9, 14, …\), illustrates this: patching first has an effect at global layer 19, the lexical share rises sharply at global layer 24, and the messiness\-4 answer is fully retrieved from layer 30, immediately after global layer 29\.

Figure 7:Layer scan for Gemma\-4 2B with the messiness\-0 \(top\) and messiness\-4 \(bottom\) counterfactuals\. Bars show the share of each mechanism when patching at each layer\.
## Appendix CMechanism Allocation Extended Results

In this section, we report the full mechanism distribution in[Table7](https://arxiv.org/html/2609.38530#A3.T7): positional, lexical, reflexive, unknown, and no effect, along with the resultingrposr\_\{\\mathrm\{pos\}\}\. While the no\-effect rate is generally low, it is substantially higher for some models, particularly Gemma\-4 12B, even when patching multiple tokens\.

[Figure8](https://arxiv.org/html/2609.38530#A3.F8)showsrposr\_\{\\mathrm\{pos\}\}separately for each task\. The ordering between RoPE and PE hybrids does not hold for every individual task, but emerges in aggregate\. Across tasks, however, there is a strong shift inrposr\_\{\\mathrm\{pos\}\}from the most positional models to the most semantic models\.

### C\.1Full Distribution Table

Table 7:Mechanism distribution per model, as a percentage of all samples \(29 task–query pairs, 29,000 samples each\), sorted byrp​o​s=P/\(L\+R\)r\_\{pos\}=P/\(L\+R\)\.Bold/underline: highest/second\-highest among Positional, Lexical, and Reflexive\.Figure 8:Full per\-task breakdown ofrposr\_\{\\mathrm\{pos\}\}across the model battery \(29 tasks\)\. Separation does not hold for every individual task, but holds when averaged across tasks\.
### C\.2Mechanism Allocation by Target Entity Position

[Figures9](https://arxiv.org/html/2609.38530#A3.F9),[10](https://arxiv.org/html/2609.38530#A3.F10)and[11](https://arxiv.org/html/2609.38530#A3.F11)show mechanism allocation as a function of target entity positionte​n​t​i​t​yt\_\{entity\}for each model\. As shown by[Gur\-Arieh et al\. \(2026\)](https://arxiv.org/html/2609.38530#bib.bib2),te​n​t​i​t​yt\_\{entity\}controls the tradeoff between lexical and reflexive mechanisms: when the target precedes the query entity in the group, models shift toward reflexive retrieval\. This breakdown confirms that our results hold consistently across all query types\.

Figure 9:Mechanism allocation by entity index forte​n​t​i​t​y=1t\_\{entity\}=1, averaged over tasks\. Reflexive retrieval dominates lexical retrieval because the target entity always precedes the query entity\.Figure 10:Mechanism allocation by entity index forte​n​t​i​t​y=2t\_\{entity\}=2, averaged over tasks\.Figure 11:Mechanism allocation by entity index forte​n​t​i​t​y=3t\_\{entity\}=3, averaged over tasks\. Lexical retrieval dominates reflexive retrieval because the target entity follows the query entity\.

## Appendix DRULER Extended Results

In[Table8](https://arxiv.org/html/2609.38530#A4.T8), we report the per\-task RULER results for the pre\-training ablation\. The per\-task results show that SWA NoPE’s gains are concentrated rather than uniform\. Its largest improvements occur on several NIAH tasks, particularly S3, MK3, MV, and MQ, as well as both QA tasks\. However, SWA NoPE harms or does not improve performance on MK1 and the variable\-tracking and aggregation tasks, motivating examining how the shift in retrieval mechanism interacts with the structure of the retrieval task\.

Table 8:Per\-task RULER scores at 16K and 32K context\.N=500N=500per task\. Best result for each task and context length is bolded\.In[Table5](https://arxiv.org/html/2609.38530#S5.T5)we showed that SWA RoPE degrades on the confusable\-key task even though its mechanism allocation stays similar to RoPE\. To study why this is, we categorize the incorrect predictions of each model at 32K context length\. As shown in[Table9](https://arxiv.org/html/2609.38530#A4.T9), 82\-94% of SWA NoPE’s errors at 32K retrieve a competing key’s value\. In contrast, roughly half of SWA RoPE’s errors return no needle value and instead consist primarily of degenerate repetitions\. Thus, SWA RoPE’s reduced accuracy at 32K reflects a distinct failure mode rather than the competing\-key retrieval errors that characterize SWA NoPE\.

Table 9:Error types on the confusable\-key task at 32K context length\. Conditioned on an incorrect answer, we report the percentage of errors that return a value bound to one of the three competing keys \(Competing key\), versus no needle value \(Neither\)\.
## Appendix ELinear Hybrids

Many recent open\-weight models increasingly use linear attention hybrids\([Kimi et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib34);[Blakeman et al\., 2025](https://arxiv.org/html/2609.38530#bib.bib35);[Qwen, 2026](https://arxiv.org/html/2609.38530#bib.bib36)\), often motivated by long\-context gains that are primarily computational\. A natural question is how these architectures affect in\-context retrieval mechanisms\.

We briefly examine three linear hybrids\. Qwen3\.5 4B interleaves Gated DeltaNet layers with full p\-RoPE attention at a 3:1 ratio, RecurrentGemma 2B interleaves gated linear recurrence with local p\-RoPE attention \(window 2048\) at a 2:1 ratio, and Granite\-4\.0\-H 1B interleaves Mamba2 layers with full NoPE attention at a 9:1 ratio\. We run the full task battery on each using the same protocol as the main experiments\.

Results\.None of the three exhibits the mechanism allocation characteristic of the PE hybrids in our main battery\. Qwen3\.5 falls within the RoPE range \(rpos=0\.72r\_\{\\mathrm\{pos\}\}=0\.72\)\. RecurrentGemma \(0\.590\.59\) and Granite \(0\.600\.60\) instead lie near the boundary: roughly matching the lowest RoPE model, Llama\-3\.1 8B \(0\.590\.59\), while remaining above the highest PE hybrid, Command R7B \(0\.530\.53\)\.[Figure12](https://arxiv.org/html/2609.38530#A5.F12)breaks down mechanism allocation by entity index\. In RecurrentGemma and Granite, reflexive retrieval accounts for under 10% of legible predictions, the lowest of any model we test\. Nearly half of the predictions for these models fall outside the three mechanisms\. Exploring whether this reflects a retrieval mechanism outside our taxonomy, model scale, or anything else, we leave for future work\.

Figure 12:Mechanism allocation across linear hybrids compared to RoPE/PE hybrids\.

相似文章

RoVE:面向相对位置依赖值路径的旋转值嵌入注意力机制

arXiv cs.LG

本文提出RoVE,一种无需参数的旋转位置嵌入改进方法,通过同时旋转值与键使值路径具备位置敏感性,将RoPE注意力转化为注意力卷积。在GPT-2模型上的实验表明,该机制在少样本上下文学习、分布外困惑度及长上下文检索方面持续提升性能。