RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways

arXiv cs.LG Papers

Summary

This paper proposes RoVE, a parameter-free modification to Rotary Position Embeddings that makes value pathways position-sensitive by rotating values simultaneously with keys, transforming RoPE attention into attentive convolution. Experiments on GPT-2 models show consistent gains in few-shot in-context learning, out-of-distribution perplexity, and long-context retrieval.

arXiv:2606.11275v1 Announce Type: new Abstract: Rotary Position Embeddings (RoPE) make attention scores position-relative but leave the value pathway position-blind: the message sent by a value token is the same regardless of its distance from the query. We propose RoVE, a parameter-free modification that makes values position-sensitive by rotating them simultaneously with keys, and show that it turns RoPE attention into attentive convolution. This new perspective unifies several independent formulations of the same operation across computer vision, robotics, and modern LLM architectures. Trained 124M and 354M GPT-2 models show consistent empirical gains over RoPE on few-shot in-context learning, out-of-distribution perplexity, and long-context retrieval, with the clearest improvements on tasks that require long-range aggregation.
Original Article
View Cached Full Text

Cached at: 06/11/26, 01:46 PM

# Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
Source: [https://arxiv.org/html/2606.11275](https://arxiv.org/html/2606.11275)
\\theorembodyfont\\theoremheaderfont\\theorempostheader

:\\theoremsep \\jmlrvolume334\\jmlryear2026\\jmlrworkshopTopology, Algebra, and Geometry in Data Science

\\NameAlejandro García\-Castellanos1\\Emaila\.garciacastellanos@uva\.nl \\NameMaurice Weiler2\\Emailm\.weiler\.ml@gmail\.com \\NameErik J\. Bekkers1\\Emaile\.j\.bekkers@uva\.nl \\addrAMLabMIT CSAIL2

###### Abstract

Rotary Position Embeddings \(RoPE\) make attention scores position\-relative but leave the value pathway position\-blind: the message sent by a value token is the same regardless of its distance from the query\. We proposeRoVE, a parameter\-free modification that makes values position\-sensitive by rotating them simultaneously with keys, and show that it turnsRoPEattention into*attentive convolution*\. This new perspective unifies several independent formulations of the same operation across computer vision, robotics, and modern LLM architectures\. Trained 124M and 354M GPT\-2 models show consistent empirical gains overRoPEon few\-shot in\-context learning, out\-of\-distribution perplexity, and long\-context retrieval, with the clearest improvements on tasks that require long\-range aggregation\.

###### keywords:

Rotary Position Embeddings, Attentive Convolution, Large Language Models

## 1Introduction

Rotary Position Embeddings \(RoPE\)\(Su et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib35)\)make attention scores*shift equivariant*: rotating queriesqiq\_\{i\}and keyskjk\_\{j\}by position dependent rotation matricesRiR\_\{i\}andRjR\_\{j\}produces the scoreqi⊤​Rj−i​kj/dq\_\{i\}^\{\\\!\\top\}R\_\{j\-i\}k\_\{j\}/\\\!\\sqrt\{d\}, which depends on positions only through the relative offsetδ=j−i\\delta=j\-i\. The value pathway, however, is untouched\. In the language of transformer circuits \(Appendix[E](https://arxiv.org/html/2606.11275#A5)andElhage et al\. \([2021](https://arxiv.org/html/2606.11275#bib.bib14)\)\),RoPEbiases the QK circuit toward relative positions while leaving the OV circuit position\-blind: unlike convolution kernels, the channel mapWVW\_\{\\\!V\}applied toxjx\_\{j\}carries no information about wherexjx\_\{j\}lies relative to the query\.

A natural completion rotates the value at positionjjbyRjR\_\{j\}before aggregation and inverts by moving the output viaRi−1R\_\{i\}^\{\-1\}to positionii\. As we show, this replaces the constant value mapWVW\_\{\\\!V\}with the offset\-dependent convolution kernelψδ=Rδ​WV\\psi\_\{\\delta\}=R\_\{\\delta\}W\_\{\\\!V\}whereδ=j−i\\delta=j\-i, endowing the OV circuit with the same relative\-position sensitivity already present in the QK circuit\.

Similar “RoPE\-on\-values” constructions have been independently discovered across several communities\.Miyato et al\. \([2024](https://arxiv.org/html/2606.11275#bib.bib27)\)introduced it for multi\-view novel\-view synthesis, encoding geometric relationships between camera frames, and subsequent work has further extended the same mechanism to computer vision\(Wu et al\.,[2026](https://arxiv.org/html/2606.11275#bib.bib38); Li et al\.,[2026](https://arxiv.org/html/2606.11275#bib.bib23)\)and robotics\(Klee et al\.,[2026](https://arxiv.org/html/2606.11275#bib.bib21)\)\. In the language\-modeling setting, DeepSeek\-V4\(DeepSeek\-AI,[2026](https://arxiv.org/html/2606.11275#bib.bib12)\)arrives at the same operation from a different direction: its compressed shared\-KV architecture causes positional information to leak from keys into values, and an inverse output rotation needs to be applied as a corrective measure to maintain the relative position\. Each work motivates the mechanism on application\-specific grounds, yet none provides a structural account of what it does to the attention operator\.

We isolate this positional embedding mechanism on the value stream, which we callRoVE, and provide a theoretical analysis of this modification\. We then evaluate it as a standalone module in standard \(non\-shared\-KV\) language models – a regime in which it has not previously been studied\. Our contributions are:

- •Structural characterisation:We show thatRoVEturnsRoPEattention into an*attentive convolution*\(Romero et al\.,[2020](https://arxiv.org/html/2606.11275#bib.bib34); Fuchs et al\.,[2020](https://arxiv.org/html/2606.11275#bib.bib17)\): the position\-blind mapWVW\_\{\\\!V\}is replaced by the offset\-dependent kernelψδ=Rδ​WV\\psi\_\{\\delta\}=R\_\{\\delta\}W\_\{\\\!V\}, and the matrix mixing operator\(Hwang et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib20)\)acquires block\-Toeplitz rather than Kronecker structure\.
- •Empirical validation:We train 124M and 354M parameter GPT\-2 language models and evaluate on in\-context learning \(ICL\) benchmarks and long\-context tasks\.RoVEconsistently improves ICL accuracy, long\-context robustness, and retrieval performance over standardRoPEattention\.

\\floatconts

fig:mixer

Figure 1:Matrix\-mixer view ofRoPEandRoVE\.RoPEfactorises into\(a\)\(a\)position\-sensitive attention weights and\(b\)\(b\)a*constant*shared value projectionWVW\_\{V\}across all offsets\(c\)\(c\)\.\(d\)\(d\)RoVEreplacesWVW\_\{V\}with the offset\-indexed familyψδ=Rδ​WV\\psi\_\{\\delta\}=R\_\{\\delta\}W\_\{V\},\(e\)\(e\)producing a block\-Toeplitz mixer whose diagonals rotate systematically with relative offset, the signature of an attentive convolution\.\\subfigure

\[Att\. weights\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x1.png)\\subfigure\[RoPEkernel\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x2.png)\\subfigure\[RoPEmixer\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x3.png)\\subfigure\[RoVEkernel\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x4.png)\\subfigure\[RoVEmixer\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x5.png)

## 2Background

#### Matrix mixers:

LetX∈ℝn×dX\\in\\mathbb\{R\}^\{n\\times d\}be the input feature tensor, wherennis the sequence length andddis the hidden dimension\. We denote byvec⁡\(X\)∈ℝn​d\\operatorname\{vec\}\(X\)\\in\\mathbb\{R\}^\{nd\}the vectorisation ofXX\. Any layer of the formvec⁡\(Y\)=ℳ​\(X\)​vec⁡\(X\)\\operatorname\{vec\}\(Y\)=\\mathcal\{M\}\(X\)\\operatorname\{vec\}\(X\), whereℳ​\(X\)\\mathcal\{M\}\(X\)is ann​d×n​dnd\\times ndmatrix partitioned intod×dd\\times dblocks, belongs to the*matrix mixer*family\(Hwang et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib20)\)\.

#### Attentive convolutions:

A classicaldd\-dimensional convolution mixes positions through a fixed offset\-dependent kernelψδ∈ℝd×d\\psi\_\{\\delta\}\\in\\mathbb\{R\}^\{d\\times d\}\.*Attentive convolutions*\(Romero et al\.,[2020](https://arxiv.org/html/2606.11275#bib.bib34); Fuchs et al\.,[2020](https://arxiv.org/html/2606.11275#bib.bib17)\)replace the fixed weights with content\-dependent scalarsA​\(X\)i​j∈ℝA\(X\)\_\{ij\}\\in\\mathbb\{R\}while keeping the kernel:

yi=∑j∈𝒩​\(i\)A​\(X\)i​j​ψj−i​xj,\\displaystyle y\_\{i\}\\ =\\ \\sum\\nolimits\_\{j\\in\\mathcal\{N\}\(i\)\}A\(X\)\_\{ij\}\\;\\psi\_\{j\-i\}\\;x\_\{j\},\(1\)whereA​\(X\)i​jA\(X\)\_\{ij\}gates the contribution of tokenjjto positionii, and𝒩​\(i\)\\mathcal\{N\}\(i\)is a neighbourhood ofii\. The corresponding mixer is block\-Toeplitz: the\(i,j\)\(i,j\)\-thd×dd\\times dblock ofℳ​\(X\)\\mathcal\{M\}\(X\)equalsA​\(X\)i​j​ψj−iA\(X\)\_\{ij\}\\psi\_\{j\-i\}, so each block diagonal carries the same kernel, scaled by a scalar gate\.

#### StandardRoPEattention:

Let𝒩​\(i\)⊆\{1,…,n\}\\mathcal\{N\}\(i\)\\subseteq\\\{1,\\dots,n\\\}denote the set of positions visible to queryii\. Typical choices include causal masking \(𝒩​\(i\)=\{j≤i\}\\mathcal\{N\}\(i\)=\\\{j\\leq i\\\}\), full attention \(𝒩​\(i\)≡\{1,…,n\}\\mathcal\{N\}\(i\)\\equiv\\\{1,\.\.\.,n\\\}\), and sparse patterns\(Child et al\.,[2019](https://arxiv.org/html/2606.11275#bib.bib8)\)\.RoPE\(Su et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib35)\)defines block\-diagonal rotation matricesRt∈SO​\(d\)R\_\{t\}\\in\\mathrm\{SO\}\(d\), where themm\-th2×22\\times 2block rotates by anglet​ωmt\\omega\_\{m\}for geometrically spaced frequenciesωm=θ0−2​m/d\\omega\_\{m\}=\\theta\_\{0\}^\{\-2m/d\}, and computes

yi=∑j∈𝒩​\(i\)A​\(X\)i​j​WV​xj,A​\(X\)i​j=softmaxj∈𝒩​\(i\)⁡\(1d​\(WQ​xi\)⊤​Rj−i​\(WK​xj\)\)\.\\displaystyle y\_\{i\}=\\sum\\nolimits\_\{j\\in\\mathcal\{N\}\(i\)\}A\(X\)\_\{ij\}\\,W\_\{\\\!V\}x\_\{j\},\\qquad A\(X\)\_\{ij\}=\\operatorname\{softmax\}\_\{j\\in\\mathcal\{N\}\(i\)\}\\\!\\left\(\\tfrac\{1\}\{\\sqrt\{d\}\}\\,\(W\_\{\\\!Q\}x\_\{i\}\)^\{\\\!\\top\}R\_\{j\-i\}\(W\_\{\\\!K\}x\_\{j\}\)\\right\)\.\(2\)The mixer factorizes asℳRoPE​\(X\)=A​\(X\)⊗WV\\mathcal\{M\}^\{\\textsc\{RoPE\}\}\(X\)=A\(X\)\\otimes W\_\{\\\!V\}across tensor axes, decoupling token routing from channel projections\(Elhage et al\.,[2021](https://arxiv.org/html/2606.11275#bib.bib14)\)\.

#### YaRN:

RoPE\-based models are typically trained at short context lengths but deployed at longer ones, creating an out\-of\-distribution problem: low\-frequency components encounter rotation angles unseen during training\(Tian et al\.,[2026](https://arxiv.org/html/2606.11275#bib.bib36)\)\. Uniform frequency stretching is a natural fix\(Chen et al\.,[2023](https://arxiv.org/html/2606.11275#bib.bib5)\), but destroys local position dependence in high\-frequency components\.YaRN\(Peng et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib30)\)resolves this via frequency\-dependent interpolation, i\.e\., preserving high frequencies while smoothly scaling the lower ones\. This rescaling can be applied post\-hoc without any additional training\. For further related work see Appendix[B](https://arxiv.org/html/2606.11275#A2)\.

## 3Method

RoPEmakes attention scores position\-relative but leaves the value pathway invariant: two tokens assigned equal attention weight contribute identically to the output regardless of their offset from the query\.RoVEextends this by additionally rotating each value into the query’s reference frame before aggregation\.

###### Definition 3\.1\(RoVE\)\.

Let𝒩​\(i\)⊆\{1,…,n\}\\mathcal\{N\}\(i\)\\subseteq\\\{1,\\dots,n\\\}be any neighbourhood function and letA​\(X\)i​jA\(X\)\_\{ij\}be theRoPEattention weights from \([2](https://arxiv.org/html/2606.11275#S2.E2)\)\.RoVEcomputes

y~i=Ri−1​∑j∈𝒩​\(i\)A​\(X\)i​j​Rj​WV​xj=∑j∈𝒩​\(i\)A​\(X\)i​j​Rj−i​WV⏟ψj−i​xj\.\\tilde\{y\}\_\{i\}\\,\\ =\\,\\ R\_\{i\}^\{\-1\}\\sum\_\{\\mathclap\{j\\in\\mathcal\{N\}\(i\)\}\}A\(X\)\_\{ij\}\\,R\_\{j\}\\,W\_\{\\\!V\}x\_\{j\}\\,\\ =\\,\\ \\sum\_\{\\mathclap\{j\\in\\mathcal\{N\}\(i\)\}\}A\(X\)\_\{ij\}\\;\\underbrace\{R\_\{j\-i\}W\_\{\\\!V\}\}\_\{\\scriptstyle\\psi\_\{j\-i\}\}\\,x\_\{j\}\.\(3\)

#### Convolution lens:

Equation \([3](https://arxiv.org/html/2606.11275#S3.E3)\) is an instance of the attentive convolution \([1](https://arxiv.org/html/2606.11275#S2.E1)\) with neighbourhood𝒩\\mathcal\{N\},RoPEattention weights, and position\-dependent kernelψδ=Rδ​WV\\psi\_\{\\delta\}=R\_\{\\delta\}W\_\{\\\!V\}\(standardRoPEis recovered by the degenerate choiceψδ≡WV\\psi\_\{\\delta\}\\equiv W\_\{\\\!V\}\)\. Crucially, the value pathway is no longer a single shared map but the tied family\{Rδ​WV\}δ\\\{R\_\{\\delta\}W\_\{\\\!V\}\\\}\_\{\\delta\}, so relative position modulates token*transformation*\(OV circuit\) as well as token*selection*\(QK circuit\)\.

#### Matrix mixer lens:

In mixer form, the\(i,j\)\(i,j\)\-thd×dd\\times dblock ofℳRoVE​\(X\)\\mathcal\{M\}^\{\\mbox\{\{R\\kern 0\.1pto\\kern\-1\.0ptVE\}\}\}\(X\)equalsA​\(X\)i​j​Rj−i​WVA\(X\)\_\{ij\}R\_\{j\-i\}W\_\{\\\!V\}, replacing the factorisedRoPEblocksA​\(X\)i​j​WVA\(X\)\_\{ij\}W\_\{\\\!V\}\. Consequently, the mixer inherits the*block\-Toeplitz structure*of attentive convolutions, up to the content\-dependent modulation byAi​jA\_\{ij\}: blocks along the same relative\-offset diagonal share the same rotated value kernel\. FigureLABEL:fig:mixervisualizes this transition from a constant value kernel to an offset\-indexed family of value kernels\. See Appendix[E](https://arxiv.org/html/2606.11275#A5)for further analysis ofRoVE’s circuit\.

#### Local\-frame lens:

Equation \([3](https://arxiv.org/html/2606.11275#S3.E3)\) admits a clean frame\-change interpretation: first, each valueWV​xjW\_\{\\\!V\}x\_\{j\}is transformed from its local frame into a shared global frame byRjR\_\{j\}, then contributions are aggregated in the global frame, and finally the result is transformed into the query’s local frame byRi−1R\_\{i\}^\{\-1\}\. The effective kernelRj−i​WVR\_\{j\-i\}W\_\{\\\!V\}is the frame\-change operator composed with the learned channel map\. This frame\-change perspective is the primary motivation inMiyato et al\. \([2024](https://arxiv.org/html/2606.11275#bib.bib27)\), where features from different views are rotated into a common reference frame before aggregation\. Moreover, in the message\-passing view, the transformed valuesRj​WV​xjR\_\{j\}W\_\{\\\!V\}x\_\{j\}can be seen as*tensorial messages*\(Lippmann et al\.,[2025](https://arxiv.org/html/2606.11275#bib.bib25)\)\.

#### Efficiency:

RoVEembeds values analogously to the query/key embeddings in RoPE\. These operations have linear complexity𝒪​\(n​d\)\\mathcal\{O\}\(nd\)since rotations act independently onnnindividual tokens andd2\\frac\{d\}\{2\}channel pairs – their computational cost is therefore negligible compared to𝒪​\(n​d2\)\\mathcal\{O\}\(nd^\{2\}\)linear and𝒪​\(n2​d\)\\mathcal\{O\}\(n^\{2\}d\)attention layers\. Like RoPE,RoVEis compatible with FlashAttention kernels\(Dao et al\.,[2022](https://arxiv.org/html/2606.11275#bib.bib11)\)since it acts on values and attention outputs before and after the kernel call\. Furthermore, it introduces no additional learned parameters\.

## 4Experiments

#### Setup:

We evaluateRoVEas a drop\-in replacement for the value pathway inRoPEattention, training GPT\-2\-style transformers\(Brown et al\.,[2020](https://arxiv.org/html/2606.11275#bib.bib4)\)at small \(∼124\{\\sim\}124M\) and medium \(∼354\{\\sim\}354M\) scale on FineWebEdu\-10B\(Lozhkov et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib26)\)with a 1024\-token context\. We evaluate on:DCLM\-Core, few\-shot ICL accuracy within the training context\(Li et al\.,[2024a](https://arxiv.org/html/2606.11275#bib.bib22)\);OOD perplexityat up to16×16\\timescontext length, with and withoutYaRN\(Peng et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib30)\); andRULER, long\-context retrieval at 4k/8k tokens scored by NLL\(Hsieh et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib19)\)\. Full details and more empirical results are shown in Appendix[A](https://arxiv.org/html/2606.11275#A1)\.

Table 1:Core ICL accuracy and perplexity for the 354M parameter model\. Core is measured within the 1024\-token training context; PPL is measured from 512 to 16384 tokens\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation, leaving Core unchanged\.Green highlightmarks column bests
#### RoVEimproves in the trained regime:

Tables[1](https://arxiv.org/html/2606.11275#S4.T1)and[3](https://arxiv.org/html/2606.11275#A1.T3)show thatRoVEimproves both Core ICL accuracy and in\-context perplexity at both scales\. ThusRoVEprovides a useful inductive bias even within the training distribution, not only at extrapolated lengths\.

#### RoVEandYaRNare complementary:

RoVEsubstantially reduces OOD perplexity, andYaRNimproves both baselines while leaving their gap largely intact\. Relative positioning on the value stream is therefore not hindered by inference\-time frequency interpolation, i\.e\., the two methods address complementary aspects of long\-context generalization\.

Table 2:RULER long\-context retrieval \(NLL; lower is better\) for the 354M parameter model\. Tasks: Common Word Extraction \(CWE\), multi\-key Needle\-in\-a\-Haystack \(NIAH\), Question Answering \(QA\), Variable Tracking \(VT\)\. Avg is the unweighted mean over all eight task/length cells\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation\.Green highlightmarks column bests\.
#### RoVEimproves long\-context retrieval:

Tables[2](https://arxiv.org/html/2606.11275#S4.T2)and[4](https://arxiv.org/html/2606.11275#A1.T4)show that perplexity gains translate to synthetic long\-context retrieval tasks, with the strongest improvements on tasks requiring information to be maintained and recombined across the context\. These results align with the attentive convolution view\. StandardRoPEmakes attention weights position\-aware but leaves value transformation position\-blind;RoVEcloses this gap by aligning selected information according to relative position before aggregation\.

## 5Conclusion

We proposeRoVE, an extension ofRoPEattention that turns it into an attentive convolution by replacing the position\-blind mapWVW\_\{\\\!V\}with the offset\-indexed kernelψδ=Rδ​WV\\psi\_\{\\delta\}=R\_\{\\delta\}W\_\{\\\!V\}, converting the attention mixer from Kronecker to block\-Toeplitz structure and endowing the OV circuit with the same relative\-position sensitivity already present in the QK circuit\.

This parameter\-free modification consistently improves uponRoPEacross model scales, in\-context evaluation, out\-of\-distribution perplexity, and RULER retrieval\. Gains are largest withYaRN, confirming the two methods are complementary\. SinceRoVEleaves attention logits unchanged and is compatible with efficient attention kernels, relative\-position\-aware values stand as a robust structural bias for LLMs\. For further discussion see Appendix[D](https://arxiv.org/html/2606.11275#A4)\.

\\acks

Alejandro García Castellanos is funded by the Hybrid Intelligence Center, a 10\-year programme funded through the research programme Gravitation which is \(partly\) financed by the Dutch Research Council \(NWO\)\. This publication is part of the project SIGN with file number VI\.Vidi\.233\.220 of the research programme Vidi which is \(partly\) financed by the Dutch Research Council \(NWO\) under the grant[https://doi\.org/10\.61686/PKQGZ71565](https://doi.org/10.61686/PKQGZ71565)\.

## References

- Arora et al\. \(2024\)Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Y Zou, Atri Rudra, and Christopher Ré\.Zoology: Measuring and improving recall in efficient language models\.In*International conference on learning representations*, volume 2024, pages 15664–15730, 2024\.
- Barbero et al\. \(2024\)Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Veličković\.Round and round we go\! what makes rotary positional encodings useful?*arXiv preprint arXiv:2410\.06205*, 2024\.
- bloc97 \(2023\)bloc97\.Ntk\-aware scaled rope allows llama models to have longer context windows\.[https://www\.reddit\.com/r/LocalLLaMA/comments/14lz7j5/ntkaware\_scaled\_rope\_allows\_llama\_models\_to\_have/](https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/), 2023\.
- Brown et al\. \(2020\)Tom B\. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert\-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M\. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei\.Language models are few\-shot learners\.*Advances in neural information processing systems*, 33:1877–1901, 2020\.
- Chen et al\. \(2023\)Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian\.Extending context window of large language models via positional interpolation\.*arXiv preprint arXiv:2306\.15595*, 2023\.
- Chen et al\. \(2025\)Yuhan Chen, Ang Lv, Jian Luan, Bin Wang, and Wei Liu\.Hope: A novel positional encoding without long\-term decay for enhanced context awareness and extrapolation\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 23044–23056, 2025\.
- Chi et al\. \(2022\)Ta\-Chung Chi, Ting\-Han Fan, Peter J Ramadge, and Alexander Rudnicky\.Kerple: Kernelized relative positional embedding for length extrapolation\.*Advances in Neural Information Processing Systems*, 35:8386–8399, 2022\.
- Child et al\. \(2019\)Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever\.Generating long sequences with sparse transformers\.*arXiv preprint arXiv:1904\.10509*, 2019\.
- Cordonnier et al\. \(2019\)Jean\-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi\.On the relationship between self\-attention and convolutional layers\.*arXiv preprint arXiv:1911\.03584*, 2019\.
- Dai et al\. \(2019\)Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov\.Transformer\-xl: Attentive language models beyond a fixed\-length context\.In*Proceedings of the 57th annual meeting of the association for computational linguistics*, pages 2978–2988, 2019\.
- Dao et al\. \(2022\)Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré\.Flashattention: Fast and memory\-efficient exact attention with io\-awareness\.*Advances in neural information processing systems*, 35:16344–16359, 2022\.
- DeepSeek\-AI \(2026\)DeepSeek\-AI\.Deepseek\-v4: Towards highly efficient million\-token context intelligence, 2026\.
- Ding et al\. \(2024\)Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang\.Longrope: Extending llm context window beyond 2 million tokens\.*arXiv preprint arXiv:2402\.13753*, 2024\.
- Elhage et al\. \(2021\)Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield\-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah\.A mathematical framework for transformer circuits\.*Transformer Circuits Thread*, 2021\.https://transformer\-circuits\.pub/2021/framework/index\.html\.
- Fu et al\. \(2023a\)Daniel Y\. Fu, Tri Dao, Khaled K\. Saab, Armin W\. Thomas, Atri Rudra, and Christopher Ré\.Hungry Hungry Hippos: Towards language modeling with state space models\.In*International Conference on Learning Representations*, 2023a\.
- Fu et al\. \(2023b\)Daniel Y\. Fu, Elliot L\. Epstein, Eric Nguyen, Armin W\. Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré\.Simple hardware\-efficient long convolutions for sequence modeling\.*arXiv preprint arXiv:2302\.06646*, 2023b\.
- Fuchs et al\. \(2020\)Fabian Fuchs, Daniel Worrall, Volker Fischer, and Max Welling\.SE\(3\)\-Transformers: 3D Roto\-Translation Equivariant Attention Networks\.In*Advances in Neural Information Processing Systems*, volume 33, pages 1970–1981\. Curran Associates, Inc\., 2020\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2020/hash/15231a7ce4ba789d13b722cc5c955834\-Abstract\.html](https://proceedings.neurips.cc/paper_files/paper/2020/hash/15231a7ce4ba789d13b722cc5c955834-Abstract.html)\.
- Gopalakrishnan et al\. \(2025\)Anand Gopalakrishnan, Robert Csordás, Jürgen Schmidhuber, and Michael C Mozer\.Decoupling the” what” and” where” with polar coordinate positional embeddings\.*arXiv preprint arXiv:2509\.10534*, 2025\.
- Hsieh et al\. \(2024\)Cheng\-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg\.Ruler: What’s the real context size of your long\-context language models?*arXiv preprint arXiv:2404\.06654*, 2024\.
- Hwang et al\. \(2024\)Sukjun Hwang, Aakash Lahoti, Tri Dao, and Albert Gu\.Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers, July 2024\.URL[http://arxiv\.org/abs/2407\.09941](http://arxiv.org/abs/2407.09941)\.arXiv:2407\.09941 \[cs\]\.
- Klee et al\. \(2026\)David Klee, Boce Hu, Andrew Cole, Heng Tian, Dian Wang, Robert Platt, and Robin Walters\.RAVEN: End\-to\-end equivariant robot learning with RGB cameras\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=z8BN7KyaPl](https://openreview.net/forum?id=z8BN7KyaPl)\.
- Li et al\. \(2024a\)Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng\-Yu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El\-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G\. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar\.Datacomp\-lm: In search of the next generation of training sets for language models\.*Advances in Neural Information Processing Systems*, 37:14200–14282, 2024a\.
- Li et al\. \(2026\)Ruilong Li, Brent Yi, Junchen Liu, Hang Gao, Yi Ma, and Angjoo Kanazawa\.Cameras as relative positional encoding\.*Advances in Neural Information Processing Systems*, 38:15984–16009, 2026\.
- Li et al\. \(2024b\)Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli\.Functional interpolation for relative positions improves long context transformers\.In*International Conference on Learning Representations*, volume 2024, pages 11303–11328, 2024b\.
- Lippmann et al\. \(2025\)Peter Lippmann, Gerrit Gerhartz, Roman Remme, and Fred A Hamprecht\.Beyond canonicalization: How tensorial messages improve equivariant message passing\.In*International Conference on Learning Representations*, volume 2025, pages 88067–88087, 2025\.
- Lozhkov et al\. \(2024\)Anton Lozhkov, Loubna Ben Allal, Leandro von Werra, and Thomas Wolf\.Fineweb\-edu: the finest collection of educational content, 2024\.URL[https://huggingface\.co/datasets/HuggingFaceFW/fineweb\-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu)\.
- Miyato et al\. \(2024\)Takeru Miyato, Bernhard Jaeger, Max Welling, and Andreas Geiger\.Gta: A geometry\-aware attention mechanism for multi\-view transformers\.In*International Conference on Learning Representations*, volume 2024, pages 8172–8208, 2024\.
- nanoGPT \(2022\)nanoGPT\.nanogpt\.[https://github\.com/karpathy/nanoGPT](https://github.com/karpathy/nanoGPT), 2022\.GitHub repository\.
- Peng et al\. \(2023\)Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, and Jiaming et al\. Kong\.Rwkv: Reinventing rnns for the transformer era\.*arXiv:2305\.13048*, 2023\.
- Peng et al\. \(2024\)Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole\.Yarn: Efficient context window extension of large language models\.In*International Conference on Learning Representations*, volume 2024, pages 31932–31951, 2024\.
- Poli et al\. \(2023\)Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y\. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré\.Hyena Hierarchy: Towards Larger Convolutional Language Models, April 2023\.URL[http://arxiv\.org/abs/2302\.10866](http://arxiv.org/abs/2302.10866)\.arXiv:2302\.10866 \[cs\]\.
- Press et al\. \(2021\)Ofir Press, Noah A Smith, and Mike Lewis\.Train short, test long: Attention with linear biases enables input length extrapolation\.*arXiv preprint arXiv:2108\.12409*, 2021\.
- Raffel et al\. \(2020\)Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu\.Exploring the limits of transfer learning with a unified text\-to\-text transformer\.*Journal of machine learning research*, 21\(140\):1–67, 2020\.
- Romero et al\. \(2020\)David W\. Romero, Erik J\. Bekkers, Jakub M\. Tomczak, and Mark Hoogendoorn\.Attentive Group Equivariant Convolutional Networks, June 2020\.URL[http://arxiv\.org/abs/2002\.03830](http://arxiv.org/abs/2002.03830)\.arXiv:2002\.03830 \[cs\]\.
- Su et al\. \(2024\)Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu\.Roformer: Enhanced transformer with rotary position embedding\.*Neurocomputing*, 568:127063, 2024\.
- Tian et al\. \(2026\)Qingyuan Tian, Wenhong Zhu, Xiaoran Liu, Xiaofeng Wang, and Rui Wang\.Mrrope: Mixed\-radix rotary position embedding\.*arXiv preprint arXiv:2601\.22181*, 2026\.
- Vaswani et al\. \(2017\)Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin\.Attention is all you need\.*Advances in neural information processing systems*, 30, 2017\.
- Wu et al\. \(2026\)Yu Wu, Minsik Jeon, Jen\-Hao Rick Chang, Oncel Tuzel, and Shubham Tulsiani\.RayRoPE: Projective Ray Positional Encoding for Multi\-view Attention, January 2026\.URL[http://arxiv\.org/abs/2601\.15275](http://arxiv.org/abs/2601.15275)\.arXiv:2601\.15275 \[cs\]\.
- Zheng et al\. \(2025\)Chuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong, Jiankai Sun, Jingyao Li, Minbin Huang, Xiaozhe Ren, Michael Ng, Xin Jiang, et al\.Dape v2: Process attention score as feature map for length extrapolation\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 10628–10666, 2025\.

## Appendix AFull Experimental Results

### A\.1Setup

#### Models:

We train two GPT\-2\-style transformers in the nanoGPT framework\(nanoGPT,[2022](https://arxiv.org/html/2606.11275#bib.bib28)\)\. The*small*model \(≈124\{\\approx\}124M parameters\) has 12 layers, 12 attention heads, and embedding dimension 768; the*medium*model \(≈354\{\\approx\}354M parameters\) has 24 layers, 16 heads, and embedding dimension 1024\. Both models share a GPT\-2 BPE vocabulary of 50 304 tokens, pre\-layer normalisation, GELU activations, and base RoPE frequencyθ0=10 000\\theta\_\{0\}=10\\,000\. The*only*architectural difference between theRoPEandRoVEconditions is the value pathway; all other architectural and training hyperparameters are held fixed\.

#### Training:

Both models are trained for one epoch on FineWebEdu\-10B \(≈10\{\\approx\}10B tokens of educational web text tokenised with the GPT\-2 tiktoken encoder\)\(Lozhkov et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib26)\), with a sequence length of 1024 tokens\. We use a total batch size of219=524 2882^\{19\}=524\\,288tokens, accumulated via gradient accumulation over micro\-batches of 32 sequences per GPU \(small\) and 16 \(medium\), distributed across four NVIDIA H100 GPUs with PyTorch DDP\. We optimise with AdamW \(β=\(0\.9,0\.95\)\\beta=\(0\.9,0\.95\), weight decay0\.10\.1, gradient clipping at norm1\.01\.0\) inbfloat16\. The learning rate follows a cosine decay from6×10−46\\times 10^\{\-4\}to6×10−56\\times 10^\{\-5\}after a 715\-step linear warm\-up, over 19 073 gradient steps in total\.

#### Evaluation:

We evaluate both models on three benchmarks\.

- •Core ICL accuracy\.We report in\-context learning accuracy on the DCLM\-Core benchmark\(Li et al\.,[2024a](https://arxiv.org/html/2606.11275#bib.bib22)\), a diverse suite of few\-shot tasks spanning multiple\-choice, Winograd schema, and language\-modelling formats\. Per\-task accuracy is centred on the random baseline and normalised to\[0,1\]\[0,1\], and the Core score is the mean across tasks\. All evaluation uses the GPT\-2 tiktoken tokeniser, matching training\.
- •Out\-of\-distribution perplexity\.We evaluate on the FineWebEdu\-10B held\-out validation split at context lengthsL∈\{512,1024,2048,4096,8192,16384\}L\\in\\\{512,1024,2048,4096,8192,16384\\\}using a sliding window with stride 512 tokens\. Only the finalmin⁡\(L,512\)\\min\(L,512\)tokens per window are scored, so each reported perplexity reflects next\-token prediction conditioned onLLtokens of preceding context\. We additionally applyYaRN\(Peng et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib30)\)positional interpolation at inference time without any fine\-tuning, where the frequency modulation is applied to all rotation matrices, covering both the QK\- and OV\-circuits\.
- •RULER long\-context retrieval\.We evaluate on four RULER synthetic tasks\(Hsieh et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib19)\), namely Common Word Extraction \(CWE\), multi\-key Needle\-in\-a\-Haystack \(NIAH\), Question Answering \(QA\), and Variable Tracking \(VT\), at context lengths 4 096 and 8 192 tokens with 500 samples each, constructed with the NVIDIA RULER pipeline using the GPT\-2 tokeniser\. Because our models are base language models without instruction tuning, candidate answers are ranked by negative log\-likelihood \(NLL, where lower values indicate higher probability assigned to the correct answer\)\. We report mean±\\pmstandard deviation over the 500 samples\.

### A\.2Results

#### In\-distribution performance:

Tables[3](https://arxiv.org/html/2606.11275#A1.T3)and[1](https://arxiv.org/html/2606.11275#S4.T1)report Core ICL accuracy and perplexity within the training context length \(≤1024\\leq 1024tokens\) for both scales\.RoVEimproves Core from0\.13750\.1375to0\.14160\.1416at 124M and from0\.16640\.1664to0\.18560\.1856at 354M\. Correspondingly, perplexity at 512 and 1024 tokens decreases from25\.2325\.23/22\.3722\.37to25\.0525\.05/22\.3022\.30\(124M\) and from17\.6817\.68/15\.6415\.64to17\.5217\.52/15\.5215\.52\(354M\)\. The consistent gains at both scales within the trained context window confirm that value\-side rotation provides a useful inductive bias independently of any length\-extrapolation effect\.

#### Out\-of\-distribution perplexity:

Beyond the training context, the advantage ofRoVEgrows substantially\. Without positional interpolation, the 354MRoVEmodel reaches perplexity311\.38311\.38and583\.84583\.84at 4k and 16k tokens, compared to840\.10840\.10and1630\.721630\.72forRoPE, i\.e\., a reduction of approximately63%63\\%and64%64\\%\. ApplyingYaRNnarrows the absolute values for both methods, but the relative gap persists:RoVE\+YaRNachieves18\.4018\.40and124\.82124\.82against48\.6148\.61and270\.98270\.98forRoPE\+YaRNat the same lengths\. The 124M model follows the same pattern \(Table[3](https://arxiv.org/html/2606.11275#A1.T3)\), withRoVE\+YaRNreaching27\.3427\.34and185\.87185\.87at 4k and 16k compared to58\.6758\.67and310\.52310\.52forRoPE\+YaRN\. These results demonstrate that value\-side relative rotation is complementary to, and not subsumed by, inference\-time frequency interpolation\.

#### Long\-context retrieval:

Tables[4](https://arxiv.org/html/2606.11275#A1.T4)and[2](https://arxiv.org/html/2606.11275#S4.T2)report RULER results withYaRNapplied\. At 354M,RoVE\+YaRNreduces the mean RULER NLL from6\.626\.62to4\.334\.33relative toRoPE\+YaRN, with the largest gains on tasks requiring long\-range information aggregation: at 4k tokens, multi\-key NIAH improves from7\.617\.61to3\.633\.63and Variable Tracking from4\.534\.53to2\.112\.11; at 8k tokens, the same tasks improve from9\.469\.46to5\.165\.16and from7\.507\.50to3\.103\.10, respectively\. At 124M, mean NLL decreases from6\.756\.75to5\.355\.35, again with the strongest gains on NIAH and Variable Tracking\. The consistent pattern across both scales and tasks indicates that positional alignment in the value pathway is especially beneficial for retrieval problems that require detecting and recombining information distributed across long contexts\.

Table 3:Core ICL accuracy and perplexity for the 124M parameter model\. Core is measured within the 1024\-token training context; PPL is measured from 512 to 16384 tokens\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation, leaving Core unchanged\.Green highlightmarks column bests\.Table 4:RULER long\-context retrieval \(NLL; lower is better\) for the 124M parameter model\. Tasks: Common Word Extraction \(CWE\), multi\-key Needle\-in\-a\-Haystack \(NIAH\), Question Answering \(QA\), Variable Tracking \(VT\)\. Avg is the unweighted mean over all eight task/length cells\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation\.Green highlightmarks column bests\.Table 5:Core ICL accuracy and perplexity for the 354M parameter model\. Core is measured within the 1024\-token training context; PPL is measured from 512 to 16384 tokens\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation, leaving Core unchanged\.Green highlightmarks column bests\.Restatement of the results presented at Table[1](https://arxiv.org/html/2606.11275#S4.T1)for easier comparison with 124M parameter model\.Table 6:RULER long\-context retrieval \(NLL; lower is better\) for the 354M parameter model\. Tasks: Common Word Extraction \(CWE\), multi\-key Needle\-in\-a\-Haystack \(NIAH\), Question Answering \(QA\), Variable Tracking \(VT\)\. Avg is the unweighted mean over all eight task/length cells\.\+YaRN\+\\textit\{YaRN\}denotes inference\-time interpolation\.Green highlightmarks column bests\.Restatement of the results presented at Table[2](https://arxiv.org/html/2606.11275#S4.T2)for easier comparison with 124M parameter model\.\.

## Appendix BRelated Work

#### Positional encodings:

Positional encodings for transformers fall into three families\.

- •*Absolute encodings*\(APE;Vaswani et al\.[2017](https://arxiv.org/html/2606.11275#bib.bib37)\) add a fixed or learned vectorpip\_\{i\}to each token embedding before projection, yielding scores Ai​jape=\(WQ​\(xi\+pi\)\)⊤​\(WK​\(xj\+pj\)\),A^\{\\mathrm\{ape\}\}\_\{ij\}=\(W\_\{\\\!Q\}\(x\_\{i\}\+p\_\{i\}\)\)^\{\\\!\\top\}\(W\_\{\\\!K\}\(x\_\{j\}\+p\_\{j\}\)\),which depend on the absolute indicesiiandjjseparately, making extrapolation to unseen lengths fragile\.
- •*Additive relative encodings*\(ARPE;Raffel et al\.[2020](https://arxiv.org/html/2606.11275#bib.bib33); Press et al\.[2021](https://arxiv.org/html/2606.11275#bib.bib32); Chi et al\.[2022](https://arxiv.org/html/2606.11275#bib.bib7); Li et al\.[2024b](https://arxiv.org/html/2606.11275#bib.bib24)\) replace the absolute positional terms with an offset\-indexed bias, Ai​jarpe=\(WQ​xi\)⊤​\(WK​xj\)\+bj−i,A^\{\\mathrm\{arpe\}\}\_\{ij\}=\(W\_\{\\\!Q\}x\_\{i\}\)^\{\\\!\\top\}\(W\_\{\\\!K\}x\_\{j\}\)\+b\_\{j\-i\},so the positional contribution depends only on the displacementδ=j−i\\delta\{=\}j\{\-\}i, which can improve length generalisation over APE but requires materialising the fulln×nn\{\\times\}nscore matrix, preventing the use of FlashAttention kernels\(Dao et al\.,[2022](https://arxiv.org/html/2606.11275#bib.bib11)\)\.
- •*Rotary position encoding*,RoPE\(Su et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib35)\), obtains the same offset\-only dependence multiplicatively: rotating queries and keys by their absolute positions before the inner product yields Ai​jrope=\(Ri​WQ​xi\)⊤​\(Rj​WK​xj\)=\(WQ​xi\)⊤​Rj−i​\(WK​xj\),A^\{\\mathrm\{rope\}\}\_\{ij\}=\(R\_\{i\}W\_\{\\\!Q\}x\_\{i\}\)^\{\\\!\\top\}\(R\_\{j\}W\_\{\\\!K\}x\_\{j\}\)=\(W\_\{\\\!Q\}x\_\{i\}\)^\{\\\!\\top\}R\_\{j\-i\}\(W\_\{\\\!K\}x\_\{j\}\),\(4\)which depends on the displacementδ=j−i\\delta\{=\}j\{\-\}irather than oniiandjjseparately, while remaining compatible with FlashAttention\. Length extrapolation withRoPEremains challenging because rotation angles at unseen positions are out of distribution; a number of methods address this via frequency rescaling, including PI, NTK, andYaRN\(see Appendix[C](https://arxiv.org/html/2606.11275#A3)for a brief overview\)\.

In all three families the value pathway carries no relative\-position signal: APE injects an absolute component viaWV​\(xj\+pj\)W\_\{\\\!V\}\(x\_\{j\}\{\+\}p\_\{j\}\), while ARPE andRoPEleaveWV​xjW\_\{\\\!V\}x\_\{j\}entirely unchanged\.RoVEcloses this gap by applying the same rotation family to the value pathway,ψδ=Rδ​WV\\psi\_\{\\delta\}\{=\}R\_\{\\delta\}W\_\{\\\!V\}, without modifying scores or sacrificing FlashAttention compatibility\.

#### Distinguishing “what” and “where” inRoPE:

A separate line of work examines how position and content interact*within*theRoPEscore itself\.Barbero et al\. \([2024](https://arxiv.org/html/2606.11275#bib.bib2)\); Chen et al\. \([2025](https://arxiv.org/html/2606.11275#bib.bib6)\)show mechanistically that low\-frequencyRoPEcomponents act as semantic channels in trained models, while high\-frequency components construct positional attention patterns\. PoPE\(Gopalakrishnan et al\.,[2025](https://arxiv.org/html/2606.11275#bib.bib18)\)formalises this through a polar decomposition of \([4](https://arxiv.org/html/2606.11275#A2.E4)\), identifying a content\-dependent phase cross\-term that entangles positional and semantic information, which is replaced by pure\-magnitude representations to yield a score that factors into a content product and a positional cosine, improving perplexity and length generalisation\. These works restructure the QK circuit to disentangle position and content in the attention mechanism;RoVEinstead introduces positional structure into the value pathway, a component none of them address\. CombiningRoVEwith semantics\-aware QK encodings such as PoPE is a natural direction for future work\.

#### Attention and convolution:

Cordonnier et al\. \([2019](https://arxiv.org/html/2606.11275#bib.bib9)\)prove that multi\-head self\-attention can express any convolutional layer under a specific relative positional encoding\. Their construction takes the ARPE score ofDai et al\. \([2019](https://arxiv.org/html/2606.11275#bib.bib10)\)and zeroes the content\-driven projection matrices, collapsing it to a position\-only term, i\.e\., a degenerate ARPE in whichbδ=v\(h\)⊤​rδb\_\{\\delta\}\{=\}v^\{\(h\)\\top\}r\_\{\\delta\}absorbs the entire score\. With the quadratic encodingrδ=\(‖δ‖2,δ1,δ2\)r\_\{\\delta\}\{=\}\(\\\|\\delta\\\|^\{2\},\\delta\_\{1\},\\delta\_\{2\}\), each head’s score peaks sharply at a single fixed offset, so its value matrixWV\(h\)W\_\{\\\!V\}^\{\(h\)\}acts as the convolutional filter for that offset\. The value pathway is thus offset\-specific, but scores become content\-independent and each filter is a separate, discrete matrix tied to one head\.RoVEreaches the same attentive\-convolution structure from the opposite side: it keeps fully content\-dependentRoPEscores \([4](https://arxiv.org/html/2606.11275#A2.E4)\) and makes the value pathway offset\-dependent through the continuous, parameter\-free familyψδ=Rδ​WV\\psi\_\{\\delta\}\{=\}R\_\{\\delta\}W\_\{\\\!V\}, a singleWVW\_\{\\\!V\}rotated through the rotation group rather than replicated per head\.

DAPE V2\(Zheng et al\.,[2025](https://arxiv.org/html/2606.11275#bib.bib39)\)takes a complementary approach, applying a narrow convolution kernel across heads over the pre\-softmax score tensor \(on top of a standard additive positional bias\), and shows that this convolution component alone provably suffices for associative recall even when the bias is zeroed out\. However, the operation requires materialising the fulln×nn\{\\times\}nscore tensor before softmax, ruling outFlashAttention\. Nevertheless, the two methods are structurally dual: DAPE V2 enriches routing while leaving values intact, whereasRoVEenriches the value transformation while leaving scores intact\.

#### Gated convolutions and recall:

Gated convolutions and state\-space models\(Fu et al\.,[2023a](https://arxiv.org/html/2606.11275#bib.bib15); Poli et al\.,[2023](https://arxiv.org/html/2606.11275#bib.bib31); Peng et al\.,[2023](https://arxiv.org/html/2606.11275#bib.bib29); Fu et al\.,[2023b](https://arxiv.org/html/2606.11275#bib.bib16)\)provide sub\-quadratic alternatives to attention\.Arora et al\. \([2024](https://arxiv.org/html/2606.11275#bib.bib1)\)show that 82% of their perplexity gap relative to attention is explained by*associative recall*: gated convolutions apply a fixed filter whose weights are set by model parameters alone and cannot adapt, based on the input, which tokens to mix, whereas attention’s input\-dependent scoressoftmax⁡\(Q​K⊤\)\\operatorname\{softmax\}\(QK^\{\\top\}\)can locate the matching token at any distance\.RoVEaddresses a complementary half of this picture\. Attention already determines*which*token to retrieve, but with standardRoPEthe value transformationWVW\_\{\\\!V\}is the same fixed map regardless of how far that token lies from the query\. ReplacingWVW\_\{\\\!V\}with the offset\-indexed familyψδ=Rδ​WV\\psi\_\{\\delta\}\{=\}R\_\{\\delta\}W\_\{\\\!V\}lets the model additionally control*how*each retrieved feature is realigned before it is recombined with the query, much as multi\-view aggregation rotates a feature into a common reference frame before fusion\(Miyato et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib27)\)\. Consistent with this view,RoVEimproves associative recall over vanillaRoPEacross the long\-context retrieval benchmarks reported in Appendix[A\.2](https://arxiv.org/html/2606.11275#A1.SS2)\.

## Appendix CAdditional Background on Frequency Scaling

At positiontt,RoPErotates frequency channelmmby anglet​ωmt\\,\\omega\_\{m\}\. Whenttexceeds the training context lengthLL, this angle falls outside the range\[0,L​ωm\]\[0,L\\,\\omega\_\{m\}\]seen during training, causing distributional shift that compounds across layers\. All practical remedies address this by rescaling the frequenciesωm=θ0−2​m/d\\omega\_\{m\}=\\theta\_\{0\}^\{\-2m/d\}so that positions up to a target lengthL′\>LL^\{\\prime\}\>Lproduce rotation angles within the training range\.

*Positional Interpolation*\(PI;Chen et al\.[2023](https://arxiv.org/html/2606.11275#bib.bib5)\) maps each positiont↦t​L/L′t\\mapsto tL/L^\{\\prime\}, equivalently applying a uniform frequency reductionωm↦ωm/s\\omega\_\{m\}\\mapsto\\omega\_\{m\}/swiths=L′/Ls=L^\{\\prime\}/L\. This guarantees all rotation angles remain in\-distribution, but does so indiscriminately: high\-frequency dimensions, which encode fine\-grained local positional distinctions, are compressed by the same factorssas low\-frequency dimensions, blurring short\-range structure\.

*NTK\-aware scaling*\(bloc97,[2023](https://arxiv.org/html/2606.11275#bib.bib3)\)corrects this imbalance by uniformly rescaling the baseθ0\\theta\_\{0\}rather than the positions directly\. Becauseωm=θ0−2​m/d\\omega\_\{m\}=\\theta\_\{0\}^\{\-2m/d\}, a single multiplicative base change has a dimension\-dependent effect on individual frequencies: the substitution

θ0↦θ0⋅sd/\(d−2\)\\theta\_\{0\}\\;\\mapsto\\;\\theta\_\{0\}\\cdot s^\{d/\(d\-2\)\}rescales each frequency asωm↦ωm⋅s−2​m/\(d−2\)\\omega\_\{m\}\\mapsto\\omega\_\{m\}\\cdot s^\{\-2m/\(d\-2\)\}, leaving the highest\-frequency dimension \(m=0m=0\) entirely unchanged while recovering the full PI factors−1s^\{\-1\}at the lowest\-frequency dimension \(m=d/2−1m=d/2\-1\), thereby preserving short\-range positional structure where it is most informative\.

YaRN\(Peng et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib30)\)takes this frequency\-dependent logic to its principled conclusion by treating each dimension according to its wavelengthλm=2​π/ωm\\lambda\_\{m\}=2\\pi/\\omega\_\{m\}relative to the training contextLL\. Dimensions withλm≪L\\lambda\_\{m\}\\ll Lcomplete many full rotations within the training window; their angles are robustly periodic and can therefore be safely*extrapolated*\(left unscaled\) at inference time\. Dimensions withλm\>L\\lambda\_\{m\}\>Lnever complete a single rotation during training: at positions beyondLLtheir rotation angles fall entirely outside the training distribution, making them the primary source of extrapolation errors\. These dimensions must therefore be*interpolated*\. Intermediate dimensions are handled by a smooth blend between the two regimes:

ωm′=\{ωmif​λm<α,\(1−γ​\(λm\)\)​ωm\+γ​\(λm\)​ωm/sif​α≤λm≤β,ωm/sif​λm\>β,\\omega\_\{m\}^\{\\prime\}=\\begin\{cases\}\\omega\_\{m\}&\\text\{if \}\\lambda\_\{m\}<\\alpha,\\\\\[2\.0pt\] \\bigl\(1\-\\gamma\(\\lambda\_\{m\}\)\\bigr\)\\,\\omega\_\{m\}\+\\gamma\(\\lambda\_\{m\}\)\\,\\omega\_\{m\}/s&\\text\{if \}\\alpha\\leq\\lambda\_\{m\}\\leq\\beta,\\\\\[2\.0pt\] \\omega\_\{m\}/s&\\text\{if \}\\lambda\_\{m\}\>\\beta,\\end\{cases\}whereγ\\gammais a smooth blending function increasing from0to11over\[α,β\]\[\\alpha,\\beta\], andα,β\\alpha,\\betaare wavelength thresholds\. To compensate for the softmax sharpness distortion introduced by frequency compression,YaRNadditionally applies an attention temperature correction1/t\\sqrt\{1/t\}to the pre\-softmax logits\. All adjustments are applied at inference time without any additional fine\-tuning\.

More advanced frequency extrapolation schemes have since been proposed\(Ding et al\.,[2024](https://arxiv.org/html/2606.11275#bib.bib13); Tian et al\.,[2026](https://arxiv.org/html/2606.11275#bib.bib36)\); however, in this work we focus onYaRNas it is among the most widely adopted techniques in production settings, as evidenced, for instance, by its use in DeepSeek\-V4\(DeepSeek\-AI,[2026](https://arxiv.org/html/2606.11275#bib.bib12)\)\.

#### Frequency scaling inRoVE:

When applyingYaRNtoRoVE, we rescale the frequencies of both the QK and OV rotation matrices uniformly, as described above\. This is a natural extension: sinceRoVEties the value pathway to the same rotation family\{Rδ\}δ\\\{R\_\{\\delta\}\\\}\_\{\\delta\}as the keys and queries, the same out\-of\-distribution problem arises in the OV circuit when positions exceedLL, and the same frequency rescaling mitigates it\. We leave as future work whether more specialised extrapolation strategies should be developed specifically for the OV pathway\. The methods reviewed above, such as, PI, NTK\-aware scaling, andYaRN, were originally designed under the assumption that rotations appear only in the QK circuit; it is therefore an open question whether their frequency thresholds and blending schedules remain optimal when the same rotation family also modulates value transformations, or whether the two pathways call for different rescaling regimes\.

## Appendix DDiscussion

#### What changes when values rotate?

In standardRoPE, relative position determines attention weights but leaves the value mapWVW\_\{\\\!V\}unchanged: position governs*which*features are aggregated, not*how*\.RoVEcloses this gap by replacingWVW\_\{\\\!V\}with the offset\-dependent kernelRδ​WVR\_\{\\delta\}W\_\{\\\!V\}, so that the layer remains shift\-equivariant while the value stream carries strictly more relative\-position information than a scalar attention coefficient can convey alone\. From a geometric perspective, attention selects which neighboring features to aggregate whileRδ​WVR\_\{\\delta\}W\_\{\\\!V\}aligns each selected feature before fusion, analogous to geometric multi\-view aggregation inMiyato et al\. \([2024](https://arxiv.org/html/2606.11275#bib.bib27)\), where features from different camera views are rotated into a common frame; in language, positions replace views and relative\-position rotations replace camera\-to\-camera transforms\. This joint conditioning of selection and aggregation explains the RULER pattern, where gains are largest on tasks requiring long\-range information to be tracked and recombined rather than merely detected\.

#### Limitations and future work:

While the experiments show a consistent advantage for rotating values, they do not fully identify the mechanism behind the improved OOD perplexity\. Our working hypothesis is thatRoVEinduces a more coherent extrapolation regime: when relative offsets exceed those seen in training, the QK and value pathways drift together because they share the same rotation family\. In standardRoPE, by contrast, the attention logits extrapolate while the value transformation remains unchanged, creating a mismatch between selection and aggregation\. A direct test would analyze errors byRoPEfrequency band, examining each term in the offset\-indexed sum of Appendix[E](https://arxiv.org/html/2606.11275#A5)separately, following the style ofChen et al\. \([2025](https://arxiv.org/html/2606.11275#bib.bib6)\), and measure whether value\-side rotations preserve the learned relationship between the QK and OV circuits at unseen offsets\.

## Appendix ECircuits Framework Analysis

#### Background:

Elhage et al\. \([2021](https://arxiv.org/html/2606.11275#bib.bib14)\)provide a notation for decomposing transformer computations into interpretable end\-to\-end paths\. We recall the elements needed here\.

The*token embedding matrix*WE∈ℝd×\|𝒱\|W\_\{E\}\\in\\mathbb\{R\}^\{d\\times\|\\mathcal\{V\}\|\}maps one\-hot token vectors to residual\-stream vectors of dimensiondd; the*unembedding matrix*WU∈ℝ\|𝒱\|×dW\_\{U\}\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}maps residual\-stream vectors back to logits over the vocabulary𝒱\\mathcal\{V\}\. Together they form the “direct path” from input tokens to output logits:Id⊗WU​WE\\mathrm\{Id\}\\otimes W\_\{U\}W\_\{E\}, whereId\\mathrm\{Id\}is the identity on the sequence dimension and⊗\\otimesdenotes the tensor product \(equivalently, a Kronecker product when written on vectorised tokens\)\.

Each attention headhhcontributes through two largely independent circuits:

- •QK circuit\.The bilinear formWQ​Kh=\(WQh\)⊤​WKh∈ℝd×dW\_\{QK\}^\{h\}=\(W\_\{\\\!Q\}^\{h\}\)^\{\\\!\\top\}W\_\{\\\!K\}^\{h\}\\in\\mathbb\{R\}^\{d\\times d\}determines the attention patternAh​\(X\)∈ℝn×nA^\{h\}\(X\)\\in\\mathbb\{R\}^\{n\\times n\}: entryAh​\(X\)i​jA^\{h\}\(X\)\_\{ij\}measures how strongly positioniiattends to positionjjbased on the residual\-stream contentXXat those positions\. It answers the question*which*tokens are attended to\.
- •OV circuit\.The matrixWO​Vh=WOh​WVh∈ℝd×dW\_\{OV\}^\{h\}=W\_\{\\\!O\}^\{h\}W\_\{\\\!V\}^\{h\}\\in\\mathbb\{R\}^\{d\\times d\}determines what is communicated when a token is attended to: it maps the residual\-stream vector at positionjjto the update written into positionii’s residual stream\. It answers the question*what*information is moved\.

The full one\-layer attention\-only transformer then expands as

T​\(X\)=Id⊗WU​WE\+∑hAh​\(X\)⊗\(WU​WO​Vh​WE\),T\(X\)=\\mathrm\{Id\}\\otimes W\_\{U\}W\_\{E\}\\;\+\\;\\sum\_\{h\}A^\{h\}\(X\)\\otimes\\bigl\(W\_\{U\}W\_\{OV\}^\{h\}W\_\{E\}\\bigr\),\(5\)where the tensor productAh​\(X\)⊗\(WU​WO​Vh​WE\)A^\{h\}\(X\)\\otimes\(W\_\{U\}W\_\{OV\}^\{h\}W\_\{E\}\)means:Ah​\(X\)A^\{h\}\(X\)routes information across sequence positions in an input\-dependent way, whileWU​WO​Vh​WEW\_\{U\}W\_\{OV\}^\{h\}W\_\{E\}transforms it in the channel dimension with fixed weights\. The two dimensions are*independent*—this is the Kronecker structure\. See Figure[2](https://arxiv.org/html/2606.11275#A5.F2)for a visual representation of Equation \([5](https://arxiv.org/html/2606.11275#A5.E5)\)\.

\\floatconts

fig:subfigex

Figure 2:Visualization of the circuits of the one\-layer attention\-only transformer forRoPE\(left; adapted fromElhage et al\. \([2021](https://arxiv.org/html/2606.11275#bib.bib14)\)\) andRoVE\(right\)\. In standardRoPE, we have as many bifurcations from the residual stream as there are attention headshih\_\{i\}at that layer\. InRoVEwe obtain new bifurcations associated with each displacementδ\\delta\.\\subfigure

\[RoPEattention circuit see Equation \([5](https://arxiv.org/html/2606.11275#A5.E5)\)\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x6.png)\\subfigure\[RoVEattention circuit; see Equation \([6](https://arxiv.org/html/2606.11275#A5.E6)\)\]![Refer to caption](https://arxiv.org/html/2606.11275v1/x7.png)

#### How RoPE fits in:

WithRoPE, the attention patternAh​\(X\)A^\{h\}\(X\)becomes shift\-equivariant \(entryAh​\(X\)i​jA^\{h\}\(X\)\_\{ij\}depends on positions only through the offsetj−ij\-iand the token content\), but the OV circuitWO​VhW\_\{OV\}^\{h\}remains a single fixed matrix\. The Kronecker structure of \([5](https://arxiv.org/html/2606.11275#A5.E5)\) is therefore preserved underRoPE: token routing and channel transformation are still decoupled\.

#### HowRoVEchanges the picture:

RoVEoperates at the attention level, replacing the constant value mapWVhW\_\{\\\!V\}^\{h\}with the offset\-dependent kernelRδ​WVhR\_\{\\delta\}W\_\{\\\!V\}^\{h\}\. From the circuits perspective, this means the effective OV matrix at offsetδ\\deltaisWOh​Rδ​WVhW\_\{\\\!O\}^\{h\}R\_\{\\delta\}W\_\{\\\!V\}^\{h\}, with the rotationRδR\_\{\\delta\}sandwiched between the two projections\. Consequently, the token\-routing and channel\-transformation dimensions are no longer independent, and \([5](https://arxiv.org/html/2606.11275#A5.E5)\) no longer holds\. Introducing shift matricesSδ∈ℝn×nS^\{\\delta\}\\in\\mathbb\{R\}^\{n\\times n\}with\(Sδ\)i​j=𝟏​\[j−i=δ\]\(S^\{\\delta\}\)\_\{ij\}=\\mathbf\{1\}\[j\-i=\\delta\]to partition the attention pattern by offset, theRoVEtransformer expands as

TRoVE​\(X\)=Id⊗WU​WE\+∑h∑δ\(Ah​\(X\)⊙Sδ\)⊗\(WU​WOh​Rδ​WVh​WE\),T^\{\\mbox\{\{R\\kern 0\.1pto\\kern\-1\.0ptVE\}\}\}\(X\)=\\mathrm\{Id\}\\otimes W\_\{U\}W\_\{E\}\\;\+\\;\\sum\_\{h\}\\sum\_\{\\delta\}\\bigl\(A^\{h\}\(X\)\\odot S^\{\\delta\}\\bigr\)\\otimes\\bigl\(W\_\{U\}W\_\{\\\!O\}^\{h\}R\_\{\\delta\}W\_\{\\\!V\}^\{h\}W\_\{E\}\\bigr\),\(6\)where⊙\\odotis the entry\-wise product\. Each term selects theδ\\delta\-offset entries ofAh​\(X\)A^\{h\}\(X\)and pairs them with the corresponding rotated end\-to\-end map, and summing overδ\\deltarecovers the full output because∑δAh​\(X\)⊙Sδ=Ah​\(X\)\\sum\_\{\\delta\}A^\{h\}\(X\)\\odot S^\{\\delta\}=A^\{h\}\(X\)\. The Kronecker structure is replaced by a*sum of Kronecker products*, one per offset diagonal, which is exactly the block\-Toeplitz structure described in Section[3](https://arxiv.org/html/2606.11275#S3)\. See Figure[2](https://arxiv.org/html/2606.11275#A5.F2)for a visual representation of Equation \([6](https://arxiv.org/html/2606.11275#A5.E6)\)\.

Similar Articles

RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

arXiv cs.CL

This paper provides a theoretical proof that Rotary Positional Embeddings (RoPE) in Transformer-based language models lose their locality bias and ability to distinguish token order in long contexts, with attention scores becoming no better than random. The authors show that increasing the RoPE base trades off position vs. token distinction and that multi-head, multi-layer architectures cannot compensate for this fundamental limitation.

RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates

arXiv cs.CL

This preliminary technical report proposes RIG-RoPE, a relation- and instance-gated rotary positional encoding with duration-aware temporal coordinates, aiming to address spatial interference and improper temporal scaling in multimodal LLMs. It introduces a gating mechanism for height/width rotations and duration-aware temporal coordinates, but leaves large-scale empirical validation to future work.