An Analysis of Residual-Stream Geometry Across Transformer Depth

arXiv cs.LG Papers

Summary

This paper proposes a geometric analysis of transformer residual streams across depth, using relative displacement and orthogonal Procrustes analysis to reveal structured regularities in six instruction-tuned models on code generation and translation tasks.

arXiv:2607.18348v1 Announce Type: new Abstract: We propose a transition-centred geometric analysis of transformer residual streams. Relative displacement measures how \emph{far} representations move between consecutive layers, and orthogonal Procrustes analysis separates each transition into a rigid rotation and a non-rigid residual. Across six instruction-tuned models, on code generation and cross-lingual translation, these measurements reveal reproducible depth regularities. Relative displacement is strongly layer-dependent; typically larger early and late, with a quieter middle third; and nearly invariant across conditions within each model. Rotation magnitude is nearly constant across depth, while Procrustes residual and angle concentration remain depth-modulated, with residual peaking at the final transition. During generation, non-English targets show larger final-layer displacement and residual than English targets. We present these as descriptive geometric regularities, not as measures of computational effort or causal explanations. The contribution is a measurement framework for residual-stream transitions and evidence that, in the settings studied here, depth curves are model-dependent and largely condition-stable.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:20 AM

# An Analysis of Residual-Stream Geometry Across Transformer Depth
Source: [https://arxiv.org/html/2607.18348](https://arxiv.org/html/2607.18348)
###### Abstract

We propose a transition\-centred geometric analysis of transformer residual streams\. Relative displacement measures how*far*representations move between consecutive layers, and orthogonal Procrustes analysis separates each transition into a rigid rotation and a non\-rigid residual\. Across six instruction\-tuned models, on code generation and cross\-lingual translation, these measurements reveal reproducible depth regularities\. Relative displacement is strongly layer\-dependent; typically larger early and late, with a quieter middle third; and nearly invariant across conditions within each model\. Rotation magnitude is nearly constant across depth, while Procrustes residual and angle concentration remain depth\-modulated, with residual peaking at the final transition\. During generation, non\-English targets show larger final\-layer displacement and residual than English targets\. We present these as descriptive geometric regularities, not as measures of computational effort or causal explanations\. The contribution is a measurement framework for residual\-stream transitions and evidence that, in the settings studied here, depth curves are model\-dependent and largely condition\-stable\.

An Analysis of Residual\-Stream Geometry Across Transformer Depth

Sunit Bhattacharya, Ravi KolliProRata AICorrespondence:[sunit@prorata\.ai](https://arxiv.org/html/2607.18348v1/mailto:[email protected])

## 1Introduction

Transformers\(Vaswaniet al\.,[2017](https://arxiv.org/html/2607.18348#bib.bib1)\)process information by repeatedly updating a residual stream\(Elhageet al\.,[2021](https://arxiv.org/html/2607.18348#bib.bib9); Olahet al\.,[2020](https://arxiv.org/html/2607.18348#bib.bib8)\)\. Most analyses of that stream ask what a layer represents or predicts: representational similarity\(Kornblithet al\.,[2019](https://arxiv.org/html/2607.18348#bib.bib15)\), vocabulary projections such as the logit lens and tuned lens\(nostalgebraist,[2020](https://arxiv.org/html/2607.18348#bib.bib14); Belroseet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib13)\), or circuit\-level accounts of specific computations\(Olahet al\.,[2020](https://arxiv.org/html/2607.18348#bib.bib8); Elhageet al\.,[2021](https://arxiv.org/html/2607.18348#bib.bib9); Bushnaqet al\.,[2024](https://arxiv.org/html/2607.18348#bib.bib7)\)\. We propose a complementary view: treat each layer transition as a geometric transformation of the token cloud in residual space\. Under this view, the basic object of study is not a static layer activation, but the map from the cloud at depthℓ\\ellto the cloud at depthℓ\+1\\ell\{\+\}1\. Relative displacement asks how far tokens move\. Orthogonal Procrustes analysis\(Schönemann,[1966](https://arxiv.org/html/2607.18348#bib.bib16)\)then separates that movement into an optimal rigid rotation and a non\-rigid residual\. Recent geometric work on transformers has studied kinematic acceleration\(Fernando and Guitchounts,[2025](https://arxiv.org/html/2607.18348#bib.bib30)\), curvature along token paths\(Damirchiet al\.,[2026](https://arxiv.org/html/2607.18348#bib.bib31)\), and the intrinsic dimension of per\-layer clouds\(Viswanathanet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib27)\)\. These approaches either snapshot a layer or follow individual tokens\. Our contribution is to measure the collective inter\-layer transformation of the full residual\-stream cloud\.

We apply this framework to six instruction\-tuned models: Qwen30\.6​B−8​B0\.6B\-8B, Gemma\-2\-2B\-IT, and StableLM\-2\-1\.6B\-Chat, on code generation and translation between English and four Indo\-European languages\. In these settings, four regularities emerge:

1. 1\.Relative displacement is strongly structured by depth: updates are typically larger early and late, with a quieter middle third\. Within each model, coding and translation conditions share essentially the same depth curve\.
2. 2\.The magnitude of the global Procrustes rotation is nearly constant across depth \(variation below a predeclared3%3\\%bound\), while Procrustes residual and angle concentration remain depth\-modulated, with residual peaking at the final transition\.
3. 3\.Conditions mainly rescale late\-layer amplitude rather than rewrite the depth curve\. During generation, non\-English targets show larger final\-layer displacement and residual than English targets\.
4. 4\.These patterns are statistically robust under hierarchical bootstrap and layer\-permutation tests with Holm correction\.

Our contribution is measurement\-first\. We introduce a transition\-centred geometric view of residual streams and document regularities that are stable across the models and conditions studied here\. We do not claim that these metrics measure computational effort, nor do we offer a causal account of why the patterns arise\. Near\-constant rotation magnitude, in particular, may partly reflect high\-dimensional Procrustes geometry; the more informative observation is that displacement and residual remain depth\-structured even when rotational scale is flat\. The middle\-layer slowdown coincides with depths previously linked to more language\-neutral processing\(Wendleret al\.,[2024](https://arxiv.org/html/2607.18348#bib.bib17); Changet al\.,[2022](https://arxiv.org/html/2607.18348#bib.bib18); Bhattacharya and Bojar,[2023](https://arxiv.org/html/2607.18348#bib.bib19)\), and the elevated non\-English residual is consistent with English\-pivot accounts\(Wendleret al\.,[2024](https://arxiv.org/html/2607.18348#bib.bib17)\)\. We report these as associations\. Taken together, the results support a compact descriptive claim: in coding and translation, residual\-stream depth curves are model\-dependent and largely condition\-stable, while content mainly rescales late\-layer amplitude\.

## 2Related Work

### 2\.1Multilingual Representations

Multilingual models exhibit structured geometry across depth: early and late layers often retain language\-specific structure, while middle layers form a more shared, language\-agnostic subspace\(Changet al\.,[2022](https://arxiv.org/html/2607.18348#bib.bib18); Bhattacharya and Bojar,[2023](https://arxiv.org/html/2607.18348#bib.bib19)\)\. Large language models also frequently process non\-English inputs by routing through an English\-centric latent space\(Wendleret al\.,[2024](https://arxiv.org/html/2607.18348#bib.bib17)\)\. Within that shared space, some computational features appear largely independent of both the input language and the final output language\(Schutet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib20); Kyriakouet al\.,[2026](https://arxiv.org/html/2607.18348#bib.bib21)\)\. These findings locate*where*language\-specific structure appears; they do not measure how residual\-stream geometry changes from one layer to the next\.

### 2\.2Residual Streams, Readout, and Depth

Mechanistic interpretability treats the residual stream as the transformer’s main communication channel, updated additively by attention and MLP blocks\(Elhageet al\.,[2021](https://arxiv.org/html/2607.18348#bib.bib9); Olahet al\.,[2020](https://arxiv.org/html/2607.18348#bib.bib8)\)\. Layer\-wise readout methods such as the logit lens and tuned lens ask what each depth predicts about the next token\(nostalgebraist,[2020](https://arxiv.org/html/2607.18348#bib.bib14); Belroseet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib13)\)\. A complementary line of work debates the functional role of depth itself: late layers have been argued to mainly refine output probabilities\(Csordáset al\.,[2026](https://arxiv.org/html/2607.18348#bib.bib44)\), while other results emphasise the need for additional computational steps rather than additional parameters\(Saunshiet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib45)\)\. Related proposals allocate extra compute through looping or longer generation\(Dehghaniet al\.,[2018](https://arxiv.org/html/2607.18348#bib.bib43); Weiet al\.,[2022](https://arxiv.org/html/2607.18348#bib.bib48)\)\. Our metrics address a different question: not what a layer predicts, but how representations move between consecutive residual states\.

### 2\.3Mechanistic Interpretability and Reasoning Dynamics

Mechanistic studies have reverse\-engineered specific algorithmic operations in transformers, such as Fourier multiplication circuits during grokking\(Nandaet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib22)\)and the flow of operand information from early attention into late MLP computations\(Stolfoet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib23)\)\. More recent work adopts geometric and kinematic views of reasoning\. Logical reasoning has been modelled as smooth geometric flows whose structure shapes trajectory velocity and curvature\(Zhouet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib24)\)\.Fernando and Guitchounts \([2025](https://arxiv.org/html/2607.18348#bib.bib30)\)treat the residual stream as a dynamical system with distinct phases of kinematic acceleration, whileDamirchiet al\.\([2026](https://arxiv.org/html/2607.18348#bib.bib31)\)use discrete curvature to study model veracity, arguing that scalar token kinematics alone can be confounded by lexical effects\. These approaches motivate transition\-centred analysis, but they primarily track individual token paths rather than global cloud geometry across depth\.

### 2\.4Geometric Analysis of Transformer Representations

Other work studies the geometry of token representations more directly\.Viswanathanet al\.\([2025](https://arxiv.org/html/2607.18348#bib.bib27)\)characterise layer\-wise token clouds using intrinsic dimension and cosine similarity, relating spatial structure to prediction loss; related efforts model cross\-layer propagation with dynamical systems\(Shanget al\.,[2026](https://arxiv.org/html/2607.18348#bib.bib26)\)\. Directional consistency has also been measured across the sequence axis:Hosseiniet al\.\([2026](https://arxiv.org/html/2607.18348#bib.bib32)\)use cosine similarity of velocity vectors to quantify trajectory straightening during in\-context learning\. Much of this literature, however, treats each layer as a static snapshot\. Kinematic approaches track transitions, but usually for isolated tokens rather than for the full token cloud\. We instead use orthogonal Procrustes analysis\(Schönemann,[1966](https://arxiv.org/html/2607.18348#bib.bib16)\)to decompose each layer transition into an optimal rigid rotation and a non\-rigid residual\. This global, transition\-centred view is what distinguishes our framework\.

### 2\.5Synthesis

The multilingual literature maps where language\-specific structure appears; mechanistic and readout methods characterise what layers compute or predict; geometric and kinematic studies describe representation shape or token trajectories\. Fewer works ask how residual\-stream geometry is organised across layer transitions\. Our contribution is to measure that organisation directly: relative displacement and Procrustes residual are strongly depth\-structured, while rotation magnitude is nearly constant across depth\. In short, prior work explains where language structure lives and what layers do; we quantify how geometric change is allocated through the stack\.

## 3Methods

### 3\.1Setup

For each example, we record residual\-stream activations for the input prompt \(*prefill*\) and for autoregressively generated tokens \(*generation*\)\. If a model hasLLlayers and hidden dimensiondd, tokenttis represented by𝐡tℓ∈ℝd\\mathbf\{h\}\_\{t\}^\{\\ell\}\\in\\mathbb\{R\}^\{d\}at depthℓ\\ell\. Prefill yields one cloud of all prompt tokens after a single forward pass; generation yields the hidden state of the newly produced token at each decoding step\. Our analysis is transition\-centered: instead of treating each layer as a static snapshot, we measure how the token cloud is transformed from depthℓ\\elltoℓ\+1\\ell\+1\.

### 3\.2Geometric Metrics

We view each layer transition as a geometric transformation of the residual stream\. The metrics below separate how far tokens move, how sharply their trajectory bends, how much of the cloud change is explained by a global rotation, and how that rotation is distributed across planes\.

#### Relative displacement\.

We measure each token’s inter\-layer movement relative to the norm of its source representation:

δℓ​\(t\)=‖𝐡tℓ\+1−𝐡tℓ‖2‖𝐡tℓ‖2\.\\delta\_\{\\ell\}\(t\)=\\frac\{\\\|\\mathbf\{h\}\_\{t\}^\{\\ell\+1\}\-\\mathbf\{h\}\_\{t\}^\{\\ell\}\\\|\_\{2\}\}\{\\\|\\mathbf\{h\}\_\{t\}^\{\\ell\}\\\|\_\{2\}\}\.\(1\)This controls for growth in residual\-stream norms with depth, which is common in Pre\-LN models: a large absolute step may still be small relative to a growing residual vector\. Largerδℓ\\delta\_\{\\ell\}means a larger relative update at that transition\. Unless noted otherwise, we report the mean ofδℓ​\(t\)\\delta\_\{\\ell\}\(t\)over tokens within a prompt, then average across prompts and conditions\.

#### Curvature\.

Treating depth as a trajectory, let𝐯ℓ​\(t\)=𝐡tℓ\+1−𝐡tℓ\\mathbf\{v\}\_\{\\ell\}\(t\)=\\mathbf\{h\}\_\{t\}^\{\\ell\+1\}\-\\mathbf\{h\}\_\{t\}^\{\\ell\}be the update at transitionℓ\\ell\. Curvature compares consecutive updates:

κℓ​\(t\)=‖\(𝐡tℓ\+2−𝐡tℓ\+1\)−\(𝐡tℓ\+1−𝐡tℓ\)‖2‖𝐡tℓ\+1−𝐡tℓ‖22\.\\kappa\_\{\\ell\}\(t\)=\\frac\{\\\|\(\\mathbf\{h\}\_\{t\}^\{\\ell\+2\}\-\\mathbf\{h\}\_\{t\}^\{\\ell\+1\}\)\-\(\\mathbf\{h\}\_\{t\}^\{\\ell\+1\}\-\\mathbf\{h\}\_\{t\}^\{\\ell\}\)\\\|\_\{2\}\}\{\\\|\\mathbf\{h\}\_\{t\}^\{\\ell\+1\}\-\\mathbf\{h\}\_\{t\}^\{\\ell\}\\\|\_\{2\}^\{2\}\}\.\(2\)Large values indicate that the update changes strongly relative to its magnitude, whether by turning direction or by changing step size\. We use curvature only as a descriptive complement to displacement\.

#### Global rotation\.

At each transition, the tokens at depthsℓ\\ellandℓ\+1\\ell\+1form two point clouds inℝd\\mathbb\{R\}^\{d\}\. We center both clouds and use orthogonal Procrustes alignment\(Schönemann,[1966](https://arxiv.org/html/2607.18348#bib.bib16)\)to find the rotationRℓ∗R\_\{\\ell\}^\{\*\}that best maps the source cloud onto the target cloud\. We summarize the size of this rigid component by

ℛℓ=‖Rℓ∗−I‖F\.\\mathcal\{R\}\_\{\\ell\}=\\\|R\_\{\\ell\}^\{\*\}\-I\\\|\_\{F\}\.\(3\)A larger value means a larger global reorientation of the token cloud\. BecauseRℓ∗R\_\{\\ell\}^\{\*\}is orthogonal,ℛℓ\\mathcal\{R\}\_\{\\ell\}depends only on how far the fitted rotation lies from the identity, not on residual\-stream scale\.

#### Procrustes residual\.

A single global rotation does not explain every token’s movement\. After applyingRℓ∗R\_\{\\ell\}^\{\*\}, we compute the Euclidean distance between each aligned source token and its centered target, then average across tokens within each prompt and transition\. The residual is the part of the cloud change that cannot be absorbed by one shared rotation: anisotropic stretching, local rearrangements, and other non\-rigid effects\. We treat it as a measure of non\-rigid mismatch, not as a direct measure of computational effort\.

#### Rotation\-angle concentration\.

A high\-dimensional rotation may be spread across many planes or dominated by a few\. LetDℓ=Rℓ∗−ID\_\{\\ell\}=R\_\{\\ell\}^\{\*\}\-Iandq=⌊d/2⌋q=\\lfloor d/2\\rfloor\. We define

𝒞ℓ=‖Dℓ‖F2q​‖Dℓ‖22\.\\mathcal\{C\}\_\{\\ell\}=\\frac\{\\\|D\_\{\\ell\}\\\|\_\{F\}^\{2\}\}\{q\\\|D\_\{\\ell\}\\\|\_\{2\}^\{2\}\}\.\(4\)This is an efficient proxy derived fromRℓ∗−IR\_\{\\ell\}^\{\*\}\-I, not a full eigendecomposition of the rotation\. Larger𝒞ℓ\\mathcal\{C\}\_\{\\ell\}means rotational change is distributed across more planes; smaller values mean fewer planes dominate\. Together withℛℓ\\mathcal\{R\}\_\{\\ell\}and the Procrustes residual, it lets us ask whether constant rotation*size*co\-occurs with structured rotation*geometry*\.

### 3\.3Statistical Validation

We analyze models and phases separately using the previously computed prompt\-level metrics; activations are not recomputed\. For each prompt, token\-level values are averaged at every transition; prompts are then averaged within condition, with conditions weighted equally\. This yields one mean depth curve per model, phase, and metric\.

#### Depth modulation\.

For relative displacement, Procrustes residual, and angle concentration, we score depth structure as the standard deviation of the mean curve divided by its absolute mean\. A large score means some transitions are consistently larger than others\. We compare the observed score with2,0002\{,\}000null scores obtained by shuffling transition labels within each prompt, which preserves observed values but removes shared depth order\. Holm\-adjustedp<0\.05p<0\.05indicates significant depth modulation\.

#### Rotation constancy\.

For rotation magnitude, we ask whether the mean value across depth is essentially constant\. We measure this with relative peak\-to\-trough variation:

V=maxℓ⁡ℛ¯ℓ−minℓ⁡ℛ¯ℓ\|meanℓ⁡\(ℛ¯ℓ\)\|\.V=\\frac\{\\max\_\{\\ell\}\\bar\{\\mathcal\{R\}\}\_\{\\ell\}\-\\min\_\{\\ell\}\\bar\{\\mathcal\{R\}\}\_\{\\ell\}\}\{\|\\operatorname\{mean\}\_\{\\ell\}\(\\bar\{\\mathcal\{R\}\}\_\{\\ell\}\)\|\}\.\(5\)Before looking at the data, we defined rotation as practically constant whenV<0\.03V<0\.03\(less than3%3\\%variation across depth\)\.

We estimate uncertainty with2,0002\{,\}000hierarchical\-bootstrap samples, resampling conditions and then prompts within conditions\. A model passes the constancy test only if both of the following hold after Holm correction: \(i\) the one\-sided 95% upper bound onVVis below0\.030\.03, and \(ii\) the bootstrap probability thatV≥0\.03V\\geq 0\.03yields an adjustedpp\-value below0\.050\.05\. The same bootstrap procedure provides confidence intervals for the other reported effects\. Holm correction is applied separately within each metric family across model–phase comparisons\.

#### Dominant transitions and middle\-layer slowdown\.

For relative displacement, Procrustes residual, and angle concentration, we identify the transition with the largest mean value across depth and normalize its index to\[0,1\]\[0,1\], where 0 is the first transition and 1 is the last\. This puts models with different depths on a common scale\. Bootstrap samples yield a 95% interval for this normalized peak depth\.

For relative displacement, we additionally test whether values are lower in the middle third of depth than in the early and late thirds:

S=12​\(μearly\+μlate\)−μmiddle\|μ\|,S=\\frac\{\\tfrac\{1\}\{2\}\(\\mu\_\{\\mathrm\{early\}\}\+\\mu\_\{\\mathrm\{late\}\}\)\-\\mu\_\{\\mathrm\{middle\}\}\}\{\|\\mu\|\},\(6\)where thirds are defined on normalized depth andμ\\muis the mean across all transitions\. The numerator compares the outer thirds with the middle; dividing by\|μ\|\|\\mu\|makes the score scale\-free\. PositiveSStherefore indicates a middle\-layer slowdown\. Coding and translation are tested separately in prefill and generation with the same bootstrap and permutation procedures\. Holm correction is applied across model–task–phase comparisons\. Slowdown is supported when the adjustedpp\-value is below0\.050\.05and the 95% bootstrap interval lies above zero\.

## 4Experiments

### 4\.1Models

Six instruction\-tuned models across three families: Qwen3 \(0\.6B, 1\.7B, 4B, 8B\), Gemma\-2\-2B\-IT, and StableLM\-2\-1\.6B\-Chat\. All use Pre\-LN architectures\. This gives us variation in both family and parameter count111Implementation details[here](https://anonymous.4open.science/r/residual-stream-geometry-ED38/)\.

### 4\.2Tasks

#### Code generation\.

Problems from LiveCodeBench\(Jainet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib41)\), stratified by difficulty \(easy, medium, hard\)\. Models generate Python solutions with greedy decoding up to 512 tokens\. Monolingual \(English in, English out\)\.

#### Translation\.

100 sentence pairs per direction from Europarl v7\(Koehn,[2005](https://arxiv.org/html/2607.18348#bib.bib34)\): French, German, Czech, and Swedish paired with English \(eight directions\)\. Sentences under ten words filtered out\. Greedy decoding up to 128 tokens\. All four languages are Indo\-European; we note this as a typological limitation\.

### 4\.3Activation Capture

Forward hooks record the residual streamHℓ∈ℝT×dH\_\{\\ell\}\\in\\mathbb\{R\}^\{T\\times d\}after each layer’s attention and MLP\. For input: allTTprompt tokens from a single prefill pass\. For generation: the hidden state of the last generated token at each autoregressive step\. All runs use PyTorch MPS infloat16on Apple M\-series hardware\. Layer indices are normalised to\[0,1\]\[0,1\]for cross\-model comparison\.

## 5Results

### 5\.1Displacement

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/relative_displacement/translation_input.png)\(a\)Translation
![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/relative_displacement/coding_input.png)\(b\)Coding

Figure 1:Relative displacement during prefill\. Within each model, the depth curve is nearly identical across conditions: the eight language directions \(a\) overlap, and coding difficulties \(b\) share the same shape, with only peak height varying\.Relative displacement follows a model\-specific depth curve that is largely condition\-invariant\. In prefill, the eight translation directions collapse onto one curve per model \(Figure[1\(a\)](https://arxiv.org/html/2607.18348#S5.F1.sf1)\), and coding difficulties reproduce the same shape \(Figure[1\(b\)](https://arxiv.org/html/2607.18348#S5.F1.sf2)\)\. Qwen models are dominated by a sharp early peak; Gemma decays more smoothly before a final rise; StableLM shows multiple early peaks and a middle valley\. Within each model, the*shape*of change through depth is therefore more stable across conditions than across models\.

This structure is statistically robust\. Depth modulation of relative displacement is significant for all six models in both prefill and generation \(Holm\-adjustedp<0\.05p<0\.05; Figure[2](https://arxiv.org/html/2607.18348#S5.F2)\)\. Prefill modulation is large \(median97\.8%97\.8\\%; model range25\.2%25\.2\\%–166\.4%166\.4\\%\)\. Depth structure persists in generation, though with smaller relative variation across layers \(median23\.8%23\.8\\%;16\.9%16\.9\\%–34\.6%34\.6\\%\)\.

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/stats/depth_structure/depth_structure_validation.png)Figure 2:Depth structure of relative displacement \(top\) and rotation magnitude \(bottom\)\. Displacement is strongly modulated by depth; rotation magnitude stays below the predeclared3%3\\%constancy bound in every model and phase\.Dominant peaks during prefill fall in the early third for five of six models \(median normalized depth0\.090\.09\); Gemma is the exception, peaking at the final transition \(Figure[3](https://arxiv.org/html/2607.18348#S5.F3)\)\. During generation, peak location is less stable for several Qwen models, with bootstrap intervals that span much of the depth axis\.

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/stats/depth_structure/dominant_peak_layers.png)Figure 3:Normalized depth of the maximizing transition, with 95% bootstrap intervals\. Displacement peaks early for most models in prefill; Procrustes residual peaks at the final transition in every model and phase\.A second regular feature is a middle\-layer slowdown \(Figure[4](https://arxiv.org/html/2607.18348#S5.F4)\)\. Relative displacement is lower in the middle depth third than in the early and late thirds for translation in all six models \(prefill median45\.0%45\.0\\%; generation19\.3%19\.3\\%\)\. Coding shows the same sign but a smaller effect \(medians≈10%\\approx 10\\%\); five of six models pass the Holm\-corrected test, with Qwen3\-1\.7B the sole exception in both coding phases\. The early\-middle\-late rhythm is thus shared across the coding and translation conditions studied here, while its strength depends on condition and phase\.

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/stats/depth_structure/middle_layer_slowdown.png)Figure 4:Middle\-layer slowdown for relative displacement\. Positive values indicate lower displacement in the middle third\. Orange markers fail the Holm\-corrected test\.![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/relative_displacement/translation_generation.png)Figure 5:Relative displacement during translation generation\. Early and middle layers remain nearly condition\-invariant; final layers separate by target language, with larger updates for non\-English targets\.Generation makes the late\-layer condition effect visible \(Figure[5](https://arxiv.org/html/2607.18348#S5.F5)\)\. Early and middle layers still share one curve across directions, but final\-layer displacement splits by target language: non\-English targets show larger updates than English targets\. Coding generation preserves the same qualitative depth rhythm, with only modest difficulty\-related magnitude differences\. As shown next, the same English/non\-English separation appears in Procrustes residual\.

### 5\.2Rotation and Non\-Rigid Structure

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/rotation_magnitude/translation_input.png)Figure 6:Rotation magnitude during translation prefill\. Within each model the depth curve is nearly flat and overlaps across language directions; absolute scale grows with model size\.Table 1:Depth\-structure tests \(medians over models\)\. Modulation = SD/mean; rotation = peak\-to\-trough variation\. All comparisons pass after Holm correction \(max⁡pHolm=0\.006\\max p\_\{\\mathrm\{Holm\}\}=0\.006\)\.Rotation magnitude is nearly constant across depth \(Figure[6](https://arxiv.org/html/2607.18348#S5.F6); Table[1](https://arxiv.org/html/2607.18348#S5.T1)\)\. Peak\-to\-trough variation is below3%3\\%for every model and phase \(prefill median0\.68%0\.68\\%; generation0\.20%0\.20\\%\), and all twelve comparisons pass the practical\-constancy test \(Figure[2](https://arxiv.org/html/2607.18348#S5.F2), bottom\)\. Absolute magnitudes grow with hidden dimensiondd, consistent with the high\-dimensional scale‖R−I‖F≈2​d\\\|R\-I\\\|\_\{F\}\\approx\\sqrt\{2d\}\. We therefore do not interpret flatness alone as evidence of a learned computational invariant; in high dimensions, Procrustes alignments can yield stable magnitudes for geometric reasons\. The depth curve itself, however, remains nearly flat across conditions\.

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/stats/depth_structure/rotation_structure_validation.png)Figure 7:Depth structure beyond rotation magnitude\. Procrustes residual and angle concentration are both significantly modulated by depth, even though rotation magnitude itself is nearly constant\.![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/procrustes_distance/translation_generation.png)Figure 8:Procrustes residual during translation generation\. Residual stays low through early and middle layers, then rises sharply near the output and separates by target language\.Near\-constant magnitude does not imply unstructured geometry\. Procrustes residual is strongly depth\-modulated in all models \(prefill median101\.5%101\.5\\%; generation114\.7%114\.7\\%; Figure[7](https://arxiv.org/html/2607.18348#S5.F7)\)\. Its maximizing transition is the final layer for every model in both phases \(normalized depth1\.01\.0\)\. During translation generation, residual curves stay low through most of the stack and then rise sharply near the output \(Figure[8](https://arxiv.org/html/2607.18348#S5.F8)\), again separating by target language: non\-English targets leave a larger non\-rigid mismatch than English targets\. As in Methods, we treat residual as geometric mismatch after global alignment, not as computational effort\.

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/principal_angles/concentration/translation_input.png)Figure 9:Rotation\-angle concentration during translation prefill\. Most models form a broad mid\-depth plateau, with sharper changes in early and final transitions; language directions largely overlap\.Angle concentration is also depth\-structured, though more weakly \(Table[1](https://arxiv.org/html/2607.18348#S5.T1)\)\. Prefill peaks fall early to mid\-depth \(median0\.300\.30\); generation peaks are more mid\-depth \(median0\.470\.47\)\. The raw curves show a broad mid\-depth plateau with sharper changes near the beginning and end of the stack \(Figure[9](https://arxiv.org/html/2607.18348#S5.F9)\)\. Together these results separate two claims: the*size*of the global rotation is nearly depth\-invariant, while the*distribution*of that rotation across planes and the residual mismatch after alignment both vary systematically with depth\. The informative signal is this dissociation, not rotational flatness alone\.

### 5\.3Curvature

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/curvature/translation_input.png)

![Refer to caption](https://arxiv.org/html/2607.18348v1/plots/curvature/translation_generation.png)

Figure 10:Curvature for translation \(descriptive\)\. Prefill is spike\-dominated for Qwen/StableLM and smoother for Gemma; generation often shows an early\-high, late\-low decay\.Curvature is used only descriptively and is not included in our statistical tests\. Language directions again overlap within each model \(Figure[10](https://arxiv.org/html/2607.18348#S5.F10)\)\. Prefill splits by model family: Qwen and StableLM show high\-variance early spikes, whereas Gemma decays more smoothly\. In generation, several models approach an early\-high, late\-low decay; Qwen3\-4B and StableLM remain spike\-dominated\. Curvature is therefore a secondary complement to displacement, with more model\-to\-model variability\.

#### Summary\.

Treating residual streams as geometric transformations reveals nearly constant rotational scale together with strongly depth\-dependent displacement and residual: early repositioning, a quieter middle third, and rising late\-layer non\-rigid mismatch\. Across the coding and translation conditions studied here, depth curves are model\-dependent and largely condition\-stable; conditions mainly rescale late\-layer amplitude rather than rewrite the depth schedule\.

## 6Discussion

#### Scope of claims\.

This paper offers a measurement framework and descriptive regularities, not a mechanistic explanation\. The metrics quantify the geometry of residual\-stream transitions; they are not validated as measures of computational effort, reasoning difficulty, or circuit activity\. Our evidence shows that, for code generation and translation, depth curves are highly consistent across conditions within each model and differ across models\. We therefore describe the schedule as*model\-dependent*and*condition\-stable*, rather than as architecture\-caused: the experiments compare prompts within fixed models and do not manipulate architecture\. Likewise, near\-constant rotation magnitude should be interpreted cautiously\. In high dimensions, orthogonal Procrustes alignments often yield‖R−I‖F\\\|R\-I\\\|\_\{F\}values near the concentration scale2​d\\sqrt\{2d\}, so flatness may partly be a geometric baseline rather than a learned computational invariant\. What is informative is the dissociation: even when rotational scale is nearly constant, relative displacement, Procrustes residual, and angle concentration remain strongly depth\-structured\.

#### What the measurements show\.

Under the transformation view, residual\-stream change is not uniform across depth\. Relative displacement is typically larger early and late, with a quieter middle third; Procrustes residual rises sharply near the output; and rotation magnitude stays nearly flat\. Within each model, coding difficulties and translation directions largely share the same depth curve\. The clearest condition effect appears late in generation: non\-English targets show larger displacement and Procrustes residual than English targets\. This is consistent with English\-centric accounts of multilingual processing\(Wendleret al\.,[2024](https://arxiv.org/html/2607.18348#bib.bib17); Conneauet al\.,[2020](https://arxiv.org/html/2607.18348#bib.bib46); Ahujaet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib47)\), but it remains an observational association restricted to the language pairs studied here\. These measurements complement residual\-stream and readout work\. Mechanistic studies treat the residual stream as a communication channel\(Elhageet al\.,[2021](https://arxiv.org/html/2607.18348#bib.bib9); Olahet al\.,[2020](https://arxiv.org/html/2607.18348#bib.bib8)\); logit\-lens methods ask what each depth predicts\(nostalgebraist,[2020](https://arxiv.org/html/2607.18348#bib.bib14); Belroseet al\.,[2023](https://arxiv.org/html/2607.18348#bib.bib13)\)\. Our question is different: how representations move between layers\. The resulting early\-middle\-late pattern is compatible with late\-layer refinement of token probabilities\(Csordáset al\.,[2026](https://arxiv.org/html/2607.18348#bib.bib44)\)and with arguments that some problems need additional compute steps\(Saunshiet al\.,[2025](https://arxiv.org/html/2607.18348#bib.bib45)\), but compatibility is not explanation\.

#### Hypotheses for future work\.

Several non\-exclusive hypotheses are consistent with the measurements\. Early displacement peaks may reflect remapping from embeddings into a working residual space\. Middle\-layer slowdown may indicate gradual refinement with smaller relative updates\. Late residual growth may reflect preparation for unembedding, where one global rotation is insufficient\. The late English/non\-English gap may reflect uneven output\-space familiarity rather than a different depth schedule\. Testing these hypotheses will require interventions that lie outside the present study\.

#### Implications\.

Within the settings studied here, two practical points follow\. First, methods that allocate extra compute through longer generation or layer looping\(Dehghaniet al\.,[2018](https://arxiv.org/html/2607.18348#bib.bib43); Weiet al\.,[2022](https://arxiv.org/html/2607.18348#bib.bib48)\)can be checked against whether they preserve or disrupt this measured depth pattern\. Second, multilingual evaluation may benefit from layer\-resolved geometric observables: shared depth curves can coexist with larger late\-layer residual for non\-English generation\. Neither point assumes that residual equals computational cost\.

## 7Conclusion

We introduced a geometric view of residual streams: each layer transition is a transformation of the token cloud, measured by relative displacement and decomposed by orthogonal Procrustes analysis into rigid rotation and non\-rigid residual\. Applied to six models on code generation and translation, this view reveals reproducible depth regularities—early and late updates, a quieter middle third, nearly constant rotation magnitude, and rising late\-layer residual— that are largely shared across conditions within each model\. Content mainly rescales late\-layer amplitude in these settings, with non\-English generation showing larger final\-layer displacement and residual than English generation\. We present these findings as descriptive geometric measurements, not as explanations of computation\. The contribution is a framework for quantifying residual\-stream transitions and evidence that, for the tasks studied here, depth curves are model\-dependent and largely condition\-stable\.

## Limitations

We evaluate only code generation and translation, with Indo\-European languages paired with English; broader tasks and language families remain untested\.

The metrics describe geometric change and are not validated as measures of computational effort\. Near\-constant rotation magnitude may partly follow from high\-dimensional Procrustes geometry, so our claim emphasises the dissociation with structured displacement and residual rather than flatness alone\. Curvature is descriptive only\.

All models are instruction\-tuned Pre\-LN decoders under greedy decoding, and because we do not manipulate architecture, model\-dependent schedules should not be read as causal architectural effects\.

## References

- K\. Ahuja, H\. Diddee, R\. Hada, M\. Ochieng, K\. Ramesh, P\. Jain, A\. Nambi, T\. Ganu, S\. Segal, M\. Ahmed,et al\.\(2023\)Mega: multilingual evaluation of generative ai\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 4232–4267\.Cited by:[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. Steinhardt \(2023\)Eliciting latent predictions from transformers with the tuned lens\.arXiv preprint arXiv:2303\.08112\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- S\. Bhattacharya and O\. Bojar \(2023\)Unveiling multilinguality in transformer models: exploring language specificity in feed\-forward networks\.InProceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, S\. Hao, J\. Jumelet, N\. Kim, A\. McCarthy, and H\. Mohebbi \(Eds\.\),Singapore,pp\. 120–126\.External Links:[Link](https://aclanthology.org/2023.blackboxnlp-1.9/),[Document](https://dx.doi.org/10.18653/v1/2023.blackboxnlp-1.9)Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p2.2),[§2\.1](https://arxiv.org/html/2607.18348#S2.SS1.p1.1)\.
- L\. Bushnaq, S\. Heimersheim, N\. Goldowsky\-Dill, D\. Braun, J\. Mendel, K\. Hänni, A\. Griffin, J\. Stöhler, M\. Wache, and M\. Hobbhahn \(2024\)The local interaction basis: identifying computationally\-relevant and sparsely interacting features in neural networks\.arXiv preprint arXiv:2405\.10928\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2)\.
- T\. A\. Chang, Z\. Tu, and B\. K\. Bergen \(2022\)The geometry of multilingual language model representations\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 119–136\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.9/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.9)Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p2.2),[§2\.1](https://arxiv.org/html/2607.18348#S2.SS1.p1.1)\.
- A\. Conneau, S\. Wu, H\. Li, L\. Zettlemoyer, and V\. Stoyanov \(2020\)Emerging cross\-lingual structure in pretrained language models\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 6022–6034\.Cited by:[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- R\. Csordás, C\. D\. Manning, and C\. Potts \(2026\)Do language models use their depth efficiently?\.Advances in Neural Information Processing Systems38,pp\. 160313–160362\.Cited by:[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- H\. Damirchi, D\. la Jara, I\. Meza, E\. Abbasnejad, A\. Shamsi, Z\. Zhang, and J\. Shi \(2026\)Truth as a trajectory: what internal representations reveal about large language model reasoning\.arXiv preprint arXiv:2603\.01326\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.3](https://arxiv.org/html/2607.18348#S2.SS3.p1.1)\.
- M\. Dehghani, S\. Gouws, O\. Vinyals, J\. Uszkoreit, and Ł\. Kaiser \(2018\)Universal transformers\.arXiv preprint arXiv:1807\.03819\.Cited by:[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px4.p1.1)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly,et al\.\(2021\)A mathematical framework for transformer circuits\.Transformer Circuits Thread1\(1\),pp\. 12\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- J\. Fernando and G\. Guitchounts \(2025\)Transformer dynamics: a neuroscientific approach to interpretability of large language models\.arXiv preprint arXiv:2502\.12131\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.3](https://arxiv.org/html/2607.18348#S2.SS3.p1.1)\.
- E\. A\. Hosseini, Y\. Li, Y\. Bahri, D\. Campbell, and A\. K\. Lampinen \(2026\)Context structure reshapes the representational geometry of language models\.arXiv preprint arXiv:2601\.22364\.Cited by:[§2\.4](https://arxiv.org/html/2607.18348#S2.SS4.p1.1)\.
- N\. Jain, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. Stoica \(2025\)Livecodebench: holistic and contamination free evaluation of large language models for code\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 58791–58831\.Cited by:[§4\.2](https://arxiv.org/html/2607.18348#S4.SS2.SSS0.Px1.p1.1)\.
- P\. Koehn \(2005\)Europarl: a parallel corpus for statistical machine translation\.InProceedings of Machine Translation Summit X: Papers,Phuket, Thailand,pp\. 79–86\.External Links:[Link](https://aclanthology.org/2005.mtsummit-papers.11/)Cited by:[§4\.2](https://arxiv.org/html/2607.18348#S4.SS2.SSS0.Px2.p1.1)\.
- S\. Kornblith, M\. Norouzi, H\. Lee, and G\. Hinton \(2019\)Similarity of neural network representations revisited\.InInternational conference on machine learning,pp\. 3519–3529\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2)\.
- A\. Kyriakou, D\. Ulmer, and I\. Titov \(2026\)Shared doubt: zero\-shot cross\-lingual confidence estimation for language models\.arXiv preprint arXiv:2605\.31220\.Cited by:[§2\.1](https://arxiv.org/html/2607.18348#S2.SS1.p1.1)\.
- N\. Nanda, L\. Chan, T\. Lieberum, J\. Smith, and J\. Steinhardt \(2023\)Progress measures for grokking via mechanistic interpretability\.InThe Eleventh International Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2607.18348#S2.SS3.p1.1)\.
- nostalgebraist \(2020\)Interpreting GPT: the logit lens\.LessWrong\.Note:[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)External Links:[Link](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- C\. Olah, N\. Cammarata, L\. Schubert, G\. Goh, M\. Petrov, and S\. Carter \(2020\)An overview of early vision in inceptionv1\.Distill5\(4\),pp\. e00024–002\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- N\. Saunshi, N\. Dikkala, Z\. Li, S\. Kumar, and S\. J Reddi \(2025\)Reasoning with latent thoughts: on the power of looped transformers\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 14855–14881\.Cited by:[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- P\. H\. Schönemann \(1966\)A generalized solution of the orthogonal procrustes problem\.Psychometrika31\(1\),pp\. 1–10\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.4](https://arxiv.org/html/2607.18348#S2.SS4.p1.1),[§3\.2](https://arxiv.org/html/2607.18348#S3.SS2.SSS0.Px3.p1.4)\.
- L\. Schut, Y\. Gal, and S\. Farquhar \(2025\)Do multilingual llms think in english?\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,Cited by:[§2\.1](https://arxiv.org/html/2607.18348#S2.SS1.p1.1)\.
- J\. Shang, G\. Kreiman, and H\. Sompolinsky \(2026\)Unraveling the geometry of visual relational reasoning\.Scientific Reports\.Cited by:[§2\.4](https://arxiv.org/html/2607.18348#S2.SS4.p1.1)\.
- A\. Stolfo, Y\. Belinkov, and M\. Sachan \(2023\)A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 7035–7052\.Cited by:[§2\.3](https://arxiv.org/html/2607.18348#S2.SS3.p1.1)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin \(2017\)Attention is all you need\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2)\.
- K\. Viswanathan, Y\. Gardinazzi, G\. Panerai, A\. Cazzaniga, and M\. Biagetti \(2025\)The geometry of tokens in internal representations of large language models\.arXiv preprint arXiv:2501\.10573\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p1.2),[§2\.4](https://arxiv.org/html/2607.18348#S2.SS4.p1.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§2\.2](https://arxiv.org/html/2607.18348#S2.SS2.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px4.p1.1)\.
- C\. Wendler, V\. Veselovsky, G\. Monea, and R\. West \(2024\)Do llamas work in english? on the latent language of multilingual transformers\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15366–15394\.Cited by:[§1](https://arxiv.org/html/2607.18348#S1.p2.2),[§2\.1](https://arxiv.org/html/2607.18348#S2.SS1.p1.1),[§6](https://arxiv.org/html/2607.18348#S6.SS0.SSS0.Px2.p1.1)\.
- Y\. Zhou, Y\. Wang, X\. Yin, S\. Zhou, and A\. R\. Zhang \(2025\)The geometry of reasoning: flowing logics in representation space\.arXiv preprint arXiv:2510\.09782\.Cited by:[§2\.3](https://arxiv.org/html/2607.18348#S2.SS3.p1.1)\.

Similar Articles

DeepLoop: Depth Scaling for Looped Transformers

arXiv cs.LG

DeepLoop introduces a residual scaling method for looped Transformers that adjusts for parameter visits, improving stability and performance when physical blocks are reused across multiple rounds.