A Mesoscopic View of Transformer Weights Through Row and Column Scale Fields

arXiv cs.LG Papers

Summary

This paper introduces 'row and column scale fields'—median-centred log-RMS profiles of Transformer weight matrices over their channels—to mesoscopically analyze how weight magnitude is distributed across functional channels, showing how these fields evolve through training, align across projections, relate to AdamW's second-moment structure, and how edits to them affect model loss.

arXiv:2609.35852v1 Announce Type: new Abstract: Pooled statistics of Transformer weights obscure how magnitude is distributed across functional channels, while individual weights are too numerous to compare directly. We study the mesoscopic level between them: row and column scale fields, the median-centred log-RMS profiles of a weight matrix over its channels, which together with a global scale and a full balanced core represent the matrix exactly. Across public Pythia checkpoints at four sizes and controlled runs from three initialization families, balancing reveals similar measured core magnitude profiles. A mixture bridge, with its form fixed before the analysis and its coefficients fitted, predicts the pooled-shape departure from field width on held-out runs and data arms of the controlled grid. The indexed fields retain further structure: they align across projections that share a functional channel, and query/key profiles follow reassigned RoPE frequencies rather than fixed matrix coordinates. Training trajectories show early field formation followed by component-dependent broadening or recession. Extending the channel-based analysis to AdamW's second moment reveals related functional organization in its log-space row and column factors. Finally, edits of a frozen checkpoint separate reciprocal scale balance, which preserves the forward computation, from relative channel gain: flattening the gain increases in-distribution loss while preserving matrix norms and the balanced core. Row and column scale fields thus connect pooled magnitude statistics to channel organization and provide coordinates for tracking and testing trained weight structure.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:37 AM

# A Mesoscopic View of Transformer WeightsThrough Row and Column Scale Fields
Source: [https://arxiv.org/html/2609.35852](https://arxiv.org/html/2609.35852)
###### Abstract

Pooled statistics of Transformer weights obscure how magnitude is distributed across functional channels, while individual weights are too numerous to compare directly\. We study the mesoscopic level between them: row and column scale fields, the median\-centred log\-RMS profiles of a weight matrix over its channels, which together with a global scale and a full balanced core represent the matrix exactly\. Across public Pythia checkpoints at four sizes and controlled runs from three initialization families, balancing reveals similar measured core magnitude profiles\. A mixture bridge, with its form fixed before the analysis and its coefficients fitted, predicts the pooled\-shape departure from field width on held\-out runs and data arms of the controlled grid\. The indexed fields retain further structure: they align across projections that share a functional channel, and query/key profiles follow reassigned RoPE frequencies rather than fixed matrix coordinates\. Training trajectories show early field formation followed by component\-dependent broadening or recession\. Extending the channel\-based analysis to AdamW’s second moment reveals related functional organization in its log\-space row and column factors\. Finally, edits of a frozen checkpoint separate reciprocal scale balance, which preserves the forward computation, from relative channel gain: flattening the gain increases in\-distribution loss while preserving matrix norms and the balanced core\. Row and column scale fields thus connect pooled magnitude statistics to channel organization and provide coordinates for tracking and testing trained weight structure\.

## 1Introduction

Transformer weights can be inspected entry by entry or summarized at the matrix level, but these views answer different questions\. Norms, spectral statistics and fitted magnitude distributions support comparison across matrices, yet do not directly identify which functional channels carry scale variation\[[Martin and Mahoney, 2021](https://arxiv.org/html/2609.35852#bib.bib20),[Ding, 2026a](https://arxiv.org/html/2609.35852#bib.bib6)\]\. In particular, pooling magnitudes can mix variation within channels with differences in scale across channels\. We study the intermediate level at which those differences can be measured, aligned and followed through training\.

Our objects are the row and column scale fields: the median\-centred log\-RMS profiles of a weight matrix over its output and input channels\. We relate these direct measurements to the balanced representationW=s​Dr​Z​DcW=sD\_\{r\}ZD\_\{c\}, wheres=RMS⁡\(W\)s=\\mathrm\{RMS\}\(W\)and every row and column ofZZhas unit RMS\. The complete representations are interconvertible when the full signed core is retained \(Section[3\.1](https://arxiv.org/html/2609.35852#S3.SS1)\)\. For statistical analysis, the fields record channel\-scale heterogeneity and arrangement, while the magnitude profile ofZZdescribes what remains after balancing\. The fitted Weibull parameters serve as pooled read\-outs rather than assumptions about every entry\. Figure[1](https://arxiv.org/html/2609.35852#S1.F1)locates these objects in the decoder and its training loop\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figures/P5v3_F1_decoder_training_loop_v8_260925.png)Figure 1:A mesoscopic view of Transformer weights through row and column scale fields\.The training loop produces a sequence of projection matrices\. At selected checkpoints, each matrix is represented by its global RMS, balancing factors and full balanced core\. Direct row and column log\-RMS fields are linked to these factors by the conversions in Section[3\.1](https://arxiv.org/html/2609.35852#S3.SS1)\. Field widths, core magnitude profiles and fitted pooled parameters are statistical read\-outs of this representation; the fitted scaleλ\\lambdatracks the global RMSssthrough the measured ratioλ/s\\lambda/s\. Coloured strips mark the empirically dominant field sides; an unmarked side need not be homogeneous\. The dashed analysis path is not part of the training update\.This framework asks what remains after channel balancing, how the fields account for pooled statistics and functional organization, and which changes in the fields affect the computation\. We address these questions with controlled LLaMA\-style\[[Touvron et al\., 2023](https://arxiv.org/html/2609.35852#bib.bib30)\]70M runs, initialization\-family and RoPE\-reassignment experiments, Pythia size comparisons\[[Biderman et al\., 2023](https://arxiv.org/html/2609.35852#bib.bib3)\], optimizer\-state records and frozen\-checkpoint edits\. Their roles and replication levels differ and are specified in Table[3](https://arxiv.org/html/2609.35852#A1.T3)\.

##### Contributions\.

1. 1\.We connect direct channel\-scale measurements to balanced\-core analysis\. Across the tested settings, balancing reveals similar measured core magnitude profiles\. A calibrated mixture bridge predicts pooled\-shape departure from field width on held\-out runs and data arms within the controlled grid, while retaining the distinction between complete fields and their scalar summaries\.
2. 2\.We identify functional and temporal organization in the fields\. Profiles match across projections sharing a channel space, and a paired RoPE reassignment movesq/kq/kprofiles with their assigned frequencies\. Training trajectories reveal broad early formation followed by component\-dependent continuation or recession, which terminal pooled statistics alone do not show\.
3. 3\.We extend the channel\-based analysis to AdamW’s second moment and test the functional role of paired weight fields\. Log\-space optimizer factors align with functional channel spaces\. Frozen\-checkpoint edits separate an invariant reciprocal balance from a loss\-sensitive relative gain: flattening the gain increases loss at fixed matrix norms and balanced core, and exceeds the tested Gaussian controls after displacement normalization\.

## 2Related work

##### Weight magnitude, direction, and channel scales\.

Weight normalization separates the length and the direction of a weight vector\[[Salimans and Kingma, 2016](https://arxiv.org/html/2609.35852#bib.bib23)\], and later work relates weight norms to effective learning rates, decoupled weight decay, and rotational equilibrium, down to the single neuron\[[van Laarhoven, 2017](https://arxiv.org/html/2609.35852#bib.bib31),[Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.35852#bib.bib18),[Wan et al\., 2021](https://arxiv.org/html/2609.35852#bib.bib33),[Kosson et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib16)\]\. Recent language\-model studies make channel magnitudes explicit optimization variables:[Velikanov et al\. \[2026\]](https://arxiv.org/html/2609.35852#bib.bib32)attach learnable per\-row and per\-column multipliers to free the weight\-decay–noise equilibrium norm and report that they broaden the distribution of row norms in the attention\-input and gate projections, and[Hägele et al\. \[2026\]](https://arxiv.org/html/2609.35852#bib.bib10)factorize each weight into a fixed\-norm direction and learnable per\-row and per\-column gains updated at separate rates\. These methods make channel magnitudes explicit training variables\. We instead measure the row and column scale structure that emerges in ordinarily trained checkpoints and relate it to pooled statistics, channel identities and controlled edits\.

##### Weight\-side and channel\-level diagnostics\.

Trained weights have been read as a record of training through spectral statistics\[[Martin and Mahoney, 2021](https://arxiv.org/html/2609.35852#bib.bib20)\]and through the training\-data mixture that can be recovered from fine\-tuned weights\[[Huang et al\., 2026](https://arxiv.org/html/2609.35852#bib.bib12)\]\. At the channel level,[Chen et al\. \[2024\]](https://arxiv.org/html/2609.35852#bib.bib5)use the variance of channel weight norms to assess layer width and to identify stages of training\. Our analysis represents that heterogeneity by complete row and column scale fields, together with the core obtained by two\-sided RMS balancing\.

##### Row–column optimizer geometry\.

Row and column structure also appears in optimizer design, from Adafactor’s row–column reconstruction of the second moment\[[Shazeer and Stern, 2018](https://arxiv.org/html/2609.35852#bib.bib26)\]to the axis\-wise preconditioners of K\-FAC and Shampoo\[[Martens and Grosse, 2015](https://arxiv.org/html/2609.35852#bib.bib19),[Gupta et al\., 2018](https://arxiv.org/html/2609.35852#bib.bib9)\], which are different decompositions; coordinatewise normalization also brings Adam’s update close to a smoothed sign direction\[[Balles and Hennig, 2018](https://arxiv.org/html/2609.35852#bib.bib1),[Kunstner et al\., 2023](https://arxiv.org/html/2609.35852#bib.bib17)\]\. Our analysis of AdamW’s second moment is a descriptive fit to the existing state, asking whether its row and column factors line up with the channel spaces found in trained weights\.

##### Transformer channel identities and shared paths\.

The architecture gives many rows and columns an identifiable role\. RoPE assigns a fixed frequency to each coordinate pair of the query and key projections\[[Su et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib29)\], and these frequencies are used non\-uniformly across heads and dependency ranges\[[Barbero et al\., 2025](https://arxiv.org/html/2609.35852#bib.bib2),[Hong et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib11)\], while context\-extension methods rescale them according to wavelength\[[Peng et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib22)\]; recent work relates frequency usage to the training data\[[Wu et al\., 2026](https://arxiv.org/html/2609.35852#bib.bib34)\]and makes the rotation frequencies themselves learnable\[[Karypis et al\., 2026](https://arxiv.org/html/2609.35852#bib.bib13)\]\. Matrix\-level analyses of the paired projections study the products rather than the individual factors:[Saponati et al\. \[2025\]](https://arxiv.org/html/2609.35852#bib.bib24)find that autoregressive training induces high\-norm columns in the query–key product \(in their token\-row convention\), and[Kobayashi et al\. \[2024\]](https://arxiv.org/html/2609.35852#bib.bib15)show that at stationary points of the weight\-decay\-regularized loss the two factors of a bilinear product have equal Gram matrices, which makes the penalty equivalent to a nuclear\-norm penalty and lowers the rank of the query–key and value–output products, and find this balance approximately satisfied, row by row, in pretrained Llama 2\. Building on these observations, we ask how the row and column scale fields of each individual projection follow architecture\-defined channel identities, how fields pair across matrices that share a channel space, and how a controlled reassignment of RoPE frequencies moves the learned scale profile\.

##### Relation to our earlier studies\.

Our earlier work introduced a protocol\-matched Weibull description of pooled weight magnitudes, decomposed AdamW scale growth, and related the resulting scale coordinate to pre\-training data statistics\[[Ding, 2026a](https://arxiv.org/html/2609.35852#bib.bib6),[Ding, 2026b](https://arxiv.org/html/2609.35852#bib.bib7),[Ding, 2026c](https://arxiv.org/html/2609.35852#bib.bib8)\]\. The effort\-versus\-quality boundary discussed in Section[6](https://arxiv.org/html/2609.35852#S6)is from a companion study that is under review, referred to below as the companion study\. The present study moves from pooled matrix coordinates to the intermediate row–column structure of the weights, using the pooled shape as one macroscopic read\-out while treating the scale fields and the core as the primary objects\.

## 3Method: Row–Column Scale Fields

Table[2](https://arxiv.org/html/2609.35852#A1.T2)in Appendix[A\.1](https://arxiv.org/html/2609.35852#A1.SS1)lists the notation used throughout the paper, with the equation or section that defines each symbol\.

### 3\.1Weight\-matrix representations and conversions

We use three complete representations of each weight matrixW∈ℝdout×dinW\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\},

W⟷\(s,Dr,Z,Dc\)⟷\(s,hr,hc,Z\),W\\;\\longleftrightarrow\\;\(s,\\,D\_\{r\},\\,Z,\\,D\_\{c\}\)\\;\\longleftrightarrow\\;\(s,\\,h^\{r\},\\,h^\{c\},\\,Z\),\(1\)built from the following quantities\.

Global scale, balancing factors and core\.The exact re\-parameterization

W=s​Dr​Z​Dc,s=RMS⁡\(W\),W=s\\,D\_\{r\}\\,Z\\,D\_\{c\},\\qquad s=\\mathrm\{RMS\}\(W\),\(2\)separates the global scaless, the positive diagonal balancing factorsDrD\_\{r\}andDcD\_\{c\}, and the balanced coreZZ,

RMSj​\(Zi​j\)=1​for every row​i,RMSi​\(Zi​j\)=1​for every column​j,\\mathrm\{RMS\}\_\{j\}\(Z\_\{ij\}\)=1\\;\\text\{ for every row \}i,\\qquad\\mathrm\{RMS\}\_\{i\}\(Z\_\{ij\}\)=1\\;\\text\{ for every column \}j,\(3\)obtained by alternating row and column RMS balancing of the Sinkhorn–Knopp type\[[Sinkhorn and Knopp, 1967](https://arxiv.org/html/2609.35852#bib.bib28),[Sinkhorn, 1967](https://arxiv.org/html/2609.35852#bib.bib27)\]\.

Scale fields\.With the channel RMSRi=RMSj​\(Wi​j\)R\_\{i\}=\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\)andCj=RMSi​\(Wi​j\)C\_\{j\}=\\mathrm\{RMS\}\_\{i\}\(W\_\{ij\}\), the row and column scale fields are their logarithms centred by the median over channels \(the centring changes neither the widths nor the profile correlations\),

hir=log⁡Ri−mediani⁡log⁡Ri,hjc=log⁡Cj−medianj⁡log⁡Cj\.h^\{r\}\_\{i\}=\\log R\_\{i\}\-\\operatorname\{median\}\_\{i\}\\log R\_\{i\},\\qquad h^\{c\}\_\{j\}=\\log C\_\{j\}\-\\operatorname\{median\}\_\{j\}\\log C\_\{j\}\.\(4\)The field widthHWH^\{W\}is the standard deviation of the field on a kind’s identity side, marked in Figure[1](https://arxiv.org/html/2609.35852#S1.F1)\.

Conversions\.Givenss, the fields determine the channel RMS: since the mean ofRi2R\_\{i\}^\{2\}over rows and the mean ofCj2C\_\{j\}^\{2\}over columns both equals2s^\{2\},

Ri=s​ehir\(dout−1​∑i′e2​hi′r\)1/2,Cj=s​ehjc\(din−1​∑j′e2​hj′c\)1/2\.R\_\{i\}=\\frac\{s\\,e^\{h^\{r\}\_\{i\}\}\}\{\\big\(d\_\{\\mathrm\{out\}\}^\{\-1\}\\textstyle\\sum\_\{i^\{\\prime\}\}e^\{2h^\{r\}\_\{i^\{\\prime\}\}\}\\big\)^\{1/2\}\},\\qquad C\_\{j\}=\\frac\{s\\,e^\{h^\{c\}\_\{j\}\}\}\{\\big\(d\_\{\\mathrm\{in\}\}^\{\-1\}\\textstyle\\sum\_\{j^\{\\prime\}\}e^\{2h^\{c\}\_\{j^\{\\prime\}\}\}\\big\)^\{1/2\}\}\.\(5\)The fields and the balancing factors describe the same channel scales but are not entrywise equal: each marginal RMS also depends on theZ2Z^\{2\}\-weighted squared factors of the opposite side\. With the core retained as a full signed matrix, positive diagonal scalings ofZZthat meet these row and column RMS targets recover the balancing factors, up to their reciprocal constant, and henceWW; Appendix[A\.2](https://arxiv.org/html/2609.35852#A1.SS2)gives the coupling between factors and fields, the iteration and its uniqueness\. These conversions hold for the complete representations; the field widths and pooled shapes used below are summaries of them, and the relation between those summaries is tested by the mixture bridge\.

##### Global scale and fitted scale\.

Letλ𝒫\\lambda\_\{\\mathcal\{P\}\}andk𝒫k\_\{\\mathcal\{P\}\}denote the scale and shape returned by the fitting protocol𝒫\\mathcal\{P\}\(Appendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)\)\. The ideal probability\-plot fit is scale equivariant,

λ𝒫​\(W\)=s​λ𝒫​\(W/s\),k𝒫​\(W\)=k𝒫​\(W/s\),\\lambda\_\{\\mathcal\{P\}\}\(W\)=s\\,\\lambda\_\{\\mathcal\{P\}\}\(W/s\),\\qquad k\_\{\\mathcal\{P\}\}\(W\)=k\_\{\\mathcal\{P\}\}\(W/s\),\(6\)and with the fixed histogram grid used here the relation holds approximately: refittingW/sW/sreproducesλ/s\\lambda/swithin 0\.2% andkkwithin 0\.004 \(Appendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4.SSS0.Px2)\)\. Every shape read\-out therefore concernsW/sW/s, andλ\\lambdaissstimes the dimensionless ratioλ/s\\lambda/s\. For an exact Weibull magnitude distributions2=λ2​Γ​\(1\+2/k\)s^\{2\}=\\lambda^\{2\}\\,\\Gamma\(1\+2/k\), so

λ/s=Γ\(1\+2/k\)−1/2,\\lambda/s=\\Gamma\(1\+2/k\)^\{\-1/2\},\(7\)which varies by only 4% betweenk=1\.17k=1\.17and1\.251\.25\(the protocol ratio itself is measured, Appendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4.SSS0.Px2)\)\. The fitted scale is therefore a read\-out of the global scale: the growth ofλ\\lambdaunder AdamW and its dependence on the training data studied earlier\[[Ding, 2026b](https://arxiv.org/html/2609.35852#bib.bib7),[Ding, 2026c](https://arxiv.org/html/2609.35852#bib.bib8)\]are growth ofss, linked toλ\\lambdathere through eq\. \([7](https://arxiv.org/html/2609.35852#S3.E7)\) at locked shape; eq\. \([6](https://arxiv.org/html/2609.35852#S3.E6)\) does not assume the Weibull form\.

### 3\.2Mixture bridge

On a kind’s identity side the field is the centred log\-scale of the one\-sided decomposition, shown here for rows,

W=diag⁡\(R\)​Zr,RMSj​\(Zi​jr\)=1​for every row​i,W=\\mathrm\{diag\}\(R\)\\,Z^\{r\},\\qquad\\mathrm\{RMS\}\_\{j\}\(Z^\{r\}\_\{ij\}\)=1\\;\\text\{ for every row \}i,\(8\)withRi=RMSj​\(Wi​j\)R\_\{i\}=\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\)\. For row\-identity kinds,kident=krowk\_\{\\mathrm\{ident\}\}=k\_\{\\mathrm\{row\}\}is the fitted shape ofZrZ^\{r\}; for column\-identity kinds,kident=kcolk\_\{\\mathrm\{ident\}\}=k\_\{\\mathrm\{col\}\}is obtained from the corresponding column\-normalized core\. Under approximate independence of the identity\-axis field and its one\-sided core the log\-variances of the two add \(Appendix[A\.5](https://arxiv.org/html/2609.35852#A1.SS5)\), which motivates the one\-axis relation, with its form fixed before the analysis,

Δkmix≡kraw−2−kident−2≈αP​\(HW\)2\.\\Delta\_\{k\}^\{\\mathrm\{mix\}\}\\equiv k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{ident\}\}^\{\-2\}\\approx\\alpha\_\{P\}\(H^\{W\}\)^\{2\}\.\(9\)The exact\-Weibull value6/π2≈0\.6086/\\pi^\{2\}\\approx 0\.608\(Appendix[A\.5](https://arxiv.org/html/2609.35852#A1.SS5)\) serves as a reference, and the protocol coefficientαP\\alpha\_\{P\}is estimated empirically\. We also test the two\-axis extension

kraw−2−kbi−2≈αr​\(HrowW\)2\+αc​\(HcolW\)2,k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{bi\}\}^\{\-2\}\\approx\\alpha\_\{r\}\(H^\{W\}\_\{\\mathrm\{row\}\}\)^\{2\}\+\\alpha\_\{c\}\(H^\{W\}\_\{\\mathrm\{col\}\}\)^\{2\},\(10\)an empirical relation, since each direct marginal also carries the opposite factors \(eq\. \([12](https://arxiv.org/html/2609.35852#A1.E12)\)\)\. Both models are fitted through the origin and evaluated by leave\-one\-run\-out and leave\-one\-data\-arm\-out prediction, in which the field width and the normalized\-core shape of each held\-out matrix are measured inputs, so the bridge is a conditional prediction of the pooled shape; identifiability checks are in Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\.

### 3\.3Optimizer\-state row–column factors

Section[4\.5](https://arxiv.org/html/2609.35852#S4.SS5)reads the same two axes off AdamW’s bias\-corrected second momentV^\\hat\{V\}\(Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5), eq\. \([16](https://arxiv.org/html/2609.35852#A2.E16)\)\), taken from the dumped state at a checkpoint\. We fit

log⁡V^i​j=μ\+ai\+bj\+εi​j,∑iai=∑jbj=0,\\log\\hat\{V\}\_\{ij\}=\\mu\+a\_\{i\}\+b\_\{j\}\+\\varepsilon\_\{ij\},\\qquad\\textstyle\\sum\_\{i\}a\_\{i\}=\\sum\_\{j\}b\_\{j\}=0,\(11\)whose least\-squares factors are the centred row and column means oflog⁡V^\\log\\hat\{V\}, withε\\varepsilonthe coordinate\-specific remainder\. We report the variance explained by the joint model and by each factor, and correlate factors that index the same channel space within a layer \(median over layers\)\. The factorization is descriptive; the update rule, the motivation of the model and the fits are in Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\.

### 3\.4Evidence design

Table[3](https://arxiv.org/html/2609.35852#A1.T3)in Appendix[A\.3](https://arxiv.org/html/2609.35852#A1.SS3)lists each evidence source with its protocol and its role\. The evidence falls into three classes: controlled multi\-seed comparisons, the paired RoPE reassignment with three seed pairs, and single\-seed state, boundary and edit analyses, which are descriptive\. The statistical unit is the training run; matrix fits within a run are not independent replications\.

## 4Results

The results show how row and column scale fields connect individual weights to pooled matrix statistics\. After two\-sided balancing, the tested components approach a common magnitude\-profile core; their pooled differences arise mainly from how scale is distributed over channels\. The fields follow architecture\-defined identities, evolve differently across components, and have a corresponding topology in AdamW’s second\-moment state\. They therefore provide a mesoscopic description of trained Transformer weights\.

### 4\.1Two\-sided balancing exposes a common core

What remains after the row and column scale fields are removed? At the public Pythia endpoints, the pooled component shapes differ, most visibly forqqandkk\. Normalizing each matrix along its architecture\-defined identity axis places the block medians ofqq,kk,ooand both FFN projections within 1\.193–1\.205 across all four model sizes\. Thevvprojection is the exception under one\-sided normalization and enters the band only after both axes are balanced \(Figure[2](https://arxiv.org/html/2609.35852#S4.F2)\(a\)\); this exception anticipates the two\-sided channel organization ofvvexamined in Section[4\.3](https://arxiv.org/html/2609.35852#S4.SS3)\. The controlled grid gives the corresponding trajectory result: after two\-sided balancing, the central\-body shape of all seven component kinds remains in component medianskbi=1\.203k\_\{\\mathrm\{bi\}\}=1\.203–1\.2121\.212at every checkpoint, while individual matrices spread wider \(Appendix[B\.1](https://arxiv.org/html/2609.35852#A2.SS1); per\-kind values in the data package\)\. Most of the component separation visible in the pooled distributions is therefore carried by the channel\-scale fields rather than by the balanced core\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_F1_normalized_core_v5.png)Figure 2:Two\-sided balancing reveals a common core across components, model sizes and initialization families\.\(a\)Final public Pythia checkpoints; grey connectors pair the raw and identity\-axis\-normalized fits of the same component kind \(filled and open markers\), stars mark two\-sided balancing ofvv, and coloured bars span the normalized block medians\.\(b\)Each quantile of the unit\-RMS core\|Z\|\|Z\|divided by the corresponding quantile of the unit\-RMS half\-normal \(the magnitude of a Gaussian weight\), minus one, for one initialization family at one checkpoint\. At 30,000 steps each family lies within 0\.26% of it throughout the displayed range \(median pooled over the three families: 0\.19%\)\.\(c\)Two\-sided core shape in nine LLaMA\-style 70M runs\. The horizontal line in \(a\) and \(c\) is the reference shape of a Gaussian weight under the middle\-80% protocol,k≈1\.20k\\approx 1\.20, derived for a half\-normal magnitude in[Ding \[2026a\]](https://arxiv.org/html/2609.35852#bib.bib6)\.The full core profiles give a stronger view of this convergence\. Gaussian, Laplace and uniform initializations begin with clearly different magnitude profiles, but their two\-sided cores approach the same protocol\-matched reference across the fitted body and through the displayed 0\.99 quantile \(Figure[2](https://arxiv.org/html/2609.35852#S4.F2)\(b\)\)\. The reference is the magnitude profile of a Gaussian weight, on which the Gaussian family starts, so the substantive evidence comes from the Laplace and uniform families: they start about 40% away in opposite directions and converge to the same profile\. Training therefore brings distinct initial core distributions into a common measured body shape\.

Figure[2](https://arxiv.org/html/2609.35852#S4.F2)\(c\) shows how this state is reached\. The Gaussian core stays near the reference, while the Laplace and uniform cores move progressively toward it and join it within the first 10,000 steps\. This motion also shows that the common core is not imposed algebraically by the balancing procedure: the balanced cores begin differently and converge during training\. Scale\-field formation proceeds at the same time\. By step 800 every identity\-axis field is already several times its initialization width \(3\.1–4\.2 times in the Gaussian runs\), while the non\-Gaussian cores are still merging \(Figure[13](https://arxiv.org/html/2609.35852#A2.F13), middle and bottom rows, Appendix[B\.4](https://arxiv.org/html/2609.35852#A2.SS4)\)\. Core convergence and field formation are therefore concurrent rather than successive\. Sections[4\.2](https://arxiv.org/html/2609.35852#S4.SS2)–[4\.4](https://arxiv.org/html/2609.35852#S4.SS4)follow these fields as they shape the pooled statistics, organize over channels and separate into distinct training trajectories\.

### 4\.2Scale\-field width predicts pooled\-shape departure

Section[4\.1](https://arxiv.org/html/2609.35852#S4.SS1)read the representation in reverse: removing the row and column scale fields exposed a common balanced core\. Here we read it forward\. Pooling channels that carry similar local core profiles at different scales moves the pooled shape away from that core, and the mixture bridge of Section[3\.2](https://arxiv.org/html/2609.35852#S3.SS2)quantifies this map back to the observed pooled statistic: the one\-axis form starts from the one\-sided core shapekidentk\_\{\\mathrm\{ident\}\}of eq\. \([8](https://arxiv.org/html/2609.35852#S3.E8)\), the two\-axis form from the balanced core shapekbik\_\{\\mathrm\{bi\}\}\.

The effect is visible directly in Figure[3](https://arxiv.org/html/2609.35852#S4.F3)\(a\): as the row\-field width grows,krawk\_\{\\mathrm\{raw\}\}departs from the reference whilekrowk\_\{\\mathrm\{row\}\}stays close to the core value\. On the controlled grid the relation \([9](https://arxiv.org/html/2609.35852#S3.E9)\) givesαP=0\.791\\alpha\_\{P\}=0\.791withR02=0\.976R\_\{0\}^\{2\}=0\.976for pooledq/kq/k\(Figure[3](https://arxiv.org/html/2609.35852#S4.F3)\(b\)\), whereasvv, gate and up sit at 0\.59–0\.64 with mean 0\.606, close to the exact\-Weibull value6/π2=0\.6086/\\pi^\{2\}=0\.608; theq/kq/kexcess is associated with a scale–shape pairing term \(Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\)\. The same direction holds in all 12 runs for each of the five row\-side kinds, with per\-kind coefficients in Figure[9](https://arxiv.org/html/2609.35852#A2.F9); the column\-side kindsooand down enter through the two\-axis form below\. Refitted with one run or one whole data arm held out, the bridge predicts the held\-outkrawk\_\{\\mathrm\{raw\}\}with a median error of 0\.2–0\.6% by kind; the worst held\-out fold reaches about 1\.6%, in the shuffled high\-repetition arm D3 \(Figure[3](https://arxiv.org/html/2609.35852#S4.F3)\(d\)\)\. The relation is therefore not confined to the fitted trajectories\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_F2_scale_field_bridge_v6.png)Figure 3:Scale\-field mixing predicts the pooled\-shape departure on the controlled grid\.Panels \(a\), \(b\) and \(d\) use the 1,440 row\-identity fits \(4 data arms×\\times3 seeds×\\times4 checkpoints×\\times5 kinds×\\times6 layers\) and panel \(c\) the two\-axis fits of all seven kinds on the same runs; colours and markers identify component kinds throughout\. Panel \(c\) uses one pooled two\-axis fit over the seven kinds\. Panel \(d\) scores the one\-axis prediction ofkrawk\_\{\\mathrm\{raw\}\}with per\-kind coefficients: bars give the in\-sample error and the medians over held\-out runs and held\-out arms of the fold\-mean relative error, and marks show the worst fold; the field width and normalized\-core shape of each held\-out matrix are measured inputs\. Absolute leave\-one\-run\-out errors of the one\- and two\-axis models are in Figure[9](https://arxiv.org/html/2609.35852#A2.F9)\(c\)\.Relative to the two\-sided core, adding both field widths extends the bridge to all seven component kinds \(Figure[3](https://arxiv.org/html/2609.35852#S4.F3)\(c\)\): the row\-plus\-column model \([10](https://arxiv.org/html/2609.35852#S3.E10)\) givesR02=0\.904R\_\{0\}^\{2\}=0\.904against 0\.803 for the row\-only model, with the gain concentrated invv,ooand down, which connects the two\-sided restoration ofvvin Figure[2](https://arxiv.org/html/2609.35852#S4.F2)\(a\) to its measured row and column fields\. This extends the bridge to the column\-identity kinds, with pooled coefficients of 0\.67 for rows and 0\.46 for columns\. The bridge retains the width of the field but not the arrangement of scales over channels \(Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\); that information is examined in Section[4\.3](https://arxiv.org/html/2609.35852#S4.SS3)\.

### 4\.3Architecture assigns the coordinates of scale fields

Figure[1](https://arxiv.org/html/2609.35852#S1.F1)distinguishes two levels of structural organization\. The computation graph determines whether a functional identity is indexed by matrix rows or columns, and architecture\-specific coordinates such as RoPE determine how that identity is arranged within the selected axis\. We test both levels through normalization, cross\-projection profile matching, and a paired RoPE reassignment\.

#### 4\.3\.1Functional identities map to matrix sides

Figure[1](https://arxiv.org/html/2609.35852#S1.F1)gives the structural map for the scale fields\. For a projectiony=W​xy=Wx, rows index output channels and columns index input channels, and the gradient contributed by each token has the outer\-product formδ​x⊤\\delta x^\{\\top\}\(Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\. A functional identity therefore appears on the row side of the matrix that produces it and on the column side of the matrix that consumes it \(Table[1](https://arxiv.org/html/2609.35852#S4.T1); the row\-side and column\-side fields are the two colours drawn on each projection in Figure[1](https://arxiv.org/html/2609.35852#S1.F1)\)\.

Table 1:Where the computation graph places each functional identity \(matrix side, a structural fact\), and which one\-sided normalization removes most of the pooled\-shape departure on the controlled grid \(an empirical result; values in Table[4](https://arxiv.org/html/2609.35852#A2.T4)\)\.Consistent with this map, row normalization removes most of the pooled\-shape departure ofqq,kk, gate and up and explains it withR02R\_\{0\}^\{2\}of 0\.89–0\.98, whereas column normalization does so forooand down \(per\-kind values in Table[4](https://arxiv.org/html/2609.35852#A2.T4), Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\)\. The value projection has the two fields closest in width \(0\.13 and 0\.11 on the grid\) and, in the public models, is the only kind that one\-sided normalization leaves clearly below the common core; two\-sided balancing restores it\. Its row field follows head/value identities, whereas its column field carries a persistent input\-side channel profile that is also visible in other read\-side projections\. Why these two fields acquire comparable strength remains unresolved\.

#### 4\.3\.2Shared computational paths carry matched fields

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v2_S10_identity_transfer_v5.png)Figure 4:Shared functional channels carry matched scale\-field profiles across projections\.Correlations are computed at step 5,000 over the 12 controlled LLaMA\-style 70M runs, per layer; pale markers are single run–layer observations, lines connect per\-layer medians over runs, and the overall median over the 72 layer×\\timesrun cells is given in each legend\. The next\-layer mismatch pairs the same projections across adjacent layers\. Up against down matches as closely as gate against down \(medianrr0\.93\)\. In \(c\), directq/kq/krow matching is weaker and layer dependent, whereas aggregating rows by RoPE frequency pair \(32 pairs per head, median over heads\) gives a consistently matched profile\. Long\-run and data\-arm contrasts are reported in Appendix[B\.3](https://arxiv.org/html/2609.35852#A2.SS3)\.Axis assignment also makes shared channels comparable: when two projections share a functional channel, one profile can be read from the producing matrix’s rows and the other from the consuming matrix’s columns, and whether training makes them match is an empirical question\. Figure[1](https://arxiv.org/html/2609.35852#S1.F1)contains three such paths,v→ov\\rightarrow o, gate/up→\\rightarrowdown andq↔kq\\leftrightarrow k, the last read on the rows of both matrices and matched by RoPE pair, and Figure[4](https://arxiv.org/html/2609.35852#S4.F4)tests them\.

All three pairings are observed: gate/up rows match down columns,vvrows matchoocolumns, andqqandkkrows match once indexed by RoPE pair\. Layer\-mismatched and within\-layer re\-paired controls sit at the null, no pairing is present at initialization, and the data arms change the strength of the matching but not its location\. The gate–down match of step 5,000 weakens by 30,000 steps, when the persistent FFN pairing is up–down \(Appendix[B\.3](https://arxiv.org/html/2609.35852#A2.SS3)\)\. The fields follow the shared functional channels\.

#### 4\.3\.3RoPE indexes and couples theq/kq/krow\-scale fields

The dominantq/kq/kscale field lies on the row axis\. We tested how this field is indexed by permuting the assignment of RoPE frequencies while preserving the frequency set, architecture, initialization, data order and seed \(Figure[5](https://arxiv.org/html/2609.35852#S4.F5)\)\. In the representative pair, the base and permuted profiles correlate at 0\.98 for bothqqandkkafter frequency re\-indexing, compared with 0\.01 and 0\.06 at the original matrix coordinates\. Across three paired seeds,Δ​rid\\Delta r\_\{\\mathrm\{id\}\}is positive and above the permutation null in all six layers for both projections\. Thevvcontrol has no reproducible positive frequency advantage \(0\.03,−0\.86\-0\.86,−0\.15\-0\.15across seeds\)\. RoPE frequency therefore indexes theq/kq/krow\-scale profiles\.

The rotary construction explains why frequency, and not the matrix coordinate, indexes these fields\. Within a head, RoPE writes the logit between positionsmmandnnas a sum over the rotary pairsppof the query and key activations,

∑p\|qp\|​\|kp\|​cos⁡\(\(n−m\)​θp\+ϕp\),\\sum\_\{p\}\|q\_\{p\}\|\\,\|k\_\{p\}\|\\cos\\\!\\big\(\(n\-m\)\\theta\_\{p\}\+\\phi\_\{p\}\\big\),soθp\\theta\_\{p\}sets how each pair’s contribution varies with relative distance, and the gradient that reaches the pair’s weight rows carries the same rotation\. A frequency thus assigns a computational role; reassigning it to other rows moves the role, and the learned profile moves with it\. The same construction sets the unit on which a scale is defined\. For distinct frequencies, the logits are unchanged by a common rotation of theqqandkkrows of a pair, which mixes the two rows, and by the reciprocal rescalingqp→a​qpq\_\{p\}\\to a\\,q\_\{p\},kp→kp/ak\_\{p\}\\to k\_\{p\}/a, the balance mode of Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6); without the rotation, any invertible mapq→A​qq\\to Aq,k→A−⁣⊤​kk\\to A^\{\-\\top\}kwithin a head leaves them unchanged\. RoPE thus reduces this freedom to one scale and one phase per pair: the scale of a single row changes under the rotation, that of the pair does not, and the pair’sqq–kkgain is fixed up to the balance\. This accounts for pair aggregation matchingqqandkkwhere single rows do not \(Figure[4](https://arxiv.org/html/2609.35852#S4.F4)\(c\); the log\-mean pair read\-out used there is not itself rotation invariant but agrees with the invariant pooled\-RMS read\-out, Appendix[B\.3](https://arxiv.org/html/2609.35852#A2.SS3)\) and for the head\-organized profile without RoPE\. The symmetry fixes the unit, not the sign or strength of the coupling: across pairs, the covariance of theqqandkkpair profiles is\[var⁡\(g\)−var⁡\(b\)\]/4\[\\mathrm\{var\}\(g\)\-\\mathrm\{var\}\(b\)\]/4, so positive matching means that training varied the gain more than the balance, a measured outcome \(Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\), andv/ov/o, with the same freedom within a head, also couples \(Figure[4](https://arxiv.org/html/2609.35852#S4.F4)\(b\)\)\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_F5_rope_v9.png)Figure 5:RoPE frequency indexes theq/kq/krow\-scale profiles\.Profiles are cumulative median\-centred changes in log row RMS\.\(a\)Base and frequency\-permuted LLaMA\-style 70M runs of the representative seed pair, aggregated over heads and layers\.\(b\)Frequency\-indexing advantageΔ​rid=rfreq−rcoord\\Delta r\_\{\\mathrm\{id\}\}=r\_\{\\mathrm\{freq\}\}\-r\_\{\\mathrm\{coord\}\}, wherercoordr\_\{\\mathrm\{coord\}\}compares the paired profiles at the same matrix coordinates andrfreqr\_\{\\mathrm\{freq\}\}at the same assigned frequencies; one point per layer and seed pair, with thevvcontrol and the per\-layer one\-sided 95th\-percentile permutation null\.\(c\)A separately trained Pythia\-protocol size series \(70M, 160M, 410M; one data arm, 8,000 steps; rotary on 16 of 64 head dimensions, 8 pairs per head\); each profile is divided by its maximum absolute value, comparing shape and peak location rather than amplitude\.The frequency organization also appears across a separately trained Pythia\-protocol size series\. Theq/kq/kprofiles correlate at 0\.88–0\.96 between 70M, 160M and 410M, and all six profiles reach their maximum at pair index 5 \(of 0–7\)\. Positive profile values denote frequencies whose rows grew more than the median row\. Both the LLaMA\-style intervention and the Pythia series place the stronger relative growth toward the lower\-frequency part of their respective rotary grids\.

The permutation separates field arrangement from its scalar summaries\. Between paired runs,krawk\_\{\\mathrm\{raw\}\}andkrowk\_\{\\mathrm\{row\}\}agree to within 0\.5%, whileHWH^\{W\}changes by different amounts across seeds and layers \(Figure[11](https://arxiv.org/html/2609.35852#A2.F11)\)\. The full profilehrh^\{r\}retains the frequency\-to\-row assignment;HWH^\{W\}summarizes its spread, and the pooled shapekksummarizes the resulting marginal shape\.

The no\-positional\-encoding arm separates row\-field formation from frequency organization\. In this run,qqandkkstill develop substantial row\-scale heterogeneity on a similar early timescale, but the profile becomes primarily head\-organized\. The pair\-levelq/kq/kcorrelation falls from 0\.90 to 0\.39, still above its re\-pairing null \(97\.5% quantile 0\.17 for this six\-layer median\), while the row\-level correlation rises from 0\.26 to 0\.85: without rotation each row coordinate is itself a unit of theq⋅kq\\cdot kproduct\. Thev/ov/oand up/down paths remain well above their re\-pairing nulls \(Figure[12](https://arxiv.org/html/2609.35852#A2.F12)\)\. Removing the rotary coordinate thus removes the frequency index and weakens the pair\-level coupling, while substantial row\-scale heterogeneity remains\. The comparison rests on one seed and reaches a higher loss, so its magnitude differences remain descriptive\.

Architecture supplies the channel coordinates of the scale fields\. Section[4\.4](https://arxiv.org/html/2609.35852#S4.SS4)next follows how the profiles written on those coordinates accumulate or recede during training\.

### 4\.4Scale fields form early in training but are retained selectively

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_F4_field_evolution_v5.png)Figure 6:Scale fields form early in every projection kind, but their later retention differs\.Each kind is read on its identity side, andkidentk\_\{\\mathrm\{ident\}\}is the fit after normalizing that side only; thevvcurve therefore reports its row\-side contribution, its two\-sided restoration being given in Section[4\.1](https://arxiv.org/html/2609.35852#S4.SS1)\. Each point is the median over 18 matrices \(six layers×\\timesthree seeds of the Gaussian family\) at seven sampled steps; a matrix enters only if both fits passR2≥0\.99R^\{2\}\\geq 0\.99\.\(a\)Identity\-axis field width; the common initial value is the finite\-sample floor when each marginal RMS is estimated from 512 entries, and triangles mark each kind’s largest sampled width\.\(b\)Shape differencekraw−kidentk\_\{\\mathrm\{raw\}\}\-k\_\{\\mathrm\{ident\}\}\.\(c\)Mixture termkraw−2−kident−2k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{ident\}\}^\{\-2\}against\(HW\)2\(H^\{W\}\)^\{2\}; in these median trajectories the receding kinds return along nearly the same path\.We follow the identity\-axis field widths from initialization to step 30,000 in the three Gaussian\-initialized runs, sampled at steps 0, 800, 3,200, 10,000, 20,000, 25,000 and 30,000; each point is the median over six layers and three seeds \(Figure[6](https://arxiv.org/html/2609.35852#S4.F6)\)\. The fields are not built monotonically toward their terminal form\. The sampled trajectories fall into three phases, and a terminal checkpoint records only the last\.Formation\(steps 0–800\): by the first sampled checkpoint every projection kind has left the initialization floor, atHW≈0\.10H^\{W\}\\approx 0\.10–0\.130\.13, while the balanced cores of the three initialization families are still converging \(Section[4\.1](https://arxiv.org/html/2609.35852#S4.SS1)\)\.Separation\(800–3,200\): the fields diverge by kind;qqandkkkeep widening,vvandooreach their largest sampled width at step 3,200, and the FFN fields stop growing, gate and down peaking already at step 800 and up nearly flat\.Selective retention\(3,200–30,000\):qqandkkwiden further, to 0\.174 and 0\.231;vvandoonarrow to about 62% of their peak and are still narrowing at step 30,000; the FFN fields narrow substantially, gate with a small late rebound and down close to its initialization width \(per\-kind values in Appendix[B\.4](https://arxiv.org/html/2609.35852#A2.SS4)\)\. Formation is common and retention is selective: the terminal differences between kinds reflect continued widening, partial recession or a return toward the initialization width, not whether a field formed\. The same qualitative sequence appears in all three initialization families\. Retention and recession here describe the width of the identity\-axis field, not the preservation of individual channel profiles or a return of weights to their initial values\.

Across the sampled trajectories, the pooled shape tracks the contemporaneous field width\. Along each kind’s median trajectory,Δkmix\\Delta\_\{k\}^\{\\mathrm\{mix\}\}against\(HW\)2\(H^\{W\}\)^\{2\}is linear through the origin with slopes 0\.59–0\.73 andR02≥0\.98R\_\{0\}^\{2\}\\geq 0\.98, and the widening and receding branches show little separation \(Figure[6](https://arxiv.org/html/2609.35852#S4.F6)\(c\)\); the slopes sit below the across\-matrix coefficient 0\.791 of the grid, so the coefficient depends on what is compared\. As the FFN fields recede, the departure of the raw from the identity\-axis fit closes to within 0\.002, inside the spread of the reference, while the retainedq/kq/kfields keep it at−0\.015\-0\.015and−0\.031\-0\.031\(Figure[6](https://arxiv.org/html/2609.35852#S4.F6)\(b\)\)\. A near\-reference terminal shape therefore does not mean that a projection stayed homogeneous: the FFN fields broadened early and later narrowed\.

The later trajectories also depend on the training condition\. Without positional encoding \(Figure[12](https://arxiv.org/html/2609.35852#A2.F12)\) theqqfield still widens, to 0\.203 against 0\.171, and becomes head\-organized; from step 3,200 to 30,000vvholds its width \(0\.185 to 0\.186\) andoodeclines only modestly \(0\.210 to 0\.192\), whereas the FFN fields recede again \(Appendix[B\.3](https://arxiv.org/html/2609.35852#A2.SS3)\)\. This run uses one seed and reaches a different loss, so it marks a boundary condition, not a cause\. FFN recession is the most reproducible trajectory across the tested conditions, and what retains the attention\-side fields remains open\.

### 4\.5Extending channel\-scale analysis to AdamW’s second moment

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v2_S19_factor_topology_v9.png)Figure 7:The row and column factors of AdamW’s second moment align with the decoder’s functional channel spaces\.Spearman correlationρ\\rhoamong the 14 row and column factors oflog⁡V^\\log\\hat\{V\}for the seven projection kinds at steps 200, 800, 3,200 and 10,000 of one LLaMA\-style 70M run \(single seed; the step\-400 dump enters the 210 fits but is not drawn\); each factor is the centred mean oflog⁡V^\\log\\hat\{V\}along its axis, and each cell the median over six layers\. Grey cells pair factors of different length and are undefined; negative values, none below−0\.10\-0\.10, are drawn as 0\. At step 10,000 the median correlation is\+0\.81\+0\.81over the 10 within\-space pairs and\+0\.01\+0\.01over the 48 comparable across\-space pairs; the permutation null \(0\.07\) is the median over cells of the 95th percentile of\|ρ\|\|\\rho\|under 200 within\-layer permutations\.The row and column axes of the analysis are not specific to the weights\. Every per\-parameter array of a projection, whether gradient, first moment or second moment, is indexed by the same output and input channels, so the channel\-scale read\-out carries over; only the statistic has to suit the array\. For AdamW’s bias\-corrected second momentV^\\hat\{V\}, the positive state from which the adaptive denominator is built, the factorsai,bja\_\{i\},b\_\{j\}of Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\) are the row and column main effects oflog⁡V^\\log\\hat\{V\}, not the log\-RMS fieldshhof the weights, but they index the same channels\. We fit them to one LLaMA\-style 70M run \(single seed\) at five steps, 210 matrix–step fits \(seven kinds×\\timessix layers×\\timesfive steps\)\. The model accounts for a median 0\.92 of the variance oflog⁡V^\\log\\hat\{V\}, with a minimum of 0\.76, and the split between the two factors follows the weight\-side identity axis: the row factor dominates forqqandkkand, less sharply, for gate and up, the column factor forooand down, and the two are comparable forvv\(Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\.

The factors align with the functional channel spaces\. Figure[7](https://arxiv.org/html/2609.35852#S4.F7)groups the fourteen row and column factors by the architecture\-defined functional channel space each one indexes, a grouping fixed before any correlation is computed\. At step 10,000 the median correlation is\+0\.81\+0\.81between factors of the same channel space and\+0\.01\+0\.01between factors of different spaces, against a permutation null of 0\.07\. The block structure is pronounced but not complete: the columns that read the FFN input and the rows that carry the residual\-output error become correlated during training, whereas the attention\-input columns do not \(Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\.

The blocks form on different time courses, marked in Figure[7](https://arxiv.org/html/2609.35852#S4.F7): the pair that reads one tensor, the gate and up columns, is already fully correlated at the first dump, the attention\-input and value\-path pairs strengthen over training, and the FFN hidden\-unit block peaks at step 3,200 and partly relaxes, the direction of the FFN weight fields in Section[4\.4](https://arxiv.org/html/2609.35852#S4.SS4)\. Factors attached to the same nominal residual coordinate but read at different normalized tensors, the attention input and the FFN input, correlate strongly at step 800 and hardly at all by step 10,000: the factors come to track the tensor at a computational position, not the coordinate index\. The optimizer state therefore carries the fixed connectivity of the architecture together with structure that accumulates path by path during training\.

The channel\-scale analysis thus extends beyond the weights: on the optimizer state the same axes show structure aligned with the functional channel spaces\. The correspondence describes the adaptive denominator, not a generator of the weight fields \(Section[6](https://arxiv.org/html/2609.35852#S6)\)\. We next test on a frozen checkpoint which changes in the paired scale fields preserve the computation\.

### 4\.6Paired scale\-field edits separate functional gain from invariant balance

For the paired identity\-axis profilesh1h\_\{1\}andh2h\_\{2\}of a shared channel, the gaing=h1\+h2g=h\_\{1\}\+h\_\{2\}is the log of the relative product of the two channel scales and the balanceb=h1−h2b=h\_\{1\}\-h\_\{2\}the log of their ratio\. On the controlled grid at step 5,000,var⁡\(g\)/var⁡\(b\)\\mathrm\{var\}\(g\)/\\mathrm\{var\}\(b\)is 14–25 on the matched paths, against about 1 at initialization and for re\-pairedv/ov/oand up/down profiles \(theq/kq/kre\-pairing null is wider, 0\.5–2\.0; Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\): the shared variation is concentrated in the gain\. Each edit rewrites all six layers of one path,q/kq/k,v/ov/oor up/down, in the 30,000\-step base checkpoint, with every other parameter fixed and no further training\. Balance edits rescale the two sides reciprocally\. Gain edits scale both sides together and then restore each matrix’s Frobenius norm, so they hold the global scale and the balanced core fixed and change only the scale fields \(Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\)\.

##### Balance preserves the computation; gain flattening changes the loss\.

Removing the balance on any path changes the loss only at the10−710^\{\-7\}level, the float32 resolution of the mean loss, whereas flattening the gain raises it by 0\.14, 0\.07 and 0\.01 nats per token \(slice 1\) onq/kq/k,v/ov/oand up/down \(Figure[8](https://arxiv.org/html/2609.35852#S4.F8)a\)\. Scale can move between the two matrices of a path without changing the network; the relative channel gain is what the computation depends on, even at fixed matrix norms and balanced core\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_F8_gain_edit_v4.png)Figure 8:Frozen\-checkpoint edits separate an invariant balance mode from a loss\-sensitive gain mode\.Edits act on one of theq/kq/k,v/ov/oand up/down paths across all six layers of the 30,000\-step base checkpoint\. Filled and open markers show two disjoint evaluation slices from the training token stream \(slice 1 and slice 2; values in Table[5](https://arxiv.org/html/2609.35852#A2.T5)\)\.\(a\)Balance removal and gain flattening\.\(b\)Loss changes relative to the flatten, divided by the corresponding ratios of squared relative Frobenius displacementΔF\\Delta\_\{F\}\.\(c\)Flattening one gain decile alone, relative to the full flatten; responses below10−310^\{\-3\}of it are drawn at the floor, and the seven non\-positive ones as open downward triangles\. Edit formulas and displacement ratios are in Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\.
##### The trained gain direction costs more than random directions\.

Per unit squared displacement, flattening the trained gain costs 3\.2–3\.3, 1\.4–1\.6 and 2\.0–3\.1 times as much as moving along a Gaussian direction in the same channel subspace on the two slices \(Figure[8](https://arxiv.org/html/2609.35852#S4.F8)b\)\. A displacement\-matched permutation of the gain across channels costs roughly what its1/21/\\sqrt\{2\}projection onto the flatten direction plus a random remainder predicts \(Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\); this control therefore gives no evidence of sensitivity to the channel assignment beyond that projection\.

##### The response depends on path and direction\.

Onq/kq/k, doubling the log\-gain profile \(τ=−1\\tau=\-1\) costs 3\.5–3\.6 times as much as flattening it, at 1\.2 times its squared displacement; onv/ov/oand up/down doubling costs 0\.6–1\.6 times flattening\. The response also concentrates at the extremes of each field \(Figure[8](https://arxiv.org/html/2609.35852#S4.F8)c\)\. Flattening only the most amplified decile produces 0\.8 of the full\-field response onq/kq/kand 0\.5 onv/ov/o, whereas on up/down the most suppressed decile produces 0\.5 against 0\.25 for the most amplified\. After displacement normalization a concentration remains at the amplified end onq/kq/kand at the suppressed end on up/down, but not onv/ov/o\. On up/down, flattening the lowest\-gain decile thus has the larger response under this intervention, which does not define a general ranking of units\.

The paired fields therefore supply functional edit coordinates\. The ranges quoted above span both evaluation slices\. The evidence concerns finite edits of one checkpoint and seed, scored on in\-distribution loss; it does not show how the gain fields were generated or how such edits act on training or generalization\.

## 5Discussion

##### What the channel\-scale description adds\.

The fields distinguish where magnitude is distributed across channels from the pooled magnitude profile of a matrix\. This distinction is useful even when pooled fits are similar: indexed fields can differ across functional channels, and a narrow terminal field can follow substantial earlier broadening\. The complete representation\(s,hr,hc,Z\)\(s,h^\{r\},h^\{c\},Z\)is invertible only when the full signed core is retained \(Appendix[A\.2](https://arxiv.org/html/2609.35852#A1.SS2)\)\. The two fields summarize channel RMS heterogeneity withdout\+dind\_\{\\mathrm\{out\}\}\+d\_\{\\mathrm\{in\}\}values; they do not replace the remaining matrix information\.

##### From a common magnitude profile to different pooled shapes\.

The balanced cores have similar measured magnitude profiles across the tested settings, while channel\-scale heterogeneity accounts for much of the variation in pooled shape\. The mixture bridge makes this connection predictive on held\-out runs and data arms within the controlled grid\. Its coefficients remain protocol\- and setting\-dependent, and equal field widths need not produce equal pooled shapes: field distributions and scale–core pairing retain information beyond width \(Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\)\. Likewise, the fitted scale tracks the global RMS through the measured ratioλ/s\\lambda/s, rather than being identical toss\. These results explain why pooled statistics can be informative while leaving channel organization unresolved\.

##### Structural constraints and training\-dependent outcomes\.

The computation graph identifies the channel spaces on which fields can be compared\. The paired RoPE reassignment further shows that the learnedq/kq/kprofiles follow assigned frequencies rather than fixed row coordinates\. Training determines the strength and trajectory of these profiles: the sampled runs show continuedq/kq/kbroadening, partialv/ov/orecession and substantial FFN narrowing\. The extension to AdamW’s second moment finds related functional organization using log\-space row and column main effects, a different read\-out on the same parameter axes\. These observations constrain explanations of field formation but do not identify its generating dynamics; neither the optimizer state nor the weight fields uniformly precede the other in the available timing comparisons \(Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\.

##### Functional tests and prospective uses\.

The paired edits distinguish scale allocation between projections from their relative channel gain\. Reciprocal balance edits preserve the forward computation on the tested paths\. Gain flattening instead increases loss while preserving each matrix’s global RMS and balanced core, and exceeds the tested Gaussian controls after displacement normalization\. The larger endpoint response occurs at the high\-gain end onq/kq/kandv/ov/o, but at the low\-gain end on up/down; these are sensitivities to the specified finite edits, not a general ranking of unit importance\. The experiments demonstrate identity\-aligned comparison, trajectory tracking and controlled checkpoint edits\. They motivate testing channel\-aware rescaling or adaptation, but establish neither gains in compression or generalization nor arrangement\-specific sensitivity beyond the reported controls\.

## 6Limitations

##### Experimental coverage\.

The main controlled trajectories and interventions use a LLaMA\-style 70M model and one corpus with transformed data conditions\. The separately trained Pythia size series provides descriptive single\-seed comparisons, while public Pythia checkpoints extend static comparisons to four sizes without optimizer or intervention records\. Optimizer\-state records, the no\-position control and the checkpoint edits each use one seed; the RoPE reassignment uses three paired seeds\. Three initialization families support the qualitative trajectory comparison, with the detailed field\-width analysis centred on Gaussian runs\. Matrices and layers within a run are not independent replications, and aggregate profiles can conceal depth\-dependent differences, as observed for learned channel multipliers\[[Velikanov et al\., 2026](https://arxiv.org/html/2609.35852#bib.bib32)\]\. Transport to substantially larger models, other corpora and other optimizers is not established\.

##### Measurement and statistical summaries\.

The core is defined by two\-sided RMS balancing; its fitted shape is measured with a middle\-80% magnitude protocol, supplemented by the displayed quantile comparisons\. Similar measured profiles do not imply identical cores, independent entries, Gaussianity or universal tail behaviour\. The bridge is a calibrated empirical relation with setting\-dependent coefficients and a residual scale–shape pairing term; field widths are not sufficient statistics for the pooled shape\. The tracking ofssbyλ\\lambdarests on the measured stability of the ratioλ/s\\lambda/s, not on their equality\. Theq/kq/kpair profiles aggregate weight\-field read\-outs, robust to the choice between the mean of the two rows’ log RMS and their pooled RMS, whereas the optimizer factors are row and column main effects oflog⁡V^\\log\\hat\{V\}\.

##### Generating dynamics\.

The RoPE intervention changes the assignment of a fixed frequency set\. It identifies frequency\-based organization, not the effect of changing the spectrum or the origin of the low\-frequency preference; rotary symmetry makes the pair the natural unit of comparison but does not determine the learned coupling strength\. Optimizer\-state correspondence and timing remain observations inside a coupled training loop and do not establish how the different field trajectories are generated, including why the attention\-side fields are retained while the FFN fields recede\.

##### Functional and practical scope\.

The edits are finite perturbations of one checkpoint, evaluated on two disjoint slices of the training token stream\. They test forward\-pass invariance and in\-distribution loss sensitivity, not subsequent training behaviour, held\-out task performance or generalization\. Control edits are compared after displacement normalization rather than at exactly equal displacements, and the permutation comparison uses a quadratic\-response approximation\. The decile responses do not define a pruning criterion, and benefits for quantization, channel selection or scale\-only adaptation remain untested\. Weight\-side read\-outs also do not measure learning quality on their own: the companion study finds data conditions with similar weight\-side read\-outs that differ widely in generalization\.

## 7Conclusion

Row and column scale fields provide a mesoscopic, channel\-indexed description between individual Transformer weights and pooled matrix statistics\. Together with the global RMS and the full balanced core they retain the matrix information, while their widths and profiles summarize channel heterogeneity and organization\. In the tested settings, similar balanced\-core magnitude profiles coexist with different pooled shapes, much of whose departure is predicted by field width\.

The fields align with functional channel spaces and evolve differently across components during training\. The same channel\-based analysis extends to AdamW’s second\-moment factors, and frozen\-checkpoint edits distinguish an invariant reciprocal balance from a loss\-sensitive gain\. These results establish a framework for comparing, tracking and experimentally probing channel\-scale structure; its generating dynamics, its transport to other settings and its benefits for model adaptation or generalization are questions for further study\.

## References

- Balles and Hennig \[2018\]Lukas Balles and Philipp Hennig\.Dissecting Adam: The sign, magnitude and variance of stochastic gradients\.In*Proceedings of the 35th International Conference on Machine Learning \(ICML\)*, volume 80 of*Proceedings of Machine Learning Research*, pages 404–413\. PMLR, 2018\.arXiv:1705\.07774\.
- Barbero et al\. \[2025\]Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Veličković\.Round and round we go\! What makes rotary positional encodings useful?In*International Conference on Learning Representations \(ICLR\)*, 2025\.arXiv:2410\.06205\.
- Biderman et al\. \[2023\]Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal\.Pythia: A suite for analyzing large language models across training and scaling\.In*Proceedings of the 40th International Conference on Machine Learning \(ICML\)*, volume 202 of*Proceedings of Machine Learning Research*, pages 2397–2430, 2023\.arXiv:2304\.01373\.
- Black et al\. \[2022\]Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach\.GPT\-NeoX\-20B: An open\-source autoregressive language model\.In*Proceedings of BigScience Episode \#5 – Workshop on Challenges & Perspectives in Creating Large Language Models*, pages 95–136, 2022\.doi:10\.18653/v1/2022\.bigscience\-1\.9\.
- Chen et al\. \[2024\]Yiting Chen, Jiazi Bu, and Junchi Yan\.Unveiling the Matthew effect across channels: Assessing layer width sufficiency via weight norm variance\.In*Advances in Neural Information Processing Systems 37 \(NeurIPS 2024\)*, 2024\.
- Ding \[2026a\]Tiexin Ding\.A two\-parameter Weibull framework for diagnosing transformer weight distributions\.*arXiv preprint arXiv:2605\.18898*, 2026a\.
- Ding \[2026b\]Tiexin Ding\.Weibull weight\-scale parameter evolution under AdamW training dynamics\.*arXiv preprint arXiv:2606\.19367*, 2026b\.
- Ding \[2026c\]Tiexin Ding\.Data predictability shapes Weibull weight\-scale growth in transformer training\.*arXiv preprint arXiv:2608\.23573*, 2026c\.
- Gupta et al\. \[2018\]Vineet Gupta, Tomer Koren, and Yoram Singer\.Shampoo: Preconditioned stochastic tensor optimization\.In*Proceedings of the 35th International Conference on Machine Learning \(ICML\)*, volume 80 of*Proceedings of Machine Learning Research*, pages 1842–1850, 2018\.
- Hägele et al\. \[2026\]Alexander Hägele, Alejandro Hernández\-Cano, Atli Kosson, and Martin Jaggi\.Improving neural network training by decoupling the magnitude and direction of weight vectors\.*arXiv preprint arXiv:2606\.25971*, 2026\.v1 24 June 2026; v2 17 July 2026\.
- Hong et al\. \[2024\]Xiangyu Hong, Che Jiang, Biqing Qi, Fandong Meng, Mo Yu, Bowen Zhou, and Jie Zhou\.On the token distance modeling ability of higher RoPE attention dimension\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 5877–5888\. Association for Computational Linguistics, 2024\.doi:10\.18653/v1/2024\.findings\-emnlp\.338\.arXiv:2410\.08703\.
- Huang et al\. \[2026\]Tzu\-Heng Huang, Aditya Goyal, John Cooper, and Frederic Sala\.WARP: Weight\-space analysis for recovering training data portfolios, 2026\.ICML 2026 Workshop on Weight\-Space Symmetries; arXiv:2607\.01686\.
- Karypis et al\. \[2026\]Petros Karypis, Sean O’Brien, Shreyas Kadekodi, Rui Zhu, and Julian McAuley\.LeRoPE: Learnable RoPE frequencies improve language modeling\.*arXiv preprint arXiv:2607\.10134*, 2026\.
- Kingma and Ba \[2015\]Diederik P\. Kingma and Jimmy Ba\.Adam: A method for stochastic optimization\.In*International Conference on Learning Representations \(ICLR\)*, 2015\.arXiv:1412\.6980\.
- Kobayashi et al\. \[2024\]Seijin Kobayashi, Yassir Akram, and Johannes von Oswald\.Weight decay induces low\-rank attention layers\.In*Advances in Neural Information Processing Systems 37 \(NeurIPS 2024\)*, 2024\.arXiv:2410\.23819\.
- Kosson et al\. \[2024\]Atli Kosson, Bettina Messmer, and Martin Jaggi\.Rotational equilibrium: How weight decay balances learning across neural networks\.In*Proceedings of the 41st International Conference on Machine Learning \(ICML\)*, volume 235 of*Proceedings of Machine Learning Research*, pages 25333–25369, 2024\.arXiv:2305\.17212\.
- Kunstner et al\. \[2023\]Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt\.Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be\.In*International Conference on Learning Representations \(ICLR\)*, 2023\.arXiv:2304\.13960\.
- Loshchilov and Hutter \[2019\]Ilya Loshchilov and Frank Hutter\.Decoupled weight decay regularization\.In*International Conference on Learning Representations \(ICLR\)*, 2019\.arXiv:1711\.05101\.
- Martens and Grosse \[2015\]James Martens and Roger Grosse\.Optimizing neural networks with Kronecker\-factored approximate curvature\.In*Proceedings of the 32nd International Conference on Machine Learning \(ICML\)*, volume 37 of*Proceedings of Machine Learning Research*, pages 2408–2417, 2015\.
- Martin and Mahoney \[2021\]Charles H\. Martin and Michael W\. Mahoney\.Implicit self\-regularization in deep neural networks: Evidence from random matrix theory and implications for learning\.*Journal of Machine Learning Research*, 22\(165\):1–73, 2021\.arXiv:1810\.01075\.
- Merity et al\. \[2017\]Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher\.Pointer sentinel mixture models\.In*International Conference on Learning Representations \(ICLR\)*, 2017\.arXiv:1609\.07843\.
- Peng et al\. \[2024\]Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole\.YaRN: Efficient context window extension of large language models\.In*International Conference on Learning Representations \(ICLR\)*, 2024\.arXiv:2309\.00071\.
- Salimans and Kingma \[2016\]Tim Salimans and Diederik P\. Kingma\.Weight normalization: A simple reparameterization to accelerate training of deep neural networks\.In*Advances in Neural Information Processing Systems 29 \(NIPS 2016\)*, 2016\.arXiv:1602\.07868\.
- Saponati et al\. \[2025\]Matteo Saponati, Pascal Sager, Pau Vilimelis Aceituno, Thilo Stadelmann, and Benjamin Grewe\.The underlying structures of self\-attention: symmetry, directionality, and emergent dynamics in transformer training\.*arXiv preprint arXiv:2502\.10927*, 2025\.v1 15 February 2025; v2 3 June 2025\.
- Shazeer \[2020\]Noam Shazeer\.GLU variants improve Transformer\.*arXiv preprint arXiv:2002\.05202*, 2020\.
- Shazeer and Stern \[2018\]Noam Shazeer and Mitchell Stern\.Adafactor: Adaptive learning rates with sublinear memory cost\.In*Proceedings of the 35th International Conference on Machine Learning \(ICML\)*, volume 80 of*Proceedings of Machine Learning Research*, pages 4596–4604, 2018\.arXiv:1804\.04235\.
- Sinkhorn \[1967\]Richard Sinkhorn\.Diagonal equivalence to matrices with prescribed row and column sums\.*The American Mathematical Monthly*, 74\(4\):402–405, 1967\.doi:10\.2307/2314570\.
- Sinkhorn and Knopp \[1967\]Richard Sinkhorn and Paul Knopp\.Concerning nonnegative matrices and doubly stochastic matrices\.*Pacific Journal of Mathematics*, 21\(2\):343–348, 1967\.
- Su et al\. \[2024\]Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu\.RoFormer: Enhanced transformer with rotary position embedding\.*Neurocomputing*, 568:127063, 2024\.doi:10\.1016/j\.neucom\.2023\.127063\.arXiv:2104\.09864\.
- Touvron et al\. \[2023\]Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie\-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample\.LLaMA: Open and efficient foundation language models\.*arXiv preprint arXiv:2302\.13971*, 2023\.
- van Laarhoven \[2017\]Twan van Laarhoven\.L2 regularization versus batch and weight normalization\.*arXiv preprint arXiv:1706\.05350*, 2017\.
- Velikanov et al\. \[2026\]Maksim Velikanov, Ilyas Chahed, Jingwei Zuo, Dhia Eddine Rhaiem, Younes Belkada, and Hakim Hacid\.Learnable multipliers: Freeing the scale of language model matrix layers\.*arXiv preprint arXiv:2601\.04890*, 2026\.8 January 2026\.
- Wan et al\. \[2021\]Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, and Jian Sun\.Spherical motion dynamics: Learning dynamics of normalized neural network using SGD and weight decay\.In*Advances in Neural Information Processing Systems 34 \(NeurIPS\)*, 2021\.arXiv:2006\.08419\.
- Wu et al\. \[2026\]Xinyi Wu, Siyuan Liu, and Ali Jadbabaie\.How data shapes RoPE frequency usage: From positional scale matching to length generalization\.*arXiv preprint arXiv:2607\.07678*, 2026\.
- Zhang and Sennrich \[2019\]Biao Zhang and Rico Sennrich\.Root mean square layer normalization\.In*Advances in Neural Information Processing Systems 32 \(NeurIPS 2019\)*, 2019\.arXiv:1910\.07467\.

## Appendix AMethods and reproducibility

This appendix describes the decomposition, the read\-outs and the training runs\. Table[3](https://arxiv.org/html/2609.35852#A1.T3)states the protocol and the role of each evidence source\.

### A\.1Notation

Table[2](https://arxiv.org/html/2609.35852#A1.T2)lists the symbols used across sections, with the equation or section that defines them; symbols local to one appendix are defined where they appear\. Throughout,log\\logdenotes the natural logarithm;log10\\log\_\{10\}is used only for the histogram grid and forHIQRWH^\{W\}\_\{\\mathrm\{IQR\}\}\. Profile correlationsrrare Pearson correlations unless stated otherwise\.

symbolmeaningdefined inMatrix and decompositionW∈ℝdout×dinW\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}one projection matrix; rows index output channels, columns input channelsSection[3\.1](https://arxiv.org/html/2609.35852#S3.SS1)W=s​Dr​Z​DcW=s\\,D\_\{r\}\\,Z\\,D\_\{c\}exact re\-parameterization: global scale, row and column balancing factors, coreeq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\)s=RMS⁡\(W\)s=\\mathrm\{RMS\}\(W\)global scale, fitting\-freeeq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\)Dr,DcD\_\{r\},\\,D\_\{c\};\(Dr\)i​i,\(Dc\)j​j\(D\_\{r\}\)\_\{ii\},\\,\(D\_\{c\}\)\_\{jj\}positive diagonal balancing factors; convertible to and from the scale fields givenssandZZeqs\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\), \([13](https://arxiv.org/html/2609.35852#A1.E13)\)ZZtwo\-sided RMS\-balanced core, unit RMS in every row and columneq\. \([3](https://arxiv.org/html/2609.35852#S3.E3)\)ZrZ^\{r\}one\-sided \(row\-normalized\) core,W=diag⁡\(R\)​ZrW=\\mathrm\{diag\}\(R\)\\,Z^\{r\}; its shape iskrowk\_\{\\mathrm\{row\}\}eq\. \([8](https://arxiv.org/html/2609.35852#S3.E8)\)χir,χjc\\chi^\{r\}\_\{i\},\\,\\chi^\{c\}\_\{j\}logZ2Z^\{2\}\-weighted RMS of the opposite\-side factors; separateshrh^\{r\}fromlog⁡Dr\\log D\_\{r\}\(andhch^\{c\}fromlog⁡Dc\\log D\_\{c\}\)Appendix[A\.2](https://arxiv.org/html/2609.35852#A1.SS2)𝒞⁡\(⋅\)\\mathcal\{C\}\(\\cdot\)median\-centring over channels,𝒞⁡\(x\)=x−median⁡\(x\)\\mathcal\{C\}\(x\)=x\-\\operatorname\{median\}\(x\)Appendix[A\.2](https://arxiv.org/html/2609.35852#A1.SS2)Ri=RMSj​\(Wi​j\)R\_\{i\}=\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\),Cj=RMSi​\(Wi​j\)C\_\{j\}=\\mathrm\{RMS\}\_\{i\}\(W\_\{ij\}\)row and column \(channel\) RMS ofWWeq\. \([4](https://arxiv.org/html/2609.35852#S3.E4)\)Scale fieldshir,hjch^\{r\}\_\{i\},\\,h^\{c\}\_\{j\}row and column scale fields: median\-centred log row and column RMSeq\. \([4](https://arxiv.org/html/2609.35852#S3.E4)\)hheither field when the axis is clear;h1,h2h\_\{1\},h\_\{2\}the identity\-axis fields of two paired projectionsSection[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)HrowW,HcolWH^\{W\}\_\{\\mathrm\{row\}\},\\,H^\{W\}\_\{\\mathrm\{col\}\}field widthssdi​\(hir\)\\mathrm\{sd\}\_\{i\}\(h^\{r\}\_\{i\}\),sdj​\(hjc\)\\mathrm\{sd\}\_\{j\}\(h^\{c\}\_\{j\}\)Appendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)HWH^\{W\}field width \(standard deviation\) on a kind’s identity side \(Table[1](https://arxiv.org/html/2609.35852#S4.T1)\)Section[3\.1](https://arxiv.org/html/2609.35852#S3.SS1)HIQRWH^\{W\}\_\{\\mathrm\{IQR\}\}interquartile row width used for the RoPE interventionAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)Weibull read\-outsshapekkWeibull shape from the middle\-80% probability\-plot fitAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)krawk\_\{\\mathrm\{raw\}\}shape of the original matrix \(pooled shape\)Appendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)krow,kcolk\_\{\\mathrm\{row\}\},\\,k\_\{\\mathrm\{col\}\}shape after row or column normalizationAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)kidentk\_\{\\mathrm\{ident\}\}krowk\_\{\\mathrm\{row\}\}orkcolk\_\{\\mathrm\{col\}\}on the identity sideeq\. \([9](https://arxiv.org/html/2609.35852#S3.E9)\)kbik\_\{\\mathrm\{bi\}\}shape of the two\-sided coreZZAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)λ\\lambda,λ𝒫\\lambda\_\{\\mathcal\{P\}\}fitted Weibull scale; the scale returned by protocol𝒫\\mathcal\{P\}eq\. \([6](https://arxiv.org/html/2609.35852#S3.E6)\)Δkmix\\Delta\_\{k\}^\{\\mathrm\{mix\}\}one\-axis mixture termkraw−2−kident−2k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{ident\}\}^\{\-2\}eq\. \([9](https://arxiv.org/html/2609.35852#S3.E9)\)αP\\alpha\_\{P\};αr,αc\\alpha\_\{r\},\\,\\alpha\_\{c\}protocol coefficient of the one\-axis bridge; row and column coefficients of the two\-axis bridgeeqs\. \([9](https://arxiv.org/html/2609.35852#S3.E9)\), \([10](https://arxiv.org/html/2609.35852#S3.E10)\)6/π26/\\pi^\{2\}reference coefficient for an exact Weibull under independenceAppendix[A\.5](https://arxiv.org/html/2609.35852#A1.SS5)R02R\_\{0\}^\{2\},R2R^\{2\}through\-origin and centred coefficients of determinationAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4)rr;ρ\\rhoprofile correlation; Spearman rank correlationAppendix[A\.4](https://arxiv.org/html/2609.35852#A1.SS4),[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)rfreq,rcoord,Δ​ridr\_\{\\mathrm\{freq\}\},\\,r\_\{\\mathrm\{coord\}\},\\,\\Delta r\_\{\\mathrm\{id\}\}paired\-profile correlation by assigned frequency, by matrix coordinate, and their differenceSection[4\.3\.3](https://arxiv.org/html/2609.35852#S4.SS3.SSS3)Optimizer stateGtG\_\{t\}mini\-batch gradient of a projection, a sum of per\-token outer productsδ​x⊤\\delta x^\{\\top\}eq\. \([15](https://arxiv.org/html/2609.35852#A2.E15)\)Mt,VtM\_\{t\},\\,V\_\{t\};M^t,V^t\\hat\{M\}\_\{t\},\\,\\hat\{V\}\_\{t\}AdamW first and second moments; bias\-correctedeqs\. \([15](https://arxiv.org/html/2609.35852#A2.E15)\), \([16](https://arxiv.org/html/2609.35852#A2.E16)\)UtU\_\{t\},ϵ\\epsilon,ηt\\eta\_\{t\},γwd\\gamma\_\{\\mathrm\{wd\}\},β1,β2\\beta\_\{1\},\\,\\beta\_\{2\}update direction, stabilizer, learning rate, decoupled weight decay, EMA coefficientseq\. \([16](https://arxiv.org/html/2609.35852#A2.E16)\)log⁡V^i​j=μ\+ai\+bj\+εi​j\\log\\hat\{V\}\_\{ij\}=\\mu\+a\_\{i\}\+b\_\{j\}\+\\varepsilon\_\{ij\}additive row–column model of the second moment;εi​j\\varepsilon\_\{ij\}the coordinate remaindereq\. \([11](https://arxiv.org/html/2609.35852#S3.E11)\)Data arms and checkpoint editsg=h1\+h2g=h\_\{1\}\+h\_\{2\},b=h1−h2b=h\_\{1\}\-h\_\{2\}path gain and balance modes of a paired identity axisSection[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)τ\\taustrength of a gain edit \(τ=1\\tau=1flattens the gain,τ=−1\\tau=\-1doubles it\)Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)Δ​L\\Delta Lchange in mean next\-token loss \(nats per token\) after an editAppendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)ΔF\\Delta\_\{F\},ΔFflat\\Delta\_\{F\}^\{\\mathrm\{flat\}\}relative squared Frobenius displacement of an edit; that of the flatten atτ=1\\tau=1Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)Table 2:Notation\. Component kinds are writtenqq,kk,vv,oo, gate, up and down; the key projectionkkalways appears alongsideqq, and the Weibull shape is always qualified as “shapekk”\.
### A\.2Conversions between representations

Equation \([1](https://arxiv.org/html/2609.35852#S3.E1)\) consists of two links, the second at fixedssandZZ:

W​⟷\(i\)​\(s,Dr,Z,Dc\)​⟷\(ii\), \(iii\)​\(s,hr,hc,Z\)\.W\\;\\overset\{\\text\{\(i\)\}\}\{\\longleftrightarrow\}\\;\(s,\\,D\_\{r\},\\,Z,\\,D\_\{c\}\)\\;\\overset\{\\text\{\(ii\), \(iii\)\}\}\{\\longleftrightarrow\}\\;\(s,\\,h^\{r\},\\,h^\{c\},\\,Z\)\.Link \(i\) is the balancing decomposition; link \(ii\) converts the balancing factors into the fields, \(iii\) inverts it, and \(iv\) shows that the inversion recoversWWuniquely\. Throughout, the matrix is entrywise nonzero, as every self\-trained matrix is \(the public Pythia checkpoints contain exact zeros in about10−610^\{\-6\}of their entries, for which the balancing also converges\), the coreZZis kept as the full signed matrix,log⁡Dr\\log D\_\{r\}andlog⁡Dc\\log D\_\{c\}denote the vectors of log diagonal entries, and𝒞⁡\(x\)=x−median⁡\(x\)\\mathcal\{C\}\(x\)=x\-\\operatorname\{median\}\(x\)denotes median\-centring\.

##### \(i\) Matrix and balancing factors,W↔\(s,Dr,Z,Dc\)W\\leftrightarrow\(s,D\_\{r\},Z,D\_\{c\}\)\.

BalancingWWgives the factors of eq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\), and their product returnsWW\. Condition \([3](https://arxiv.org/html/2609.35852#S3.E3)\) is a matrix\-scaling problem on the entrywise squaresW∘WW\\circ W, solved by alternating row and column RMS normalization; its solution is unique up to\(Dr,Dc\)→\(c​Dr,Dc/c\)\(D\_\{r\},D\_\{c\}\)\\to\(cD\_\{r\},D\_\{c\}/c\)\(see \(iv\)\)\. The constantccis fixed by the iteration itself, in which both factors start at one and rows are normalized before columns in each sweep; the fields, their widths andZZdo not depend on it\. The iteration stops at a relative row\- and column\-RMS spread of10−610^\{\-6\}\(at most 200 sweeps\), and every reported self\-trained matrix passes the checks spread<10−3<10^\{\-3\}and reconstructionW=s​Dr​Z​DcW=s\\,D\_\{r\}ZD\_\{c\}within10−910^\{\-9\}of the global RMS\. The public Pythia matrices and two early extraction scripts used a looser routine \(60 sweeps, tolerance10−510^\{\-5\}\); re\-balancing the 348 public Pythia matrices with the routine above changeskbik\_\{\\mathrm\{bi\}\}by at most10−610^\{\-6\}\.

##### \(ii\) Factors to fields,\(Dr,Dc\)→\(hr,hc\)\(D\_\{r\},D\_\{c\}\)\\rightarrow\(h^\{r\},h^\{c\}\)\.

Substituting eq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\) into the row RMS of eq\. \([4](https://arxiv.org/html/2609.35852#S3.E4)\) givesRMSj​\(Wi​j\)=s​\(Dr\)i​i​eχir\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\)=s\\,\(D\_\{r\}\)\_\{ii\}\\,e^\{\\chi^\{r\}\_\{i\}\}, so

hr\\displaystyle h^\{r\}=𝒞⁡\(log⁡Dr\+χr\),\\displaystyle=\\mathcal\{C\}\(\\log D\_\{r\}\+\\chi^\{r\}\),χir\\displaystyle\\qquad\\chi^\{r\}\_\{i\}=12​log⁡\(1din​∑jZi​j2​\(Dc\)j​j2\),\\displaystyle=\\tfrac\{1\}\{2\}\\log\\Big\(\\tfrac\{1\}\{d\_\{\\mathrm\{in\}\}\}\\textstyle\\sum\_\{j\}Z\_\{ij\}^\{2\}\\,\(D\_\{c\}\)\_\{jj\}^\{2\}\\Big\),\(12\)hc\\displaystyle h^\{c\}=𝒞⁡\(log⁡Dc\+χc\),\\displaystyle=\\mathcal\{C\}\(\\log D\_\{c\}\+\\chi^\{c\}\),χjc\\displaystyle\\qquad\\chi^\{c\}\_\{j\}=12​log⁡\(1dout​∑iZi​j2​\(Dr\)i​i2\)\.\\displaystyle=\\tfrac\{1\}\{2\}\\log\\Big\(\\tfrac\{1\}\{d\_\{\\mathrm\{out\}\}\}\\textstyle\\sum\_\{i\}Z\_\{ij\}^\{2\}\\,\(D\_\{r\}\)\_\{ii\}^\{2\}\\Big\)\.Each field thus carries its own factor and a termχ\\chifrom the opposite factor\. Since the rows ofZZhave unit RMS \(eq\. \([3](https://arxiv.org/html/2609.35852#S3.E3)\)\),e2​χire^\{2\\chi^\{r\}\_\{i\}\}is aZ2Z^\{2\}\-weighted mean of\(Dc\)j​j2\(D\_\{c\}\)\_\{jj\}^\{2\}along rowii, andhr=𝒞⁡\(log⁡Dr\)h^\{r\}=\\mathcal\{C\}\(\\log D\_\{r\}\)holds exactly whenχr\\chi^\{r\}is constant across rows, for instance whenDcD\_\{c\}is a multiple of the identity\. This is not the case in general, so the fields and the centred log factors are related exactly but need not coincide entrywise\. Median centring is not additive:𝒞⁡\(log⁡Dr\+χr\)\\mathcal\{C\}\(\\log D\_\{r\}\+\\chi^\{r\}\)and𝒞⁡\(log⁡Dr\)\+𝒞⁡\(χr\)\\mathcal\{C\}\(\\log D\_\{r\}\)\+\\mathcal\{C\}\(\\chi^\{r\}\)differ by a constant, which leaves widths and correlations unchanged\.

##### \(iii\) Fields to factors,\(hr,hc\)→\(Dr,Dc\)\(h^\{r\},h^\{c\}\)\\rightarrow\(D\_\{r\},D\_\{c\}\)\.

Solving eq\. \([12](https://arxiv.org/html/2609.35852#A1.E12)\) for the factors, with the unknown centring constants fixed byRMS⁡\(Dr​Z​Dc\)=1\\mathrm\{RMS\}\(D\_\{r\}ZD\_\{c\}\)=1, gives

log⁡Dr=hr−χr−12​log⁡\(1dout​∑ie2​hir\),log⁡Dc=hc−χc−12​log⁡\(1din​∑je2​hjc\)\.\\log D\_\{r\}=h^\{r\}\-\\chi^\{r\}\-\\tfrac\{1\}\{2\}\\log\\Big\(\\tfrac\{1\}\{d\_\{\\mathrm\{out\}\}\}\\textstyle\\sum\_\{i\}e^\{2h^\{r\}\_\{i\}\}\\Big\),\\qquad\\log D\_\{c\}=h^\{c\}\-\\chi^\{c\}\-\\tfrac\{1\}\{2\}\\log\\Big\(\\tfrac\{1\}\{d\_\{\\mathrm\{in\}\}\}\\textstyle\\sum\_\{j\}e^\{2h^\{c\}\_\{j\}\}\\Big\)\.\(13\)Becauseχr\\chi^\{r\}depends onDcD\_\{c\}andχc\\chi^\{c\}onDrD\_\{r\}, the two equations are solved by updating them in turn, which is the balancing scheme of \(i\) with row and column targets set by the fields\. The global scale does not enter; eq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\) then returnsW=s​Dr​Z​DcW=s\\,D\_\{r\}ZD\_\{c\}\.

##### \(iv\) Uniqueness\.

At a fixed point of eq\. \([13](https://arxiv.org/html/2609.35852#A1.E13)\) the matrixΩ=Dr2​\(Z∘Z\)​Dc2\\Omega=D\_\{r\}^\{2\}\\,\(Z\\circ Z\)\\,D\_\{c\}^\{2\}, withZ∘ZZ\\circ Zthe entrywise square, has row sumsdin​e2​hird\_\{\\mathrm\{in\}\}e^\{2h^\{r\}\_\{i\}\}and column sumsdout​e2​hjcd\_\{\\mathrm\{out\}\}e^\{2h^\{c\}\_\{j\}\}, each divided by the channel mean ofe2​he^\{2h\}on its side as in eq\. \([13](https://arxiv.org/html/2609.35852#A1.E13)\), so that both totals equaldout​dind\_\{\\mathrm\{out\}\}d\_\{\\mathrm\{in\}\}\. For a positive matrix the diagonally scaled matrix with prescribed row and column sums is unique\[[Sinkhorn, 1967](https://arxiv.org/html/2609.35852#bib.bib27)\], which covers the self\-trained matrices since they have no exact zeros\. By \(ii\),\(W∘W\)/s2\(W\\circ W\)/s^\{2\}has this form and these sums, soΩ=\(W∘W\)/s2\\Omega=\(W\\circ W\)/s^\{2\}at any fixed point: the factors are determined up to\(Dr,Dc\)→\(c​Dr,Dc/c\)\(D\_\{r\},D\_\{c\}\)\\to\(cD\_\{r\},D\_\{c\}/c\), andWi​j=s​sgn⁡\(Zi​j\)​Ωi​j1/2W\_\{ij\}=s\\,\\operatorname\{sgn\}\(Z\_\{ij\}\)\\,\\Omega\_\{ij\}^\{1/2\}, the positive factors preserving signs\.

The conversions need the full signed core with its channel indices\. Field widths and pooled magnitude profiles are summaries, and the relation between them is the empirical mixture bridge of Section[3\.2](https://arxiv.org/html/2609.35852#S3.SS2)\.

### A\.3Training configurations

All LLaMA\-style\[[Touvron et al\., 2023](https://arxiv.org/html/2609.35852#bib.bib30)\]runs train the same 70M\-parameter decoder \(6 layers, width 512, FFN width 1,376, 8 heads of dimension 64, SwiGLU\[[Shazeer, 2020](https://arxiv.org/html/2609.35852#bib.bib25)\], RoPE\[[Su et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib29)\], RMSNorm\[[Zhang and Sennrich, 2019](https://arxiv.org/html/2609.35852#bib.bib35)\]withϵ=10−5\\epsilon=10^\{\-5\}, untied output layer, initialization scale 0\.02\) in fp32 with AdamW\[[Kingma and Ba, 2015](https://arxiv.org/html/2609.35852#bib.bib14),[Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.35852#bib.bib18)\]\(β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.999\\beta\_\{2\}=0\.999,ϵ=10−8\\epsilon=10^\{\-8\}, decoupled weight decay 0\.1 applied to all parameters, including embeddings, the output layer and the RMSNorm gains\) and no gradient clipping, at peak learning rate10−310^\{\-3\}with 200 warm\-up steps and a cosine schedule of a 30,000\-step horizon that decays to 10% of peak, batch size 24 and sequence length 512 unless Table[3](https://arxiv.org/html/2609.35852#A1.T3)states otherwise\. The corpus is the first10810^\{8\}tokens of the WikiText\-103 training split with empty lines removed\[[Merity et al\., 2017](https://arxiv.org/html/2609.35852#bib.bib21)\], tokenized with the Pythia \(GPT\-NeoX\) tokenizer \(vocabulary padded to 50,304\)\. Two sampling schemes are used\. The controlled grid reads a transformed stream sequentially in a single pass: for a token budgetTT, the firstT/nrepT/n\_\{\\mathrm\{rep\}\}tokens form its unique block, a fractionffof the positions in that block have their tokens permuted among themselves, and the block is tilednrepn\_\{\\mathrm\{rep\}\}times\. All 30,000\-step runs instead sample 512\-token windows uniformly with replacement from the10810^\{8\}\-token stream, about 3\.7 expected passes\. Weight read\-outs come from stored checkpoints; optimizer\-state read\-outs and field\-edit evaluations use their own stored states and evaluation streams\. Matrix fits are nested within runs\.

##### Protocols that differ\.

In the paired RoPE reassignment, the base of pair 1 repeats the seed\-1 Gaussian initialization run \(same initialization and batch sequence\), and pairs 2–3 use the seed\-2 and seed\-3 Gaussian runs as base arms and are compared at the three checkpoints 3,200, 10,000 and 30,000 that both arms store\. The controlled grid uses a token budget ofT=61\.44T=61\.44M, is checkpointed at steps 800, 1,600, 3,200 and 5,000, and stops at 5,000 steps of the 30,000\-step schedule, at about 94% of the peak learning rate\. The optimizer\-state records come from two runs\. A dump run in the seed\-1 Gaussian configuration \(batch 24, same batch sequence as the base run\) stops at 10,000 steps of the 30,000\-step schedule and stores the fullWW,M^\\hat\{M\}andV^\\hat\{V\}of every projection at steps 200, 400, 800, 3,200 and 10,000 \(Figure[7](https://arxiv.org/html/2609.35852#S4.F7)and the residual comparison of Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\. A separate full\-length Gaussian run with batch size 16 and its own window sequence logs axis\-wiseWW,M^\\hat\{M\}andV^\\hat\{V\}statistics over a 20\-step window every 200 steps \(lock\-in timing of Appendix[B\.5](https://arxiv.org/html/2609.35852#A2.SS5)\)\. The Pythia size series uses batch24×51224\\times 512, 200 warm\-up steps and cosine decay to 10% of peak over its 8,000\-step budget, with checkpoints at 512, 2,000, 5,000 and 8,000; its windows are sampled with replacement from a 98\.3M\-token arm stream built from the full 118\.7M\-token WikiText\-103 training split, pre\-tokenized separately\.

Table 3:Evidence sources, their roles and protocol\-specific details; all LLaMA\-style runs share the architecture, optimizer and tokenizer stated above, while sampling, schedule length and batch size differ by row\. Rows denote evidence streams rather than independent sample counts: the Gaussian initialization runs are the base arms for RoPE seed pairs 2 and 3\. Fit counts and the mapping to figures are in the data documentation\.

### A\.4Weibull fitting protocol and the spread statistic

We use the middle\-80% probability\-plot fit of[Ding \[2026a\]](https://arxiv.org/html/2609.35852#bib.bib6), implemented on a histogram:\|W\|\|W\|is binned into a 1,024\-binlog10\\log\_\{10\}histogram on\[−12,2\]\[\-12,2\]and fitted by a probability\-plot regression oflog⁡\(−log⁡\(1−F\)\)\\log\(\-\\log\(1\-F\)\)onlog⁡x\\log xat the bin centresxx, whereFFis the empirical cumulative distribution of\|W\|\|W\|evaluated at the mid\-point of each bin \(for a Weibull distribution this plot is linear with slopekk\); the regression is weighted by bin count and restricted to the bins withFFin the middle 80%; the slope is the shapekkand the intercept givesλ=exp\(−c0/k\)\\lambda=\\exp\(\-c\_\{0\}/k\)withc0c\_\{0\}the intercept\. The same protocol is applied to the original, row\-normalized, column\-normalized and two\-sided balanced matrices, givingkrawk\_\{\\mathrm\{raw\}\},krowk\_\{\\mathrm\{row\}\},kcolk\_\{\\mathrm\{col\}\}andkbik\_\{\\mathrm\{bi\}\}; unless stated otherwiseλ\\lambdais the fitted scale of the original matrix\. TheR2R^\{2\}of the regression is recorded for every fit with a gate of0\.990\.99; matrices with a fit below the gate are retained and flagged, and figures that apply the gate say so\. They occur only in the uniform initialization family at steps 0 and 800 \(133 of 2,646 matrix records\) and in two rawqqfits of the public Pythia endpoints, so the gate removes no Gaussian\-family matrix\. No entry of the audited self\-trained checkpoints lies below10−1210^\{\-12\}and none is exactly zero; the public Pythia checkpoints contain 1,326 exact zeros \(about10−610^\{\-6\}of their entries\), which fall outside the grid and are excluded\. Every coefficient relatingHWH^\{W\}to the shapekkis a coefficient of this window and estimator\.

The field widths are the population standard deviations \(normalized by the channel count\) of the natural\-log fields,HrowW=sdi​\(hir\)H^\{W\}\_\{\\mathrm\{row\}\}=\\mathrm\{sd\}\_\{i\}\(h^\{r\}\_\{i\}\)andHcolW=sdj​\(hjc\)H^\{W\}\_\{\\mathrm\{col\}\}=\\mathrm\{sd\}\_\{j\}\(h^\{c\}\_\{j\}\), andHWH^\{W\}without a subscript denotes the width on a kind’s identity axis \(Table[1](https://arxiv.org/html/2609.35852#S4.T1)\)\. The paired RoPE comparison is read on the row side for every kind with a second, interquartile statistic fixed before the intervention,HIQRW=IQRi​\(log10⁡RMSj​\(Wi​j\)2\)H^\{W\}\_\{\\mathrm\{IQR\}\}=\\mathrm\{IQR\}\_\{i\}\\big\(\\log\_\{10\}\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\)^\{2\}\\big\), which weights the central rows wheresd\\mathrm\{sd\}weights the tails; the two can move differently under the permutation \(Appendix[B\.3](https://arxiv.org/html/2609.35852#A2.SS3)\)\.

##### Bridge statistics\.

All bridge coefficients are fitted through the origin and scored by the through\-originR02=1−∑n\(yn−y^n\)2/∑nyn2R\_\{0\}^\{2\}=1\-\\sum\_\{n\}\(y\_\{n\}\-\\widehat\{y\}\_\{n\}\)^\{2\}/\\sum\_\{n\}y\_\{n\}^\{2\}over the fitted matricesnn, withy^n=αP​xr,n\\widehat\{y\}\_\{n\}=\\alpha\_\{P\}x\_\{r,n\}for the one\-axis model andy^n=αr​xr,n\+αc​xc,n\\widehat\{y\}\_\{n\}=\\alpha\_\{r\}x\_\{r,n\}\+\\alpha\_\{c\}x\_\{c,n\}for the two\-axis model; this is not the centredR2R^\{2\}\. The statistical unit is the training run \(one seed at one data condition, an arm\)\. Leave\-one\-run\-out refits the coefficient on the other runs and predicts the held\-out run; leave\-one\-data\-arm\-out refits on three arms and predicts the fourth\.

##### Global scale and fitted scale\.

The global scales=RMS⁡\(W\)s=\\mathrm\{RMS\}\(W\)is computed without fitting and scales in proportion toWW\. For an ideal probability\-plot fit the fitted scaleλ\\lambdadoes the same, because a common factor shifts the plot sideways without changing its slope\. With the fixed log grid, a shift that is not a whole number of bins \(0\.0137 decades\) changes the binning slightly: refittingW/sW/sfor the 378 matrices of the base run of the paired RoPE reassignment reproducesλ/s\\lambda/swithin 0\.18% andkkwithin 0\.0040 \(medians 0\.06% and 0\.001\), which is the approximate form of eq\. \([6](https://arxiv.org/html/2609.35852#S3.E6)\)\. On that run \(42 matrices, nine stored steps\),log⁡s\\log srises over training by 1\.26 onqqandkk, 1\.04 on gate, 0\.94 on up and down, and 0\.74–0\.77 onvvandoo, while the layer medians ofλ/s\\lambda/sstay within 0\.845–0\.887 on every kind \(individual matrices 0\.786–0\.890\), soλ\\lambdafollows the global scale\. The upper end matches the valueλ/s≈0\.8875\\lambda/s\\approx 0\.8875derived for a Gaussian initialization under the middle\-80% protocol\[[Ding, 2026a](https://arxiv.org/html/2609.35852#bib.bib6)\]\. The measured ratio lies 3–9% above the Weibull value 0\.817 of eq\. \([7](https://arxiv.org/html/2609.35852#S3.E7)\) atk≈1\.20k\\approx 1\.20, comparable to the≈\\approx4\.6% bridge residual between the RMS\-derived and the fittedλ\\lambdathat[Ding \[2026b\]](https://arxiv.org/html/2609.35852#bib.bib7)attribute to the nonlinearity of the Weibull fit\.

### A\.5The full\-Weibull reference coefficient

ForX∼Weibull⁡\(k,λ\)X\\sim\\mathrm\{Weibull\}\(k,\\lambda\),E=\(X/λ\)kE=\(X/\\lambda\)^\{k\}is standard exponential, solog⁡X=log⁡λ\+k−1​log⁡E\\log X=\\log\\lambda\+k^\{\-1\}\\log EandVar⁡\(log⁡X\)=k−2​Var​\(log⁡E\)\\mathrm\{Var\}\(\\log X\)=k^\{\-2\}\\,\\mathrm\{Var\}\(\\log E\)\. The moments𝔼⁡\[Em\]=Γ⁡\(m\+1\)\\mathbb\{E\}\[E^\{m\}\]=\\Gamma\(m\+1\)giveVar⁡\(log⁡E\)=ψ′​\(1\)=∑n≥1n−2=π2/6\\mathrm\{Var\}\(\\log E\)=\\psi^\{\\prime\}\(1\)=\\sum\_\{n\\geq 1\}n^\{\-2\}=\\pi^\{2\}/6, withψ\\psithe digamma function \(equivalently,−log⁡E\-\\log Eis standard Gumbel\)\. Hencek−2=\(6/π2\)​Var​\(log⁡X\)k^\{\-2\}=\(6/\\pi^\{2\}\)\\,\\mathrm\{Var\}\(\\log X\)\. For the one\-sided decompositionWi​j=RMSj​\(Wi​j\)​Zi​jrW\_\{ij\}=\\mathrm\{RMS\}\_\{j\}\(W\_\{ij\}\)\\,Z^\{r\}\_\{ij\}, pooling all entries uniformly gives exactlyVar⁡\(log⁡\|W\|\)=\(HrowW\)2\+Var⁡\(log⁡\|Zr\|\)\+2​Covi​\(hir,meanj​log​\|Zi​jr\|\)\\mathrm\{Var\}\(\\log\|W\|\)=\(H^\{W\}\_\{\\mathrm\{row\}\}\)^\{2\}\+\\mathrm\{Var\}\(\\log\|Z^\{r\}\|\)\+2\\,\\mathrm\{Cov\}\_\{i\}\\big\(h^\{r\}\_\{i\},\\mathrm\{mean\}\_\{j\}\\log\|Z^\{r\}\_\{ij\}\|\\big\)\. Because every row ofZrZ^\{r\}has unit RMS, the row mean oflog⁡\|Zi​jr\|\\log\|Z^\{r\}\_\{ij\}\|is a statistic of the within\-row shape, so the last term pairs each row’s scale with its shape; it is the log\-variance counterpart of the pairing term of Appendix[B\.2](https://arxiv.org/html/2609.35852#A2.SS2)\. Neglecting it \(approximate independence of the row scale andZrZ^\{r\}\) and applying the log\-variance identity to the raw and the row\-normalized magnitudes gives

kraw−2−krow−2≈6π2​\(HrowW\)2,k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{row\}\}^\{\-2\}\\approx\\frac\{6\}\{\\pi^\{2\}\}\\,\(H^\{W\}\_\{\\mathrm\{row\}\}\)^\{2\},\(14\)with equality when both the raw and the row\-normalized magnitudes are exactly Weibull and the pairing term vanishes; the column case follows by transposition\.6/π26/\\pi^\{2\}is therefore the reference coefficient of the mixture bridge under those two conditions, not a value that the middle\-80% protocol coefficientαP\\alpha\_\{P\}must attain\.

### A\.6Code and data

The release is the directoryscale\_fieldat tagp5\-scale\-field\-v1of the repository

[https://github\.com/tiexinding/NPM\-Weibull\-public](https://github.com/tiexinding/NPM-Weibull-public)

It contains the training code and configurations of the self\-trained runs, the script that builds the pre\-tokenized training stream from the public corpus, the analysis code that computes the data tables from the checkpoints and optimizer states and the reported statistics from the tables, the data tables, and documentation of their contents\. Model checkpoints and optimizer dumps are not released because of their size\.

## Appendix BSupporting results

The supporting results follow the order of Section[4](https://arxiv.org/html/2609.35852#S4): the common core, the mixture bridge, the functional axes and RoPE, the weight\-field trajectories, the AdamW second\-moment factors, and the paired gain modes and frozen\-checkpoint edits\.

### B\.1Common core

This appendix bounds the common\-core statement of Section[4\.1](https://arxiv.org/html/2609.35852#S4.SS1)\. Figure[2](https://arxiv.org/html/2609.35852#S4.F2)reports component\-level medians, not an identical core in every matrix\. Across the 2,016 two\-sided fits of the controlled grid the individualkbik\_\{\\mathrm\{bi\}\}values span 1\.18–1\.27, with the upper end in the two shuffled arms\. Thevvprojection of the shuffled high\-repetition arm shows the largest one\-sided departure: two\-sided balancing moves its lowest block fromkrow=1\.09k\_\{\\mathrm\{row\}\}=1\.09tokbi=1\.19k\_\{\\mathrm\{bi\}\}=1\.19\. Among the normalized public\-Pythia blocks, a fewqqblocks overshoot the reference \(up to 1\.28 at 410M\) and a fewkkblocks fall below it \(down to 1\.125 at 70M\), although the block medians of both kinds lie in 1\.193–1\.205\. Per\-kind and per\-checkpoint values are in the data package\.

### B\.2Mixture bridge

This appendix supports the bridge of Section[4\.2](https://arxiv.org/html/2609.35852#S4.SS2)and the axis choice of Section[4\.3](https://arxiv.org/html/2609.35852#S4.SS3): hold\-out prediction, axis\-specific normalization, the two\-axis form, and what the width\-only bridge leaves out\. Fits are nested in their training runs and are never treated as independent samples\.

To check that the bridge of Figure[3](https://arxiv.org/html/2609.35852#S4.F3)is not carried by many matrices of the same training trajectory, its coefficient is refitted with a whole run or a whole data arm withheld \(definitions in Appendix[A](https://arxiv.org/html/2609.35852#A1)\)\. The held\-out errors stay small on average, and the shuffled high\-repetition arm D3 is the worst fold for every kind under both schemes \(Figure[3](https://arxiv.org/html/2609.35852#S4.F3)d; absolute leave\-one\-run\-out errors in Figure[9](https://arxiv.org/html/2609.35852#A2.F9)c\)\. The coefficient is positive in all 12 runs for the five row\-side kinds \(Figure[9](https://arxiv.org/html/2609.35852#A2.F9)a\), while its value moves from 0\.791 on the grid to 0\.654 on pooledq/kq/kof the paired RoPE\-reassignment runs and drifts in time by about a factor 1\.3 on the FFN side\. The relation is therefore predictive within the grid; its numerical coefficient is a property of the protocol and the setting\.

#### B\.2\.1Axis\-specific normalization and the two\-sidedvvexception

Table[4](https://arxiv.org/html/2609.35852#A2.T4)gives, per kind, the raw fit, the fits after row, column and two\-sided normalization, and the through\-originR02R\_\{0\}^\{2\}on each axis\. Read along a row, the side whose normalization returns the fit to the reference band and whose width explains the departure is the identity side of Table[1](https://arxiv.org/html/2609.35852#S4.T1)\. Which side carries the identity is fixed by the computation graph and tested by the cross\-projection pairing of Section[4\.3](https://arxiv.org/html/2609.35852#S4.SS3); which side’s measured field is the wider one usually follows, but that is empirical: it holds throughout training forqq,kkandoo, the two sides stay comparable forvv, and in the FFN kinds, whose fields are the narrowest, the wider side is not stable over training\.vvis the exception: removing only one side leaves the other field mixed into the pooled distribution, which is whyvvalone needs two\-sided normalization to recover the common core\. An independent check on the paired RoPE\-reassignment runs gives the same direction: the pooled shapekkofq/kq/kis returned to the initialization value by row and not by column normalization, while that ofWoW\_\{o\}is returned by column and not by row normalization\.

Table 4:Medians over the 288 fits per kind on the controlled grid \(12 runs×\\times4 checkpoints×\\times6 layers\): the raw fit, the fit after row, column and two\-sided normalization, the two heterogeneities, andR02R\_\{0\}^\{2\}of the through\-origin bridgekraw−2−kbi−2k\_\{\\mathrm\{raw\}\}^\{\-2\}\-k\_\{\\mathrm\{bi\}\}^\{\-2\}against\(HrowW\)2\(H^\{W\}\_\{\\mathrm\{row\}\}\)^\{2\},\(HcolW\)2\(H^\{W\}\_\{\\mathrm\{col\}\}\)^\{2\}, and both\.
#### B\.2\.2The two\-axis form of the bridge and its identifiability

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_SB_bridge_coefficients_v3.png)Figure 9:Bridge coefficients per kind, on one axis and on two\.\(a\)The one\-axis protocol coefficientαP\\alpha\_\{P\}on the 1,440 grid fits; all 12 runs share the sign in every kind\.\(b\)Two\-axis coefficients fitted per kind through the origin on the 2,016 two\-axis fits\. On the axis that carries each kind’s identity the coefficient is 0\.75–0\.87 \(qq,kk,oo, gate, up, down\);vv, whose two fields are of similar size, hasαr=0\.42\\alpha\_\{r\}=0\.42andαc=0\.88\\alpha\_\{c\}=0\.88\. The minor\-axis coefficient is poorly identified by these data: the minor\-axis width has per\-kind medians of 0\.06–0\.12, its square is 2–7 times below the major term, and the fitted minor coefficients include negative values \(gateαc=−1\.14\\alpha\_\{c\}=\-1\.14, downαr=−1\.13\\alpha\_\{r\}=\-1\.13\); ranges beyond±2\.5\\pm 2\.5are truncated\.\(c\)Each model is fitted per kind\. The column term is what bringsvv,ooand down into the same relation; it changesq/kq/klittle and slightly worsens gate/up, and the cross term changes the error by at most 0\.001 in every kind\. The worst fold is a high\-repetition arm in every kind and model, the shuffled one \(D3\) in all but two cases\.The balancing factors of eq\. \([2](https://arxiv.org/html/2609.35852#S3.E2)\) give a consistency check on the additive form: the variance oflog\|Wi​j\|=log⁡s\+log⁡\(Dr\)i​i\+log⁡\(Dc\)j​j\+log⁡\|Zi​j\|\\log\|W\_\{ij\}\|=\\log s\+\\log\(D\_\{r\}\)\_\{ii\}\+\\log\(D\_\{c\}\)\_\{jj\}\+\\log\|Z\_\{ij\}\|splits into the two balancing\-factor terms and the core term with a cross\-covariance share whose per\-kind median is at most 0\.3% on the 2,016 grid matrices \(per\-matrix maximum 3\.3%; per\-matrix values in the data release, Appendix[A\.6](https://arxiv.org/html/2609.35852#A1.SS6)\)\. The bridge itself is fitted on the direct marginal widths \([4](https://arxiv.org/html/2609.35852#S3.E4)\), which differ from these factors by the opposite\-side term of eq\. \([12](https://arxiv.org/html/2609.35852#A1.E12)\), and remains an empirical relation\.

The bridge on both axes, eq\. \([10](https://arxiv.org/html/2609.35852#S3.E10)\), is an empirical through\-origin approximation and not a theorem about Weibull mixtures\. The one\-axis form \([9](https://arxiv.org/html/2609.35852#S3.E9)\), whose baseline iskrowk\_\{\\mathrm\{row\}\}rather thankbik\_\{\\mathrm\{bi\}\}, is its one\-axis counterpart on the identity axis, which is the form the data identify\. The second\-axis term is not identifiable where the second field is small — its sign agrees in only 2–3 of 12 runs forq/kq/k, gate and up — and it is required where the second field is not small: adding it lowers the leave\-one\-run\-out error from 0\.0060 to 0\.0032 forvv, from 0\.0109 to 0\.0042 forooand from 0\.0235 to 0\.0071 for down, and changesq/kq/klittle\. A single pooled two\-axis fit over all seven kinds reaches onlyR02=0\.90R\_\{0\}^\{2\}=0\.90with unequal coefficients \(0\.67 rows, 0\.46 columns\), so the one\-axis bridge remains the main result and the two\-axis form is its extension to the column\-identity kinds\. Adding a cross termγ​HrowW​HcolW\\gamma\\,H^\{W\}\_\{\\mathrm\{row\}\}H^\{W\}\_\{\\mathrm\{col\}\}as a robustness check changes the leave\-one\-run\-out error by at most 0\.001 in every kind \(Figure[9](https://arxiv.org/html/2609.35852#A2.F9)c\)\.

On synthetic scale fields injected into real matrices the exact\-Weibull coefficient6/π26/\\pi^\{2\}predicts the refitted shapekkbetter than 0\.791 does, and on the grid the excess ofq/kq/kover6/π26/\\pi^\{2\}is carried by the scale\-times\-within\-row\-shape pairing term; that is an association within the protocol, not a cause\.

The width does not fix the departure: a scale\-times\-within\-row\-shape pairing term carries 12–28% ofΔkmix\\Delta\_\{k\}^\{\\mathrm\{mix\}\}and varies with the data condition, and the local coefficient bends withHWH^\{W\}in a way we report and do not explain\.

#### B\.2\.3Information retained beyond field width

Field width explains most of the pooled\-shape departure but does not specify the field\. Two synthetic checks on the controlled grid \(Table[3](https://arxiv.org/html/2609.35852#A1.T3)\) separate what it leaves out\. Shape: four synthetic field distributions of equal width \(log\-normal, two\-point, sparse, arcsine\) applied to the sameq/kq/kcores give fitted shape values that spread increasingly with width, by 0\.0015, 0\.007 and 0\.023 atHW=0\.10H^\{W\}=0\.10, 0\.20 and 0\.35, against a width\-driven drop of 0\.005, 0\.022 and 0\.068\. Order: reordering a fixed synthetic field over the channels changes the shapekkby at most 0\.0002 in this test; that does not remove the scale–core pairing term observed in real matrices above, which is a different quantity\. The quantitative evidence is therefore synthetic\. Figure[10](https://arxiv.org/html/2609.35852#A2.F10)adds the empirical fact on real weights: the skewness and tails of the fields vary with the data condition, particularly in the FFN projections: the FFN fields change the sign of their skewness and become heavy\-tailed with the data arm, theq/kq/kfields are heavy\-tailed mainly in the structured arm D1, and thev/ov/ofields stay mostly light\-tailed\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v2_S12_field_distribution_v6.png)Figure 10:Scale\-field distributions retain information beyond field width\.Controlled data grid \(Table[3](https://arxiv.org/html/2609.35852#A1.T3)\), 504 matrices at step 5,000, identity\-axis field of each matrix\. The plotted quantity ish~\\tilde\{h\}, the fieldhhof \([4](https://arxiv.org/html/2609.35852#S3.E4)\) re\-centred by its mean rather than by its median; the shift changes neither width, skewness nor kurtosis, andh~=0\\tilde\{h\}=0is the geometric\-mean channel RMS\.\(a\)Positive skewness is a tail of amplified channels, negative a tail of suppressed channels\.\(b\)Channel histograms of five example matrices, chosen to show the range of shapes and not matched in width\.

### B\.3Functional axes and RoPE

This section gives the controls for the cross\-projection field matching and the RoPE reassignment and no\-positional\-encoding comparisons of Section[4\.3](https://arxiv.org/html/2609.35852#S4.SS3)\.

#### B\.3\.1When cross\-projection matching persists or weakens

Figure[4](https://arxiv.org/html/2609.35852#S4.F4)reports the grid\-endpoint pairings and their null controls; this section adds what the figure does not show\. Before training, the pairings are absent: at step 0 every row\-level correlation across the three initialization families lies within 0\.09 of zero\. The longer initialization\-family runs show which matches persist\. At 30,000 steps, up rows and down columns correlate at about 0\.90 in each family, andvvrows andoocolumns at 0\.95–0\.96, whereas gate–down and gate–up correlations fall to about 0\.35 and 0\.4, so the persistent FFN pairing is specifically up–down\. Within the 5,000\-step data grid, matching is weakest in D2 \(fully shuffled, single pass\): gate–down is 0\.66 andv/ov/ois 0\.26\. The equally shuffled but highly repeated D3 arm reaches 0\.95 and 0\.91\. The strength of a shared\-channel match therefore changes with the data condition, and the D2 medians still exceed their within\-layer re\-pairing controls\.

#### B\.3\.2RoPE assignment and removal: field organization versus formation

Section[4\.3\.3](https://arxiv.org/html/2609.35852#S4.SS3.SSS3)establishes on one representative pair that theq/kq/krow profiles follow the reassigned frequencies\. Pair\-level profiles average the log row RMS of the two rows of a rotary pair \(rowjjand rowjjplus half the head dimension\) and take the median over heads\. This average is not exactly invariant under the common rotation of Section[4\.3\.3](https://arxiv.org/html/2609.35852#S4.SS3.SSS3), whereas the log of the pair’s pooled RMS is; on the grid endpoints of Figure[4](https://arxiv.org/html/2609.35852#S4.F4)\(c\) the two give medianq/kq/kcorrelations of 0\.929 and 0\.936 and gain\-to\-balance ratios of 22\.2 and 22\.0\. Figure[11](https://arxiv.org/html/2609.35852#A2.F11)extends this to every paired seed and shared checkpoint and separates what the reassignment leaves unchanged from what it moves\. Reassigning the same RoPE frequencies to different rows changes the coordinate\-indexedq/kq/kprofiles, while indexing by the assigned frequencies restores their agreement across the three paired seeds;vvshows no comparable separation \(Figure[11](https://arxiv.org/html/2609.35852#A2.F11)a\), although at seed 3 its pooled value \(−0\.15\-0\.15\) lies above its one\-sided null \(−0\.17\-0\.17\), so the pre\-specified negative\-control criterion is not met at that seed\. Of the scalar read\-outs, the pooled shapekkandkrowk\_\{\\mathrm\{row\}\}change by less than 0\.5%, whereas the row width is not equally stable \(Figure[11](https://arxiv.org/html/2609.35852#A2.F11)b\): its change depends on the statistic \(sd 1–4% against the interquartile read\-out 7–10% forq/kq/k\), on the aggregation \(vv:−9\.7%\-9\.7\\%pooled over 18 layer–checkpoint matrices,−5\.1%\-5\.1\\%as the median over checkpoints of layer medians,−2\.5%\-2\.5\\%at 30,000 steps alone\) and on the seed \(qqat 30,000 steps:\+2%\+2\\%,\+26%\+26\\%,\+14%\+14\\%\)\. The full field thus retains the frequency\-to\-row assignment that the scalar read\-outs largely omit\. The held\-out loss differs between the two arms by 0\.0013 nat\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_SB_rope_readout_v6.png)Figure 11:The RoPE permutation changes field arrangement while the pooled shape remains stable\.Paired runs with base and reassigned RoPE frequencies, three seeds, same data and initialization\.\(a\)Similarity between the permuted and the base row profile at each paired seed, pooled over heads and layers on the checkpoints shared by each pair \(steps 3,200, 10,000 and 30,000\); the pooledΔ​rid\\Delta r\_\{\\mathrm\{id\}\}is the difference of the two bars\.\(b\)Change under the permutation at seed 1 \(the width change differs across seeds, see text\), each bar a pooled median over 18 matrices per kind \(6 layers×\\timessteps 20,000, 25,000 and 30,000\);oois not part of the row\-side width comparison because its identity axis is the column axis\.Removing RoPE gives a complementary comparison \(Figure[12](https://arxiv.org/html/2609.35852#A2.F12)\)\. In this run,qqandkkrow fields still develop, but the pair\-indexedqq–kkcoupling weakens from 0\.90 to 0\.39, above its re\-pairing null \(97\.5% quantile 0\.17 for the six\-layer median; three of six layers exceed their per\-layer null\), and the pair gain\-to\-balance ratio falls from 19 to 2\.3, just above its null of 2\.0; the row\-levelqq–kkcorrelation rises from 0\.26 to 0\.85\. Thev/ov/oand up/down paths remain well above their nulls\. RoPE therefore organizes the frequency\-levelq/kq/krelationship in these runs, while formation of a row\-scale field does not require that coordinate\. The identity\-axis widths of the other kinds in this run at steps 3,200 and 30,000 \(vv0\.185/0\.186,oo0\.210/0\.192, gate 0\.216/0\.097, up 0\.202/0\.066, down 0\.205/0\.056\) show thev/ov/ofields holding their level whereas the FFN fields recede from a higher peak, as on the base run\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v3_SB_nope_v7.png)Figure 12:Aq/kq/krow field forms without RoPE, but its rotary\-pair organization and theqq–kkpair coupling weaken\.One run trained without positional rotation \(rotarycos=1\\cos=1,sin=0\\sin=0\) against the base run: single seed, same data and initialization, terminal loss 3\.174 against 2\.977\. In \(d\),g=h1\+h2g=h\_\{1\}\+h\_\{2\}andb=h1−h2b=h\_\{1\}\-h\_\{2\}are the sum and difference of the two paired profiles on each shared path \(Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)\), used here only as a pairing read\-out; its grey marks are the upper 97\.5% quantile of 100 within\-layer re\-pairings\.##### Boundaries\.

The reassignment keeps the frequency set fixed, whereas the run without positional encoding is one higher\-loss seed; neither comparison identifies the dynamics that generate the fields\.

### B\.4Weight\-field trajectories

This appendix checks the field trajectories of Section[4\.4](https://arxiv.org/html/2609.35852#S4.SS4)across Gaussian, Laplace and uniform initializations\. Figure[6](https://arxiv.org/html/2609.35852#S4.F6)follows three seeds of one family; Figure[13](https://arxiv.org/html/2609.35852#A2.F13)places the same nine runs of Figure[2](https://arxiv.org/html/2609.35852#S4.F2)\(c\) side by side in three read\-outs, so that field formation can be seen against core convergence in every family\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v2_S11_timing_chain_v8.png)Figure 13:Core convergence and scale\-field formation unfold together across initializations\.Nine LLaMA\-style 70M runs on one data setting \(Gaussian, Laplace and uniform initialization×\\timesthree seeds\), in three columns of kinds \(q/kq/k,v/ov/o, gate/up/down\)\. Top row: pooled shapekrawk\_\{\\mathrm\{raw\}\}; middle row: identity\-axis field width \(HrowWH^\{W\}\_\{\\mathrm\{row\}\}forqq,kk, gate, up;HcolWH^\{W\}\_\{\\mathrm\{col\}\}foroo, down; both forvv\); bottom row: two\-sided core shapekbik\_\{\\mathrm\{bi\}\}\. Lines are family medians over 18 matrices \(three seeds×\\timessix layers\)\.In the Gaussian runs of Figure[6](https://arxiv.org/html/2609.35852#S4.F6), gate falls from its step\-800 peak to 0\.065 by step 10,000 and rebounds slightly to 0\.072, while up and down narrow by2\.8×2\.8\\timesand3\.9×3\.9\\timesfrom their peaks, down to within 4% of its initialization width and up to 27% above it\. The bottom row of Figure[13](https://arxiv.org/html/2609.35852#A2.F13)gives the two\-sided core; forq/kq/kit is 1\.21 preserved under Gaussian initialization, 1\.00 and 1\.37 at step 0 under Laplace and uniform, within 0\.01 of the Gaussian value by step 3,200, and 1\.195–1\.218 per matrix at 30,000 steps for all three families\. Once the cores have merged, the terminal pooled shapekkof each kind agrees to within 0\.01 across the three initializations\. The middle row shows that the field is already well formed while the cores are still merging, and that the formation–retention split of Figure[6](https://arxiv.org/html/2609.35852#S4.F6)\(accumulatingq/kq/k, receding FFN,v/ov/obetween\) appears in each family\.

### B\.5AdamW second\-moment factors

This appendix gives the update rule, the motivation and fits of Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\), and the timing comparison behind Section[4\.5](https://arxiv.org/html/2609.35852#S4.SS5)\.

##### The AdamW update\.

For a projectionWt∈ℝdout×dinW\_\{t\}\\in\\mathbb\{R\}^\{d\_\{\\mathrm\{out\}\}\\times d\_\{\\mathrm\{in\}\}\}with mini\-batch gradientGtG\_\{t\}, the AdamW state\[[Kingma and Ba, 2015](https://arxiv.org/html/2609.35852#bib.bib14),[Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.35852#bib.bib18)\]is

Mt=β1​Mt−1\+\(1−β1\)​Gt,Vt=β2​Vt−1\+\(1−β2\)​Gt⊙2,M\_\{t\}=\\beta\_\{1\}M\_\{t\-1\}\+\(1\-\\beta\_\{1\}\)\\,G\_\{t\},\\qquad V\_\{t\}=\\beta\_\{2\}V\_\{t\-1\}\+\(1\-\\beta\_\{2\}\)\\,G\_\{t\}^\{\\odot 2\},\(15\)withM0=V0=0M\_\{0\}=V\_\{0\}=0, and the weights advance as

M^t=Mt1−β1t,V^t=Vt1−β2t,Ut=M^tV^t\+ϵ,Wt\+1=\(1−ηt​γwd\)​Wt−ηt​Ut,\\hat\{M\}\_\{t\}=\\frac\{M\_\{t\}\}\{1\-\\beta\_\{1\}^\{t\}\},\\qquad\\hat\{V\}\_\{t\}=\\frac\{V\_\{t\}\}\{1\-\\beta\_\{2\}^\{t\}\},\\qquad U\_\{t\}=\\frac\{\\hat\{M\}\_\{t\}\}\{\\sqrt\{\\hat\{V\}\_\{t\}\}\+\\epsilon\},\\qquad W\_\{t\+1\}=\(1\-\\eta\_\{t\}\\gamma\_\{\\mathrm\{wd\}\}\)\\,W\_\{t\}\-\\eta\_\{t\}\\,U\_\{t\},\(16\)where all operations are elementwise,ϵ\\epsilonis the stabilizer \(distinct from the fit residualεi​j\\varepsilon\_\{ij\}below\),ηt\\eta\_\{t\}the scheduled learning rate andγwd\\gamma\_\{\\mathrm\{wd\}\}the decoupled weight decay; the values used are in Appendix[A\.3](https://arxiv.org/html/2609.35852#A1.SS3)\. Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\) is fitted to the dumpedV^t\\hat\{V\}\_\{t\}; the bias correction multiplies every coordinate by the same number and therefore shiftsμ\\muonly\.

##### Why an additive model in log space\.

For a linear projectiony=W​xy=Wxthe gradient contributed by a single token is the outer productδ​x⊤\\delta x^\{\\top\}, whose squared entriesδi2​xj2\\delta\_\{i\}^\{2\}x\_\{j\}^\{2\}separate into a row and a column factor\. The mini\-batch gradient is a sum of such outer products,Gi​j=∑uδi\(u\)​xj\(u\)G\_\{ij\}=\\sum\_\{u\}\\delta\_\{i\}^\{\(u\)\}x\_\{j\}^\{\(u\)\}, soGi​j2G\_\{ij\}^\{2\}also contains cross\-token termsδi\(u\)​δi\(v\)​xj\(u\)​xj\(v\)\\delta\_\{i\}^\{\(u\)\}\\delta\_\{i\}^\{\(v\)\}x\_\{j\}^\{\(u\)\}x\_\{j\}^\{\(v\)\}, andV^\\hat\{V\}further averages these squares over time\. The per\-token structure motivates row and column effects in the accumulated second moment; if persistent row\- and column\-dependent parts dominate,V^i​j≈CV​Ai​Bj\\hat\{V\}\_\{ij\}\\approx C\_\{V\}A\_\{i\}B\_\{j\}, which is Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\) after taking logarithms and centring\. The cross\-token terms, the coupling of errors and activations, the attention mixing, the nonlinearities and the coordinate\-specific history remain inεi​j\\varepsilon\_\{ij\}, and the adequacy of the model is judged by itsR2R^\{2\}\. The factorsaia\_\{i\}andbjb\_\{j\}are thus the output\- and input\-channel factors of the accumulated squared\-gradient state, indexed by channel, which is why they group by tensor \(Figure[7](https://arxiv.org/html/2609.35852#S4.F7)\)\. With one\-sidedR2R^\{2\}valuesdin​∑iai2/SStotd\_\{\\mathrm\{in\}\}\\sum\_\{i\}a\_\{i\}^\{2\}/\\mathrm\{SS\}\_\{\\mathrm\{tot\}\}anddout​∑jbj2/SStotd\_\{\\mathrm\{out\}\}\\sum\_\{j\}b\_\{j\}^\{2\}/\\mathrm\{SS\}\_\{\\mathrm\{tot\}\}\(per\-kind medians over the 30 fits of six layers and five dump steps\), the full model reaches 0\.86–0\.98; the row factor alone explains 0\.80 and 0\.83 forqqandkkand 0\.68 and 0\.69 for gate and up, the column factor alone 0\.61 forooand 0\.85 for down, andvvsplits at 0\.40 and 0\.44, so the dominant factor is the identity axis of each kind\.

##### The remainder shows weak coordinate\-level association with the weights\.

The remainderεi​j\\varepsilon\_\{ij\}holds 0\.2–10% of the variance oflog⁡V^\\log\\hat\{V\}at step 10,000 and becomes heavy\-tailed as training proceeds in every kind exceptoo\(excess kurtosis below 1 at step 200; 7–22 at step 10,000, whileoostays near 1\)\. Double\-centringlog⁡\|Wi​j\|\\log\|W\_\{ij\}\|by its row and column means removes the same two axes from the weights, so the two residuals can be compared coordinate by coordinate within one matrix\. Their Spearman correlation, measured at the same checkpoint, is weak: over the 210 matrix–step fits of the five\-dump run of Figure[7](https://arxiv.org/html/2609.35852#S4.F7)the median\|ρ\|\|\\rho\|is 0\.017 and the maximum 0\.089, against a coordinate\-permutation null below 0\.006; 21 fits exceed 0\.05, all negative, on gate,kk,vvandooat intermediate steps\. In this contemporaneous comparison the correspondence between optimizer state and weight fields is carried by the row and column axes; a lagged influence through the update is not tested\.

##### What the factors mean for the update\.

Ignoring the stabilizer, the AdamW step isUi​j=M^i​j/V^i​jU\_\{ij\}=\\hat\{M\}\_\{ij\}/\\sqrt\{\\hat\{V\}\_\{ij\}\}, so under Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\) the denominator factorizes aseμ/2​eai/2​ebj/2e^\{\\mu/2\}e^\{a\_\{i\}/2\}e^\{b\_\{j\}/2\}up toeεi​j/2e^\{\\varepsilon\_\{ij\}/2\}: for a fixed first moment, a positiveaia\_\{i\}damps every update in rowiiand a positivebjb\_\{j\}every update in columnjj\. The realized update also depends onM^i​j\\hat\{M\}\_\{ij\}, so the factors describe the geometry of the denominator, not the complete step\. Adafactor\[[Shazeer and Stern, 2018](https://arxiv.org/html/2609.35852#bib.bib26)\]shares the row–column viewpoint but reconstructsV^\\hat\{V\}from original\-space marginals to save memory; Equation \([11](https://arxiv.org/html/2609.35852#S3.E11)\) is a descriptive fit in log space, and a high log\-spaceR2R^\{2\}does not makeV^\\hat\{V\}rank one\.

##### Cells of Figure[7](https://arxiv.org/html/2609.35852#S4.F7)\.

The gate and up columns, which read the same FFN input, correlate at 0\.998 from the first dump\. The attention\-input and FFN\-input columns, attached to the same residual coordinate, correlate at 0\.53–0\.76 at step 800 and at 0\.00–0\.07 by step 10,000\. The FFN\-input columns and the rows carrying the residual\-output error rise from 0\.26–0\.29 at step 200 to 0\.45–0\.54 at step 10,000, whereas the attention\-input columns stay within±0\.10\\pm 0\.10of both\. The permutation null of 0\.07 is a median over cells of cell\-specific 95th percentiles, not a single threshold for the whole matrix\.

The dump\-based factors above are not interchangeable with the logger profilelog⁡mean​V^\\log\\sqrt\{\\mathrm\{mean\}\\,\\hat\{V\}\}along an axis, used for the time\-resolved analyses, which differs most for the row\-levelq/kq/kpair; the timing below refers to logger profiles\. They come from the separate logger run of Table[3](https://arxiv.org/html/2609.35852#A1.T3)\(batch size 16, its own window sequence\), which records axis\-wiseWW,M^\\hat\{M\}andV^\\hat\{V\}statistics over a 20\-step window every 200 steps to 30,000 steps\. Taking as lock\-in the first step at which a row profile reaches a Spearman correlation of 0\.7 with its profile in the final window \(steps 29,801–29,820\), theV^\\hat\{V\}row profile locks in before the weight row field on six of seven kinds, by 0\.7k–1\.5k steps onqq,vvandooand by 3\.7k–5\.9k onkk, up and down\. Gate is the exception: its weight field locks in at 4\.8k steps while itsV^\\hat\{V\}profile locks in at 9\.5k\. The reversal is not specific toV^\\hat\{V\}: every gradient\-side quantity feeding gate \(the error rows, the gradient,M^\\hat\{M\}andV^\\hat\{V\}\) locks in between 9\.5k and 16\.7k steps, so the gradient statistics reaching gate keep reorganizing long after gate’s own row profile has settled\. A candidate reading is the SwiGLU product\[[Shazeer, 2020](https://arxiv.org/html/2609.35852#bib.bib25)\]: the gradient reaching a gate row is weighted by the paired up activation and inherits the up field, which recedes until about 10k steps, whereas up’s gradient inherits the early\-settling gate field and locks in at 5k–6k steps, close to gate’s weight lock\-in\. Cross\-matrix pairing runs the other way:vvrows andoocolumns correlate above 0\.7 at 676 steps inWWbut only at 3,528 steps inV^\\hat\{V\}\. Neither state therefore uniformly precedes the other, and the gate reading is a single\-run candidate\.

### B\.6Paired gain and frozen\-checkpoint edits

This appendix gives the gain–balance decomposition, the edit protocol and the controls behind Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)\.

##### Paired gain and balance modes\.

For two projections sharing a unit identity, with centred log\-scale profilesh1,h2h\_\{1\},h\_\{2\}, each unit has a gain modeg=h1\+h2g=h\_\{1\}\+h\_\{2\}and a balance modeb=h1−h2b=h\_\{1\}\-h\_\{2\}\. Sincevar⁡\(g\)−var⁡\(b\)=4​cov​\(h1,h2\)\\mathrm\{var\}\(g\)\-\\mathrm\{var\}\(b\)=4\\,\\mathrm\{cov\}\(h\_\{1\},h\_\{2\}\), their variance ratio reads the cross\-channel matching; the edits below test function\. At 5,000 steps the ratiovar⁡\(g\)/var⁡\(b\)\\mathrm\{var\}\(g\)/\\mathrm\{var\}\(b\)is 22 forq/kq/kby RoPE pair, 14 forvvrows againstoocolumns and 25 for up rows against down columns, against 0\.97–0\.99 at step 0\. Within\-layer re\-pairing gives 95% null bands of 0\.85–1\.18 onv/ov/oand 0\.91–1\.11 on up/down, and pairing with the adjacent layer gives 1\.00 on both; onq/kq/kthe re\-pairing band is wider \(0\.50–2\.01\) and the adjacent\-layer pairing stays at 4\.6, consistent with a frequency profile shared across layers\. The two sides of a matched channel move together \(Figure[14](https://arxiv.org/html/2609.35852#A2.F14)a\)\. On the initialization\-family runs at 30,000 steps the gate/down and gate/up ratios fall to 1\.7–2\.3 while up/down stays at 12–14, so the persistent FFN matching is on up/down, and the edits leave gate fixed\. Where the unit has an architectural coordinate the gain mode also aligns across runs: forq/kq/kby RoPE pair its correlation is 0\.79 across seeds and 0\.92 across initialization families, against 0\.27 and 0\.43 for the balance mode \(Figure[14](https://arxiv.org/html/2609.35852#A2.F14)b\); thev/ov/ohead index does not align, whereas the sorted profiles of both modes correlate at 0\.97–1\.00 per channel on every path \(0\.93–0\.94 at the head level\): the shape of the gain profile is shared across runs, the assignment to heads is not\. For theq/kq/kandv/ov/opairs the coupling echoes the balancing condition derived for regularized bilinear factors\[[Kobayashi et al\., 2024](https://arxiv.org/html/2609.35852#bib.bib15)\]; the gated FFN is not a bilinear factorization, and its up/down matching is an empirical observation\.

![Refer to caption](https://arxiv.org/html/2609.35852v1/figure_v2/figures/P5v2_S13_path_gain_balance_v7.png)Figure 14:Matched scale fields accumulate predominantly in a shared gain mode\.Gainggand balancebbof a paired channel are defined in Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6); their variance ratio exceeds 1 when the two profiles move together\. Re\-paired profiles give about 1 \(0\.5–2\.0 onq/kq/k\); the adjacent\-layerq/kq/kpairing stays near 4\.6\. Both run sources are defined in Table[3](https://arxiv.org/html/2609.35852#A1.T3)\.\(a\)The re\-pairing null \(shaded bars\) and the next\-layer pairings \(crosses\) are computed at step 5,000 of the data grid only and are drawn there, one per path\.\(b\)Head order is a permutation\-symmetric label, sov/ov/ohas no reproducible index; up/down is shown as a sorted comparison only\.
##### What is edited\.

The edits of Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)act on the final checkpoint of the base run of the paired RoPE reassignment \(LLaMA\-style 70M, Gaussian initialization, seed 1, 30,000 steps\)\. No parameter is trained; every condition is a rewrite of the stored weights followed by forward evaluation, with fixed random seeds for the randomized controls\. A path is a pair of projections that share a channel:vvrows withoocolumns \(512 head/value channels per layer\), up rows with down columns \(1,376 FFN hidden units per layer\), andqqrows withkkrows taken at the RoPE\-pair level \(rowsddandd\+32d\+32of each head form one channel with scale\(Rd2\+Rd\+322\)/2\\sqrt\{\(R\_\{d\}^\{2\}\+R\_\{d\+32\}^\{2\}\)/2\}, the rotation\-invariant pooled RMS; 256 pairs per layer\)\. For each layer and channeliithe two identity\-axis fieldsh1​\(i\)h\_\{1\}\(i\)andh2​\(i\)h\_\{2\}\(i\)are the median\-centred log scales of the two projections on the shared side, as in eq\. \([4](https://arxiv.org/html/2609.35852#S3.E4)\);ggandbbare not re\-centred, and flattening sets the selected channels to the median levelg=0g=0\(the median ofgglies within 0\.006 of zero on every path and layer\)\. The gain isg=h1\+h2g=h\_\{1\}\+h\_\{2\}and the balance isb=h1−h2b=h\_\{1\}\-h\_\{2\}\. An edit multiplies the row \(or column\)iiof side 1 byϕ1​\(i\)\\phi\_\{1\}\(i\)and of side 2 byϕ2​\(i\)\\phi\_\{2\}\(i\):

editϕ1​\(i\)\\phi\_\{1\}\(i\)ϕ2​\(i\)\\phi\_\{2\}\(i\)balancee−b\(i\)/2e^\{\-b\(i\)/2\}e\+b\(i\)/2e^\{\+b\(i\)/2\}flatten at strengthτ\\taue−τg\(i\)/2e^\{\-\\tau\\,g\(i\)/2\}e−τg\(i\)/2e^\{\-\\tau\\,g\(i\)/2\}shuffle by permutationπ\\pie\(g⁡\(π⁡\(i\)\)−g⁡\(i\)\)/2e^\{\(g\(\\pi\(i\)\)\-g\(i\)\)/2\}e\(g⁡\(π⁡\(i\)\)−g⁡\(i\)\)/2e^\{\(g\(\\pi\(i\)\)\-g\(i\)\)/2\}shuffle, displacement\-matchede\(g⁡\(π⁡\(i\)\)−g⁡\(i\)\)/\(2​2\)e^\{\(g\(\\pi\(i\)\)\-g\(i\)\)/\(2\\sqrt\{2\}\)\}samedecilemme−g\(i\)/2e^\{\-g\(i\)/2\}ifg⁡\(i\)g\(i\)in decilemm, else 1sameGaussian directione−ξ\(i\)/2e^\{\-\\xi\(i\)/2\},ξ∼𝒩⁡\(0,var⁡\(g\)\)\\xi\\sim\\mathcal\{N\}\(0,\\mathrm\{var\}\(g\)\)i\.i\.d\.samerandom signse−σ\(i\)g\(i\)/2e^\{\-\\sigma\(i\)\\,g\(i\)/2\},σ⁡\(i\)=±1\\sigma\(i\)=\\pm 1i\.i\.d\.sameOn theq/kq/kpathϕ\\phiis applied to both rows of a pair, so the path is edited at pair level and a within\-pair asymmetry is never touched; the gate projection is not edited\. Every edit acts on all six layers of one path at once and leaves the other five kinds untouched\. After every gain edit each matrix is rescaled to its original Frobenius norm, which holds the global scalessfixed; positive diagonal scalings preserve the balanced coreZZ, while the direct marginals on either axis can change \(Appendix[A\.2](https://arxiv.org/html/2609.35852#A1.SS2)\)\. Balance edits are not renormalized, since their invariance rests on the exact reciprocal factors\. At strengthτ\\tauthe gain becomes\(1−τ\)​g\(1\-\\tau\)gup to channel\-independent constants:τ=1\\tau=1flattens it,τ=−1\\tau=\-1doubles it andτ=2\\tau=2reverses it\. The size of an edit is its displacement, the relative squared Frobenius change summed over the twelve matricesWmW\_\{m\}of the path,ΔF=∑m‖Δ​Wm‖F2/‖Wm‖F2\\Delta\_\{F\}=\\sum\_\{m\}\\\|\\Delta W\_\{m\}\\\|\_\{F\}^\{2\}/\\\|W\_\{m\}\\\|\_\{F\}^\{2\};ΔFflat\\Delta\_\{F\}^\{\\mathrm\{flat\}\}is that of the flatten atτ=1\\tau=1\. The loss read\-out of an edit isΔ​L\\Delta L, the change in mean next\-token loss \(nats per token\) relative to the unedited model on the same tokens\.

##### Why the balance edit is an identity\.

On thev/ov/opath, the attention\-weighted value in each head is linear in its value channels\. Scaling rowiiofWvW\_\{v\}byccscales that channel at every token; the attention weights are unchanged becauseq/kq/kare not edited\. Scaling the corresponding column ofWoW\_\{o\}by1/c1/cthen cancels the change\. The same holds for up/down through the FFN hidden unit, except that the nonlinearity sits between the two sides: SwiGLU\[[Shazeer, 2020](https://arxiv.org/html/2609.35852#bib.bib25)\]multiplies the up activation by the gated activation, so scaling up byccand down by1/c1/cis exact only because the gate is not edited and the activation is linear in the up side\. Forq/kq/kthe scoreq⊤​kq^\{\\top\}kof a head is a sum over RoPE pairs; rotation by the pair’s frequency commutes with a scalar applied to both rows of the pair, so scalingqq’s pair byccandkk’s pair by1/c1/cleaves every score unchanged\. Biases are absent in this architecture\. The balance edits change the loss at the10−710^\{\-7\}level on every path and slice \(at most1\.2×10−71\.2\\times 10^\{\-7\}\), the float32 resolution of the mean loss, which verifies the pairing, the pair\-level treatment of RoPE and the implementation for the forward pass\.

##### Evaluation\.

The loss is the mean next\-token cross\-entropy over 48 sequences of 512 tokens, in batches of 8, read from the pre\-tokenized stream of10810^\{8\}tokens that the run was trained on\. Slice 1 starts at offset9×1079\\times 10^\{7\}and slice 2 at offset6×1076\\times 10^\{7\}, without overlap; Figure[8](https://arxiv.org/html/2609.35852#S4.F8)shows both, and the absolute costs quoted in Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)are slice\-1 values\. Training sampled the stream with replacement, so both slices are in\-distribution; each edit is compared with the unedited model on the same tokens\.

##### Controls and reproducibility across slices\.

The flatten atτ=1\\tau=1is compared with four controls: the gain permutation scaled by1/21/\\sqrt\{2\}\(at its own width it displaces about twice as much as the flatten\), a Gaussian direction in the same channel subspace with the variance ofgg, the flatten with an independent random sign per channel, and the gain doubled \(τ=−1\\tau=\-1\)\. The controls lie at displacement ratios 0\.87–1\.52 of the flatten, the scaled permutation at 0\.98–1\.00 \(Table[5](https://arxiv.org/html/2609.35852#A2.T5)\)\. Figure[8](https://arxiv.org/html/2609.35852#S4.F8)\(b\) reports the displacement\-normalized response\(Δ​L/Δ​Lflat\)/\(ΔF/ΔFflat\)\(\\Delta L/\\Delta L^\{\\mathrm\{flat\}\}\)/\(\\Delta\_\{F\}/\\Delta\_\{F\}^\{\\mathrm\{flat\}\}\), a descriptive quantity for finite edits; the table gives its two factors separately\.

After this normalization, flattening costs more than the Gaussian direction on both slices, by 3\.2–3\.3 onq/kq/k, 1\.4–1\.6 onv/ov/oand 2\.0–3\.1 on up/down\. The scaled permutation is its1/21/\\sqrt\{2\}projection onto−g\-gplus a remainder of half the variance ofgg; to second order its cost relative to the flatten is therefore12\+12×\\tfrac\{1\}\{2\}\+\\tfrac\{1\}\{2\}\\timesthe raw Gaussian cost ratio, which predicts 0\.65, 0\.81 and 0\.75 on slice 1 and 0\.64, 0\.77 and 0\.67 on slice 2, against the observed 0\.64/0\.66, 0\.87/0\.85 and 0\.65/0\.69\. The flatten costs themselves depart from a quadratic dependence onτ\\tau, most onq/kq/k, so the agreement is approximate; the control gives no evidence of sensitivity to the channel assignment beyond that projection\.

Across the two slices the balance identity, the positive flatten costs, the advantage over the Gaussian direction and the endpoint decile with the larger response all reproduce\. Absolute costs are higher on slice 2, and theq/kq/kdoubling asymmetry reproduces \(3\.64 and 3\.48\), whereas the FFN doubling response is slice\-dependent \(1\.59 and 0\.64\)\. The ten decile edits ofv/ov/osum to 0\.68–0\.70 of the full flatten, so these finite responses are not additive\. The four middle deciles change the loss by at most6×10−46\\times 10^\{\-4\}nats, against2×10−32\\times 10^\{\-3\}to 0\.14 for the two extreme deciles, and the ratio of flatten cost to squared gain width moves with the slice and is kept only as a candidate\.

##### Localization\.

Part of the decile concentration of Section[4\.6](https://arxiv.org/html/2609.35852#S4.SS6)follows from the larger displacement at the extremes; after displacement normalization a concentration of 1\.6 remains at the amplified end ofq/kq/kand 1\.2–1\.4 at the suppressed end of up/down, while the amplified end ofv/ov/ofalls to 0\.8, below its share of the displacement\. Of the predictions recorded before the edit runs, the identity and the monotone cost hold, the permutation’s arrangement reading was withdrawn after the displacement comparisons, and the amplified\-end prediction reverses on the FFN path\.

Table 5:Control\-edit displacements and loss responses\.Unless marked as an absolute quantity, the slice columns report the change in loss of the edit relative to flattening the gain \(τ=1\\tau=1\) on the same path and slice, averaged over the five random draws where marked; size is the squared displacement ratioΔF/ΔFflat\\Delta\_\{F\}/\\Delta\_\{F\}^\{\\mathrm\{flat\}\}of Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\. The shuffle at the histogram’s width is the unscaled permutation of the edit table in Appendix[B\.6](https://arxiv.org/html/2609.35852#A2.SS6)\. The lower block reads the 63\-condition set of the same two slices: the flatten cost in nats \(base loss 3\.347 on slice 1, 3\.093 on slice 2\), and, relative to it, the reversalτ=2\\tau=2, the two extreme deciles ofggedited alone, the sum of all ten decile edits, and the flatten cost divided by the squared gain width \(0\.275, 0\.209, 0\.068 on the three paths\)\.

Similar Articles

Geometric and Behavioral Stratification in Transformer Residual Streams

arXiv cs.LG

This paper investigates the residual stream geometry of trained transformers, showing that the prediction direction acts as a privileged anchor that stratifies residual-stream variation into narrow, readout-relevant and broad, computational regions across 18 models. The findings have implications for interpretability and evaluation.