Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
摘要
This paper introduces Power Law Graph Attention (PLGA) and the PLDR-LLM architecture, an exact generalization of scaled dot-product attention using input-generated bilinear operators. It presents theoretical results including an inference-collapse theorem, empirical stability measurements, and machine-checked proofs in Lean 4.
查看缓存全文
缓存时间: 2026/08/12 08:28
# exact generalization of scaled dot-product attention, empirical collapse at inference
Source: [https://arxiv.org/html/2608.10288](https://arxiv.org/html/2608.10288)
\(Date: August 10, 2026\)
###### Abstract\.
The Large Language Model from Power Law Decoder Representations \(PLDR\-LLM\) and its attention, Power Law Graph Attention \(PLGA\), replace the fixed bilinear form of scaled dot\-product attention \(SDPA\) with a*learned, input\-generated bilinear operator*GLMG\_\{\\mathrm\{LM\}\}, built from a positive tensorALMA\_\{\\mathrm\{LM\}\}by elementwise power laws\. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture\. Unconditionally: PLGA contains SDPA exactly atGLM=IG\_\{\\mathrm\{LM\}\}=I;ALMA\_\{\\mathrm\{LM\}\}andAPA\_\{P\}are strictly entrywise positive, with Perron–Frobenius structure onALMA\_\{\\mathrm\{LM\}\}; the DAG regularizer has the NOTEARS walk\-counting form and positivity obstructs exact acyclicity; and, under nonresonance \(satisfied by standard rotary frequencies\), a commutant criterion identifies which operators preserve relative\-position dependence\. An inference\-collapse theorem: exact input invariance of deductive outputs collapses inference to generalized SDPA with a constant operator\. Measured invariance: relative fluctuations of10−610^\{\-6\}and below; perturbation bounds quantify but do not certify cached inference; the assembled proxy misses the decoding margin\. A conditional three\-stage mechanism \(rotary twirl, concentration, row\-map contraction\) is measured on a released checkpoint\. Blockwise training and scoring under the global Gram are stated with explicit target exposure; on tested samples, block and sequential scoring select identical answers and agree on the published TruthfulQA probability\-mass metric within5×10−55\\times 10^\{\-5\}per item\. Self\-organized criticality enters as a phenomenological framework with an intrinsic order parameter; open claims become falsifiable conjectures\. Selected proof cores are machine\-checked in Lean 4\.
###### Key words and phrases:
Large language models, attention mechanisms, power law graph attention, learned bilinear operators, Perron–Frobenius theory, self\-organized criticality
###### 2020 Mathematics Subject Classification:
Primary 68T07; Secondary 15B48, 05C50, 82C27
###### Contents
1. [1Introduction](https://arxiv.org/html/2608.10288#S1)
2. [2Preliminaries and Notation](https://arxiv.org/html/2608.10288#S2)1. [2\.1Tokens, graphs, and the quantization/manifold duality](https://arxiv.org/html/2608.10288#S2.SS1) 2. [2\.2Matrix conventions](https://arxiv.org/html/2608.10288#S2.SS2)
3. [3The Power Law Graph Attention Operator](https://arxiv.org/html/2608.10288#S3)1. [3\.1Formal definition](https://arxiv.org/html/2608.10288#S3.SS1) 2. [3\.2Basic structural properties](https://arxiv.org/html/2608.10288#S3.SS2) 3. [3\.3PLGA as attention with a learned bilinear form](https://arxiv.org/html/2608.10288#S3.SS3) 4. [3\.4Multi\-head structure as a local decomposition](https://arxiv.org/html/2608.10288#S3.SS4)
4. [4The PLDR\-LLM Architecture](https://arxiv.org/html/2608.10288#S4)1. [4\.1Rotary position embeddings and a commutant characterization](https://arxiv.org/html/2608.10288#S4.SS1) 2. [4\.2The full decoder map](https://arxiv.org/html/2608.10288#S4.SS2) 3. [4\.3Training objective and the DAG regularizer](https://arxiv.org/html/2608.10288#S4.SS3) 4. [4\.4Inference: KV\-cache and G\-cache](https://arxiv.org/html/2608.10288#S4.SS4) 5. [4\.5The online generation contract and historical\-row prefix consistency](https://arxiv.org/html/2608.10288#S4.SS5)
5. [5Deductive Outputs as Invariant Operators](https://arxiv.org/html/2608.10288#S5)1. [5\.1The steady state and the order parameter](https://arxiv.org/html/2608.10288#S5.SS1) 2. [5\.2The inference\-collapse theorem](https://arxiv.org/html/2608.10288#S5.SS2) 3. [5\.3Quantitative stability of the collapse](https://arxiv.org/html/2608.10288#S5.SS3) 4. [5\.4The origin of the invariance: how the metric learner generates an invariant metric generator](https://arxiv.org/html/2608.10288#S5.SS4)1. [5\.4\.1Stage 1: the rotary twirl projects the density operator onto a commutant](https://arxiv.org/html/2608.10288#S5.SS4.SSS1) 2. [5\.4\.2Stage 2: statistical concentration of the normalized density operator](https://arxiv.org/html/2608.10288#S5.SS4.SSS2) 3. [5\.4\.3Stage 3: a conditional contraction bound for the trained row map](https://arxiv.org/html/2608.10288#S5.SS4.SSS3) 4. [5\.4\.4Why training at criticality might select the constant map: a hypothesis](https://arxiv.org/html/2608.10288#S5.SS4.SSS4)
6. [6Power Laws, Scale Invariance, and Self\-Organized Criticality](https://arxiv.org/html/2608.10288#S6)1. [6\.1Why power laws: the unique scale\-equivariant interaction](https://arxiv.org/html/2608.10288#S6.SS1) 2. [6\.2Criticality of the attention dynamics: a conditional spectral dictionary](https://arxiv.org/html/2608.10288#S6.SS2) 3. [6\.3The SOC training picture as a phenomenological framework](https://arxiv.org/html/2608.10288#S6.SS3)
7. [7Advantages of PLDR\-LLM over SDPA\-LLM](https://arxiv.org/html/2608.10288#S7)1. [7\.1A larger parameterized family with distinct training dynamics](https://arxiv.org/html/2608.10288#S7.SS1) 2. [7\.2An inspectable law representation](https://arxiv.org/html/2608.10288#S7.SS2) 3. [7\.3An intrinsic evaluation diagnostic \(prospective\)](https://arxiv.org/html/2608.10288#S7.SS3) 4. [7\.4Phase\-aware, self\-instrumented training](https://arxiv.org/html/2608.10288#S7.SS4) 5. [7\.5Inference efficiency and a deployment asymmetry](https://arxiv.org/html/2608.10288#S7.SS5) 6. [7\.6Inductive bias and transfer](https://arxiv.org/html/2608.10288#S7.SS6) 7. [7\.7Costs, trade\-offs, and honest limits](https://arxiv.org/html/2608.10288#S7.SS7)
8. [8Conjectures](https://arxiv.org/html/2608.10288#S8)
9. [9Discussion and Open Problems](https://arxiv.org/html/2608.10288#S9)
10. [ANotation](https://arxiv.org/html/2608.10288#A1)
11. [BCode, Models, and Verification Resources](https://arxiv.org/html/2608.10288#A2)
12. [CLean Formalization: Exact Coverage](https://arxiv.org/html/2608.10288#A3)
13. [DNumerical Audit on a Released Checkpoint](https://arxiv.org/html/2608.10288#A4)1. [D\.1Setup and provenance](https://arxiv.org/html/2608.10288#A4.SS1) 2. [D\.2Singular values and numerical rank \(in place of determinants\), trained versus initialized](https://arxiv.org/html/2608.10288#A4.SS2) 3. [D\.3The LayerNorm scale error, measured directly](https://arxiv.org/html/2608.10288#A4.SS3) 4. [D\.4Twirl energy against pre\-rotation data](https://arxiv.org/html/2608.10288#A4.SS4) 5. [D\.5Commutant residual of the trained operators](https://arxiv.org/html/2608.10288#A4.SS5) 6. [D\.6Row\-map contraction, measured on the composition](https://arxiv.org/html/2608.10288#A4.SS6) 7. [D\.7The invariance budget as a sample\-extrema proxy, against measured margins](https://arxiv.org/html/2608.10288#A4.SS7) 8. [D\.8DAG\-loss values](https://arxiv.org/html/2608.10288#A4.SS8) 9. [D\.9Order parameter, two normalizations](https://arxiv.org/html/2608.10288#A4.SS9) 10. [D\.10The online contract, historical\-row movement, and padding](https://arxiv.org/html/2608.10288#A4.SS10) 11. [D\.11Sequential versus one\-pass block scores](https://arxiv.org/html/2608.10288#A4.SS11) 12. [D\.12Held\-out sequential NLL and real benchmark items under both scoring protocols](https://arxiv.org/html/2608.10288#A4.SS12) 13. [D\.13Scope](https://arxiv.org/html/2608.10288#A4.SS13)
14. [References](https://arxiv.org/html/2608.10288#bib)
## 1\.Introduction
Modern large language models are overwhelmingly built on the transformer\[[43](https://arxiv.org/html/2608.10288#bib.bib6)\]with scaled dot\-product attention \(SDPA\): attention scores are Euclidean inner products of linearly projected token representations\. The Power Law Graph Attention \(PLGA\) mechanism\[[10](https://arxiv.org/html/2608.10288#bib.bib1),[11](https://arxiv.org/html/2608.10288#bib.bib2)\]and the decoder\-only architecture built from it, the Large Language Model from Power Law Decoder Representations \(PLDR\-LLM\)\[[12](https://arxiv.org/html/2608.10288#bib.bib3),[13](https://arxiv.org/html/2608.10288#bib.bib4),[14](https://arxiv.org/html/2608.10288#bib.bib5)\], generalize this design in a specific and mathematically meaningful way:*the bilinear form used to compare queries and keys is itself learned, nonlinearly, from the input*, through a chain
input⟶query Gram \(“density”\) operator⟶positive tensorALM\\displaystyle\\text\{input\}\\;\\longrightarrow\\;\\text\{query Gram \(\`\`density''\) operator\}\\;\\longrightarrow\\;\\text\{positive tensor \}A\_\{\\mathrm\{LM\}\}⟶potential tensorAP⟶bilinear score operatorGLM⟶attentionELM,\\displaystyle\\;\\longrightarrow\\;\\text\{potential tensor \}A\_\{P\}\\;\\longrightarrow\\;\\text\{bilinear score operator \}G\_\{\\mathrm\{LM\}\}\\;\\longrightarrow\\;\\text\{attention \}E\_\{\\mathrm\{LM\}\},in which the stepALM↦AP=ALM⊙PA\_\{\\mathrm\{LM\}\}\\mapsto A\_\{P\}=A\_\{\\mathrm\{LM\}\}^\{\\odot P\}imposes an element\-wise*power law*with learned exponents, and the stepAP↦GLM=aAP\+baA\_\{P\}\\mapsto G\_\{\\mathrm\{LM\}\}=a\\,A\_\{P\}\+b\_\{a\}superposes the resulting interaction profiles\. The intermediate tensors are exposed as*deductive outputs*of the model, alongside the usual next\-token probabilities \(the*inductive output*\)\. The title’s two claims are glossed here once\. “Exact generalization” means precisely this algebraic relationship: SDPA is contained in the PLGA family exactly, as itsGLM=IG\_\{\\mathrm\{LM\}\}=Ipoint \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(i\), machine\-checked\); containment does not assert strict function\-class separation at matched resources, which is explicitly not proved \(Remark[5\.5](https://arxiv.org/html/2608.10288#S5.Thmtheorem5)\), nor performance superiority\. “Empirical collapse at inference” names the phenomenon of Section[5](https://arxiv.org/html/2608.10288#S5)at its actual epistemic level: the collapse theorem is proved conditionally on exact operator invariance, and the invariance itself is a measured hypothesis, not a theorem\. Three empirical discoveries make this architecture mathematically distinctive:
1. \(1\)Operator invariance\.After pretraining, the deductive outputs are, on the tested workloads, invariant under change of input up to a minute perturbation \(relative fluctuations of order10−610^\{\-6\}–10−1110^\{\-11\}, and0at floating\-point resolution for the best models\), so the entire deep nonlinear PLGA subnetwork can be replaced at inference by a single cached tensor operatorGLMG\_\{\\mathrm\{LM\}\}; the published cached\-versus\-uncached benchmark evaluations are unchanged\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]\(scored by one\-pass block scoring, a path in which no generation\-time cache participates by construction, see Appendix[B](https://arxiv.org/html/2608.10288#A2); the substantive fidelity evidence is the measured operator and logit agreement of Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\.
2. \(2\)A learned singularity condition\.At convergence the generatorAAis numerically singular in the sharpest sense: rank one under an explicit singular\-value tolerance, with identical rows \(identical, moreover,*across attention heads*within a layer\), so that its single nonzero eigenvalue \(the common row sum\) is its spectral radius\. The interaction tensorALMA\_\{\\mathrm\{LM\}\}built from it remains strictly entry\-wise positive with Perron–Frobenius spectral structure\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\], and is generically of near\-full numerical rank: the float\-zero determinants reported for both tensors\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]are confirmed as rank statements forAAby its singular\-value spectrum, and traced forALMA\_\{\\mathrm\{LM\}\}to determinant underflow with no rank implication \(Remark[3\.10](https://arxiv.org/html/2608.10288#S3.Thmtheorem10), Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\.
3. \(3\)Critical\-like training phenomenology\.In the published experiments, whether the model generalizes \(produces coherent language and transfers to reasoning benchmarks\) tracks whether pretraining is carried out in a critical\-like regime of its driving/dissipation dynamics, with maximum learning rate and warm\-up steps as control parameters; an intrinsic*order parameter*built from deductive\-output fluctuations separates the observed phases in that sample\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The source papers read this as self\-organized criticality; in this paper that reading is treated as a phenomenological framework and hypothesis, not an established result \(Section[6\.3](https://arxiv.org/html/2608.10288#S6.SS3)\)\.
The goal of this paper is a full analytical description of PLDR\-LLM and PLGA\. We \(a\) give formal definitions of every component and the end\-to\-end model map, each verified against the reference implementations; \(b\) prove the structural properties underlying the three discoveries above, citing existing proofs where they exist and supplying proofs where they do not, including a conditional mechanistic analysis of how invariance of the generatorAAcan arise inside the metric learner, with hypotheses stated and their measurable diagnostics computed directly \(Section[5\.4](https://arxiv.org/html/2608.10288#S5.SS4), Appendix[D](https://arxiv.org/html/2608.10288#A4)\); \(c\) consolidate these results into a comparison of PLDR\-LLM with its SDPA base point in which each advantage is tied to a result at its actual epistemic strength, stating the costs with equal explicitness \(Section[7](https://arxiv.org/html/2608.10288#S7)\); and \(d\) collect the program’s open claims as precise, falsifiable conjectures \(rigidity of the invariant operator, a spectral form of the order parameter, and operator transfer across domains; Section[8](https://arxiv.org/html/2608.10288#S8)\)\.
##### Epistemic conventions\.
Every claim in this paper carries one of four labels, and the categories are never blurred\. \(1\)*Algebraic theorems*\(theorem, proposition, lemma, corollary environments with no unverified hypotheses\) are proved in place or cited\. \(2\)*Conditional theorems*are proved under hypotheses, stated inside the formal statement itself, that are*not*verified for trained models \(e\.g\. Lipschitz or stationarity assumptions\); their conclusions inherit that conditionality wherever they are used\. \(3\)*Empirical observations*are cited to the specific source experiment and never restated as theorems; floating\-point zeros are reported as underflow\-level observations, not exact identities\. \(4\)*Analogies and conjectures*are flagged as such in their statements \(conjectures in a dedicated environment\) and are never used as premises\. A correct elementary identity is not treated as evidence for a stronger interpretive claim in whose direction it points\. Where a hypothesis is measurable, we measure it on a released checkpoint \(Appendix[D](https://arxiv.org/html/2608.10288#A4)\) instead of assuming it\.
##### Related work\.
The ingredients of PLGA have distinct lineages, and priority is claimed narrowly\. Learned and input\-generated attention patterns predate this program: Synthesizer\[[39](https://arxiv.org/html/2608.10288#bib.bib12)\]learns or generates score matrices directly, and talking\-heads attention\[[34](https://arxiv.org/html/2608.10288#bib.bib13)\]learns linear maps across heads; SDPA itself carries a learned constant bilinear formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}in pre\-projection coordinates\. The broader dynamic\-parameter lineage is also necessary context: hypernetworks\[[16](https://arxiv.org/html/2608.10288#bib.bib35)\]generate the weights of one network with another; bilinear attention networks\[[21](https://arxiv.org/html/2608.10288#bib.bib36)\]learn explicit bilinear attention structure; the fast\-weight\-programmer line\[[31](https://arxiv.org/html/2608.10288#bib.bib41),[19](https://arxiv.org/html/2608.10288#bib.bib42)\]develops the view of attention as input\-updated operator/associative memory, with linear attention\[[20](https://arxiv.org/html/2608.10288#bib.bib40)\]as a fast\-weight update rule\. That line supplies a construction worth contrasting directly: its*cumulative*outer\-product stateSt=∑n≤tϕ\(kn\)vn⊤S\_\{t\}=\\sum\_\{n\\leq t\}\\phi\(k\_\{n\}\)v\_\{n\}^\{\\top\}is updated causally token by token, so every row of one parallel pass is its own sequential conditional \(historical\-row prefix consistency, Definition[4\.8](https://arxiv.org/html/2608.10288#S4.Thmtheorem8), holds by construction\), whereas PLGA’s*global*query Gram deliberately spends that property to let all known tokens shape a nonlinear learned operator \(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\), with the masked cumulative Gram \([4\.16](https://arxiv.org/html/2608.10288#S4.E16)\) as the interpolating option; no equivalence between the two mechanisms is claimed; and dynamic bilinear low\-rank attention \(DBA\)\[[29](https://arxiv.org/html/2608.10288#bib.bib18)\]is the closest in name, generating*input\-sensitive low\-rank projection matrices*that compress sequence length for efficient attention, whereas PLGA generates a strictly positive, power\-law\-deformeddk×dkd\_\{k\}\\times d\_\{k\}*head\-space score operator*whose learned tensor is itself the object of regularization, measurement, and caching: the shared idea is input\-conditioned bilinear structure, the object and purpose differ\. Closest at the level of the score bilinear form, PaTH attention\[[46](https://arxiv.org/html/2608.10288#bib.bib48)\]also places a data\-dependent matrix inside the query–key logit, writingqi⊤Hijkjq\_\{i\}^\{\\top\}H\_\{ij\}k\_\{j\}withHijH\_\{ij\}an accumulated product of data\-dependent identity\-plus\-rank\-one \(Householder\-like\) transitions along the positional path fromjjtoii: a content\-conditioned position encoding with one transition per query–key pair, built causally link by link, so that one parallel pass keeps the cumulative prefix semantics of the fast\-weight line\. PLGA instead generates a*single*sequence\-globaldk×dkd\_\{k\}\\times d\_\{k\}operator per layer and head from the query Gram through a deep row\-wise metric learner and a strictly positive elementwise power law chain, shares it across every row of the call, and exposes it as a tensor for regularization, spectral measurement, transfer, and caching; the global Gram spends the historical\-row prefix consistency \(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\) that the accumulated construction keeps, and neither construction reproduces the other\. Closest in the operator’s generating statistic, Curvature\-Conditioned Query \(CCQ\)\[[22](https://arxiv.org/html/2608.10288#bib.bib50)\]likewise inserts a context\-generateddk×dkd\_\{k\}\\times d\_\{k\}operator into the bilinear key–query score, deriving it from a second\-moment statistic of the context: the centered causal running key covarianceΣt\\Sigma\_\{t\}of the prefix enters the prescribed affine contractionI−λtΣtI\-\\lambda\_\{t\}\\Sigma\_\{t\}, with one learned scalar gateλt\\lambda\_\{t\}per token, so that the cleaned read isqt⊤\(I−λtΣt\)kjq\_\{t\}^\{\\top\}\(I\-\\lambda\_\{t\}\\Sigma\_\{t\}\)\\,k\_\{j\}, a per\-token causal correction to the read of a linear\-attention memory that keeps the prefix semantics of its recurrent state\. PLGA also generates its operator from a second\-moment statistic, the global rotated\-query Gram, but maps it through a deep learned strictly positive power law generator, shares the one resulting operator across every row of the call, and exposes it as a deductive tensor for regularization, spectral measurement, transfer, and caching, at the cost of the historical\-row prefix consistency \(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\) that the running\-covariance state keeps; a prescribed contraction with a scalar gate differs from a learned deep positive power law generator, a per\-token causal state from one sequence\-global row\-shared operator, and a linear\-attention read correction from an operator inside ordinary softmax scores, so neither construction reproduces the other\. Closest in feature\-space mechanics, cross\-covariance attention \(XCiT\)\[[8](https://arxiv.org/html/2608.10288#bib.bib45)\]also contracts the token axis into adk×dkd\_\{k\}\\times d\_\{k\}feature\-space matrix \(the key–query cross\-covariance\), but uses that matrix, normalized and softmax\-ed,*as the attention map itself*, mixing feature channels in place of token–token attention for linear\-in\-tokens efficiency; PLGA instead feeds its query Gram to a deep nonlinear row\-wise metric learner whose outputGLMG\_\{\\mathrm\{LM\}\}is a bilinear operator applied*inside*ordinary token–token attention scores, which remain the mixing mechanism, and the operator \(not the mixing\) is what is regularized, measured, and cached\. The two lines developed in parallel: XCiT and the encoder–decoder PLGA of\[[11](https://arxiv.org/html/2608.10288#bib.bib2)\]appeared within weeks of each other in mid\-2021, the feature\-space lineage here descending from the graph\-attention construction of\[[10](https://arxiv.org/html/2608.10288#bib.bib1)\]\. Closest in interpretive frame,*Attention as a Hypernetwork*\[[32](https://arxiv.org/html/2608.10288#bib.bib19)\]reads standard multi\-head attention itself as an implicit hypernetwork producing key–query\-conditioned operations, sharpening the operator\-generation view of attention; PLGA differs in object and mechanism, generating an*explicit*, strictly positive power\-law operator on head space that is exposed as a tensor, regularized, and empirically cacheable\. Closest in the operational use of a fixed head\-space operator, Autonomy\-of\-Heads \(AoH\)\[[47](https://arxiv.org/html/2608.10288#bib.bib49)\]extracts the frozen bilinear form that each trained SDPA head carries in pre\-projection coordinates and uses its effective rank as a data\-free spectral diagnostic, classifying retrieval heads \(concentrated spectra\) against streaming heads \(diffuse spectra\) and driving sparse\-attention and KV\-cache policy from the classification alone; the fixed pre\-projection form thus supports direct, operationally consequential spectral diagnostics without any input\. PLGA’s operator differs in being generated anew from each input’s rotated\-query Gram in head space through a deep strictly positive power law generator, shared across every row of the call, and exposed as a deductive tensor for fluctuation, spectral, transfer, and collapse analysis; AoH has no input\-conditioned generator and no invariant\-operator removal theorem, and neither construction reproduces the other\. None of these anticipates the specific PLGA construction, but a claim centered on input\-conditioned operator generation and removable fast/deductive weights belongs inside that lineage\. What is distinctive in PLGA is the*input\-conditioned generation of the head\-space operator*through a positive elementwise power law chain \(Section[3\.3](https://arxiv.org/html/2608.10288#S3.SS3)\), together with the removable/cacheable deductive network of Section[5](https://arxiv.org/html/2608.10288#S5)\. On the training side, a contemporaneous gradient\-flow analysis of standard attention\[[42](https://arxiv.org/html/2608.10288#bib.bib51)\]shows that factorizing the query–key and output–value circuits implicitly rescales their relative learning rates, with faster query–key movement sharpening attention at comparable loss; the analysis concerns fixed SDPA parameterizations, with no input\-conditioned operator, and it is independent precedent for the reading adopted in Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(iii\) and Remark[5\.5](https://arxiv.org/html/2608.10288#S5.Thmtheorem5): a parameterization changes the gradient geometry of training without, by itself, changing the realized function class\. Softmax attention as an averaging \(Markov\) operator and the rank\-collapse/oversmoothing phenomenology have an established literature\[[7](https://arxiv.org/html/2608.10288#bib.bib14),[30](https://arxiv.org/html/2608.10288#bib.bib15)\]\. The rotary\-embedding structure we analyze originates with RoFormer\[[38](https://arxiv.org/html/2608.10288#bib.bib10)\]\. Its design line generalizes the rotations through commuting/group\-structured generators\[[28](https://arxiv.org/html/2608.10288#bib.bib16),[48](https://arxiv.org/html/2608.10288#bib.bib17)\]; its analysis line has characterized the surviving symmetries: the gauge\-symmetry characterization of\[[44](https://arxiv.org/html/2608.10288#bib.bib43)\]identifies the invertible query/key reparameterizations of rotary attention exactly with the commutant of the position rotations \(typically\(GL\(1,ℂ\)\)dk/2\(\\mathrm\{GL\}\(1,\\mathbb\{C\}\)\)^\{d\_\{k\}/2\}for standard RoPE\), and the functional\-equivalence analysis of\[[40](https://arxiv.org/html/2608.10288#bib.bib44)\]ties the symmetries of rotary attention to the same commutant\. The rotation commutant is thus an established organizing principle of this literature, and its mathematical content \(the commutant of a semisimple matrix with distinct eigenvalues\) is classical linear algebra; Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)contributes the PLGA\-specific instance: an*arbitrary inserted operator*GGbetween rotated queries and keys preserves offset\-only score dependence exactly when it lies in that commutant, under an explicit nonresonance hypothesis\. From this instance follow the absorption step of Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(ii\), the head\-space codimension count of Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1), and the commutant\-residual diagnostic of Appendix[D](https://arxiv.org/html/2608.10288#A4)\. The DAG regularizer is NOTEARS\[[50](https://arxiv.org/html/2608.10288#bib.bib7)\]\. The remaining mathematical tools \(Perron–Frobenius theory, multiplicative Cauchy equations, geometric sums, Lipschitz composition, rank\-one algebra\) are classical; this paper’s mathematical contribution is application and synthesis, not new general theorems\.
##### How to read this paper\.
Section[2](https://arxiv.org/html/2608.10288#S2)fixes notation\. Sections[3](https://arxiv.org/html/2608.10288#S3)and[4](https://arxiv.org/html/2608.10288#S4)are the formal core: the PLGA operator and the full PLDR\-LLM architecture, with basic propositions\. Section[5](https://arxiv.org/html/2608.10288#S5)contains the operator\-invariance theory \(inference collapse, perturbation bounds, order parameter\), culminating in the new mechanism analysis of Section[5\.4](https://arxiv.org/html/2608.10288#S5.SS4)\. Section[6](https://arxiv.org/html/2608.10288#S6)formalizes scale invariance and the SOC training picture\. Section[7](https://arxiv.org/html/2608.10288#S7)consolidates, with formal backing, the advantages of PLDR\-LLM over SDPA\-LLM and their costs\. Section[8](https://arxiv.org/html/2608.10288#S8)collects the program’s conjectures\. Section[9](https://arxiv.org/html/2608.10288#S9)collects open problems\. Proofs are short and given in place\. A notation table is provided in Appendix[A](https://arxiv.org/html/2608.10288#A1); code, model, and verification resources in Appendix[B](https://arxiv.org/html/2608.10288#A2)\.
##### Dependency structure of the main claims\.
The load\-bearing chain of the paper, with the epistemic status of each link made explicit, is:
No result in a lower row is used as a premise for a claim in a higher row\.
##### Machine\-checked proofs\.
The elementary algebraic and analytic cores of the proofs in this paper have been formalized in Lean 4 over mathlib and verified by the Lean proof checker; the library builds with no unproved obligations, and a continuous\-integration axiom audit confirms that every exported theorem depends only on mathlib’s standard classical principles \(Appendix[C](https://arxiv.org/html/2608.10288#A3)\)\. The Lean source is available at[https://github\.com/burcgokden/PLDR\-LLM\-Math\-Foundations](https://github.com/burcgokden/PLDR-LLM-Math-Foundations)\(Appendix[B](https://arxiv.org/html/2608.10288#A2)\)\. Appendix[C](https://arxiv.org/html/2608.10288#A3)gives the exact claim\-by\-claim coverage table: which part of each numbered result is kernel\-checked and which is not\. Formalization of a proof core is*not*presented as validation of surrounding unformalized claims; wherever a result has both a prose and a formalized statement, the prose is kept no stronger than what is formally proved \(see Proposition[3\.9](https://arxiv.org/html/2608.10288#S3.Thmtheorem9)\)\.
## 2\.Preliminaries and Notation
### 2\.1\.Tokens, graphs, and the quantization/manifold duality
Fix a finite vocabulary𝒱\\mathcal\{V\}with\|𝒱\|=V\\left\\lvert\\mathcal\{V\}\\right\\rvert=V\(in the reference implementations, a SentencePiece unigram vocabulary withV=32,000V=32\{,\}000\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\]\)\. A*context*is a finite sequencex=\(x1,…,xS\)∈𝒱Sx=\(x\_\{1\},\\dots,x\_\{S\}\)\\in\\mathcal\{V\}^\{S\}withS≤SmaxS\\leq S\_\{\\max\}\(context lengthSmax=1024S\_\{\\max\}=1024in\[[12](https://arxiv.org/html/2608.10288#bib.bib3),[13](https://arxiv.org/html/2608.10288#bib.bib4),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\)\. An embedding map
ι:𝒱→ℝdmodel\\iota:\\mathcal\{V\}\\to\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\}assigns each token a dense feature vector; a context is identified with the matrixX∈ℝS×dmodelX\\in\\mathbb\{R\}^\{S\\times d\_\{\\mathrm\{model\}\}\}whoseii\-th row isι\(xi\)⊤\\iota\(x\_\{i\}\)^\{\\top\}\.
Following\[[11](https://arxiv.org/html/2608.10288#bib.bib2),[12](https://arxiv.org/html/2608.10288#bib.bib3)\], a context is regarded as a weighted graphG=\(𝒱x,Ex\)G=\(\\mathcal\{V\}\_\{x\},E\_\{x\}\)whose nodes are the tokens with feature vectorsι\(xi\)\\iota\(x\_\{i\}\)\. Two dual descriptions coexist:
- •the*quantization set*: the discrete vocabulary𝒱\\mathcal\{V\}, from which contexts are sampled as graph instances \(a “local” object\);
- •the*language\-model manifold*: admodeld\_\{\\mathrm\{model\}\}\-dimensional continuous feature space whose interaction structure is learned globally from the ensemble of all instances\.
PLGA is designed so that dataset\-level \(*global*\) structure is carried by learned parameters\(P,a,ba,W,bW\)\(P,a,b\_\{a\},W,b\_\{W\}\), while instance\-level \(*local*\) structure is carried by the inferred tensorsAA,ALMA\_\{\\mathrm\{LM\}\},APA\_\{P\},GLMG\_\{\\mathrm\{LM\}\},ELME\_\{\\mathrm\{LM\}\}\[[11](https://arxiv.org/html/2608.10288#bib.bib2)\]\. A central empirical result of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\], formalized in Section[5](https://arxiv.org/html/2608.10288#S5), is that after pretraining at criticality the “local” tensors become global: they are \(numerically\) independent of the instance\.
### 2\.2\.Matrix conventions
For matricesM,NM,Nof equal shape,M⊙NM\\odot Nis the Hadamard \(element\-wise\) product, and forMMwith positive entries and any real matrixPPof the same shape, the*element\-wise power*is
\(2\.1\)\(M⊙P\)ij=MijPij=exp\(PijlogMij\)\.\\bigl\(M^\{\\odot P\}\\bigr\)\_\{ij\}\\;=\\;M\_\{ij\}^\{\\,P\_\{ij\}\}\\;=\\;\\exp\\\!\\bigl\(P\_\{ij\}\\,\\log M\_\{ij\}\\bigr\)\.Plain juxtapositionaMaMof two matrices always denotes the ordinary matrix product\.𝟏\\mathbf\{1\}is the all\-ones column vector \(its dimension clear from context\), so𝟏𝟏⊤\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}is the all\-ones matrix;IIis the identity\. Rows of a matrix are writtenMi,:M\_\{i,:\}or, where subscripts would crowd,M\[i,⋅\]M\[i,\\cdot\]; likewiseV\[j\]V\[j\]is rowjjofVV\.softmax\\operatorname\{softmax\}acts row\-wise:softmax\(z\)i=ezi/∑jezj\\operatorname\{softmax\}\(z\)\_\{i\}=e^\{z\_\{i\}\}/\\sum\_\{j\}e^\{z\_\{j\}\}\. For a nonempty index set𝒥⊆\{1,…,S\}\\mathcal\{J\}\\subseteq\\\{1,\\dots,S\\\}the*ideal masked softmax*is
softmax𝒥\(z\)i=\{ezi/∑j∈𝒥ezj,i∈𝒥,0,i∉𝒥,\\operatorname\{softmax\}\_\{\\mathcal\{J\}\}\(z\)\_\{i\}\\;=\\;\\begin\{cases\}e^\{z\_\{i\}\}\\big/\\sum\_\{j\\in\\mathcal\{J\}\}e^\{z\_\{j\}\},&i\\in\\mathcal\{J\},\\\\ 0,&i\\notin\\mathcal\{J\},\\end\{cases\}equivalently the ordinary softmax with masked entries set to−∞\-\\inftyin the extended reals\. All causal\-mask statements in this paper are proved for the ideal masked softmax; the implementations realize it by adding a finite negative constant to disallowed scores \(−109\-10^\{9\}in the native TensorFlow and PyTorch repositories; the dtype minimumtorch\.finfo\(dtype\)\.min, via the Transformers causal\-mask utility, in the Hugging Face port\), whose exact real\-valued softmax is strictly positive everywhere and which reproduces the ideal operator only through floating\-point underflow of the masked entries under ordinary score ranges\. We keep the two semantics explicitly separate \(Remark[3\.13](https://arxiv.org/html/2608.10288#S3.Thmtheorem13)\)\.LN\\operatorname\{LN\}denotes LayerNorm\[[2](https://arxiv.org/html/2608.10288#bib.bib11)\]acting on the last \(feature\) axis; throughout the reference implementations LayerNorm usesεLN=10−6\\varepsilon\_\{\\operatorname\{LN\}\}=10^\{\-6\}and learned affine parameters; theεLN\\varepsilon\_\{\\operatorname\{LN\}\}\-dependence of its invariances is made explicit in Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\.∥⋅∥2\\left\\lVert\\cdot\\right\\rVert\_\{2\}is the spectral norm for matrices and the Euclidean norm for vectors;∥⋅∥F\\left\\lVert\\cdot\\right\\rVert\_\{F\}the Frobenius norm;∥⋅∥∞\\left\\lVert\\cdot\\right\\rVert\_\{\\infty\}the operator norm induced by the sup\-norm \(maximum absolute row sum\)\. The Swish/SiLU function isσs\(u\)=uς\(u\)\\sigma\_\{s\}\(u\)=u\\,\\varsigma\(u\)withς\\varsigmathe logistic sigmoid\. The source papers writeDQ=Q⊤QD\_\{Q\}=Q^\{\\top\}Qin Dirac notation as\|Q⊤⟩⟨Q⊤\|\\left\|Q^\{\\top\}\\right\\rangle\\\!\\left\\langle Q^\{\\top\}\\right\|\[[11](https://arxiv.org/html/2608.10288#bib.bib2)\]; in this paper we use plain matrix notation throughout and record the correspondence here once for cross\-reference\.
## 3\.The Power Law Graph Attention Operator
Throughout this section we work inside a single attention head of widthdk=dmodel/hd\_\{k\}=d\_\{\\mathrm\{model\}\}/h, wherehhis the number of heads\. All statements extend head\-wise; the multi\-head structure is discussed in Section[3\.4](https://arxiv.org/html/2608.10288#S3.SS4)\.
### 3\.1\.Formal definition
###### Definition 3\.1\(iSwiGLU\)\.
The*identity\-SwiGLU*activation is the mapiSwiGLU:ℝ→ℝ≥0\\operatorname\{iSwiGLU\}:\\mathbb\{R\}\\to\\mathbb\{R\}\_\{\\geq 0\},
\(3\.1\)iSwiGLU\(u\)=σs\(u\)⋅u=u2ς\(u\)≥0,\\operatorname\{iSwiGLU\}\(u\)\\;=\\;\\sigma\_\{s\}\(u\)\\cdot u\\;=\\;u^\{2\}\\,\\varsigma\(u\)\\;\\geq\\;0,applied element\-wise; it is SwiGLU\[[35](https://arxiv.org/html/2608.10288#bib.bib8),[5](https://arxiv.org/html/2608.10288#bib.bib9)\]with both weight matrices set to the identity and no bias\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\]\. It is smooth, non\-negative, and vanishes only atu=0u=0\.
###### Definition 3\.2\(Power Law Graph Attention\[[11](https://arxiv.org/html/2608.10288#bib.bib2),[12](https://arxiv.org/html/2608.10288#bib.bib3),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\)\.
LetQ,K,V∈ℝS×dkQ,K,V\\in\\mathbb\{R\}^\{S\\times d\_\{k\}\}be query, key, and value matrices for a context of lengthSS\. LetΦres:ℝdk×dk→ℝdk×dk\\Phi\_\{\\mathrm\{res\}\}:\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}\\to\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}be a deep residual network,*shared by all heads of the decoder layer*, consisting in the reference design ofNres=8N\_\{\\mathrm\{res\}\}=8residual units, each applyingnA=2n\_\{A\}=2SwiGLU blocks in succession, where each block is a complete gated mapℝdk→ℝdk\\mathbb\{R\}^\{d\_\{k\}\}\\to\\mathbb\{R\}^\{d\_\{k\}\}with hidden widthAdffA\_\{d\\\!f\\\!f\}\(two linear mapsdk→Adffd\_\{k\}\\to A\_\{d\\\!f\\\!f\}whose outputs are multiplied elementwise, followed by a linear mapAdff→dkA\_\{d\\\!f\\\!f\}\\to d\_\{k\}, acting on the last axis\), followed by a residual sum and LayerNorm; it therefore acts*row\-wise*on its matrix argument \(Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\)\. LetW,bW,P,a,ba∈ℝdk×dkW,b\_\{W\},P,a,b\_\{a\}\\in\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}be five full parameter matrices per head, andϵ\>0\\epsilon\>0a small constant \(ϵ=10−9\\epsilon=10^\{\-9\}\)\. PLGA is the composite map defined by
\(3\.2\)DQ\\displaystyle D\_\{Q\}=Q⊤Q\\displaystyle\\;=\\;Q^\{\\top\}Q\(“density operator”: query Gram matrix\)\(3\.3\)A\\displaystyle A=Φres\(LN\(DQ\)\)\\displaystyle\\;=\\;\\Phi\_\{\\mathrm\{res\}\}\\bigl\(\\operatorname\{LN\}\(D\_\{Q\}\)\\bigr\)\(generator\)\(3\.4\)ALM\\displaystyle A\_\{\\mathrm\{LM\}\}=iSwiGLU\(WA\+bW\)\+ϵ\\displaystyle\\;=\\;\\operatorname\{iSwiGLU\}\\bigl\(WA\+b\_\{W\}\\bigr\)\+\\epsilon\(“metric tensor”: positive interaction tensor\)\(3\.5\)AP\\displaystyle A\_\{P\}=ALM⊙P\\displaystyle\\;=\\;A\_\{\\mathrm\{LM\}\}^\{\\odot P\}\(potential tensor\)\(3\.6\)GLM\\displaystyle G\_\{\\mathrm\{LM\}\}=aAP\+ba\\displaystyle\\;=\\;a\\,A\_\{P\}\+b\_\{a\}\(“energy–curvature”: bilinear score operator\)\(3\.7\)E\\displaystyle E=QGLMK⊤dk\\displaystyle\\;=\\;\\frac\{Q\\,G\_\{\\mathrm\{LM\}\}\\,K^\{\\top\}\}\{\\sqrt\{d\_\{k\}\}\}\(scores\)\(3\.8\)ELM\\displaystyle E\_\{\\mathrm\{LM\}\}=softmax\[mask\(E\)\]\\displaystyle\\;=\\;\\operatorname\{softmax\}\\bigl\[\\operatorname\{mask\}\(E\)\\bigr\]\(attention operator\)\(3\.9\)VLM\\displaystyle V\_\{\\mathrm\{LM\}\}=ELMV\\displaystyle\\;=\\;E\_\{\\mathrm\{LM\}\}V\(inductive head output\)\.In \([3\.4](https://arxiv.org/html/2608.10288#S3.E4)\) and \([3\.6](https://arxiv.org/html/2608.10288#S3.E6)\),WAWAandaAPa\\,A\_\{P\}are ordinary matrix products \(WWandaaact on the left\), and the biases are full matrices added entry\-wise; component\-wise,
\(3\.10\)\(ALM\)ij=iSwiGLU\(∑kWikAkj\+\(bW\)ij\)\+ϵ,\(GLM\)ij=∑kaik\(AP\)kj\+\(ba\)ij\.\(A\_\{\\mathrm\{LM\}\}\)\_\{ij\}=\\operatorname\{iSwiGLU\}\\Bigl\(\\textstyle\\sum\_\{k\}W\_\{ik\}A\_\{kj\}\+\(b\_\{W\}\)\_\{ij\}\\Bigr\)\+\\epsilon,\\qquad\(G\_\{\\mathrm\{LM\}\}\)\_\{ij\}=\\textstyle\\sum\_\{k\}a\_\{ik\}\\,\(A\_\{P\}\)\_\{kj\}\+\(b\_\{a\}\)\_\{ij\}\.RowiiofGLMG\_\{\\mathrm\{LM\}\}is thus a coupling\-weighted*superposition of the interaction profiles*\(rows\) of the potential tensor; this is the precise sense of “superposition of potentials” in\[[11](https://arxiv.org/html/2608.10288#bib.bib2),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The tuple\(A,ALM,AP,GLM,ELM\)\(A,A\_\{\\mathrm\{LM\}\},A\_\{P\},G\_\{\\mathrm\{LM\}\},E\_\{\\mathrm\{LM\}\}\)constitutes the*deductive outputs*;VLMV\_\{\\mathrm\{LM\}\}is the*inductive output*of the head\.mask\\operatorname\{mask\}restricts the*score support*to the causal index sets𝒥i=\{j:j≤i\}\\mathcal\{J\}\_\{i\}=\\\{j:j\\leq i\\\}: in this paper’s theorems it is the ideal masked softmax of Section[2](https://arxiv.org/html/2608.10288#S2); the implementations add a large finite negative constant to disallowed scores before the softmax, with the provenance\-specific values as recorded in Section[2](https://arxiv.org/html/2608.10288#S2)\(Remark[3\.13](https://arxiv.org/html/2608.10288#S3.Thmtheorem13)\)\. The mask constrains which keys each row may attend to; what score\-support masking does and does not imply for the row\-wise dependence structure of the full model is stated precisely in Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\.
### 3\.2\.Basic structural properties
###### Proposition 3\.6\(Density operator\)\.
DQ=Q⊤QD\_\{Q\}=Q^\{\\top\}Qis symmetric positive semi\-definite withrankDQ≤min\(S,dk\)\\operatorname\{rank\}D\_\{Q\}\\leq\\min\(S,d\_\{k\}\), and1SDQ\\tfrac\{1\}\{S\}D\_\{Q\}is the \(uncentered\) second moment of the query token vectors\. In particular, for short contexts \(S<dkS<d\_\{k\}\) the raw instance geometry seen by the metric learner is rank\-deficient, and any full\-rank structure inALMA\_\{\\mathrm\{LM\}\}is necessarily*generated*byΦres\\Phi\_\{\\mathrm\{res\}\}and the downstream maps rather than present as rank in the single\-instance Gram matrix\. \(Generation downstream does not by itself attribute the structure to training data: biases, the nonlinearity, or random initialization can produce full\-rankALMA\_\{\\mathrm\{LM\}\}from a rank\-deficient input\. Attribution requires a trained\-versus\-initialized comparison, which Appendix[D](https://arxiv.org/html/2608.10288#A4)carries out: the same spectral battery run at random initialization\.\)
###### Proof\.
v⊤Q⊤Qv=‖Qv‖22≥0v^\{\\top\}Q^\{\\top\}Qv=\\left\\lVert Qv\\right\\rVert\_\{2\}^\{2\}\\geq 0andrank\(Q⊤Q\)=rankQ≤min\(S,dk\)\\operatorname\{rank\}\(Q^\{\\top\}Q\)=\\operatorname\{rank\}Q\\leq\\min\(S,d\_\{k\}\)\. The moment statement is the definition of the empirical second moment\. ∎
###### Proposition 3\.7\(Strict positivity and well\-posedness of the power law\)\.
ALM≥ϵ\>0A\_\{\\mathrm\{LM\}\}\\geq\\epsilon\>0entry\-wise\. Consequently the element\-wise power \([2\.1](https://arxiv.org/html/2608.10288#S2.E1)\) in \([3\.5](https://arxiv.org/html/2608.10288#S3.E5)\) is well defined and jointly smooth in\(ALM,P\)\(A\_\{\\mathrm\{LM\}\},P\), and for fixedALMA\_\{\\mathrm\{LM\}\}the family\{ALM⊙tP\}t∈ℝ\\\{A\_\{\\mathrm\{LM\}\}^\{\\odot tP\}\\\}\_\{t\\in\\mathbb\{R\}\}is a one\-parameter multiplicative group:
ALM⊙\(s\+t\)P=ALM⊙sP⊙ALM⊙tP,ALM⊙0=𝟏𝟏⊤\.A\_\{\\mathrm\{LM\}\}^\{\\odot\(s\+t\)P\}=A\_\{\\mathrm\{LM\}\}^\{\\odot sP\}\\odot A\_\{\\mathrm\{LM\}\}^\{\\odot tP\},\\qquad A\_\{\\mathrm\{LM\}\}^\{\\odot 0\}=\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}\.Equivalently,L\(t\)=logALM⊙tP=t\(P⊙logALM\)L\(t\)=\\log A\_\{\\mathrm\{LM\}\}^\{\\odot tP\}=t\\,\(P\\odot\\log A\_\{\\mathrm\{LM\}\}\)solves the linear flowL˙=P⊙logALM\\dot\{L\}=P\\odot\\log A\_\{\\mathrm\{LM\}\}: the potential tensor is the time\-11point of a linear dynamical system in log\-space\.
###### Proof\.
iSwiGLU≥0\\operatorname\{iSwiGLU\}\\geq 0by Definition[3\.1](https://arxiv.org/html/2608.10288#S3.Thmtheorem1), soALM≥ϵA\_\{\\mathrm\{LM\}\}\\geq\\epsilonentry\-wise by \([3\.4](https://arxiv.org/html/2608.10288#S3.E4)\)\. The remaining statements follow from \([2\.1](https://arxiv.org/html/2608.10288#S2.E1)\) and elementary properties ofexp\\expandlog\\logapplied entry\-wise\. ∎
###### Theorem 3\.8\(Perron–Frobenius structure of the positive interaction tensor\)\.
Each head’s tensorALM∈ℝdk×dkA\_\{\\mathrm\{LM\}\}\\in\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}is entry\-wise positive; hence its spectral radiusρ\(ALM\)\\rho\(A\_\{\\mathrm\{LM\}\}\)is a simple, positive eigenvalue \(the Perron root\) with entry\-wise positive left and right eigenvectors, and every other eigenvalueλ\\lambdasatisfies\|λ\|<ρ\(ALM\)\\left\\lvert\\lambda\\right\\rvert<\\rho\(A\_\{\\mathrm\{LM\}\}\)\.
###### Proof\.
This is the classical Perron–Frobenius theorem for positive matrices; see\[[18](https://arxiv.org/html/2608.10288#bib.bib37), Ch\. 8\]or\[[33](https://arxiv.org/html/2608.10288#bib.bib20)\]\. Positivity ofALMA\_\{\\mathrm\{LM\}\}is Proposition[3\.7](https://arxiv.org/html/2608.10288#S3.Thmtheorem7)\. ∎
The empirical finding of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]is sharper: at convergence, the generatorAAhas \(to high precision\)*identical rows*, with the same row profile appearing on every head of a layer, and the determinants of bothAAandALMA\_\{\\mathrm\{LM\}\}evaluate to zero at floating\-point resolution on every head of every layer\. \(What these float\-zero determinants do and do not imply is quantified in Remark[3\.10](https://arxiv.org/html/2608.10288#S3.Thmtheorem10)and Appendix[D](https://arxiv.org/html/2608.10288#A4): forAA, the singular\-value spectrum confirms numerical rank one; forALMA\_\{\\mathrm\{LM\}\}, it does not\.\) The following elementary proposition is the exact statement of that limit configuration for the generator, which\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]calls the*learned singularity condition*\.
###### Proposition 3\.9\(Rank\-one singularity condition\)\.
SupposeA=𝟏α⊤A=\\mathbf\{1\}\\alpha^\{\\top\}for some row vectorα⊤∈ℝ1×dk\\alpha^\{\\top\}\\in\\mathbb\{R\}^\{1\\times d\_\{k\}\}\(identical rows\), and writes=α⊤𝟏=∑jαjs=\\alpha^\{\\top\}\\mathbf\{1\}=\\sum\_\{j\}\\alpha\_\{j\}for the common row sum\. Then:
1. \(i\)rankA≤1\\operatorname\{rank\}A\\leq 1, with equality if and only ifα≠0\\alpha\\neq 0, anddetA=0\\det A=0wheneverdk≥2d\_\{k\}\\geq 2;
2. \(ii\)A𝟏=s𝟏A\\mathbf\{1\}=s\\mathbf\{1\}and the characteristic polynomial ofAAisλdk−1\(λ−s\)\\lambda^\{d\_\{k\}\-1\}\(\\lambda\-s\), so the spectrum is\{s\}∪\{0\}\\\{s\\\}\\cup\\\{0\\\}fordk≥2d\_\{k\}\\geq 2and\{s\}\\\{s\\\}fordk=1d\_\{k\}=1; ifs≠0s\\neq 0, thenssis the unique nonzero eigenvalue,\|s\|=ρ\(A\)\\left\\lvert s\\right\\rvert=\\rho\(A\), andA/sA/sis idempotent, a rank\-one \(oblique\) projection; ifs=0s=0andα≠0\\alpha\\neq 0, thenAAis nonzero and nilpotent of index22\(A2=0A^\{2\}=0\), with spectral radius0;
3. \(iii\)Ak=sk−1AA^\{k\}=s^\{k\-1\}Afor allk≥1k\\geq 1\.
4. \(iv\)\(Perturbation, singular\-value form\.\) IfA′=𝟏α⊤\+ΔA^\{\\prime\}=\\mathbf\{1\}\\alpha^\{\\top\}\+\\Deltawith‖Δ‖2≤δ\\left\\lVert\\Delta\\right\\rVert\_\{2\}\\leq\\delta, thenσj\(A′\)≤δ\\sigma\_\{j\}\(A^\{\\prime\}\)\\leq\\deltafor everyj≥2j\\geq 2, and hence \|detA′\|≤\(‖𝟏α⊤‖2\+δ\)δdk−1=\(dk‖α‖2\+δ\)δdk−1\.\\left\\lvert\\det A^\{\\prime\}\\right\\rvert\\;\\leq\\;\\bigl\(\\left\\lVert\\mathbf\{1\}\\alpha^\{\\top\}\\right\\rVert\_\{2\}\+\\delta\\bigr\)\\,\\delta^\{d\_\{k\}\-1\}\\;=\\;\\bigl\(\\sqrt\{d\_\{k\}\}\\,\\left\\lVert\\alpha\\right\\rVert\_\{2\}\+\\delta\\bigr\)\\,\\delta^\{d\_\{k\}\-1\}\.No conclusion about the locations of individual*eigenvalues*ofA′A^\{\\prime\}is drawn:A′A^\{\\prime\}is in general nonnormal, and its eigenvalues can move at rateδ\\sqrt\{\\delta\}under a perturbation of sizeδ\\delta\(take𝟏α⊤=\(1−11−1\)\\mathbf\{1\}\\alpha^\{\\top\}=\\begin\{pmatrix\}1&\-1\\\\ 1&\-1\\end\{pmatrix\}andΔ=\(00δ0\)\\Delta=\\begin\{pmatrix\}0&0\\\\ \\delta&0\\end\{pmatrix\}: the eigenvalues ofA′A^\{\\prime\}are±iδ\\pm i\\sqrt\{\\delta\}anddetA′=δ\\det A^\{\\prime\}=\\delta\)\.111Weyl’s eigenvalue inequality is valid only for Hermitian \(more generally normal\) matrices and does not apply toA′A^\{\\prime\}; the counterexample displayed here refutes the eigenvalue and determinant bounds it would otherwise suggest\. Eigenvalue\-location statements for the nonnormalA′A^\{\\prime\}would require normality or explicit pseudospectral/eigenbasis\-conditioning hypotheses, which we do not impose\.
###### Proof\.
\(i\)–\(ii\):Av=𝟏\(α⊤v\)Av=\\mathbf\{1\}\(\\alpha^\{\\top\}v\), so the range is contained inspan\{𝟏\}\\operatorname\{span\}\\\{\\mathbf\{1\}\\\}, proper iffα≠0\\alpha\\neq 0;A𝟏=s𝟏A\\mathbf\{1\}=s\\mathbf\{1\}; sincerankA≤1\\operatorname\{rank\}A\\leq 1, the eigenvalue0has geometric \(hence algebraic\) multiplicity at leastdk−1d\_\{k\}\-1, and the tracessaccounts for the remaining root, givingλdk−1\(λ−s\)\\lambda^\{d\_\{k\}\-1\}\(\\lambda\-s\)\. Ifs=0s=0andα≠0\\alpha\\neq 0thenA≠0A\\neq 0whileA2=𝟏\(α⊤𝟏\)α⊤=0A^\{2\}=\\mathbf\{1\}\(\\alpha^\{\\top\}\\mathbf\{1\}\)\\alpha^\{\\top\}=0\. \(iii\):A2=sAA^\{2\}=sAand induction\. \(iv\):σj\(A′\)≤σj\(𝟏α⊤\)\+‖Δ‖2≤0\+δ\\sigma\_\{j\}\(A^\{\\prime\}\)\\leq\\sigma\_\{j\}\(\\mathbf\{1\}\\alpha^\{\\top\}\)\+\\left\\lVert\\Delta\\right\\rVert\_\{2\}\\leq 0\+\\deltaforj≥2j\\geq 2by Weyl’s inequality*for singular values*\[[18](https://arxiv.org/html/2608.10288#bib.bib37), Cor\. 7\.3\.5\(a\) and eq\. \(7\.3\.13\)\], since𝟏α⊤\\mathbf\{1\}\\alpha^\{\\top\}has rank≤1\\leq 1; then\|detA′\|=∏jσj\(A′\)≤σ1\(A′\)δdk−1\\left\\lvert\\det A^\{\\prime\}\\right\\rvert=\\prod\_\{j\}\\sigma\_\{j\}\(A^\{\\prime\}\)\\leq\\sigma\_\{1\}\(A^\{\\prime\}\)\\,\\delta^\{d\_\{k\}\-1\}andσ1\(A′\)≤‖𝟏α⊤‖2\+δ\\sigma\_\{1\}\(A^\{\\prime\}\)\\leq\\left\\lVert\\mathbf\{1\}\\alpha^\{\\top\}\\right\\rVert\_\{2\}\+\\delta, with‖𝟏α⊤‖2=‖𝟏‖2‖α‖2=dk‖α‖2\\left\\lVert\\mathbf\{1\}\\alpha^\{\\top\}\\right\\rVert\_\{2\}=\\left\\lVert\\mathbf\{1\}\\right\\rVert\_\{2\}\\left\\lVert\\alpha\\right\\rVert\_\{2\}=\\sqrt\{d\_\{k\}\}\\,\\left\\lVert\\alpha\\right\\rVert\_\{2\}\. ∎
To state the next result precisely, we make the internal structure of the metric learner explicit\. Each of itsNresN\_\{\\mathrm\{res\}\}residual units acts on a single rowr∈ℝdkr\\in\\mathbb\{R\}^\{d\_\{k\}\}of its matrix argument as
\(3\.11\)uj\(r\)=LN\(r\+gj,2\(gj,1\(r\)\)\),j=1,…,Nres,u\_\{j\}\(r\)\\;=\\;\\operatorname\{LN\}\\bigl\(r\+g\_\{j,2\}\(g\_\{j,1\}\(r\)\)\\bigr\),\\qquad j=1,\\dots,N\_\{\\mathrm\{res\}\},wheregj,1,gj,2:ℝdk→ℝdkg\_\{j,1\},g\_\{j,2\}:\\mathbb\{R\}^\{d\_\{k\}\}\\to\\mathbb\{R\}^\{d\_\{k\}\}are the two successive complete SwiGLU blocks of unitjj, each a gated map with hidden widthAdffA\_\{d\\\!f\\\!f\}\(Definition[3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2)\); in the reference code each block is aGLUVariantmodule with twodk→Adffd\_\{k\}\\to A\_\{d\\\!f\\\!f\}linear maps multiplied elementwise and adkd\_\{k\}\-dimensional linear output map\.222The typing matters: the natural\-looking alternativegj,1:ℝdk→ℝAdffg\_\{j,1\}:\\mathbb\{R\}^\{d\_\{k\}\}\\to\\mathbb\{R\}^\{A\_\{d\\\!f\\\!f\}\},gj,2:ℝAdff→ℝdkg\_\{j,2\}:\\mathbb\{R\}^\{A\_\{d\\\!f\\\!f\}\}\\to\\mathbb\{R\}^\{d\_\{k\}\}would describe the two internal halves of a single feed\-forward block, whereasResLayerAin the reference implementations applies two full gated blocks in succession; the typing above matches the code\.When the dependence on the input sequencexxmatters, we writeQ\(x\)Q\(x\)for the query matrix produced on inputxxand, correspondingly,D\(x\)=Q\(x\)⊤Q\(x\)D\(x\)=Q\(x\)^\{\\top\}Q\(x\)andA\(x\)A\(x\)for the resulting density operator and metric generator\.
###### Proposition 3\.11\(Row\-factorization and permutation equivariance of the metric learner\)\.
In the reference implementations, every layer ofΦres\\Phi\_\{\\mathrm\{res\}\}\(SwiGLU blocks acting on the last axis, residual sums, and LayerNorm over the last axis\) acts on each row of itsdk×dkd\_\{k\}\\times d\_\{k\}argument independently and identically\. Consequently there is a single mapφ=uNres∘⋯∘u1:ℝdk→ℝdk\\varphi=u\_\{N\_\{\\mathrm\{res\}\}\}\\circ\\cdots\\circ u\_\{1\}:\\mathbb\{R\}^\{d\_\{k\}\}\\to\\mathbb\{R\}^\{d\_\{k\}\}\(the*row map*, the composition of the residual units \([3\.11](https://arxiv.org/html/2608.10288#S3.E11)\)\) such that
Φres\(M\)i,:=φ\(Mi,:\),i=1,…,dk,\\Phi\_\{\\mathrm\{res\}\}\(M\)\_\{i,:\}\\;=\\;\\varphi\\bigl\(M\_\{i,:\}\\bigr\),\\qquad i=1,\\dots,d\_\{k\},for every inputMM\. Hence:
1. \(i\)\(Equivariance\.\)Φres\(ΠM\)=ΠΦres\(M\)\\Phi\_\{\\mathrm\{res\}\}\(\\Pi M\)=\\Pi\\,\\Phi\_\{\\mathrm\{res\}\}\(M\)for every row permutationΠ\\Pi\.
2. \(ii\)\(Collapse criterion\.\) The generatorA\(x\)A\(x\)has identical rows with one*common*row valueα∗⊤\\alpha^\{\\ast\\top\}shared across all inputsxxin a set𝒳\\mathcal\{X\}if and only ifφ\\varphiis constant on the union of all rows ofLN\(D\(x\)\)\\operatorname\{LN\}\(D\(x\)\),x∈𝒳x\\in\\mathcal\{X\}\. \(Identical rows within each single input separately is the weaker condition thatφ\\varphiis constant on each input’s row set; constancy on the union is what the cross\-input invariance of Section[5](https://arxiv.org/html/2608.10288#S5)requires\.\)
3. \(iii\)\(Cross\-head identity\.\) Sinceφ\\varphiis shared by all heads of a layer, constancy ofφ\\varphiforcesA\(i\)=𝟏α∗⊤A^\{\(i\)\}=\\mathbf\{1\}\\,\\alpha^\{\\ast\\top\}with the*same*α∗\\alpha^\{\\ast\}for every headii, even though the heads’ density operators differ\.
4. \(iv\)\(Lipschitz transfer\.\)‖A\(x\)−A\(x′\)‖F≤Lip\(φ\)‖LN\(D\(x\)\)−LN\(D\(x′\)\)‖F\\left\\lVert A\(x\)\-A\(x^\{\\prime\}\)\\right\\rVert\_\{F\}\\leq\\operatorname\{Lip\}\(\\varphi\)\\,\\left\\lVert\\operatorname\{LN\}\(D\(x\)\)\-\\operatorname\{LN\}\(D\(x^\{\\prime\}\)\)\\right\\rVert\_\{F\}wheneverφ\\varphiisLip\(φ\)\\operatorname\{Lip\}\(\\varphi\)\-Lipschitz on the visited row set\.
###### Proof\.
Each SwiGLU block is a map applied to the last axis, i\.e\. to each row separately with shared weights; LayerNorm over the last axis normalizes each row separately; residual addition is row\-wise\. A composition of row\-wise maps with shared parameters is row\-wise with a single shared row mapφ\\varphi\. \(i\) follows since applying the same function to permuted rows permutes the outputs\. \(ii\): a common output value across all rows of all inputs in𝒳\\mathcal\{X\}is exactly constancy ofφ\\varphion the union of the row sets; per\-input identical rows alone constrainsφ\\varphionly on each row set separately\. \(iii\): apply \(ii\) per head and noteφ\\varphiis the same map for all heads \(the parameter tensorsW,bW,P,a,baW,b\_\{W\},P,a,b\_\{a\}downstream are per\-head, butΦres\\Phi\_\{\\mathrm\{res\}\}is instantiated once per layer\)\. \(iv\) is the definition of a Lipschitz constant applied row\-wise and summed in Frobenius norm\. ∎
Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\(iii\) turns the observed cross\-head identity ofAA\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]from a curiosity into*diagnostic evidence*: the heads feed different density operators to the sameφ\\varphi, and a locally constantφ\\varphiwould force exactly the observed agreement\. Observed equality on the visited set does not by itself identify local constancy \(equal or symmetry\-related head inputs, or pointwise agreement without flatness, are alternatives\); the direct discriminating measurements, composite Jacobians and pairwise contraction ratios of the trainedφ\\varphion visited rows, are reported in Appendix[D](https://arxiv.org/html/2608.10288#A4)and support the locally\-constant reading on the audited checkpoint\. The mechanism is analyzed in Section[5\.4](https://arxiv.org/html/2608.10288#S5.SS4)\.
###### Proposition 3\.12\(The ideal attention operator is Markov\)\.
With the ideal masked softmax of Section[2](https://arxiv.org/html/2608.10288#S2),ELME\_\{\\mathrm\{LM\}\}is row\-stochastic with exact causal support:ELM≥0E\_\{\\mathrm\{LM\}\}\\geq 0,ELM𝟏=𝟏E\_\{\\mathrm\{LM\}\}\\mathbf\{1\}=\\mathbf\{1\}, and\(ELM\)ij=0\(E\_\{\\mathrm\{LM\}\}\)\_\{ij\}=0forj\>ij\>i, with\(ELM\)ij\>0\(E\_\{\\mathrm\{LM\}\}\)\_\{ij\}\>0forj≤ij\\leq i\. Consequently:
1. \(i\)‖ELM‖∞=1\\left\\lVert E\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{\\infty\}=1and every eigenvalue ofELME\_\{\\mathrm\{LM\}\}lies in the closed unit disk, with11always attained \(lower\-triangular stochastic structure gives spectrum equal to the diagonal entries in the causal case\);
2. \(ii\)rowiiofVLM=ELMVV\_\{\\mathrm\{LM\}\}=E\_\{\\mathrm\{LM\}\}Vis a convex combination \(an expectation\) of value vectors of positionsj≤ij\\leq i:VLM\[i\]=𝔼j∼ELM\[i,⋅\]V\[j\]V\_\{\\mathrm\{LM\}\}\[i\]=\\mathbb\{E\}\_\{j\\sim E\_\{\\mathrm\{LM\}\}\[i,\\cdot\]\}\\,V\[j\]\.
###### Proof\.
Rows of the ideal masked softmax are probability vectors supported exactly on the allowed set\{j≤i\}\\\{j\\leq i\\\}\. \(i\) is standard for stochastic matrices \(Gershgorin or the sub\-multiplicativity of∥⋅∥∞\\left\\lVert\\cdot\\right\\rVert\_\{\\infty\}\); a causal \(lower\-triangular\) matrix has its eigenvalues on the diagonal\. \(ii\) is the definition of matrix multiplication with stochastic rows\. ∎
Proposition[3\.12](https://arxiv.org/html/2608.10288#S3.Thmtheorem12)is a statement about the*support*of the attention operator: rowiimixes values of positionsj≤ij\\leq ionly\. It does not by itself constrain how the scores and the operatorGLMG\_\{\\mathrm\{LM\}\}entering rowiidepend on the input, and in PLGA they depend on*all*supplied rows through the density operator \([3\.2](https://arxiv.org/html/2608.10288#S3.E2)\)\. The row\-wise dependence property that lower\-triangular support does and does not deliver is stated precisely in Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5); the deployed final\-row generation interface \(Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)\) does not require the stronger property\.
Proposition[3\.12](https://arxiv.org/html/2608.10288#S3.Thmtheorem12)is elementary but load\-bearing: PLGA’s inductive action is that of an input\-adapted*averaging operator*on the token graph, the same structural role played by transfer operators in ergodic theory and by graph Laplacian smoothers\.
### 3\.3\.PLGA as attention with a learned bilinear form
SDPA computes scoresQK⊤/dkQK^\{\\top\}/\\sqrt\{d\_\{k\}\}: the bilinear form comparing queries with keys is the Euclidean one,B\(q,k\)=q⊤IkB\(q,k\)=q^\{\\top\}Ik\. PLGA computesQGLMK⊤/dkQG\_\{\\mathrm\{LM\}\}K^\{\\top\}/\\sqrt\{d\_\{k\}\}: the form isBG\(q,k\)=q⊤GLMkB\_\{G\}\(q,k\)=q^\{\\top\}G\_\{\\mathrm\{LM\}\}k, withGLMG\_\{\\mathrm\{LM\}\}produced by the nonlinear chain \([3\.2](https://arxiv.org/html/2608.10288#S3.E2)\)–\([3\.6](https://arxiv.org/html/2608.10288#S3.E6)\) from the input itself\. Three regimes must be distinguished:
1. \(1\)GLM≡IG\_\{\\mathrm\{LM\}\}\\equiv I\(frozen identity\)\.PLGA is exactly SDPA \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(i\)\)\.
2. \(2\)GLM=G∗G\_\{\\mathrm\{LM\}\}=G^\{\\ast\}frozen but arbitrary\.A generalized SDPA with a learned constant bilinear operator\. Without position\-dependent transforms this is linearly equivalent to SDPA by absorbingG∗G^\{\\ast\}into the query projection; with rotary embeddings the equivalence generically fails \(Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\), so even the frozen head realizes score functions outside the SDPA\-realizable family, in the conditional, head\-level sense counted by Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1)\.
3. \(3\)GLM=GLM\(x\)G\_\{\\mathrm\{LM\}\}=G\_\{\\mathrm\{LM\}\}\(x\)input\-generated \(training\-time PLGA\)\.The score mapQ↦QGLM\(Q\)K⊤Q\\mapsto Q\\,G\_\{\\mathrm\{LM\}\}\(Q\)\\,K^\{\\top\}is generically nonlinear in the input \(degenerate parameter choices linearize it:a=0a=0makesGLM≡baG\_\{\\mathrm\{LM\}\}\\equiv b\_\{a\}constant and the head falls back to regime 2\), and gradients flow throughΦres,W,P,a\\Phi\_\{\\mathrm\{res\}\},W,P,a; this regime is*not*reducible to SDPA, which is the content of the training/inference asymmetry \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(iii\)\)\.
### 3\.4\.Multi\-head structure as a local decomposition
Withhhheads, the model spaceℝdmodel\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\}is split as an internal direct sum⨁i=1hℝdk\\bigoplus\_\{i=1\}^\{h\}\\mathbb\{R\}^\{d\_\{k\}\}and each head learns its own tuple\(A\(i\),ALM\(i\),AP\(i\),GLM\(i\),ELM\(i\)\)\(A^\{\(i\)\},A\_\{\\mathrm\{LM\}\}^\{\(i\)\},A\_\{P\}^\{\(i\)\},G\_\{\\mathrm\{LM\}\}^\{\(i\)\},E\_\{\\mathrm\{LM\}\}^\{\(i\)\}\); outputs are concatenated and mixed by a linear map\. Formally, the deductive state of one decoder layer is the block\-diagonal operator
\(3\.12\)GLMlayer=⨁i=1hGLM\(i\)∈ℝdmodel×dmodel\(in the head\-aligned basis\),G\_\{\\mathrm\{LM\}\}^\{\\mathrm\{layer\}\}\\;=\\;\\bigoplus\_\{i=1\}^\{h\}G\_\{\\mathrm\{LM\}\}^\{\(i\)\}\\;\\in\\;\\mathbb\{R\}^\{d\_\{\\mathrm\{model\}\}\\times d\_\{\\mathrm\{model\}\}\}\\quad\\text\{\(in the head\-aligned basis\)\},so the layer’s score computation is assembled from mutually non\-interacting local components, each acting on its own subspace\[[11](https://arxiv.org/html/2608.10288#bib.bib2)\]\. The output projectionWOW\_\{O\}then mixes the head*outputs*linearly, after the scores are formed; it does not act on, or conjugate, the per\-head score operators themselves\. The division of parameters respects this structure: the metric\-learner networkΦres\\Phi\_\{\\mathrm\{res\}\}is shared across the heads of a layer \(Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\), while the five tensorsW,bW,P,a,baW,b\_\{W\},P,a,b\_\{a\}are per\-head: heads are differentiated*after*the common metric generator, not before\.
## 4\.The PLDR\-LLM Architecture
### 4\.1\.Rotary position embeddings and a commutant characterization
PLDR\-LLM applies rotary position embeddings \(RoPE\)\[[38](https://arxiv.org/html/2608.10288#bib.bib10)\]to queries and keys after head\-splitting\. RoPE at positionnnis the orthogonal block\-diagonal rotation
\(4\.1\)Rn=⨁j=1dk/2Rot\(nθj\),Rot\(ϕ\)=\(cosϕ−sinϕsinϕcosϕ\),θj=Θ−2\(j−1\)/dk,R\_\{n\}\\;=\\;\\bigoplus\_\{j=1\}^\{d\_\{k\}/2\}\\operatorname\{Rot\}\(n\\theta\_\{j\}\),\\qquad\\operatorname\{Rot\}\(\\phi\)=\\begin\{pmatrix\}\\cos\\phi&\-\\sin\\phi\\\\ \\sin\\phi&\\cos\\phi\\end\{pmatrix\},\\qquad\\theta\_\{j\}=\\Theta^\{\-2\(j\-1\)/d\_\{k\}\},withΘ=104\\Theta=10^\{4\}in the reference configuration, son↦Rnn\\mapsto R\_\{n\}is a group homomorphismℤ→SO\(dk\)\\mathbb\{Z\}\\to SO\(d\_\{k\}\)into a maximal torusT≅\(S1\)dk/2T\\cong\(S^\{1\}\)^\{d\_\{k\}/2\}\(the*ambient RoPE torus*\), withRn⊤=R−nR\_\{n\}^\{\\top\}=R\_\{\-n\}andRnRm=Rn\+mR\_\{n\}R\_\{m\}=R\_\{n\+m\}\. Scores become
\(4\.2\)Enm=\(Rnqn\)⊤GLM\(Rmkm\)dk=qn⊤R−nGLMRmkmdk\.E\_\{nm\}\\;=\\;\\frac\{\(R\_\{n\}q\_\{n\}\)^\{\\top\}\\,G\_\{\\mathrm\{LM\}\}\\,\(R\_\{m\}k\_\{m\}\)\}\{\\sqrt\{d\_\{k\}\}\}\\;=\\;\\frac\{q\_\{n\}^\{\\top\}\\,R\_\{\-n\}\\,G\_\{\\mathrm\{LM\}\}\\,R\_\{m\}\\,k\_\{m\}\}\{\\sqrt\{d\_\{k\}\}\}\.For SDPA \(GLM=IG\_\{\\mathrm\{LM\}\}=I\) one hasR−nIRm=Rm−nR\_\{\-n\}IR\_\{m\}=R\_\{m\-n\}: scores depend on positions only through the offsetm−nm\-n\(the celebrated relative\-position property\)\. With a learned metric the situation is characterized exactly:
###### Proposition 4\.1\(Commutant characterization of relative\-position invariance\)\.
Assume the rotation anglesθ1,…,θdk/2\\theta\_\{1\},\\dots,\\theta\_\{d\_\{k\}/2\}satisfy the*nonresonance conditions*
θa≢±θb\(mod2π\)\(a≠b\),θa≢0\(modπ\)\(alla\),\\theta\_\{a\}\\not\\equiv\\pm\\,\\theta\_\{b\}\\pmod\{2\\pi\}\\quad\(a\\neq b\),\\qquad\\theta\_\{a\}\\not\\equiv 0\\pmod\{\\pi\}\\quad\(\\text\{all \}a\),i\.e\. the eigenvaluese±iθae^\{\\pm i\\theta\_\{a\}\}ofR1R\_\{1\}aredkd\_\{k\}distinct non\-real numbers\. The standard frequenciesθj=Θ−2\(j−1\)/dk\\theta\_\{j\}=\\Theta^\{\-2\(j\-1\)/d\_\{k\}\}withΘ=104\\Theta=10^\{4\}satisfy these conditions: they lie in\(0,1\]⊂\(0,π\)\(0,1\]\\subset\(0,\\pi\)and are strictly decreasing, hence pairwise distinct with no pair summing to a multiple of2π2\\pi\.333Nonresonance is the correct hypothesis here\. The natural\-looking alternative, rational independence of theθa\\theta\_\{a\}with a density argument for the cyclic orbit in the torus, is both insufficient for density \(which requires rational independence of1,θ1/2π,…1,\\theta\_\{1\}/2\\pi,\\dots; e\.g\.θ=π\\theta=\\piis irrational yet has a finite orbit\) and false for the standard frequencies, which obey the exact rational relationsθj\+dk/8=θj/10\\theta\_\{j\+d\_\{k\}/8\}=\\theta\_\{j\}/10for base10410^\{4\}\. The nonresonance route avoids density altogether\.Then the map\(n,m\)↦R−nGRm\(n,m\)\\mapsto R\_\{\-n\}\\,G\\,R\_\{m\}depends only onm−nm\-n\(for all offsets\) if and only ifGGcommutes with every rotationRnR\_\{n\}\(equivalently, with the ambient RoPE torusTT\), which holds if and only ifGGis block\-diagonal with2×22\\times 2blocks of the form\(cj−sjsjcj\)\\begin\{pmatrix\}c\_\{j\}&\-s\_\{j\}\\\\ s\_\{j\}&c\_\{j\}\\end\{pmatrix\}, i\.e\.G∈⨁j\{cjI2\+sjJ2\}≅ℂdk/2G\\in\\bigoplus\_\{j\}\\\{c\_\{j\}I\_\{2\}\+s\_\{j\}J\_\{2\}\\\}\\cong\\mathbb\{C\}^\{d\_\{k\}/2\}, whereI2I\_\{2\}is the2×22\\times 2identity andJ2=\(0−110\)J\_\{2\}=\\begin\{pmatrix\}0&\-1\\\\ 1&0\\end\{pmatrix\}is the rotation byπ/2\\pi/2\(the complex structure of the plane\):GGacts as the complex scalarcj\+isjc\_\{j\}\+is\_\{j\}on thejj\-th rotation plane\. For generic learnedGLMG\_\{\\mathrm\{LM\}\}this fails, and PLGA scores carry*absolute*positional information through the conjugation orbitR−nGLMRnR\_\{\-n\}G\_\{\\mathrm\{LM\}\}R\_\{n\}\.
###### Proof\.
Offset dependence for all\(n,m\)\(n,m\)is equivalent toR−nGRm=R−n′GRm′R\_\{\-n\}GR\_\{m\}=R\_\{\-n^\{\\prime\}\}GR\_\{m^\{\\prime\}\}wheneverm−n=m′−n′m\-n=m^\{\\prime\}\-n^\{\\prime\}; takingn′=0n^\{\\prime\}=0,m′=m−nm^\{\\prime\}=m\-ngivesR−nGRn=GR\_\{\-n\}GR\_\{n\}=Gfor allnn, i\.e\.GGcommutes with the cyclic group generated byR1R\_\{1\}, which holds iffGGcommutes with the single matrixR1R\_\{1\}\(every element is a power ofR1R\_\{1\}\)\.R1R\_\{1\}is real semisimple; after complexification it is diagonal with eigenvaluese±iθae^\{\\pm i\\theta\_\{a\}\}, which under the nonresonance hypothesis aredkd\_\{k\}*distinct*numbers\. The commutant of a diagonalizable matrix with distinct eigenvalues is the algebra of matrices diagonal in the same eigenbasis; regrouping the conjugate eigenpairs into their real2×22\\times 2planes, the real matrices in this commutant are exactly the block\-diagonal matrices whosejj\-th block commutes withRot\(θj\)\\operatorname\{Rot\}\(\\theta\_\{j\}\), i\.e\. \(forθj∉πℤ\\theta\_\{j\}\\notin\\pi\\mathbb\{Z\}\) lies in\{cI2\+sJ2\}\\\{cI\_\{2\}\+sJ\_\{2\}\\\}, the algebra generated by the rotation itself \(isomorphic toℂ\\mathbb\{C\}\)\. Since everyRtR\_\{t\}is block\-diagonal with blocks in these algebras, any suchGGcommutes with everyRtR\_\{t\}, giving offset dependence; density of the orbit in the torus is not needed\. ∎
### 4\.2\.The full decoder map
###### Definition 4\.3\(PLDR\-LLM\[[12](https://arxiv.org/html/2608.10288#bib.bib3),[13](https://arxiv.org/html/2608.10288#bib.bib4)\]\)\.
Fix depthLL, headshh, widthdmodel=hdkd\_\{\\mathrm\{model\}\}=hd\_\{k\}, feed\-forward widthdffd\_\{f\\\!f\}, metric\-learner widthAdffA\_\{d\\\!f\\\!f\}, and vocabulary𝒱\\mathcal\{V\}\. The PLDR\-LLM is the length\-indexed family of mapsFθ:⋃S≤Smax𝒱S→⋃S≤SmaxΔ\(𝒱\)SF\_\{\\theta\}:\\bigcup\_\{S\\leq S\_\{\\max\}\}\\mathcal\{V\}^\{S\}\\to\\bigcup\_\{S\\leq S\_\{\\max\}\}\\Delta\(\\mathcal\{V\}\)^\{S\}\(row\-wise probability simplices\), with one output row per input row: an input of lengthSSis mapped intoΔ\(𝒱\)S\\Delta\(\\mathcal\{V\}\)^\{S\}, so\|Fθ\(x\)\|=\|x\|\\left\\lvert F\_\{\\theta\}\(x\)\\right\\rvert=\\left\\lvert x\\right\\rvertalways \(this length\-preservation statement is carried as an explicit predicate in the Lean development, Appendix[C](https://arxiv.org/html/2608.10288#A3)\)\. The computation is as follows\. WithX\(0\)=LN\(dmodel⋅ι\(x\)\)X^\{\(0\)\}=\\operatorname\{LN\}\\bigl\(\\sqrt\{d\_\{\\mathrm\{model\}\}\}\\,\\cdot\\,\\iota\(x\)\\bigr\), forℓ=1,…,L\\ell=1,\\dots,Land headsi=1,…,hi=1,\\dots,h:
\(4\.3\)Q\(ℓ,i\)\\displaystyle Q^\{\(\\ell,i\)\}=X\(ℓ−1\)WQ\(ℓ,i\)\+𝟏bQ\(ℓ,i\)⊤,K\(ℓ,i\)=X\(ℓ−1\)WK\(ℓ,i\)\+𝟏bK\(ℓ,i\)⊤,\\displaystyle=X^\{\(\\ell\-1\)\}W\_\{Q\}^\{\(\\ell,i\)\}\+\\mathbf\{1\}\\,b\_\{Q\}^\{\(\\ell,i\)\\top\},\\qquad K^\{\(\\ell,i\)\}=X^\{\(\\ell\-1\)\}W\_\{K\}^\{\(\\ell,i\)\}\+\\mathbf\{1\}\\,b\_\{K\}^\{\(\\ell,i\)\\top\},\(4\.4\)V\(ℓ,i\)\\displaystyle V^\{\(\\ell,i\)\}=X\(ℓ−1\)WV\(ℓ,i\)\+𝟏bV\(ℓ,i\)⊤,Q~\(ℓ,i\)=RoPE\(Q\(ℓ,i\)\),K~\(ℓ,i\)=RoPE\(K\(ℓ,i\)\),\\displaystyle=X^\{\(\\ell\-1\)\}W\_\{V\}^\{\(\\ell,i\)\}\+\\mathbf\{1\}\\,b\_\{V\}^\{\(\\ell,i\)\\top\},\\qquad\\widetilde\{Q\}^\{\(\\ell,i\)\}=\\operatorname\{RoPE\}\(Q^\{\(\\ell,i\)\}\),\\qquad\\widetilde\{K\}^\{\(\\ell,i\)\}=\\operatorname\{RoPE\}\(K^\{\(\\ell,i\)\}\),\(4\.5\)VLM\(ℓ,i\)\\displaystyle V\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\}=PLGA\(ℓ,i\)\(Q~\(ℓ,i\),K~\(ℓ,i\),V\(ℓ,i\)\),\\displaystyle=\\operatorname\{PLGA\}^\{\(\\ell,i\)\}\\bigl\(\\widetilde\{Q\}^\{\(\\ell,i\)\},\\widetilde\{K\}^\{\(\\ell,i\)\},V^\{\(\\ell,i\)\}\\bigr\),\(4\.6\)U\(ℓ\)\\displaystyle U^\{\(\\ell\)\}=LN\(X\(ℓ−1\)\+\[VLM\(ℓ,1\)‖⋯‖VLM\(ℓ,h\)\]WO\(ℓ\)\+𝟏bO\(ℓ\)⊤\),\\displaystyle=\\operatorname\{LN\}\\Bigl\(X^\{\(\\ell\-1\)\}\+\\bigl\[V\_\{\\mathrm\{LM\}\}^\{\(\\ell,1\)\}\\\|\\cdots\\\|V\_\{\\mathrm\{LM\}\}^\{\(\\ell,h\)\}\\bigr\]W\_\{O\}^\{\(\\ell\)\}\+\\mathbf\{1\}\\,b\_\{O\}^\{\(\\ell\)\\top\}\\Bigr\),\(4\.7\)X\(ℓ\)\\displaystyle X^\{\(\\ell\)\}=LN\(U\(ℓ\)\+SwiGLU\-FFN\(ℓ\)\(U\(ℓ\)\)\),\\displaystyle=\\operatorname\{LN\}\\Bigl\(U^\{\(\\ell\)\}\+\\operatorname\{SwiGLU\\text\{\-\}FFN\}^\{\(\\ell\)\}\\bigl\(U^\{\(\\ell\)\}\\bigr\)\\Bigr\),and finallyFθ\(x\)=softmax\(X\(L\)Wvocab\+𝟏bvocab⊤\)F\_\{\\theta\}\(x\)=\\operatorname\{softmax\}\\bigl\(X^\{\(L\)\}W\_\{\\mathrm\{vocab\}\}\+\\mathbf\{1\}\\,b\_\{\\mathrm\{vocab\}\}^\{\\top\}\\bigr\), with𝟏∈ℝS\\mathbf\{1\}\\in\\mathbb\{R\}^\{S\}the all\-ones column, so each bias adds its row to every position\. In \([4\.5](https://arxiv.org/html/2608.10288#S4.E5)\) each head applies the PLGA operator per Definition[3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2)withD~\(ℓ,i\)=Q~\(ℓ,i\)⊤Q~\(ℓ,i\)\\widetilde\{D\}^\{\(\\ell,i\)\}=\\widetilde\{Q\}^\{\(\\ell,i\)\\top\}\\widetilde\{Q\}^\{\(\\ell,i\)\}in place ofDQD\_\{Q\}: the density operator is built from the*rotated*queryQ~\\widetilde\{Q\}, so positional phases enter the metric learner \(Remark[4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2)\)\. The displayed affine terms are those of the released configuration \(biases enabled on all attention projections and on the unembedding\); the unembedding is not tied to the embeddingι\\iota, and the gated feed\-forward and metric\-learner blocks carry their own internal biases\. The deployed generative interface is*final\-row online generation*: at stepttthe model is invoked on exactly the known prefixx1:tx\_\{1:t\}, so the call hasS=tS=trows \(one\-based; in zero\-based code indexing the rows are0,…,S−10,\\dots,S\{\-\}1and the selected row isS−1S\{\-\}1\), and only the final output row is consumed,
\(4\.8\)pθ\(⋅∣x1:t\)=Fθ\(x1:t\)t,p\_\{\\theta\}\(\\,\\cdot\\mid x\_\{1:t\}\)\\;=\\;F\_\{\\theta\}\(x\_\{1:t\}\)\_\{t\},the final row of the call on the prefix itself\. Every tensor of the call, including the density operatorD~=Q~⊤Q~\\widetilde\{D\}=\\widetilde\{Q\}^\{\\top\}\\widetilde\{Q\}and the deductive family derived from it, is a function ofx1:tx\_\{1:t\}alone, so the generator defined by \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\) is causal: no token beyond positionttenters the conditional that predictsxt\+1x\_\{t\+1\}\. The causal mask in \([3\.8](https://arxiv.org/html/2608.10288#S3.E8)\) constrains the*score support*of every row; it does not make internal rowsr<Sr<Sof a longer call functions of their own prefixes, and they are not so interpreted \(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\)\.444One initialization detail distinguishes the released model generations: in the v510 code accompanying\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]the value projectionWVW\_\{V\}keeps PyTorch’s default uniform initialization whileWQ,WKW\_\{Q\},W\_\{K\}are Glorot\-uniform with zero bias \(a duplicated initializer call onWKW\_\{K\}; the detail is shared by the base v510 model and its DAG and ablation variants in that repository\); the code accompanying\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]initializesWVW\_\{V\}identically toWQ,WKW\_\{Q\},W\_\{K\}\. This shifts the near\-critical\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)region slightly\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\.
The deductive output of the full model is the indexed family
\(4\.9\)𝒟\(x;θ\)=\{A\(ℓ,i\),ALM\(ℓ,i\),AP\(ℓ,i\),GLM\(ℓ,i\)\}ℓ≤L,i≤h,\\mathcal\{D\}\(x;\\theta\)\\;=\\;\\bigl\\\{\\,A^\{\(\\ell,i\)\},\\ A\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\},\\ A\_\{P\}^\{\(\\ell,i\)\},\\ G\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\}\\,\\bigr\\\}\_\{\\ell\\leq L,\\,i\\leq h\},a point inℝ4Lhdk2\\mathbb\{R\}^\{4Lhd\_\{k\}^\{2\}\}, computed alongside the inductive output\.
### 4\.3\.Training objective and the DAG regularizer
Pretraining minimizes a*blockwise \(global\-context\) cross\-entropy*\. For a training block\(x0,…,xT−1\)\(x\_\{0\},\\dots,x\_\{T\-1\}\)the reference implementations run*one*forward pass on the shifted input\(x0,…,xT−2\)\(x\_\{0\},\\dots,x\_\{T\-2\}\)and apply cross\-entropy at every aligned output row:
\(4\.10\)ℒblock\(θ\)=−𝔼∑t=0T−2logqθ,t\(xt\+1;x0:T−2\),\\mathcal\{L\}\_\{\\mathrm\{block\}\}\(\\theta\)\\;=\\;\-\\,\\mathbb\{E\}\\sum\_\{t=0\}^\{T\-2\}\\log q\_\{\\theta,t\}\\bigl\(x\_\{t\+1\};\\,x\_\{0:T\-2\}\\bigr\),whereqθ,tq\_\{\\theta,t\}denotes the row\-ttconditional of the single full\-block call \(rowttofFθ\(x0:T−2\)F\_\{\\theta\}\(x\_\{0:T\-2\}\), evaluated atxt\+1x\_\{t\+1\}\)\. As implemented in both reference trainers, the objective is the masked, normalized form
ℒ^block\(θ\)=−∑b,tmbtlogqθ,t\(xb,t\+1;xb,0:T−2\)∑b,tmbt,mbt=𝟏\{xb,t\+1≠0\},\\widehat\{\\mathcal\{L\}\}\_\{\\mathrm\{block\}\}\(\\theta\)\\;=\\;\-\\,\\frac\{\\displaystyle\\sum\_\{b,t\}m\_\{bt\}\\,\\log q\_\{\\theta,t\}\\bigl\(x\_\{b,t\+1\};\\,x\_\{b,0:T\-2\}\\bigr\)\}\{\\displaystyle\\sum\_\{b,t\}m\_\{bt\}\},\\qquad m\_\{bt\}=\\mathbf\{1\}\\\{x\_\{b,t\+1\}\\neq 0\\\},with the sums over the batchb≤Bb\\leq Band the aligned rowstt: padding labels \(token id0\) are excluded, and the normalizer is the batch’s nonpadding\-label count\. For fixed\-length un\-padded blocks this equals \([4\.10](https://arxiv.org/html/2608.10288#S4.E10)\) up to the positive constant1/\(T−1\)1/\(T\{\-\}1\), reading the expectation as the batch mean; in padded batches the two differ, and the mask removes padded*labels*from the loss while padded*rows*still enter the forward pass and its query Gram \(see the padding paragraph of Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\)\. The final summand \(t=T−2t=T\{\-\}2\) is exactly the deployed online conditional−logpθ\(xT−1∣x0:T−2\)\-\\log p\_\{\\theta\}\(x\_\{T\-1\}\\mid x\_\{0:T\-2\}\)of \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\), since the held\-out target is absent from the input; the block’s Gram aggregates only tokens that precede it\. The first summand \(t=0t=0\) is also structurally guaranteed: under the ideal causal mask row0has a single admissible key, so every attention layer assigns it the weight vector\(1\)\(1\)independently of the operatorGLMG\_\{\\mathrm\{LM\}\}and of every later supplied row, and all other sublayers act row\-wise; by induction the row\-0output depends only onx0x\_\{0\}, and the summand equals the online conditional−logpθ\(x1∣x0\)\-\\log p\_\{\\theta\}\(x\_\{1\}\\mid x\_\{0\}\)\(exactly so in the implemented finite\-mask arithmetic whenever the masked exponentials underflow to zero, as on the audited calls; Appendix[D](https://arxiv.org/html/2608.10288#A4)reports the structural insensitivity of prefix position0on both released checkpoints\)\. For the intermediate rows1≤t≤T−31\\leq t\\leq T\{\-\}3the summands are auxiliary historical\-row predictions conditioned on the*entire*supplied block through the deductive operator, and the conditioning must be stated at its sharpest: for each suchttthe labelxt\+1x\_\{t\+1\}is*itself among the supplied rows*: it enters the query Gram, hence the operatorGLMG\_\{\\mathrm\{LM\}\}, used to produce the row\-ttprediction, as do all tokens later than the target\. The intermediate summands are therefore*target\-exposed*auxiliary scores, generically target\-dependent absent operator collapse or another degeneracy of the score, value, or output path \(Remark[5\.3](https://arxiv.org/html/2608.10288#S5.Thmtheorem3)exhibits such degeneracies; Appendix[D](https://arxiv.org/html/2608.10288#A4)measures a collapsed checkpoint on which historical rows do not move\), and they are not next\-token log probabilities as a protocol; forT≥3T\\geq 3exactly two of theT−1T\{\-\}1summands per block, the first and the final, are structurally guaranteed to equal deployed conditionals \(forT=2T=2the single summand is both\), while the remainingT−3T\{\-\}3need not be \(though every summand updates shared parameters\), and a low value ofℒblock\\mathcal\{L\}\_\{\\mathrm\{block\}\}cannot be read as a low autoregressive perplexity without a sequential measurement \(Appendix[D](https://arxiv.org/html/2608.10288#A4)reports one on held\-out text\)\. Historical\-row prefix consistency \(Definition[4\.8](https://arxiv.org/html/2608.10288#S4.Thmtheorem8)\) at the visited inputs guarantees that the summands coincide with the chain\-rule factors−logpθ\(xt\+1∣x0:t\)\-\\log p\_\{\\theta\}\(x\_\{t\+1\}\\mid x\_\{0:t\}\); the property is equivalent to equality of the full row conditionals there \(equality of every possible next\-token factor\), and it is strictly stronger than equality of the realized quantities alone, since probability mass can move among untargeted tokens without changing a realized factor, and summed losses can agree through cancellation across positions\. The global Gram does not supply prefix consistency in general \(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5); measured on the released checkpoints in Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\.ℒblock\\mathcal\{L\}\_\{\\mathrm\{block\}\}is a coherent surrogate in the limited sense of a well\-defined differentiable criterion that trains the causal final\-row generator while letting all past tokens interact through the operator being learned; it is not a proper scoring rule for the chain\-rule language distribution, and a sufficiently input\-sensitive metric learner could in principle exploit the target exposure during training, precisely when the learner is meant to be input\-sensitive \(Section[9](https://arxiv.org/html/2608.10288#S9)lists the along\-training measurement this motivates\)\. In this paper the terms*autoregressive NLL*and*perplexity*are reserved for the sequential final\-row quantity−∑tlogpθ\(xt\+1∣x0:t\)\-\\sum\_\{t\}\\log p\_\{\\theta\}\(x\_\{t\+1\}\\mid x\_\{0:t\}\), and every likelihood or benchmark number is labeled with the protocol that produced it \(Appendices[B](https://arxiv.org/html/2608.10288#A2)and[D](https://arxiv.org/html/2608.10288#A4)\)\.
The objective is optionally augmented by the*DAG regularizer*of the deductive outputs\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\], built on the NOTEARS characterization\[[50](https://arxiv.org/html/2608.10288#bib.bib7)\]\. As implemented in the reference code, the per\-tensor DAG loss averages the log of the*normalized*heat trace over the batch elementsb≤Bb\\leq B\(BBthe batch size\), the layersℓ≤L\\ell\\leq L, and the headsi≤hi\\leq h, withM\(b,ℓ,i\)M^\{\(b,\\ell,i\)\}denoting the instance of the tensorMMinferred for batch elementbbat layerℓ\\ell, headii:
\(4\.11\)DL\(M\)=1BLh∑b,ℓ,i\|log\(1dktreM\(b,ℓ,i\)⊙M\(b,ℓ,i\)\)\|,ℒ=ℒblock\+λ1DL\(ALM\)\+λ2DL\(AP\)\+λ3DL\(GLM\)\.\\begin\{split\}D\_\{L\}\(M\)&=\\frac\{1\}\{BLh\}\\sum\_\{b,\\ell,i\}\\,\\Bigl\\lvert\\,\\log\\Bigl\(\\tfrac\{1\}\{d\_\{k\}\}\\,\\operatorname\{tr\}\\,e^\{\\,M^\{\(b,\\ell,i\)\}\\odot M^\{\(b,\\ell,i\)\}\}\\Bigr\)\\Bigr\\rvert,\\\\ \\mathcal\{L\}&=\\mathcal\{L\}\_\{\\mathrm\{block\}\}\+\\lambda\_\{1\}D\_\{L\}\(A\_\{\\mathrm\{LM\}\}\)\+\\lambda\_\{2\}D\_\{L\}\(A\_\{P\}\)\+\\lambda\_\{3\}D\_\{L\}\(G\_\{\\mathrm\{LM\}\}\)\.\\end\{split\}By Theorem[4\.4](https://arxiv.org/html/2608.10288#S4.Thmtheorem4)below,treM⊙M≥dk\\operatorname\{tr\}e^\{M\\odot M\}\\geq d\_\{k\}always, so the absolute value in \([4\.11](https://arxiv.org/html/2608.10288#S4.E11)\) is analytically redundant \(it guards the numerics\) andDL\(M\)≥0D\_\{L\}\(M\)\\geq 0, with equality iff the support graph of each summand is acyclic\. For the two entrywise\-positive tensorsALMA\_\{\\mathrm\{LM\}\}andAPA\_\{P\}that equality is unattainable \(Remark[4\.5](https://arxiv.org/html/2608.10288#S4.Thmtheorem5)\); the mixed tensorGLM=aAP\+baG\_\{\\mathrm\{LM\}\}=aA\_\{P\}\+b\_\{a\}is not sign\-constrained, and no positive floor is asserted forDL\(GLM\)D\_\{L\}\(G\_\{\\mathrm\{LM\}\}\), whose value is an empirical matter \(measured on a released checkpoint in Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\.
###### Theorem 4\.4\(Walk\-counting characterization of acyclicity\[[50](https://arxiv.org/html/2608.10288#bib.bib7)\]\)\.
ForM∈ℝd×dM\\in\\mathbb\{R\}^\{d\\times d\}, leth\(M\)=treM⊙M−dh\(M\)=\\operatorname\{tr\}e^\{M\\odot M\}\-d\. Thenh\(M\)≥0h\(M\)\\geq 0always, andh\(M\)=0h\(M\)=0if and only if the weighted directed graph with adjacencyMM\(edgei→ji\\to jiffMij≠0M\_\{ij\}\\neq 0\) has no directed cycle\.
###### Proof\.
LetN=M⊙MN=M\\odot M, which has entriesNij=Mij2≥0N\_\{ij\}=M\_\{ij\}^\{2\}\\geq 0\. Then\(Nk\)ii=∑Mij12Mj1j22⋯Mjk−1i2\(N^\{k\}\)\_\{ii\}=\\sum M\_\{ij\_\{1\}\}^\{2\}M\_\{j\_\{1\}j\_\{2\}\}^\{2\}\\cdots M\_\{j\_\{k\-1\}i\}^\{2\}sums non\-negative weights over closed walks of lengthkkthroughii\. Hence
\(4\.12\)treN=d\+∑k≥1trNkk\!=d\+∑k≥11k\!\(total squared\-weight of closedk\-walks\)≥d,\\operatorname\{tr\}e^\{N\}\\;=\\;d\+\\sum\_\{k\\geq 1\}\\frac\{\\operatorname\{tr\}N^\{k\}\}\{k\!\}\\;=\\;d\+\\sum\_\{k\\geq 1\}\\frac\{1\}\{k\!\}\\bigl\(\\text\{total squared\-weight of closed $k$\-walks\}\\bigr\)\\;\\geq\\;d,with equality iff every term vanishes, i\.e\. iff there is no closed walk of any length, i\.e\. iff the graph is acyclic\. \(A directed cycle yields a closed walk and conversely any closed walk contains a cycle\.\) ∎
### 4\.4\.Inference: KV\-cache and G\-cache
At inference, PLDR\-LLM admits two nested caches, implemented as follows in the reference code\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]:
- •KV\-cache=\(Kcached,Vcached,Acached\)=\(K\_\{\\mathrm\{cached\}\},V\_\{\\mathrm\{cached\}\},A\_\{\\mathrm\{cached\}\}\): after the prompt pass, the rotated keys, the values,*and the metric generatorAA*are stored; for each generated token, the newk,vk,vrows are rotated \(with tracked positions\) and appended, whileAAis*not*recomputed; the prompt\-inferredAAis reused\.
- •G\-cache=\(ALM,GLM\)=\(A\_\{\\mathrm\{LM\}\},G\_\{\\mathrm\{LM\}\}\): additionally, the outputs of \([3\.4](https://arxiv.org/html/2608.10288#S3.E4)\)–\([3\.6](https://arxiv.org/html/2608.10288#S3.E6)\) are stored and the parameter mapsA↦ALM↦AP↦GLMA\\mapsto A\_\{\\mathrm\{LM\}\}\\mapsto A\_\{P\}\\mapsto G\_\{\\mathrm\{LM\}\}are skipped\.
New tokens are then scored by
\(4\.13\)E=q~GLMcachedKcached⊤dk,ELM=softmax\(E\),vnext=ELMVcached\.E=\\frac\{\\widetilde\{q\}\\;G\_\{\\mathrm\{LM\}\}^\{\\mathrm\{cached\}\}\\;K\_\{\\mathrm\{cached\}\}^\{\\top\}\}\{\\sqrt\{d\_\{k\}\}\},\\qquad E\_\{\\mathrm\{LM\}\}=\\operatorname\{softmax\}\(E\),\\qquad v\_\{\\mathrm\{next\}\}=E\_\{\\mathrm\{LM\}\}V\_\{\\mathrm\{cached\}\}\.Three inference semantics must be kept separate, and this paper uses them under the following fixed names\.
- •*Exact prefix recomputation*: the uncached online loop of \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\); every step is a freshS=tS=tcall on the grown prefix, and every tensor \(includingAAandGLMG\_\{\\mathrm\{LM\}\}\) is recomputed\. This is the reference semantics; it defines the model’s conditionals\.
- •*Prompt\-frozen conditional generation*: what KV\-cache and G\-cache literally compute;AA\(henceALM,GLMA\_\{\\mathrm\{LM\}\},G\_\{\\mathrm\{LM\}\}\) is frozen at the prompt and reused for every generated token\. This is a*definition*of the cached conditional, not an approximation claim: given the cachedAA, the maps toALMA\_\{\\mathrm\{LM\}\}andGLMG\_\{\\mathrm\{LM\}\}are deterministic functions ofAAand the fixed parameters, so*G\-cache is exact relative to KV\-cache*; it saves compute and introduces no additional approximation beyond the freeze itself\.
- •*Empirical freeze after collapse*: the measured statement that the first two semantics coincide on a given checkpoint \(the invariance property of Section[5](https://arxiv.org/html/2608.10288#S5)\)\. On the audited checkpoint the recomputedGLMG\_\{\\mathrm\{LM\}\}equals the cached value bitwise at every compared step and the cached\-versus\-recomputed logit deviation stays below the decoding margins \(Appendix[D](https://arxiv.org/html/2608.10288#A4)\), which is what makes the cached conditional a faithful implementation of the reference semantics there\.
Empirically, caching yields a∼3×\\sim 3\\timesspeedup\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\], with the reported aggregate deductive\-output statistics \(cross\-head RMSE values and maximum determinant magnitudes\) agreeing to1515printed decimal digits on the tested prompt and greedy decoding run\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]; agreement of these aggregates is evidence of, but not identical to, uniform element\-wise equality of all tensors on all prompts\. The published benchmark evaluations of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]score answer candidates by*one\-pass block scoring*\(one full call on context plus candidate, summing candidate log\-probabilities; Appendix[B](https://arxiv.org/html/2608.10288#A2)pins the evaluation wrapper\), a protocol in which no generation\-time cache participates by construction; the sequential\-versus\-block score comparison on the released checkpoints is reported in Appendix[D](https://arxiv.org/html/2608.10288#A4)\.
### 4\.5\.The online generation contract and historical\-row prefix consistency
The global density operator makes PLDR\-LLM’s row\-wise dependence structure different from SDPA’s in one specific, easily misread way\. This subsection states the deployed contract and the stronger property it deliberately does not supply, and fixes the terminology used for training and evaluation semantics throughout the paper\.
The deployed interface is \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\): supply exactly the known prefixx1:tx\_\{1:t\}\(S=tS=t\), read the final row\. Two properties must be distinguished\.
###### Definition 4\.8\(Historical\-row prefix consistency\)\.
A decoder mapFθF\_\{\\theta\}is*historical\-row prefix consistent*if for every inputx1:Sx\_\{1:S\}and every rowrrwith1≤r≤S1\\leq r\\leq S\(one\-based, as in \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\)\),
\(4\.14\)Fθ\(x1:S\)r=Fθ\(x1:r\)r,F\_\{\\theta\}\(x\_\{1:S\}\)\_\{r\}\\;=\\;F\_\{\\theta\}\(x\_\{1:r\}\)\_\{r\},i\.e\. every row of one parallel call equals the final row of the call on its own prefix\.
Property \([4\.14](https://arxiv.org/html/2608.10288#S4.E14)\) is what permits a single forward pass to be read simultaneously as all the sequential conditionals; it is the bridge between the blockwise objective \([4\.10](https://arxiv.org/html/2608.10288#S4.E10)\) and the chain\-rule factorization, and it is exactly what one\-pass likelihood evaluation assumes\.
##### An optional cumulative construction\.
A masked running Gram restores \([4\.14](https://arxiv.org/html/2608.10288#S4.E14)\) by construction, at a cost: with validity maskmn∈\{0,1\}m\_\{n\}\\in\\\{0,1\\\},
\(4\.16\)D~tcum=∑n≤tmnq~nq~n⊤,GLM\(t\)=Ψ\(D~tcum\),ztj=q~t⊤GLM\(t\)k~jdk\(j≤t\),\\widetilde\{D\}\_\{t\}^\{\\,\\mathrm\{cum\}\}\\;=\\;\\sum\_\{n\\leq t\}m\_\{n\}\\,\\widetilde\{q\}\_\{n\}\\widetilde\{q\}\_\{n\}^\{\\top\},\\qquad G\_\{\\mathrm\{LM\}\}^\{\(t\)\}=\\Psi\\bigl\(\\widetilde\{D\}\_\{t\}^\{\\,\\mathrm\{cum\}\}\\bigr\),\\qquad z\_\{tj\}=\\frac\{\\widetilde\{q\}\_\{t\}^\{\\top\}\\,G\_\{\\mathrm\{LM\}\}^\{\(t\)\}\\,\\widetilde\{k\}\_\{j\}\}\{\\sqrt\{d\_\{k\}\}\}\\quad\(j\\leq t\),computed layerwise\. The rank\-one updates to the Gram are cheap, but rowttneedsΨ\\Psievaluated at its own prefix Gram: a parallel all\-row pass evaluates the deep metric learner at up toSSprefixes per layer and head \(or in blocks, trading granularity for cost\), against*one*O\(Sdk2\)O\(Sd\_\{k\}^\{2\}\)contraction plus*one*Ψ\\Psievaluation per layer and head for the global Gram\. Equation \([4\.16](https://arxiv.org/html/2608.10288#S4.E16)\) is an implementation option for callers who want one\-pass all\-row autoregressive semantics or padded batching; it is not required for sequential generation, and the reference implementations and released checkpoints use the global Gram\.
##### Padding at the Gram boundary\.
The reference implementations formD~\\widetilde\{D\}from*every*supplied query row: rows excluded by the attention mask still enter the sequence contraction, and the measured final\-real\-row deviations under appended attention\-masked padding \(including dependence on the padding*content*\) are reported in Appendix[D](https://arxiv.org/html/2608.10288#A4)\. The deployed contract is therefore the*unpadded*S=tS=tinterface: sequential generation presents only genuine prompt tokens, and callers batching with right padding should strip padding before scoring \(stripping restores the unpadded call bitwise\)\. A Gram\-side validity mask as in \([4\.16](https://arxiv.org/html/2608.10288#S4.E16)\) is an implementation option for future releases; the pinned released checkpoints and their remote code are kept byte\-stable instead\.
x1:Sx\_\{1:S\}X\(0\)X^\{\(0\)\}Q~,K~,V\\widetilde\{Q\},\\widetilde\{K\},Vaffine \+ RoPE \(row\-wise\)E=Q~GLMK~⊤/dkE=\\widetilde\{Q\}\\,G\_\{\\mathrm\{LM\}\}\\widetilde\{K\}^\{\\top\}\\\!/\\sqrt\{d\_\{k\}\}mask \+ softmaxscore supportj≤ij\\leq iVLMV\_\{\\mathrm\{LM\}\}D~=Q~⊤Q~\\widetilde\{D\}=\\widetilde\{Q\}^\{\\top\}\\widetilde\{Q\}:sequence axis contracted;*all*supplied rows enterΨ\\Psi: LN, shared row mapφ\\varphi, iSwiGLU\+ϵ\{\}\+\\epsilon,elementwise power⊙P\\odot P, couplinga\(⋅\)\+baa\(\\cdot\)\+b\_\{a\}GLMG\_\{\\mathrm\{LM\}\}\(dk×dkd\_\{k\}\\times d\_\{k\}\)shared by all rowspadding rows included unless stripped\(Appendix[D](https://arxiv.org/html/2608.10288#A4)\)causal mask acts here onlyFigure 1\.Information flow in one PLGA head\. Solid path: row\-local operations \(rowiiuses rows≤i\\leq ithrough the masked softmax\)\. Dashed path: the global density operator; the sequence axis is contracted*before*any mask acts, so the learned operatorGLMG\_\{\\mathrm\{LM\}\}depends on every supplied row\. Under the online contract \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\) all supplied rows are known context; historical rows of a longer call are recomputed under the enlarged context \(Remark[4\.9](https://arxiv.org/html/2608.10288#S4.Thmtheorem9)\)\.
##### Terminology for training and evaluation semantics\.
Five distinct semantics appear in this paper, under these fixed names:*blockwise \(global\-context\) cross\-entropy*\([4\.10](https://arxiv.org/html/2608.10288#S4.E10)\), the implemented training objective \(one full\-block pass, cross\-entropy at every aligned row\);*sequential autoregressive NLL*\(equivalently*perplexity*\),−∑tlogpθ\(xt\+1∣x0:t\)\-\\sum\_\{t\}\\log p\_\{\\theta\}\(x\_\{t\+1\}\\mid x\_\{0:t\}\)computed by repeated final\-row prefix calls, the chain\-rule quantity, equal to the blockwise value whenever \([4\.14](https://arxiv.org/html/2608.10288#S4.E14)\) holds at the visited inputs \(equality of the two realized scalar totals alone does not imply \([4\.14](https://arxiv.org/html/2608.10288#S4.E14)\)\);*one\-pass \(whole\-candidate\) block scoring*, the published evaluation protocol \(one call on context plus candidate, candidate log\-probabilities summed; Appendix[B](https://arxiv.org/html/2608.10288#A2)\);*sequential generation*, the deployed online loop of \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\); and*prompt\-frozen cache inference*, the KV/G\-cache semantics of Section[4\.4](https://arxiv.org/html/2608.10288#S4.SS4)\. The measured gaps between one\-pass and sequential scoring on both released checkpoints are reported in Appendix[D](https://arxiv.org/html/2608.10288#A4)\.
## 5\.Deductive Outputs as Invariant Operators
### 5\.1\.The steady state and the order parameter
###### Definition 5\.1\(ε\\varepsilon\-invariance and the order parameter\)\.
Letμ\\mube a distribution over admissible inputs \(prompt plus stochastic continuation\)\. A deductive outputT\(x\)∈ℝνT\(x\)\\in\\mathbb\{R\}^\{\\nu\}\(ν\\nuthe number of entries ofTT\) is*ε\\varepsilon\-invariant*if there exists a constant tensorT∗T^\{\\ast\}with‖T\(x\)−T∗‖rms≤ε\\left\\lVert T\(x\)\-T^\{\\ast\}\\right\\rVert\_\{\\mathrm\{rms\}\}\\leq\\varepsilonforμ\\mu\-almost allxx, where∥⋅∥rms=∥⋅∥F/ν\\left\\lVert\\cdot\\right\\rVert\_\{\\mathrm\{rms\}\}=\\left\\lVert\\cdot\\right\\rVert\_\{F\}/\\sqrt\{\\nu\}\. Following\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]in substance, define the*order parameter*of a trained modelθ\\thetafrom two independent generation runsx\(1\),x\(2\)x^\{\(1\)\},x^\{\(2\)\}\(or a run against the cached pass\):
\(5\.1\)m\(θ\)=‖𝒟\(x\(1\);θ\)−𝒟\(x\(2\);θ\)‖rmsrms\(𝒟\),rms\(𝒟\)=12\(‖𝒟\(x\(1\);θ\)‖rms\+‖𝒟\(x\(2\);θ\)‖rms\),m\(\\theta\)\\;=\\;\\frac\{\\left\\lVert\\,\\mathcal\{D\}\(x^\{\(1\)\};\\theta\)\-\\mathcal\{D\}\(x^\{\(2\)\};\\theta\)\\,\\right\\rVert\_\{\\mathrm\{rms\}\}\}\{\\operatorname\{rms\}\(\\mathcal\{D\}\)\},\\qquad\\operatorname\{rms\}\(\\mathcal\{D\}\)=\\tfrac\{1\}\{2\}\\bigl\(\\left\\lVert\\mathcal\{D\}\(x^\{\(1\)\};\\theta\)\\right\\rVert\_\{\\mathrm\{rms\}\}\+\\left\\lVert\\mathcal\{D\}\(x^\{\(2\)\};\\theta\)\\right\\rVert\_\{\\mathrm\{rms\}\}\\bigr\),the RMSE between deductive outputs across runs normalized by the RMS magnitude of their entries; it is computed and reported*per tensor type*, withGLMG\_\{\\mathrm\{LM\}\}\(orAA, the most sensitive\) used as representative\. The statistic is defined*piecewise*:m=0m=0whenever the numerator vanishes \(equal tensors; this covers the all\-zero pair, for which the denominator also vanishes and the plain quotient would be0/00/0\), and by the displayed quotient otherwise, in which case the RMS denominator is strictly positive, since it can vanish only if both tensors, hence the numerator, vanish\.555The source paper\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]normalizes by the absolute value of the signed global mean,\|μ𝒟\|\\left\\lvert\\mu\_\{\\mathcal\{D\}\}\\right\\rvert, which can be arbitrarily small by cancellation for tensors centered near zero; for that normalization the case “denominator zero, numerator positive” does occur \(means cancel\) and the statistic is\+∞\+\\inftythere\. The RMS denominator used here is stable; both normalizations are computed side by side in Appendix[D](https://arxiv.org/html/2608.10288#A4)\. A reported valuem=0m=0means agreement at floating\-point resolution, not exact invariance\.
###### Proposition 5\.2\(Invariance suffices for exact cacheability\)\.
1. \(i\)\(Sufficiency\.\) If, for every layer and head,GLM\(ℓ,i\)\(x\)=G∗\(ℓ,i\)G\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\}\(x\)=G^\{\\ast\(\\ell,i\)\}for all admissible inputsxx, then the G\-cache \([4\.13](https://arxiv.org/html/2608.10288#S4.E13)\) computes the identical function to the full forward pass for every prompt and continuation\. Invariance of the upstream tensorsA,ALM,APA,A\_\{\\mathrm\{LM\}\},A\_\{P\}is*not*necessary for this conclusion\.
2. \(ii\)\(Definitional equivalence\.\) For a fixed tensor typeTT,m\(θ\)=0m\(\\theta\)=0in exact arithmetic over all input pairs if and only ifTTis exactly input\-invariant \(with the piecewise convention of Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1), which makes this hold without exception, including for identically zero tensors\)\.
3. \(iii\)\(Two\-level characterization\.\)*Row level \(iff\):*two finite allowed score rows induce the same softmax distribution if and only if they differ by a common additive scalar\.*Model level \(if\):*if on every reachable decoding state every recomputed allowed\-score row differs from its cached counterpart by a row\-wise additive constant, then cached and full attention rows, head outputs, and decoding distributions agree for all admissible continuations\. Constancy ofGLMG\_\{\\mathrm\{LM\}\}is sufficient but not necessary for the row condition, and*no converse from full\-model output equality to score\-row equality is asserted*: output equality does not imply row equality \(Remark[5\.3](https://arxiv.org/html/2608.10288#S5.Thmtheorem3)\)\.
###### Proof\.
\(i\) At every decoding step the cached operator equals the operator the full network would recompute, so scores, attention rows, and outputs agree token\-by\-token\. \(ii\) is immediate from \([5\.1](https://arxiv.org/html/2608.10288#S5.E1)\) with an exact\-arithmetic reading of0and the piecewise convention\. \(iii\) Row level: sufficiency of a common shift is invariance of the softmax under adding a constant to an allowed row; necessity follows by taking logarithms of the ratio of the two distributions, which shows the score differencezi′−zi=log∑jezj′−log∑jezjz^\{\\prime\}\_\{i\}\-z\_\{i\}=\\log\\sum\_\{j\}e^\{z^\{\\prime\}\_\{j\}\}\-\\log\\sum\_\{j\}e^\{z\_\{j\}\}is the same for every allowedii\(both directions are machine\-checked; Appendix[C](https://arxiv.org/html/2608.10288#A3)\)\. Model level: equal attention rows on every reachable state give equal head outputs, hence equal states and logits, by induction over decoding steps; this direction only\. ∎
Empirically\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]: models pretrained at near\-criticality havem∼10−6m\\sim 10^\{\-6\}to10−1110^\{\-11\}\(and0at float resolution forAP,GLMA\_\{P\},G\_\{\\mathrm\{LM\}\}of the model trained on4141B tokens\), while sub\-critical models havem∼1m\\sim 1–5050\. The order parameter thus separates the two phases sharply in the published sample and agrees with benchmark rankings at the phase level; within the near\-critical group the relationship is not strictly monotone \(small reversals occur between models whose benchmark averages differ by fractions of a point\), and no uncertainty estimates are available, so we do not claim thatm\(θ\)m\(\\theta\)precisely ranks reasoning ability\. Establishing \(or refuting\) a finer\-grained relationship requires the multi\-seed protocol of Section[9](https://arxiv.org/html/2608.10288#S9)\.
### 5\.2\.The inference\-collapse theorem
###### Theorem 5\.4\(SDPA as a special case; inference collapse; training–inference asymmetry\)\.
Consider a PLDR\-LLMFθF\_\{\\theta\}per Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)\.
1. \(i\)IfGLM\(ℓ,i\)≡IG\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\}\\equiv Ifor all layers and heads, thenFθF\_\{\\theta\}is exactly a decoder\-only transformer with scaled dot\-product attention \(with RoPE, SwiGLU FFN, and the stated normalizations\): SDPA\-LLM is the point of PLDR\-LLM model space with identity energy–curvature tensor\.
2. \(ii\)If every deductive output is exactly input\-invariant with valuesG∗\(ℓ,i\)G^\{\\ast\(\\ell,i\)\}, then the inference map ofFθF\_\{\\theta\}coincides with the map of the architecture in which the subnetwork \([3\.3](https://arxiv.org/html/2608.10288#S3.E3)\)–\([3\.6](https://arxiv.org/html/2608.10288#S3.E6)\) is deleted and replaced by the constantsG∗\(ℓ,i\)G^\{\\ast\(\\ell,i\)\}, a*generalized SDPA*with learned bilinear forms\. If additionally RoPE is absent, this is exactly an SDPA\-LLM with query projectionsWQ\(ℓ,i\)G∗\(ℓ,i\)W\_\{Q\}^\{\(\\ell,i\)\}\\,G^\{\\ast\(\\ell,i\)\}\. With RoPE:*\(sufficiency, for the model at hand\)*if eachG∗\(ℓ,i\)G^\{\\ast\(\\ell,i\)\}lies in the ambient torus commutant of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1), the constant operator is absorbable into pre\-rotation projections and the model is exactly an SDPA\-LLM;*\(necessity, at the unrestricted operator level\)*if the head’s bilinear score family\(n,m,q,k\)↦q⊤R−nG∗Rmk\(n,m,q,k\)\\mapsto q^\{\\top\}R\_\{\-n\}G^\{\\ast\}R\_\{m\}kis required to be SDPA\-realizable for*all*positions and allq,k∈ℝdkq,k\\in\\mathbb\{R\}^\{d\_\{k\}\}, thenG∗G^\{\\ast\}must lie in the commutant; this is Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1), which carries the quantifiers\. For a fixed trained model, whose reachable queries and keys span restricted subspaces, commutant membership is sufficient but*not*necessary for SDPA\-realizability of the inference map \(WQ=0W\_\{Q\}=0*and*bQ=0b\_\{Q\}=0, so that the affine query map, and with it every score, vanishes identically regardless ofG∗G^\{\\ast\}, is the degenerate witness\); no model\-level necessity is claimed\.
3. \(iii\)\(Asymmetry, structural\.\) The gradient of the loss in the full parameterization decomposes as the collapsed\-parameterization terms plus the chain\-rule term∑ℓ,i⟨∂ℒ/∂GLM\(ℓ,i\),∂GLM\(ℓ,i\)/∂θ⟩\\sum\_\{\\ell,i\}\\bigl\\langle\\partial\\mathcal\{L\}/\\partial G\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\},\\;\\partial G\_\{\\mathrm\{LM\}\}^\{\(\\ell,i\)\}/\\partial\\theta\\bigr\\rangleflowing through∂GLM/∂\(Φres,W,P,a,ba,WQ\)\\partial G\_\{\\mathrm\{LM\}\}/\\partial\(\\Phi\_\{\\mathrm\{res\}\},W,P,a,b\_\{a\},W\_\{Q\}\)\(the products denote adjoint\-Jacobian, i\.e\. vector–Jacobian, actions, not scalar multiplication\); this term is absent by construction from the collapsed model, whereGLMG\_\{\\mathrm\{LM\}\}is a constant\.*Under the explicit hypothesis that this term is nonzero at the evaluation point*, the instantaneous gradients of the two parameterizations differ\. \(Comparing*training dynamics*further requires fixing an optimizer and a correspondence between the parameter spaces; the empirical separation of loss curves and benchmark scores between learned, transferred, random, and identityGLMG\_\{\\mathrm\{LM\}\}\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]is cited as evidence that the difference is realized in practice, not as part of the statement\.\)
###### Proof\.
\(i\) SubstitutingGLM=IG\_\{\\mathrm\{LM\}\}=Iinto \([3\.7](https://arxiv.org/html/2608.10288#S3.E7)\)–\([3\.9](https://arxiv.org/html/2608.10288#S3.E9)\) givesE=Q~K~⊤/dkE=\\widetilde\{Q\}\\widetilde\{K\}^\{\\top\}/\\sqrt\{d\_\{k\}\},ELM=softmax\(mask\(E\)\)E\_\{\\mathrm\{LM\}\}=\\operatorname\{softmax\}\(\\operatorname\{mask\}\(E\)\),VLM=ELMVV\_\{\\mathrm\{LM\}\}=E\_\{\\mathrm\{LM\}\}V: the SDPA equations\[[43](https://arxiv.org/html/2608.10288#bib.bib6)\]; all remaining blocks in Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)are shared\. \(ii\) Under exact invariance, at every decoding step the recomputedGLM\(x\)G\_\{\\mathrm\{LM\}\}\(x\)equalsG∗G^\{\\ast\}; replacing the computation by the constant yields the same scores, hence the same distribution over continuations\. Without RoPE,QG∗K⊤=\(QG∗\)K⊤=Q′K⊤QG^\{\\ast\}K^\{\\top\}=\(QG^\{\\ast\}\)K^\{\\top\}=Q^\{\\prime\}K^\{\\top\}with
Q′=QG∗=X\(WQG∗\)\+𝟏\(bQ⊤G∗\),Q^\{\\prime\}\\;=\\;Q\\,G^\{\\ast\}\\;=\\;X\\,\\bigl\(W\_\{Q\}G^\{\\ast\}\\bigr\)\+\\mathbf\{1\}\\,\\bigl\(b\_\{Q\}^\{\\top\}G^\{\\ast\}\\bigr\),an SDPA parameterization in which the affine query map transforms as a whole:WQ′=WQG∗W\_\{Q\}^\{\\prime\}=W\_\{Q\}G^\{\\ast\}*and*bQ′⊤=bQ⊤G∗b\_\{Q\}^\{\\prime\\,\\top\}=b\_\{Q\}^\{\\top\}G^\{\\ast\}\(the projection of Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)is bias\-bearing, so absorbingG∗G^\{\\ast\}into the weight alone would change the model; the affine identity is machine\-checked, Appendix[C](https://arxiv.org/html/2608.10288#A3)\)\. With RoPE, sufficiency: ifG∗G^\{\\ast\}commutes with everyRtR\_\{t\}, thenR−nG∗Rm=G∗Rm−nR\_\{\-n\}G^\{\\ast\}R\_\{m\}=G^\{\\ast\}R\_\{m\-n\}and the score \([4\.2](https://arxiv.org/html/2608.10288#S4.E2)\) is the pure offset form realized by SDPA with the same absorbed affine projection applied before rotation: a commutingG∗G^\{\\ast\}also commutes with eachRt⊤=R−tR\_\{t\}^\{\\top\}=R\_\{\-t\}, so, position by position,
RoPE\(QG∗\)=RoPE\(Q\)G∗\.\\operatorname\{RoPE\}\(Q\\,G^\{\\ast\}\)\\;=\\;\\operatorname\{RoPE\}\(Q\)\\,G^\{\\ast\}\.Necessity at the unrestricted operator level is Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1), whose proof is self\-contained \(equality of the score*functions*for all positions and allq,kq,kforces commutant membership\); for a fixed model only the restrictions ofR−nG∗RmR\_\{\-n\}G^\{\\ast\}R\_\{m\}to the reachable query/key subspaces are observable, so no necessity is claimed there\. \(iii\) Write the loss asℒ\(θ\)=ℒ\(blocks;GLM\(⋅;θ\)\)\\mathcal\{L\}\(\\theta\)=\\mathcal\{L\}\\bigl\(\\text\{blocks\};\\,G\_\{\\mathrm\{LM\}\}\(\\cdot;\\theta\)\\bigr\); the chain rule gives the stated decomposition, and the extra term is identically absent from the collapsed parameterization\. If the term is nonzero atθ\\theta, the two gradients differ atθ\\thetaby exactly that term\. ∎
### 5\.3\.Quantitative stability of the collapse
In practice invariance isε\\varepsilon\-exact rather than exact\. The following bounds show that the observedε\\varepsilon\(down to10−1110^\{\-11\}relative\) propagates to a perturbation of the output distribution that is bounded by explicit constants\. Whether the resulting bound is small enough to force bit\-identical decoding is a separate, quantitative question: it requires comparing the evaluated end\-to\-end constant with the realized logit margins, which we do on a released checkpoint in Appendix[D](https://arxiv.org/html/2608.10288#A4)\. The bit\-identical cached\-versus\-uncached benchmark scores reported in\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]were obtained under one\-pass block scoring, whose scoring path performs a single uncached pass by construction \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\); the informative empirical facts here are the measured cached\-versus\-recomputed operator and logit deviations of Appendix[D](https://arxiv.org/html/2608.10288#A4), which are consistent with these bounds but are an empirical observation, not a corollary of them\.
###### Lemma 5\.6\(Softmax is11\-Lipschitz\[[9](https://arxiv.org/html/2608.10288#bib.bib21)\]\)\.
Forz,z′∈ℝSz,z^\{\\prime\}\\in\\mathbb\{R\}^\{S\},‖softmax\(z\)−softmax\(z′\)‖2≤‖z−z′‖2\\left\\lVert\\operatorname\{softmax\}\(z\)\-\\operatorname\{softmax\}\(z^\{\\prime\}\)\\right\\rVert\_\{2\}\\leq\\left\\lVert z\-z^\{\\prime\}\\right\\rVert\_\{2\}\.
###### Proposition 5\.7\(Perturbation bound for G\-caching\)\.
Fix a head with rotated inputsQ~,K~\\widetilde\{Q\},\\widetilde\{K\}and valuesVV, and letELM,VLME\_\{\\mathrm\{LM\}\},V\_\{\\mathrm\{LM\}\}andELM′,VLM′E\_\{\\mathrm\{LM\}\}^\{\\prime\},V\_\{\\mathrm\{LM\}\}^\{\\prime\}be computed with operatorsGGandG′G^\{\\prime\}where‖G−G′‖2≤ε\\left\\lVert G\-G^\{\\prime\}\\right\\rVert\_\{2\}\\leq\\varepsilon\. Then, row\-wise for each positiontt,
\(5\.2\)‖ELM\[t,⋅\]−ELM′\[t,⋅\]‖2≤‖q~t‖2‖K~‖2dkε,‖VLM\[t,⋅\]−VLM′\[t,⋅\]‖2≤‖q~t‖2‖K~‖2‖V‖2dkε\.\\left\\lVert E\_\{\\mathrm\{LM\}\}\[t,\\cdot\]\-E\_\{\\mathrm\{LM\}\}^\{\\prime\}\[t,\\cdot\]\\right\\rVert\_\{2\}\\;\\leq\\;\\frac\{\\left\\lVert\\widetilde\{q\}\_\{t\}\\right\\rVert\_\{2\}\\,\\left\\lVert\\widetilde\{K\}\\right\\rVert\_\{2\}\}\{\\sqrt\{d\_\{k\}\}\}\\;\\varepsilon,\\qquad\\left\\lVert V\_\{\\mathrm\{LM\}\}\[t,\\cdot\]\-V\_\{\\mathrm\{LM\}\}^\{\\prime\}\[t,\\cdot\]\\right\\rVert\_\{2\}\\;\\leq\\;\\frac\{\\left\\lVert\\widetilde\{q\}\_\{t\}\\right\\rVert\_\{2\}\\,\\left\\lVert\\widetilde\{K\}\\right\\rVert\_\{2\}\\,\\left\\lVert V\\right\\rVert\_\{2\}\}\{\\sqrt\{d\_\{k\}\}\}\\;\\varepsilon\.Consequently, if every non\-PLGA block of the network is Lipschitz on the relevant compact set \(true of theε\\varepsilon\-LayerNorm globally, Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(iv\), and of linear maps, SwiGLU, and softmax\), the final logits of the cached and uncached models differ by at mostC\(θ\)εC\(\\theta\)\\,\\varepsilonfor a constant depending on operator norms of the trained weights, uniformly over inputs\.
###### Proof\.
The score rows differ byΔet=q~t⊤\(G−G′\)K~⊤/dk\\Delta e\_\{t\}=\\widetilde\{q\}\_\{t\}^\{\\top\}\(G\-G^\{\\prime\}\)\\widetilde\{K\}^\{\\top\}/\\sqrt\{d\_\{k\}\}, so by sub\-multiplicativity
‖Δet‖2≤‖q~t‖2‖G−G′‖2‖K~‖2/dk;\\left\\lVert\\Delta e\_\{t\}\\right\\rVert\_\{2\}\\;\\leq\\;\\left\\lVert\\widetilde\{q\}\_\{t\}\\right\\rVert\_\{2\}\\,\\left\\lVert G\-G^\{\\prime\}\\right\\rVert\_\{2\}\\,\\left\\lVert\\widetilde\{K\}\\right\\rVert\_\{2\}/\\sqrt\{d\_\{k\}\};under the ideal masked softmax \(Remark[3\.13](https://arxiv.org/html/2608.10288#S3.Thmtheorem13)\) the restriction to the common allowed support only removes coordinates ofΔet\\Delta e\_\{t\}and cannot increase the norm\. Apply Lemma[5\.6](https://arxiv.org/html/2608.10288#S5.Thmtheorem6)row\-wise for the first bound\. For the second,ΔVLM\[t,⋅\]=\(ΔELM\[t,⋅\]\)V\\Delta V\_\{\\mathrm\{LM\}\}\[t,\\cdot\]=\(\\Delta E\_\{\\mathrm\{LM\}\}\[t,\\cdot\]\)Vand‖u⊤V‖2≤‖u‖2‖V‖2\\left\\lVert u^\{\\top\}V\\right\\rVert\_\{2\}\\leq\\left\\lVert u\\right\\rVert\_\{2\}\\left\\lVert V\\right\\rVert\_\{2\}\. The network\-level statement is the composition of Lipschitz maps, each perturbation entering additively with the product of downstream Lipschitz constants\. ∎
### 5\.4\.The origin of the invariance: how the metric learner generates an invariant metric generator
The results so far characterize the invariant steady state and its consequences; this subsection analyzes*how*the metric generatorA=Φres\(LN\(D~\)\)A=\\Phi\_\{\\mathrm\{res\}\}\(\\operatorname\{LN\}\(\\widetilde\{D\}\)\)becomes input\-invariant\. The analysis rests on three verified structural facts: \(F1\) the density operator is built from*rotary\-rotated*queries \(Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)\); \(F2\) LayerNorm is applied to the density operator row\-wise before the residual network; \(F3\) the metric learner is a single shared row mapφ\\varphiapplied independently to every row and every head \(Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\)\. The proposed mechanism decomposes into three stages: a phase\-averaging at the source, statistical concentration of what remains, and contraction inside the learned row map\. The epistemic status of each stage differs and is flagged in place: Stage 1 is an exact bound whose constant is honest but large at the reference context length; Stage 2 is a conditional theorem under an independence idealization; Stage 3 is a conditional theorem whose contraction hypothesis is*measured*, not proved \(Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\. The mechanism as a whole is therefore a quantitative hypothesis with proved ingredients, not a theorem\.
#### 5\.4\.1\.Stage 1: the rotary twirl projects the density operator onto a commutant
Because queries are rotated before the Gram product, the density operator of a head is a*position\-twisted*second moment:
\(5\.3\)D~=Q~⊤Q~=∑n=1SRnqnqn⊤Rn⊤\.\\widetilde\{D\}\\;=\\;\\widetilde\{Q\}^\{\\top\}\\widetilde\{Q\}\\;=\\;\\sum\_\{n=1\}^\{S\}R\_\{n\}\\,q\_\{n\}q\_\{n\}^\{\\top\}\\,R\_\{n\}^\{\\top\}\.The position\-dependent conjugation acts as a*twirl*\(an average over a group orbit\) and suppresses every component of the summands that does not commute with the ambient RoPE torusTT:
###### Lemma 5\.9\(Quantitative RoPE twirl\)\.
LetRnR\_\{n\}be RoPE rotations with anglesθ1,…,θdk/2\\theta\_\{1\},\\dots,\\theta\_\{d\_\{k\}/2\}such that none of the finitely many frequenciesω∈\{2θa\}∪\{θa±θb:a≠b\}\\omega\\in\\\{2\\theta\_\{a\}\\\}\\cup\\\{\\theta\_\{a\}\\pm\\theta\_\{b\}:a\\neq b\\\}is a multiple of2π2\\pi\. LetPTP\_\{T\}denote the orthogonal projection ofℝdk×dk\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}onto the commutant of the ambient RoPE torusTT\(the block\-diagonal algebra of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\)\. Then for every fixedMM,
‖1S∑n=1SRnMRn⊤−PT\(M\)‖F≤CΘS‖M‖F,CΘ=maxω≠01\|sin\(ω/2\)\|,\\Bigl\\lVert\\,\\frac\{1\}\{S\}\\sum\_\{n=1\}^\{S\}R\_\{n\}MR\_\{n\}^\{\\top\}\\;\-\\;P\_\{T\}\(M\)\\,\\Bigr\\rVert\_\{F\}\\;\\leq\\;\\frac\{C\_\{\\Theta\}\}\{S\}\\,\\left\\lVert M\\right\\rVert\_\{F\},\\qquad C\_\{\\Theta\}=\\max\_\{\\omega\\neq 0\}\\ \\frac\{1\}\{\\,\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert\\,\},the maximum over the nonzero frequencies above\.
###### Proof\.
Complexify each rotation plane: in the eigenbasis of the torus,RnR\_\{n\}is diagonal with entriese±iθane^\{\\pm i\\theta\_\{a\}n\}, and conjugation multiplies the\(u,v\)\(u,v\)matrix entry ofMM\(in this basis\) byei\(ωu−ωv\)ne^\{i\(\\omega\_\{u\}\-\\omega\_\{v\}\)n\}withωu,ωv∈\{±θa\}\\omega\_\{u\},\\omega\_\{v\}\\in\\\{\\pm\\theta\_\{a\}\\\}\. The difference frequencies are exactly0,±2θa\\pm 2\\theta\_\{a\}, and±\(θa∓θb\)\\pm\(\\theta\_\{a\}\\mp\\theta\_\{b\}\)\. The zero\-frequency entries span precisely the commutant \(thecI2\+sJ2cI\_\{2\}\+sJ\_\{2\}blocks of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\), and are left fixed by the average\. Every nonzero\-frequency entry is multiplied by1S∑n=1Seiωn\\frac\{1\}\{S\}\\sum\_\{n=1\}^\{S\}e^\{i\\omega n\}, whose modulus is\|sin\(Sω/2\)\|/\(S\|sin\(ω/2\)\|\)≤1/\(S\|sin\(ω/2\)\|\)\\left\\lvert\\sin\(S\\omega/2\)\\right\\rvert/\(S\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert\)\\leq 1/\(S\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert\)by the geometric sum\. Since the basis change is unitary, the entry\-wise bounds assemble into the stated Frobenius bound\. ∎
Applied to \([5\.3](https://arxiv.org/html/2608.10288#S5.E3)\)*under the stationarity idealization of Proposition[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10)*\(a common second momentΣ=𝔼\[qq⊤\]\\Sigma=\\mathbb\{E\}\[qq^\{\\top\}\]across positions; the lemma itself concerns one fixed matrixMM, and passing to the empirical sum requires this additional assumption\), Lemma[5\.9](https://arxiv.org/html/2608.10288#S5.Thmtheorem9)shows that the deterministic part ofD~/S\\widetilde\{D\}/Sconverges at rateO\(1/S\)O\(1/S\)toPT\(Σ\)P\_\{T\}\(\\Sigma\)\. Two qualifications keep this honest\. First, since every summandqnqn⊤q\_\{n\}q\_\{n\}^\{\\top\}is symmetric, the projection ofΣ\\Sigmaonto the commutant has vanishingJ2J\_\{2\}\-components: the effective target is the smaller algebra of*scalar*blockscjI2c\_\{j\}I\_\{2\}\. Second, the constant matters\. For the standard frequencies atdk=64d\_\{k\}=64, base10410^\{4\}, the exact worst constant isCΘ=maxω≠0\|sin\(ω/2\)\|−1≈4\.4968×104C\_\{\\Theta\}=\\max\_\{\\omega\\neq 0\}\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert^\{\-1\}\\approx 4\.4968\\times 10^\{4\}\(attained by the smallest cross\-plane difference frequency\), so at the reference context lengthS=1024S=1024the uniform boundCΘ/S≈43\.9C\_\{\\Theta\}/S\\approx 43\.9is vacuous\. The suppression is frequency\-resolved: the exact per\-frequency multiplier is\|sin\(Sω/2\)\|/\(S\|sin\(ω/2\)\|\)\\left\\lvert\\sin\(S\\omega/2\)\\right\\rvert/\(S\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert\), which is strongly suppressing for the many fast frequencies but approaches11for the slowest ones \(the worst component has multiplier≈0\.99991\\approx 0\.99991atS=1024S=1024; even the slowest within\-plane frequency2θdk/22\\theta\_\{d\_\{k\}/2\}has multiplier≈0\.9969\\approx 0\.9969\)\. The twirl therefore erases position\-of\-occurrence structure carried by the fast rotary planes but leaves the slowest planes essentially untouched at practical context lengths; how much instance structure it removes*in practice*depends on how the energy ofD~\\widetilde\{D\}distributes across twirl frequencies, which we measure on a released checkpoint in Appendix[D](https://arxiv.org/html/2608.10288#A4)rather than infer from the asymptotic\. With these qualifications, rotary embeddings, adopted in\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\]for their training benefits, also act as a partial self\-averaging channel for the deductive path\.
#### 5\.4\.2\.Stage 2: statistical concentration of the normalized density operator
What the twirl does not remove, the fluctuation of the empirical second moment around its mean, concentrates statistically:
###### Proposition 5\.10\(Concentration of the metric\-learner input; conditional\)\.
Model the \(rotated\-frame\) query vectorsq1,…,qSq\_\{1\},\\dots,q\_\{S\}as independent,‖qn‖2≤Bq\\left\\lVert q\_\{n\}\\right\\rVert\_\{2\}\\leq B\_\{q\}, with common second momentΣ\\Sigma\. This is an*idealization*: contextual transformer queries are causally dependent, non\-identically distributed across position, and trained jointly with RoPE, so independence and a common second moment are modeling assumptions, not architectural facts; the conclusion is conditional on them\. Then for anyδ∈\(0,1\)\\delta\\in\(0,1\), with probability at least1−δ1\-\\delta,
‖D~S−PT\(Σ\)‖2≤CΘS‖Σ‖F⏟twirl remainder\+Bq22log\(2dk/δ\)S\+2Bq2log\(2dk/δ\)3S⏟matrix Bernstein\.\\Bigl\\lVert\\,\\frac\{\\widetilde\{D\}\}\{S\}\-P\_\{T\}\(\\Sigma\)\\Bigr\\rVert\_\{2\}\\;\\leq\\;\\underbrace\{\\frac\{C\_\{\\Theta\}\}\{S\}\\left\\lVert\\Sigma\\right\\rVert\_\{F\}\}\_\{\\text\{twirl remainder\}\}\\;\+\\;\\underbrace\{B\_\{q\}^\{2\}\\sqrt\{\\frac\{2\\log\(2d\_\{k\}/\\delta\)\}\{S\}\}\+\\frac\{2B\_\{q\}^\{2\}\\log\(2d\_\{k\}/\\delta\)\}\{3S\}\}\_\{\\text\{matrix Bernstein\}\}\.In particular the metric\-learner input fluctuates around a single dataset\-level object at rateO\(S−1/2\)O\(S^\{\-1/2\}\)\.
###### Proof\.
WriteD~/S−PT\(Σ\)=\(1S∑nRnΣRn⊤−PT\(Σ\)\)\+1S∑nRn\(qnqn⊤−Σ\)Rn⊤\\widetilde\{D\}/S\-P\_\{T\}\(\\Sigma\)=\\bigl\(\\frac\{1\}\{S\}\\sum\_\{n\}R\_\{n\}\\Sigma R\_\{n\}^\{\\top\}\-P\_\{T\}\(\\Sigma\)\\bigr\)\+\\frac\{1\}\{S\}\\sum\_\{n\}R\_\{n\}\(q\_\{n\}q\_\{n\}^\{\\top\}\-\\Sigma\)R\_\{n\}^\{\\top\}\. The first term is Lemma[5\.9](https://arxiv.org/html/2608.10288#S5.Thmtheorem9)\. The second is an average of independent, mean\-zero, symmetric random matrices; from the Löwner\-order bounds0⪯qnqn⊤⪯Bq2I0\\preceq q\_\{n\}q\_\{n\}^\{\\top\}\\preceq B\_\{q\}^\{2\}Iand0⪯Σ⪯Bq2I0\\preceq\\Sigma\\preceq B\_\{q\}^\{2\}Ione gets−Bq2I⪯qnqn⊤−Σ⪯Bq2I\-B\_\{q\}^\{2\}I\\preceq q\_\{n\}q\_\{n\}^\{\\top\}\-\\Sigma\\preceq B\_\{q\}^\{2\}I, hence‖Rn\(qnqn⊤−Σ\)Rn⊤‖2≤Bq2\\left\\lVert R\_\{n\}\(q\_\{n\}q\_\{n\}^\{\\top\}\-\\Sigma\)R\_\{n\}^\{\\top\}\\right\\rVert\_\{2\}\\leq B\_\{q\}^\{2\}\(conjugation by orthogonal matrices preserves norms\), and the summand variance is bounded byBq4B\_\{q\}^\{4\}; the matrix Bernstein inequality\[[41](https://arxiv.org/html/2608.10288#bib.bib38), Thm\. 1\.6\.2\]applied to the average \(summand boundBq2/SB\_\{q\}^\{2\}/S, variance proxyBq4/SB\_\{q\}^\{4\}/S\) gives exactly the displayed tail\. ∎
We emphasize a bookkeeping point:SShere is the*context length*\. It governs concentration within one forward pass and must not be substituted for the pretraining token countNtokensN\_\{\\mathrm\{tokens\}\}, which controls a different limit \(how well the learned parameters approximate dataset\-level statistics\); the two enter the invariance question through different mechanisms\.
The LayerNorm stage then*standardizes*this concentrating input\. The implemented map carries the regularization constantεLN=10−6\\varepsilon\_\{\\operatorname\{LN\}\}=10^\{\-6\}, and its invariances must be stated with that constant in place:666Exact positive\-scale invariance and a spherical range hold only atεLN=0\\varepsilon\_\{\\operatorname\{LN\}\}=0and are false for the implemented map \(e\.g\. forr=\(10−4,−10−4\)r=\(10^\{\-4\},\-10^\{\-4\}\),γ=𝟏\\gamma=\\mathbf\{1\},β=0\\beta=0, the normalized coordinates are±0\.0995\\pm 0\.0995atc=1c=1and±0\.1961\\pm 0\.1961atc=2c=2\)\. The lemma below therefore keeps exact shift invariance, quantifies the scale\-invariance error, and replaces the sphere by a ball\.
###### Lemma 5\.11\(ε\\varepsilon\-LayerNorm: exact shift invariance, approximate scale invariance, compact range\)\.
LetρLNε\(r\)=γ⊙r−r¯𝟏v\(r\)\+ε\+β\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(r\)=\\gamma\\odot\\dfrac\{r\-\\bar\{r\}\\mathbf\{1\}\}\{\\sqrt\{v\(r\)\+\\varepsilon\}\}\+\\betabe the implemented LayerNorm row map with learned\(γ,β\)\(\\gamma,\\beta\), wherer¯\\bar\{r\}andv\(r\)v\(r\)are the mean and \(biased\) variance of the entries ofr∈ℝdkr\\in\\mathbb\{R\}^\{d\_\{k\}\}andε=εLN\>0\\varepsilon=\\varepsilon\_\{\\operatorname\{LN\}\}\>0\. Then:
1. \(i\)\(Exact shift invariance\.\)ρLNε\(r\+d𝟏\)=ρLNε\(r\)\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(r\+d\\mathbf\{1\}\)=\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(r\)for alld∈ℝd\\in\\mathbb\{R\}\.
2. \(ii\)\(Approximate scale invariance\.\) Forc\>0c\>0and nonconstantrr, ρLNε\(cr\)−ρLNε\(r\)=γ⊙r−r¯𝟏v\(r\)\(11\+ε/\(c2v\(r\)\)−11\+ε/v\(r\)\),\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(cr\)\-\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(r\)=\\gamma\\odot\\frac\{r\-\\bar\{r\}\\mathbf\{1\}\}\{\\sqrt\{v\(r\)\}\}\\left\(\\frac\{1\}\{\\sqrt\{1\+\\varepsilon/\(c^\{2\}v\(r\)\)\}\}\-\\frac\{1\}\{\\sqrt\{1\+\\varepsilon/v\(r\)\}\}\\right\),so that‖ρLNε\(cr\)−ρLNε\(r\)‖2≤‖γ‖∞dkε2min\(1,c2\)v\(r\)\\left\\lVert\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(cr\)\-\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}\(r\)\\right\\rVert\_\{2\}\\leq\\left\\lVert\\gamma\\right\\rVert\_\{\\infty\}\\sqrt\{d\_\{k\}\}\\,\\dfrac\{\\varepsilon\}\{2\\min\(1,c^\{2\}\)\\,v\(r\)\}\. Forγ≠0\\gamma\\neq 0, scale invariance for*all*c\>0c\>0and all nonconstant rows holds iffε=0\\varepsilon=0\(atc=1c=1, or whenγ=0\\gamma=0, the two sides coincide trivially for anyε\\varepsilon\)\. In particularLN\(D~\)=LN\(D~/S\)\\operatorname\{LN\}\(\\widetilde\{D\}\)=\\operatorname\{LN\}\(\\widetilde\{D\}/S\)holds only approximately, with per\-row relative error in the centered component bounded by the*upper proxy*ε/\(2vi\)\\varepsilon/\(2v\_\{i\}\), whereviv\_\{i\}is the row variance ofD~/S\\widetilde\{D\}/S\(equivalentlyεS2/\(2Vi\)\\varepsilon S^\{2\}/\(2V\_\{i\}\)in terms of the row varianceViV\_\{i\}ofD~\\widetilde\{D\}\); the exact first\-order coefficient for this pair of scales is\(1−S−2\)ε/\(2vi\)\(1\-S^\{\-2\}\)\\,\\varepsilon/\(2v\_\{i\}\), so the proxy overestimates already at first order by the \(small\) factor\(1−S−2\)−1\(1\-S^\{\-2\}\)^\{\-1\}, is valid as an expansion only forε≪vi\\varepsilon\\ll v\_\{i\}, and overestimates grossly outside that regime \(on an exactly constant row the true discrepancy is0while the proxy diverges\)\. The discrepancy itself is measured directly, row by row with the checkpoint’sγ,β,ε\\gamma,\\beta,\\varepsilon, in Appendix[D](https://arxiv.org/html/2608.10288#A4)\.
3. \(iii\)\(Compact range: a ball, not a sphere\.\) Every output row lies in the fixed compact setΣLN=\{γ⊙u\+β:𝟏⊤u=0,‖u‖2≤dk\}\\Sigma\_\{\\operatorname\{LN\}\}=\\\{\\gamma\\odot u\+\\beta:\\mathbf\{1\}^\{\\top\}u=0,\\ \\left\\lVert u\\right\\rVert\_\{2\}\\leq\\sqrt\{d\_\{k\}\}\\\}; indeed the normalized vectoru=\(r−r¯𝟏\)/v\(r\)\+εu=\(r\-\\bar\{r\}\\mathbf\{1\}\)/\\sqrt\{v\(r\)\+\\varepsilon\}satisfies exactly‖u‖22=dkv\(r\)/\(v\(r\)\+ε\)<dk\\left\\lVert u\\right\\rVert\_\{2\}^\{2\}=d\_\{k\}\\,v\(r\)/\(v\(r\)\+\\varepsilon\)<d\_\{k\}, approaching the sphere only asv\(r\)/ε→∞v\(r\)/\\varepsilon\\to\\infty, andu=0u=0for constant rows\.
4. \(iv\)\(Global Lipschitz on variance\-floored rows\.\) On rows withv\(r\)≥v0≥0v\(r\)\\geq v\_\{0\}\\geq 0,ρLNε\\rho\_\{\\operatorname\{LN\}\}^\{\\varepsilon\}is Lipschitz with constant at most‖γ‖∞/v0\+ε\\left\\lVert\\gamma\\right\\rVert\_\{\\infty\}/\\sqrt\{v\_\{0\}\+\\varepsilon\}\(an exact Jacobian bound\); theε\\varepsilonfloor makes the constant finite on*all*rows \(v0=0v\_\{0\}=0\), removing the null\-set caveat of the idealized map\.
###### Proof\.
\(i\) Centering removesd𝟏d\\mathbf\{1\}andvvis shift\-invariant\. \(ii\) Both terms share the unit vector\(r−r¯𝟏\)/v\(r\)\(r\-\\bar\{r\}\\mathbf\{1\}\)/\\sqrt\{v\(r\)\}scaled by\(1\+ε/\(c2v\)\)−1/2\(1\+\\varepsilon/\(c^\{2\}v\)\)^\{\-1/2\}and\(1\+ε/v\)−1/2\(1\+\\varepsilon/v\)^\{\-1/2\}respectively; the difference of the scalars is bounded, via1−\(1\+x\)−1/2≤x/21\-\(1\+x\)^\{\-1/2\}\\leq x/2forx≥0x\\geq 0, byε/\(2min\(1,c2\)v\)\\varepsilon/\(2\\min\(1,c^\{2\}\)v\), and‖\(r−r¯𝟏\)/v‖2=dk\\left\\lVert\(r\-\\bar\{r\}\\mathbf\{1\}\)/\\sqrt\{v\}\\right\\rVert\_\{2\}=\\sqrt\{d\_\{k\}\}\. \(iii\)‖r−r¯𝟏‖22=dkv\(r\)\\left\\lVert r\-\\bar\{r\}\\mathbf\{1\}\\right\\rVert\_\{2\}^\{2\}=d\_\{k\}\\,v\(r\)gives the exact norm identity\. \(iv\) WithP=I−𝟏𝟏⊤/dkP=I\-\\mathbf\{1\}\\mathbf\{1\}^\{\\top\}/d\_\{k\}\(centering\),u\(r\)=Pr/v\(r\)\+εu\(r\)=Pr/\\sqrt\{v\(r\)\+\\varepsilon\}, a direct computation gives the Jacobian factorizationJu\(r\)=\(I−uu⊤/dk\)P/v\(r\)\+εJ\_\{u\}\(r\)=\\bigl\(I\-uu^\{\\top\}/d\_\{k\}\\bigr\)\\,P\\big/\\sqrt\{v\(r\)\+\\varepsilon\}\. Both matrix factors are symmetric with spectrum in\[0,1\]\[0,1\]\(for the first,‖u‖22≤dk\\left\\lVert u\\right\\rVert\_\{2\}^\{2\}\\leq d\_\{k\}by \(iii\)\), so‖Ju\(r\)‖2≤1/v\(r\)\+ε\\left\\lVert J\_\{u\}\(r\)\\right\\rVert\_\{2\}\\leq 1/\\sqrt\{v\(r\)\+\\varepsilon\}at every floored row \(the pointwise bound used in the sample\-extrema budget proxy of Appendix[D](https://arxiv.org/html/2608.10288#A4)\), and with the affine output map contributing at most‖γ‖∞\\left\\lVert\\gamma\\right\\rVert\_\{\\infty\}, the mean value inequality gives the Lipschitz claim globally forv0=0v\_\{0\}=0, where the domain is all ofℝdk\\mathbb\{R\}^\{d\_\{k\}\}\. Forv0\>0v\_\{0\}\>0the floored set\{v\(r\)≥v0\}\\\{v\(r\)\\geq v\_\{0\}\\\}is the complement of a convex set and is not convex, so segments between floored rows may leave it; the global claim follows instead from a radial argument\. Writew=Prw=Pr, sov\(r\)=‖w‖22/dkv\(r\)=\\left\\lVert w\\right\\rVert\_\{2\}^\{2\}/d\_\{k\}andu=g\(w\)u=g\(w\)withg\(w\)=w\(‖w‖22/dk\+ε\)−1/2g\(w\)=w\\,\(\\left\\lVert w\\right\\rVert\_\{2\}^\{2\}/d\_\{k\}\+\\varepsilon\)^\{\-1/2\}, a radial mapg\(w\)=h\(ρ\)w^g\(w\)=h\(\\rho\)\\,\\hat\{w\}withρ=‖w‖2\\rho=\\left\\lVert w\\right\\rVert\_\{2\}andh\(ρ\)=ρ\(ρ2/dk\+ε\)−1/2h\(\\rho\)=\\rho\\,\(\\rho^\{2\}/d\_\{k\}\+\\varepsilon\)^\{\-1/2\}\. Setρ0=dkv0\\rho\_\{0\}=\\sqrt\{d\_\{k\}v\_\{0\}\}andL=\(v0\+ε\)−1/2L=\(v\_\{0\}\+\\varepsilon\)^\{\-1/2\}\. On\[ρ0,∞\)\[\\rho\_\{0\},\\infty\):h\(ρ\)/ρ=\(ρ2/dk\+ε\)−1/2≤Lh\(\\rho\)/\\rho=\(\\rho^\{2\}/d\_\{k\}\+\\varepsilon\)^\{\-1/2\}\\leq L, and0≤h′\(ρ\)=ε\(ρ2/dk\+ε\)−3/2≤ε\(v0\+ε\)−3/2≤L0\\leq h^\{\\prime\}\(\\rho\)=\\varepsilon\\,\(\\rho^\{2\}/d\_\{k\}\+\\varepsilon\)^\{\-3/2\}\\leq\\varepsilon\\,\(v\_\{0\}\+\\varepsilon\)^\{\-3/2\}\\leq L, so\|h\(ρ1\)−h\(ρ2\)\|≤L\|ρ1−ρ2\|\\left\\lvert h\(\\rho\_\{1\}\)\-h\(\\rho\_\{2\}\)\\right\\rvert\\leq L\\left\\lvert\\rho\_\{1\}\-\\rho\_\{2\}\\right\\rvertby the mean value theorem on the*interval*\[ρ0,∞\)\[\\rho\_\{0\},\\infty\), which is convex\. Forw1,w2w\_\{1\},w\_\{2\}withρi≥ρ0\\rho\_\{i\}\\geq\\rho\_\{0\}and angleθ\\thetabetween them,‖g\(w1\)−g\(w2\)‖22=h12\+h22−2h1h2cosθ\\left\\lVert g\(w\_\{1\}\)\-g\(w\_\{2\}\)\\right\\rVert\_\{2\}^\{2\}=h\_\{1\}^\{2\}\+h\_\{2\}^\{2\}\-2h\_\{1\}h\_\{2\}\\cos\\thetaand‖w1−w2‖22=ρ12\+ρ22−2ρ1ρ2cosθ\\left\\lVert w\_\{1\}\-w\_\{2\}\\right\\rVert\_\{2\}^\{2\}=\\rho\_\{1\}^\{2\}\+\\rho\_\{2\}^\{2\}\-2\\rho\_\{1\}\\rho\_\{2\}\\cos\\theta, soΦ\(c\)=L2‖w1−w2‖22−‖g\(w1\)−g\(w2\)‖22\\Phi\(c\)=L^\{2\}\\left\\lVert w\_\{1\}\-w\_\{2\}\\right\\rVert\_\{2\}^\{2\}\-\\left\\lVert g\(w\_\{1\}\)\-g\(w\_\{2\}\)\\right\\rVert\_\{2\}^\{2\}is affine inc=cosθc=\\cos\\thetawith slope2\(h1h2−L2ρ1ρ2\)≤02\(h\_\{1\}h\_\{2\}\-L^\{2\}\\rho\_\{1\}\\rho\_\{2\}\)\\leq 0by the first inequality;Φ\\Phiis therefore minimized atc=1c=1, whereΦ\(1\)=L2\(ρ1−ρ2\)2−\(h1−h2\)2≥0\\Phi\(1\)=L^\{2\}\(\\rho\_\{1\}\-\\rho\_\{2\}\)^\{2\}\-\(h\_\{1\}\-h\_\{2\}\)^\{2\}\\geq 0by the second\. HenceggisLL\-Lipschitz on the whole floored set, and composing with the11\-Lipschitz centeringPPand the‖γ‖∞\\left\\lVert\\gamma\\right\\rVert\_\{\\infty\}diagonal gives the stated constant, convexity of the floored set nowhere used\. ∎
Stages 1–2 together say,*under their stated idealizations*: the input to the learned row mapφ\\varphiis confined to the compact setΣLN\\Sigma\_\{\\operatorname\{LN\}\}and, across inputsxx, concentrates in anO\(S−1/2\)O\(S^\{\-1/2\}\)\-neighborhood of a single point set, the normalized rows ofPT\(Σ\)P\_\{T\}\(\\Sigma\)\. Invariance ofAAis now a property of whatφ\\varphidoes on that neighborhood\.
#### 5\.4\.3\.Stage 3: a conditional contraction bound for the trained row map
###### Proposition 5\.12\(Conditional contraction bound, collapse, and the converse\)\.
Let𝒮⊂ΣLN\\mathcal\{S\}\\subset\\Sigma\_\{\\operatorname\{LN\}\}be the set of normalized density\-operator rows visited across inputs \(all heads pooled\), and letφ=uNres∘⋯∘u1\\varphi=u\_\{N\_\{\\mathrm\{res\}\}\}\\circ\\cdots\\circ u\_\{1\}be the row map of Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\.*Hypothesis \(not proved; sampled diagnostics for it, at the level of the full composition, are reported in Appendix[D](https://arxiv.org/html/2608.10288#A4)\): each unituju\_\{j\}isLjL\_\{j\}\-Lipschitz on the relevant tube around the set it receives\.*
1. \(i\)diamφ\(𝒮\)≤\(∏jLj\)diam𝒮\\operatorname\{diam\}\\varphi\(\\mathcal\{S\}\)\\leq\\bigl\(\\prod\_\{j\}L\_\{j\}\\bigr\)\\operatorname\{diam\}\\mathcal\{S\}\. If additionallyLj≤κ<1L\_\{j\}\\leq\\kappa<1for alljj\(the*contractive regime*, an assumption about the trained weights that residual connections and LayerNorm do not automatically deliver\), the output diameter decays exponentially in depth,diamφ\(𝒮\)≤κNresdiam𝒮\\operatorname\{diam\}\\varphi\(\\mathcal\{S\}\)\\leq\\kappa^\{N\_\{\\mathrm\{res\}\}\}\\operatorname\{diam\}\\mathcal\{S\}, and the output ofφ\\varphion𝒮\\mathcal\{S\}is confined to a small set\. We emphasize that theuju\_\{j\}are*distinct*learned maps: this is a diameter bound for a composition, not a Banach fixed\-point iteration, and no fixed point or attractor dynamics is asserted\.
2. \(ii\)In that regime,A\(x\)≈𝟏α∗⊤A\(x\)\\approx\\mathbf\{1\}\\alpha^\{\\ast\\top\}simultaneously for every head and layer sharingφ\\varphi, whereα∗\\alpha^\{\\ast\}may be taken to beφ\(r0\)\\varphi\(r\_\{0\}\)for any fixedr0∈𝒮r\_\{0\}\\in\\mathcal\{S\}\(consistent with the rank\-one, cross\-head\-identical singularity condition of Proposition[3\.9](https://arxiv.org/html/2608.10288#S3.Thmtheorem9)and the observations of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]\)\. The transfer from the Stage\-2 concentration statement \(about the normalized rowsrxr\_\{x\}ofD~x/Sx\\widetilde\{D\}\_\{x\}/S\_\{x\}\) to the*implemented*row\-map inputLN\(D~x\)\\operatorname\{LN\}\(\\widetilde\{D\}\_\{x\}\)passes through theε\\varepsilon\-LayerNorm scale discrepancy: writingδLN\(x\)=‖LN\(D~x\)−LN\(rx\)‖\\delta\_\{\\operatorname\{LN\}\}\(x\)=\\left\\lVert\\operatorname\{LN\}\(\\widetilde\{D\}\_\{x\}\)\-\\operatorname\{LN\}\(r\_\{x\}\)\\right\\rVertper row, on a domain where the constantsLip\(φ\)\\operatorname\{Lip\}\(\\varphi\)\(row\-map\) andLLNL\_\{\\operatorname\{LN\}\}\(Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(iv\)\) are valid, ‖φ\(LN\(D~x\)\)−φ\(LN\(D~y\)\)‖≤Lip\(φ\)\[δLN\(x\)\+LLN‖rx−ry‖\+δLN\(y\)\],\\left\\lVert\\varphi\(\\operatorname\{LN\}\(\\widetilde\{D\}\_\{x\}\)\)\-\\varphi\(\\operatorname\{LN\}\(\\widetilde\{D\}\_\{y\}\)\)\\right\\rVert\\;\\leq\\;\\operatorname\{Lip\}\(\\varphi\)\\,\\bigl\[\\delta\_\{\\operatorname\{LN\}\}\(x\)\\;\+\\;L\_\{\\operatorname\{LN\}\}\\left\\lVert r\_\{x\}\-r\_\{y\}\\right\\rVert\\;\+\\;\\delta\_\{\\operatorname\{LN\}\}\(y\)\\bigr\],whose three ingredients are exactly the directly measured quantities of Appendix[D](https://arxiv.org/html/2608.10288#A4): the scale discrepanciesδLN\\delta\_\{\\operatorname\{LN\}\}\(small at the median, with a disclosed nonuniform tail\), the concentration radius of Proposition[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10)controlling‖rx−ry‖\\left\\lVert r\_\{x\}\-r\_\{y\}\\right\\rVert, and, absent a tube certificate forLip\(φ\)\\operatorname\{Lip\}\(\\varphi\), the*sampled*composite Jacobians, which are local linearizations, not a uniform constant\.777A seemingly natural alternative adds a transferred\-input term to an output\-diameter term over the*actual*visited rows, which the diameter already bounds: such a sum is not false, but it double\-counts rather than derives the transfer, and it hides the role of the measuredδLN\\delta\_\{\\operatorname\{LN\}\}tail\. Consequently Stage 2 concentration and the sampled Stage\-3 diagnostics are two separate pieces of evidence for the collapse mechanism; they combine into a single quantitative budget only under a tube\-uniform bound onLip\(φ\)\\operatorname\{Lip\}\(\\varphi\), which is not claimed\.
3. \(iii\)\(Converse, pairwise\.\) If for a pair of visited rowsr,r′r,r^\{\\prime\}one has the lower bound‖φ\(r\)−φ\(r′\)‖≥c0‖r−r′‖\\left\\lVert\\varphi\(r\)\-\\varphi\(r^\{\\prime\}\)\\right\\rVert\\geq c\_\{0\}\\left\\lVert r\-r^\{\\prime\}\\right\\rVert, then the corresponding rows ofAAdiffer by at leastc0‖r−r′‖c\_\{0\}\\left\\lVert r\-r^\{\\prime\}\\right\\rVert: that selected pair is not collapsed\. A lower bound on the*sampled*order parameter of Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1)requires, in addition, a distributional assumption \(that the two stochastic continuations produce separated row pairs with positive probability, together with control of the RMS denominator\); wherever such a bound is invoked it is under that stated assumption\. The pairwise statement is consistent with the sub\-critical phenomenology of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\.
###### Proof\.
\(i\) is the sub\-multiplicativity of Lipschitz constants under composition applied to the diameter of the image\. \(ii\): every row of every head’s normalized density operator lies in𝒮\\mathcal\{S\}, and for any fixedr0∈𝒮r\_\{0\}\\in\\mathcal\{S\}every output lies withindiamφ\(𝒮\)≤κNresdiam𝒮\\operatorname\{diam\}\\varphi\(\\mathcal\{S\}\)\\leq\\kappa^\{N\_\{\\mathrm\{res\}\}\}\\operatorname\{diam\}\\mathcal\{S\}ofα∗=φ\(r0\)\\alpha^\{\\ast\}=\\varphi\(r\_\{0\}\)\(a set of diameterDDneed not lie in a ball of radiusD/2D/2in dimension\>1\>1, but any of its points serves as a center at radiusDD\); the sameφ\\varphiserves all heads of the layer \(Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\(iii\)\)\. The transfer display is the triangle inequality‖LN\(D~x\)−LN\(D~y\)‖≤δLN\(x\)\+‖LN\(rx\)−LN\(ry\)‖\+δLN\(y\)\\left\\lVert\\operatorname\{LN\}\(\\widetilde\{D\}\_\{x\}\)\-\\operatorname\{LN\}\(\\widetilde\{D\}\_\{y\}\)\\right\\rVert\\leq\\delta\_\{\\operatorname\{LN\}\}\(x\)\+\\left\\lVert\\operatorname\{LN\}\(r\_\{x\}\)\-\\operatorname\{LN\}\(r\_\{y\}\)\\right\\rVert\+\\delta\_\{\\operatorname\{LN\}\}\(y\)followed by‖LN\(rx\)−LN\(ry\)‖≤LLN‖rx−ry‖\\left\\lVert\\operatorname\{LN\}\(r\_\{x\}\)\-\\operatorname\{LN\}\(r\_\{y\}\)\\right\\rVert\\leq L\_\{\\operatorname\{LN\}\}\\left\\lVert r\_\{x\}\-r\_\{y\}\\right\\rVert\(Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(iv\)\) and the Lipschitz bound onφ\\varphi\. \(iii\) is immediate from the lower bound\. ∎
Whether the trained map is in fact contractive on the visited tube is an empirical question about the trained weights, and the quantity that controls diameters is the*composition*, not the individual units: per\-unit Lipschitz numbers do not compose informatively \(singular\-vector alignment matters\), and no product of per\-unit summaries is a measurement of the composite\. Appendix[D](https://arxiv.org/html/2608.10288#A4)therefore reports, for a released checkpoint, the largest singular values of the Jacobian of the*full*row mapφ\\varphiat rows visited by every audit prompt, together with empirical pairwise contraction ratios‖φ\(r\)−φ\(r′\)‖/‖r−r′‖\\left\\lVert\\varphi\(r\)\-\\varphi\(r^\{\\prime\}\)\\right\\rVert/\\left\\lVert r\-r^\{\\prime\}\\right\\rVertwithin and across prompts\. These are sampled pointwise statistics on visited data: they can support or refute the contraction hypothesis on the sample, and they do not constitute a tube\-uniform Lipschitz certificate, which would require a maximum \(or a justified high\-probability bound\) over a specified tube and remains future work\.
###### Corollary 5\.13\(End\-to\-end invariance budget, with evaluable constants\)\.
Suppose the visited inputs confine the entries of the preactivationWA\+bWWA\+b\_\{W\}to a compact interval\[−U,U\]\[\-U,U\]and the entries ofALMA\_\{\\mathrm\{LM\}\}to\[mA,MA\]\[m\_\{A\},M\_\{A\}\]withmA≥ϵm\_\{A\}\\geq\\epsilon\. \(SuchUUexists on the visited domain: rows ofAAlie in the image under the continuousφ\\varphiof the compact LayerNorm range of Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(iii\), and a prioriU≤‖W‖∞supvisited‖A‖maxdk\+‖bW‖maxU\\leq\\left\\lVert W\\right\\rVert\_\{\\infty\}\\sup\_\{\\text\{visited\}\}\\left\\lVert A\\right\\rVert\_\{\\max\}d\_\{k\}\+\\left\\lVert b\_\{W\}\\right\\rVert\_\{\\max\}; sharper, measured values are used in Appendix[D](https://arxiv.org/html/2608.10288#A4)\.\) Then the parameter maps \([3\.4](https://arxiv.org/html/2608.10288#S3.E4)\)–\([3\.6](https://arxiv.org/html/2608.10288#S3.E6)\) satisfy an explicit Lipschitz estimate in two*named*regimes, distinguished by the lower endpoint used formAm\_\{A\}: with the*measured*minimum entry ofALMA\_\{\\mathrm\{LM\}\}on the visited states the constant isCGsampC\_\{G\}^\{\\mathrm\{samp\}\}, valid for every*pair of visited states*\(stagewise scalar mean\-value bounds at the endpoint values, whose connecting segments stay inside the measured ranges\); with the architectural floormA=ϵm\_\{A\}=\\epsilonit isCGtubeC\_\{G\}^\{\\mathrm\{tube\}\}, which bounds the derivative along*arbitrary*perturbation paths within the stated bounded domain \(the upper boundsMAM\_\{A\}andUUare hypotheses along the path\), the only certificate valid off the visited sample\. Both share the expression
CG=‖a‖2⋅maxij\(\|Pij\|max\(mAPij−1,MAPij−1\)\)⋅sup\|u\|≤U\|iSwiGLU′\(u\)\|⋅‖W‖2,C\_\{G\}\\;=\\;\\left\\lVert a\\right\\rVert\_\{2\}\\cdot\\max\_\{ij\}\\Bigl\(\\left\\lvert P\_\{ij\}\\right\\rvert\\max\\bigl\(m\_\{A\}^\{P\_\{ij\}\-1\},M\_\{A\}^\{P\_\{ij\}\-1\}\\bigr\)\\Bigr\)\\cdot\\sup\_\{\\left\\lvert u\\right\\rvert\\leq U\}\\left\\lvert\\operatorname\{iSwiGLU\}^\{\\prime\}\(u\)\\right\\rvert\\cdot\\left\\lVert W\\right\\rVert\_\{2\},in which every factor is finite,888The restriction of theiSwiGLU\\operatorname\{iSwiGLU\}\-derivative supremum to\[−U,U\]\[\-U,U\]is essential:iSwiGLU′\(u\)→2u\\operatorname\{iSwiGLU\}^\{\\prime\}\(u\)\\to 2uasu→\+∞u\\to\+\\infty, so the unrestricted supremum is infinite and a constant built from it would be vacuous\. The zero\-crossing caveat attaches toCGtubeC\_\{G\}^\{\\mathrm\{tube\}\}*only*: along a continuous input\-space path between two visited states theiSwiGLU\\operatorname\{iSwiGLU\}preactivation can cross zero, where the only lower bound onALMA\_\{\\mathrm\{LM\}\}entries is the architectural floorϵ=10−9\\epsilon=10^\{\-9\}, at which exponentsPij<1P\_\{ij\}<1makemAPij−1m\_\{A\}^\{P\_\{ij\}\-1\}astronomically large\. It does*not*invalidateCGsampC\_\{G\}^\{\\mathrm\{samp\}\}for pairs of visited states: the stagewise argument compares endpoint values within each stage’s measured range, not along any input\-space path\. Appendix[D](https://arxiv.org/html/2608.10288#A4)evaluates both named constants and labels each accordingly\.with the norms of the conclusion fixed explicitly\. The derivative argument proves the Frobenius\-to\-Frobenius inequality
‖ΔGLM‖F≤CG‖ΔA‖F\\left\\lVert\\Delta G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}\\;\\leq\\;C\_\{G\}\\,\\left\\lVert\\Delta A\\right\\rVert\_\{F\}\(an entry\-wise scalar map with derivative bounded byLLisLL\-Lipschitz in Frobenius norm, and left multiplication byWWor byaahas induced Frobenius operator norm at most the multiplier’s spectral norm\)\. If the compared contexts satisfy the uniform per\-row boundmaxi‖ΔAi,:‖2≤εA\\max\_\{i\}\\left\\lVert\\Delta A\_\{i,:\}\\right\\rVert\_\{2\}\\leq\\varepsilon\_\{A\}\(the form delivered by the left side of Proposition[5\.12](https://arxiv.org/html/2608.10288#S5.Thmtheorem12)\(ii\)’s transfer display\), row aggregation gives‖ΔA‖F≤dkεA\\left\\lVert\\Delta A\\right\\rVert\_\{F\}\\leq\\sqrt\{d\_\{k\}\}\\,\\varepsilon\_\{A\}, hence
∥ΔGLM∥2≤∥ΔGLM∥F≤CGdkεA=:εG,\\left\\lVert\\Delta G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{2\}\\;\\leq\\;\\left\\lVert\\Delta G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}\\;\\leq\\;C\_\{G\}\\,\\sqrt\{d\_\{k\}\}\\;\\varepsilon\_\{A\}\\;=:\\;\\varepsilon\_\{G\},and by Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)\(whose hypothesis is the spectral norm\) the logits of cached and uncached inference differ by at mostC\(θ\)εGC\(\\theta\)\\,\\varepsilon\_\{G\}\.999Thedk\\sqrt\{d\_\{k\}\}aggregation cannot be dropped in general: forΔA=𝟏v⊤\\Delta A=\\mathbf\{1\}v^\{\\top\}with‖v‖2=1\\left\\lVert v\\right\\rVert\_\{2\}=1, every row obeys the per\-row bound withεA=1\\varepsilon\_\{A\}=1while‖ΔA‖2=‖ΔA‖F=dk\\left\\lVert\\Delta A\\right\\rVert\_\{2\}=\\left\\lVert\\Delta A\\right\\rVert\_\{F\}=\\sqrt\{d\_\{k\}\}\(an all\-rows\-equal perturbation that a row\-combiningWWtransports undiminished\)\. A sharper constant would require a provedℓ2,∞\\ell\_\{2,\\infty\}\-to\-spectral mixed\-norm analysis of the composite map, not undertaken here\. Appendix[D](https://arxiv.org/html/2608.10288#A4)carries the factordk=8\\sqrt\{d\_\{k\}\}=8explicitly wherever this conversion is used; its assembled end\-to\-end chain enters at a different point, a directly hypothesized spectral radius‖ΔGLM‖2=εspec\\left\\lVert\\Delta G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{2\}=\\varepsilon\_\{\\mathrm\{spec\}\}, and is unaffected by the conversion\.Two further conversions are required before comparing with measured quantities: reported relative\-RMS invariancesεrms\\varepsilon\_\{\\mathrm\{rms\}\}enter the spectral hypothesis of Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)only through‖G−G′‖2≤‖G−G′‖F=dk‖G−G′‖rms\\left\\lVert G\-G^\{\\prime\}\\right\\rVert\_\{2\}\\leq\\left\\lVert G\-G^\{\\prime\}\\right\\rVert\_\{F\}=d\_\{k\}\\,\\left\\lVert G\-G^\{\\prime\}\\right\\rVert\_\{\\mathrm\{rms\}\}times the normalization scale; and a boundBBon the logit perturbation forces bit\-identical*greedy*decoding under the sufficient condition2B<Δ2B<\\Delta, whereΔ\\Deltais the minimum realized top\-two logit margin along the decoded paths: the factor22is necessary in this criterion, since the leading logit can move down byBBwhile the runner\-up moves up byBB\(for anℓ2\\ell\_\{2\}bound the two\-coordinate sufficient comparison is2B<Δ\\sqrt\{2\}\\,B<\\Delta; for stochastic sampling, only a distributional bound under coupled randomness follows\)\. This budget is a certificate only to the extent that its evaluated value closes the halved margin*and*its factors are certified on a domain containing every compared state; on the released checkpoint audited in Appendix[D](https://arxiv.org/html/2608.10288#A4)the chain, with every factor evaluated as a sample\-extrema proxy \(certified analytic derivative envelopes; all other extrema sampled over the audit prompts\), assembles to a coefficient tens of orders of magnitude too large to certify decoding decisions, and the honest summary of cache fidelity remains*empirically indistinguishable under tested decoding*, with the budget quantifying the mechanism rather than certifying the outcome\.
###### Proof\.
Compose the Lipschitz constants of the three maps on the stated domain:A↦WA\+bWA\\mapsto WA\+b\_\{W\}contributes‖W‖2\\left\\lVert W\\right\\rVert\_\{2\}; the entry\-wiseiSwiGLU\\operatorname\{iSwiGLU\}contributes its derivative bound on\[−U,U\]\[\-U,U\]; the entry\-wise power \([2\.1](https://arxiv.org/html/2608.10288#S2.E1)\) has derivativePijuPij−1P\_\{ij\}u^\{P\_\{ij\}\-1\}, extremized at the ends of\[mA,MA\]\[m\_\{A\},M\_\{A\}\]; the left action ofaacontributes‖a‖2\\left\\lVert a\\right\\rVert\_\{2\}; biases drop out of differences\. Finiteness of each factor is by compactness of the stated domain\. ∎
#### 5\.4\.4\.Why training at criticality might select the constant map: a hypothesis
The preceding results reduce the invariance question to:*why does gradient training drive the row mapφ\\varphiinto the contractive, constant\-output regime?*A derivation from the training dynamics is open \(Section[9](https://arxiv.org/html/2608.10288#S9), Problem 1\)\. We state the proposed selection mechanism explicitly as a hypothesis, with the epistemic status of each ingredient tagged; it is consistent with the observations cited but is not derived from them\.
###### Hypothesis 5\.14\(Selection of the constant map at criticality\)\.
1. \(1\)Expressivity permits it\(proved, for the collapsed instance\)\. By Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(ii\) and Remark[5\.5](https://arxiv.org/html/2608.10288#S5.Thmtheorem5), replacing an \(approximately\) constantGLMG\_\{\\mathrm\{LM\}\}by its constant value preserves the trained model’s inference map: all instance\-specific information can be carried by the linearQ,K,VQ,K,Vpath\. The constant\-map manifold is therefore loss\-competitive*for models near it*\.
2. \(2\)The gradient signal favors it\(heuristic\)\. The loss reachesφ\\varphithrough batch averages of the concentrated input \(Proposition[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10), under its idealization\): the instance\-specific component of∂ℒ/∂φ\\partial\\mathcal\{L\}/\\partial\\varphiis argued to haveO\(S−1/2\)O\(S^\{\-1/2\}\)leverage relative to its dataset\-mean component\. This is a plausibility argument: concentrated inputs do not by themselves force instance\-specific gradients to be negligible, and no measurement of the gradient decomposition has been made\.
3. \(3\)Stability at criticality enforces it\(heuristic\)\. An input\-sensitiveφ\\varphi\(Proposition[5\.12](https://arxiv.org/html/2608.10288#S5.Thmtheorem12)\(iii\)\) is argued to have larger Jacobians through the deductive path, hence larger loss curvature in the PLGA parameters, so that at the large maximum learning rates of the near\-critical regime such directions are annealed away or trigger dragon\-king events\[[14](https://arxiv.org/html/2608.10288#bib.bib5),[36](https://arxiv.org/html/2608.10288#bib.bib32)\], while at small learning rates \(sub\-critical\) input\-sensitive solutions survive and the order parameter grows, as observed\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The curvature claim is not established; input sensitivity does not in general imply larger loss curvature\.
Three empirical signatures discriminate this hypothesis from alternatives: \(a\) the*cross\-head identity*ofAA\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]is naturally explained by collapse in the shared row map \(Stage 3\) and not by head\-specific statistics, and the Stage\-3 reading is directly supported on the audited checkpoint by the composite\-Jacobian and pairwise measurements of Appendix[D](https://arxiv.org/html/2608.10288#A4); \(b\) invariance*improves with training data*\(m→0m\\to 0at float resolution for the4141B\-token model\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\), consistent with longer annealing of Stage 3 \(note that training\-token count is distinct from the context lengthSSthat drives Stage 1–2 concentration\); and \(c\) DAG regularization, which constrains the downstream tensors and hence deforms the attractor, measurably*increases*the caching perturbation\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\], a trade\-off expected if invariance is an attractor property rather than an architectural identity\.
## 6\.Power Laws, Scale Invariance, and Self\-Organized Criticality
### 6\.1\.Why power laws: the unique scale\-equivariant interaction
###### Proposition 6\.1\(Scale covariance forces power laws\)\.
Letf:\(0,∞\)→\(0,∞\)f:\(0,\\infty\)\\to\(0,\\infty\)be measurable and suppose there existsggwithf\(λu\)=g\(λ\)f\(u\)f\(\\lambda u\)=g\(\\lambda\)f\(u\)for allλ,u\>0\\lambda,u\>0\(*scale covariance*: rescaling the input rescales the output independently ofuu\)\. Then there existc\>0c\>0andp∈ℝp\\in\\mathbb\{R\}withf\(u\)=cupf\(u\)=c\\,u^\{p\}andg\(λ\)=λpg\(\\lambda\)=\\lambda^\{p\}\.
###### Proof\.
Settingu=1u=1:f\(λ\)=g\(λ\)f\(1\)f\(\\lambda\)=g\(\\lambda\)f\(1\), sog=f/f\(1\)g=f/f\(1\)andggsatisfiesg\(λμ\)=g\(λ\)g\(μ\)g\(\\lambda\\mu\)=g\(\\lambda\)g\(\\mu\)\. Every measurable solution of the multiplicative Cauchy equation on\(0,∞\)\(0,\\infty\)has the formg\(λ\)=λpg\(\\lambda\)=\\lambda^\{p\}for some realpp\[[1](https://arxiv.org/html/2608.10288#bib.bib39), Ch\. 2\]\. Thenf\(u\)=f\(1\)upf\(u\)=f\(1\)u^\{p\}\. ∎
###### Corollary 6\.2\(Elementwise scale covariance\)\.
Among measurable*element\-wise*maps of a single positive entry, the familyu↦cupu\\mapsto c\\,u^\{p\}used in \([3\.5](https://arxiv.org/html/2608.10288#S3.E5)\) is, up to the learned constants, the only one that transforms covariantly under rescaling of its argument\. In this element\-wise sense the potential stage is the minimal scale\-covariant interaction ansatz\.
### 6\.2\.Criticality of the attention dynamics: a conditional spectral dictionary
The claim “the correlation length diverges at criticality”\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]suggests an operator reading through the Markov structure of Proposition[3\.12](https://arxiv.org/html/2608.10288#S3.Thmtheorem12)\. Making it precise requires care on three points\. \(a\) A fixed finite matrix has an atomic spectral measure and cannot carry a density accumulating at11, so the critical case can only concern an infinite\-dimensional operator or a limit of a family\. \(b\) A row\-stochastic matrix is generally nonnormal, so eigen\-expansions require eigenbasis conditioning and the stationary inner product\. \(c\) Thett\-th spectral moment is the autocorrelation⟨f,Etf⟩\\left\\langle f,E^\{t\}f\\right\\rangle, not the norm‖Etf‖\\left\\lVert E^\{t\}f\\right\\rVert, whose decay exponent differs by a factor of two\. The statement below builds all three into its hypotheses: it is a*dictionary lemma*for a family of reversible operators, and its application to PLDR\-LLM is Conjecture[8\.2](https://arxiv.org/html/2608.10288#S8.Thmtheorem2)\.
###### Proposition 6\.4\(Spectral gap vs\. critical slowing down, for reversible families\)\.
Let\(Ed\)d\(E\_\{d\}\)\_\{d\}be a family of Markov operators, each reversible with respect to a stationary distributionπd\\pi\_\{d\}\(hence self\-adjoint onL2\(πd\)L^\{2\}\(\\pi\_\{d\}\)with real spectrum in\[−1,1\]\[\-1,1\]\), and letfdf\_\{d\}be observables withπd\\pi\_\{d\}\-mean zero, normalized inL2\(πd\)L^\{2\}\(\\pi\_\{d\}\), with spectral measuresμfd\\mu\_\{f\_\{d\}\}\(with respect toEdE\_\{d\}\)\.
1. \(i\)\(Gapped/off\-critical\.\) If the spectral edges are uniformly controlled,suppμfd⊆\[−1\+δ′,1−δ\]\\operatorname\{supp\}\\mu\_\{f\_\{d\}\}\\subseteq\[\-1\+\\delta^\{\\prime\},1\-\\delta\]withδ,δ′\>0\\delta,\\delta^\{\\prime\}\>0, then the autocorrelation obeys \|⟨fd,Edtfd⟩πd\|≤rt,r=max\(1−δ,1−δ′\),\\left\\lvert\\left\\langle f\_\{d\},E\_\{d\}^\{t\}f\_\{d\}\\right\\rangle\_\{\\pi\_\{d\}\}\\right\\rvert\\;\\leq\\;r^\{t\},\\qquad r=\\max\(1\-\\delta,\\;1\-\\delta^\{\\prime\}\),with correlation “length”ξ=−1/logr\\xi=\-1/\\log r: exponential decay with a finite scale\.*Both*edges enter: reversibility does not exclude spectrum near−1\-1, and an eigenvalue there produces slowly decaying alternating correlations even under a large gap at\+1\+1\.101010Omitting the lower edge leads to a genuine error, not a technicality: for the reversible two\-state chainE=\(0\.050\.950\.950\.05\)E=\\bigl\(\\begin\{smallmatrix\}0\.05&0\.95\\\\ 0\.95&0\.05\\end\{smallmatrix\}\\bigr\)with uniformπ\\pi, the normalized mean\-zero observable has eigenvalue−0\.9\-0\.9and autocorrelation\(−0\.9\)t\(\-0\.9\)^\{t\}, so withδ′=0\.1\\delta^\{\\prime\}=0\.1,δ=0\.5\\delta=0\.5the one\-sided bound\(1−δ\)t\(1\-\\delta\)^\{t\}fails already att=1t=1\(0\.9≰0\.50\.9\\not\\leq 0\.5\)\. This is near\-periodicity, not nonnormality\.
2. \(ii\)\(Critical\.\) Assume additionally a uniform gap at the lower edge:suppμfd⊆\[−1\+δ0,1\]\\operatorname\{supp\}\\mu\_\{f\_\{d\}\}\\subseteq\[\-1\+\\delta\_\{0\},1\]for a fixedδ0\>0\\delta\_\{0\}\>0independent ofdd\(automatic for lazy or positive semidefinite families, whose spectrum lies in\[0,1\]\[0,1\]\), and thatμfd→μ\\mu\_\{f\_\{d\}\}\\to\\muweakly, so that*in the iterated sense that the family limit is taken first at each fixedttand thet→∞t\\to\\inftyasymptotics are those of the limit measure*,limd→∞⟨fd,Edtfd⟩=∫λt𝑑μ\\lim\_\{d\\to\\infty\}\\left\\langle f\_\{d\},E\_\{d\}^\{t\}f\_\{d\}\\right\\rangle=\\int\\lambda^\{t\}\\,d\\mu\. Gap closing with power law accumulation at the upper edge \(necessarily along the family: each finite\-ddmeasure is atomic\) then yields power law decay, at two levels of hypothesis: 1. \(ii\.a\)\(Comparability\.\) Ifdμ\(1−s\)≍sϑ−1dsd\\mu\(1\-s\)\\asymp s^\{\\vartheta\-1\}\\,dsass↓0s\\downarrow 0\(two\-sided bounds with unspecified positive constants\), then ∫λt𝑑μ=Θ\(t−ϑ\),\(∫λ2t𝑑μ\)1/2=Θ\(t−ϑ/2\),t→∞:\\int\\lambda^\{t\}\\,d\\mu\\;=\\;\\Theta\\bigl\(t^\{\-\\vartheta\}\\bigr\),\\qquad\\Bigl\(\\int\\lambda^\{2t\}\\,d\\mu\\Bigr\)^\{1/2\}\\;=\\;\\Theta\\bigl\(t^\{\-\\vartheta/2\}\\bigr\),\\qquad t\\to\\infty:the decay*exponent*is determined, the coefficient is not, and no asymptotic equivalent is implied \(comparability constants may oscillate\)\. 2. \(ii\.b\)\(Exact edge density\.\) If moreoverdμ\(1−s\)=\(c\+o\(1\)\)sϑ−1dsd\\mu\(1\-s\)=\(c\+o\(1\)\)\\,s^\{\\vartheta\-1\}\\,dsass↓0s\\downarrow 0for some constantc\>0c\>0, with no singular component at the edge, then ∫λt𝑑μ∼cΓ\(ϑ\)t−ϑ,\(∫λ2t𝑑μ\)1/2∼cϑt−ϑ/2,cϑ=cΓ\(ϑ\)2−ϑ/2\.\\int\\lambda^\{t\}\\,d\\mu\\;\\sim\\;c\\,\\Gamma\(\\vartheta\)\\,t^\{\-\\vartheta\},\\qquad\\Bigl\(\\int\\lambda^\{2t\}\\,d\\mu\\Bigr\)^\{1/2\}\\;\\sim\\;c\_\{\\vartheta\}\\,t^\{\-\\vartheta/2\},\\qquad c\_\{\\vartheta\}=\\sqrt\{c\\,\\Gamma\(\\vartheta\)\}\\;2^\{\-\\vartheta/2\}\. In either case correlations decay as a*power law*with no characteristic scale;ξ=∞\\xi=\\infty\. \(The autocorrelation and the norm decay with different exponents; both are recorded to prevent conflation\. No uniformity inttof the family convergence is claimed; without the lower\-edge gap the conclusion fails, e\.g\. for an atom at−1\-1\.\)111111The two\-level split is forced: under comparability alone the exact equivalent is false \(a normalized measure whose edge density is2sϑ−12s^\{\\vartheta\-1\}satisfiesdμ\(1−s\)≍sϑ−1dsd\\mu\(1\-s\)\\asymp s^\{\\vartheta\-1\}dsyet contributes2Γ\(ϑ\)t−ϑ2\\Gamma\(\\vartheta\)t^\{\-\\vartheta\}, and oscillating comparability constants can prevent any asymptotic equivalent from existing\), so \(ii\.a\) claims onlyΘ\\Theta, and the exact constants live in \(ii\.b\), where the density hypothesis supports them\.
###### Proof\.
Both parts are the spectral calculus of a self\-adjoint contraction:⟨f,Etf⟩=∫λt𝑑μf\(λ\)\\left\\langle f,E^\{t\}f\\right\\rangle=\\int\\lambda^\{t\}\\,d\\mu\_\{f\}\(\\lambda\)and‖Etf‖2=∫λ2t𝑑μf\(λ\)\\left\\lVert E^\{t\}f\\right\\rVert^\{2\}=\\int\\lambda^\{2t\}\\,d\\mu\_\{f\}\(\\lambda\)\. \(i\): on the support,\|λ\|≤max\(1−δ,1−δ′\)=r\\left\\lvert\\lambda\\right\\rvert\\leq\\max\(1\-\\delta,1\-\\delta^\{\\prime\}\)=r, so\|∫λt𝑑μf\|≤rt\\left\\lvert\\int\\lambda^\{t\}d\\mu\_\{f\}\\right\\rvert\\leq r^\{t\}\. \(ii\): for fixedtt,λ↦λt\\lambda\\mapsto\\lambda^\{t\}is bounded and continuous on\[−1,1\]\[\-1,1\], so weak convergence of the compactly supportedμfd\\mu\_\{f\_\{d\}\}gives the limit∫λt𝑑μ\\int\\lambda^\{t\}d\\mu\. Splittingμ\\muat1−η1\-\\etafor any fixedη∈\(0,min\{δ0,1\}\)\\eta\\in\(0,\\min\\\{\\delta\_\{0\},1\\\}\): the contribution of\[−1\+δ0,1−η\]\[\-1\+\\delta\_\{0\},1\-\\eta\]is bounded in absolute value bymax\(1−η,1−δ0\)t\\max\(1\-\\eta,1\-\\delta\_\{0\}\)^\{t\}, exponentially small\. Near the upper edge the substitutionλ=1−s\\lambda=1\-sgives Beta\-type integrals:∫0η\(1−s\)tsϑ−1𝑑s∼Γ\(ϑ\)t−ϑ\\int\_\{0\}^\{\\eta\}\(1\-s\)^\{t\}s^\{\\vartheta\-1\}ds\\sim\\Gamma\(\\vartheta\)t^\{\-\\vartheta\}\(viaB\(t\+1,ϑ\)=Γ\(t\+1\)Γ\(ϑ\)/Γ\(t\+1\+ϑ\)∼Γ\(ϑ\)t−ϑB\(t\+1,\\vartheta\)=\\Gamma\(t\+1\)\\Gamma\(\\vartheta\)/\\Gamma\(t\+1\+\\vartheta\)\\sim\\Gamma\(\\vartheta\)\\,t^\{\-\\vartheta\}, Stirling\)\. \(ii\.a\): the hypothesis sandwiches the edge contribution betweenc1c\_\{1\}andc2c\_\{2\}times this integral for some0<c1≤c20<c\_\{1\}\\leq c\_\{2\}, giving theΘ\\Thetabounds and nothing stronger\. \(ii\.b\): writing the edge density as\(c\+o\(1\)\)sϑ−1\(c\+o\(1\)\)s^\{\\vartheta\-1\}and splitting off theo\(1\)o\(1\)factor at a radius where it is uniformly small, dominated convergence carries the constant through,∫λt𝑑μ∼cΓ\(ϑ\)t−ϑ\\int\\lambda^\{t\}d\\mu\\sim c\\,\\Gamma\(\\vartheta\)t^\{\-\\vartheta\}\. The norm computations replacettby2t2t; in \(ii\.b\) this givescΓ\(ϑ\)2−ϑt−ϑc\\,\\Gamma\(\\vartheta\)2^\{\-\\vartheta\}t^\{\-\\vartheta\}inside the square root, i\.e\. the displayedcϑc\_\{\\vartheta\}\. ∎
Whether this dictionary describes PLDR\-LLM is an open question, for two reasons: trained attention matrices are neither reversible nor constant across layers and tokens \(repeated powers of a singleEEdo not model a depth\-LLnetwork\), and no spectral measurement of trained operators near criticality has been published\. We therefore record the intended application as Conjecture[8\.2](https://arxiv.org/html/2608.10288#S8.Thmtheorem2)and use the following correspondence only as interpretive language:*sub\-critical training*↔\\leftrightarrowgapped learned operators, exponential decay of influence;*critical training*↔\\leftrightarrowgap closing with power law spectral accumulation, scale\-free propagation of constraints across the context\.
### 6\.3\.The SOC training picture as a phenomenological framework
Following\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\], pretraining is modeled as a slowly driven dissipative system in the sense of\[[3](https://arxiv.org/html/2608.10288#bib.bib25),[6](https://arxiv.org/html/2608.10288#bib.bib26)\]\. We emphasize the epistemic status before the definitions: what the published experiments establish is*critical\-like optimizer phenomenology associated with low sampled deductive\-output fluctuation*, observed in single training runs per condition with externally tuned schedules\. Establishing self\-organized criticality in the technical sense would additionally require an identified self\-tuning feedback mechanism, separation of drive and relaxation scales, and standard discriminants \(finite\-size scaling and data collapse, susceptibility, avalanche or1/f1/fstatistics\), none of which has been measured; see the empirical program in Section[9](https://arxiv.org/html/2608.10288#S9)\. The definitions below therefore fix the*vocabulary*of the source papers as a phenomenological framework, not as established physics\.
###### Definition 6\.5\(Control and order parameters of PLDR\-LLM pretraining; phenomenological\)\.
Letηmax\\eta\_\{\\max\}be the maximum learning rate andTwT\_\{w\}the linear warm\-up step count of the schedule \(cosine annealing to0\.1ηmax0\.1\\,\\eta\_\{\\max\}\)\. The pair\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)are the*control parameters*: token batches under forward propagation are the slow external drive; gradient updates under backward propagation are the dissipation\. A trained model is:
- •*near\-critical*if its order parameter \([5\.1](https://arxiv.org/html/2608.10288#S5.E1)\) satisfiesm\(θ\)≈0m\(\\theta\)\\approx 0\(empirically≲10−2\\lesssim 10^\{\-2\}, typically≤10−5\\leq 10^\{\-5\}\) while text generation is non\-degenerate;
- •*sub\-critical*ifm\(θ\)=O\(1\)m\(\\theta\)=O\(1\)or larger; loss is lower \(overfit\-like\) but generation degenerates and benchmark scores drop to near\-chance\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\];
- •subject to*dragon\-king events*\[[36](https://arxiv.org/html/2608.10288#bib.bib32)\]\(sharp loss spikes from self\-amplifying drive/dissipation imbalance\), which mark departures from power law criticality and degrade the final state even when the loss trajectory recovers\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\.
The phenomenology reported in\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]is*consistent with*a second\-order\-transition reading: near\-critical models across a range of\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)collapse onto nearly identical loss trajectories; the deductive outputs reach a metastable steady state, theε\\varepsilon\-invariant operators of Section[5](https://arxiv.org/html/2608.10288#S5); and proximity ofm\(θ\)m\(\\theta\)to zero separates the observed phases in agreement with benchmark rankings at the phase level \(not strictly monotonically within phases; see after Remark[5\.3](https://arxiv.org/html/2608.10288#S5.Thmtheorem3)\)\. Three caveats bound what this shows\. The near/sub\-critical labels are substantially defined through the order parameter and generation quality they are then used to explain, so an independent phase criterion fixed in advance is needed to break the circularity; the\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)pairs are externally selected, so the evidence shows schedule\-tuned critical\-like behavior rather than self\-organization; and each condition has a single training run with no uncertainty quantification\. Within those limits, the parallel with criticality hypotheses for cortical dynamics\[[4](https://arxiv.org/html/2608.10288#bib.bib33),[17](https://arxiv.org/html/2608.10288#bib.bib34)\]and with the ubiquity of SOC in natural systems\[[24](https://arxiv.org/html/2608.10288#bib.bib31)\]motivates, but does not support beyond motivation, the conjecture that invariant operators transfer across domains within a universality class \(Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3)\)\.
## 7\.Advantages of PLDR\-LLM over SDPA\-LLM
Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(i\) places the two architectures inside a single family: an SDPA\-LLM is the point of PLDR\-LLM model space at which the operatorGLMG\_\{\\mathrm\{LM\}\}is pinned to the identity\. The comparison in this section is therefore not between rival designs but between a family and its base point, and the one\-sentence form of the whole comparison is:*PLDR\-LLM learns and exposes an operator sector that SDPA fixes in head space, at additional training cost*\. Every advantage below is a statement about what the*learned*operator sector buys relative to the*frozen*one, and every claim is tied to a result of this paper*at that result’s epistemic level*\(several are conditional or prospective, and are so marked\) and, where available, to an experimental anchor in the source papers\. We state the costs with equal explicitness \(§[7\.7](https://arxiv.org/html/2608.10288#S7.SS7)\)\. The comparison is summarized first:
### 7\.1\.A larger parameterized family with distinct training dynamics
Three claims must be kept separate here\. \(1\)*Algebraic inclusion*\(proved, machine\-checked\): SDPA is exactly theGLM=IG\_\{\\mathrm\{LM\}\}=Ipoint of the family \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(i\)\)\. \(2\)*Distinct training dynamics*\(generic\): at training time PLGA operates in regime 3 of §[3\.3](https://arxiv.org/html/2608.10288#S3.SS3), the score map is generically nonlinear in the input \(degenerate choices such asa=0a=0linearize it, §[3\.3](https://arxiv.org/html/2608.10288#S3.SS3)\), and gradients flow through the metric learner and the power law parameters, so by Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(iii\),*where its nonvanishing hypothesis holds*, the two parameterizations induce different gradient flows\. \(3\)*Strict function\-class containment at matched resources*\(not proved\): different training dynamics do not show that no SDPA\-LLM realizes the same input–output function; a nonrepresentability theorem at fixed depth/width/positional scheme would be required, and we do not have one \(Remark[5\.5](https://arxiv.org/html/2608.10288#S5.Thmtheorem5)\)\. The claims below rest on \(1\) and \(2\) only\. The ablations of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]instantiate \(2\) experimentally: under matched data and schedules, the model with a learnableGLMG\_\{\\mathrm\{LM\}\}attains the best average benchmark scores \(under the published one\-pass block protocol, Appendix[B](https://arxiv.org/html/2608.10288#A2)\), ahead of models trained with the operator frozen to a transferred, identity \(i\.e\. SDPA\-equivalent\), or random constant, and its loss trajectory is distinct, tracking the transferred\-operator model early in training and the identity\-operator model late, a signature of the learned operator moving through the family rather than sitting at a fixed point of it\.
Even at*inference*, after the collapse of Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(ii\), the cached model’s raw score functions generically remain outside those realizable by an SDPA head, and the difference can be counted at the level of head\-space operators:
###### Corollary 7\.1\(Head\-level positional operator codimension; conditional\)\.
Assume the nonresonance hypothesis of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1), and consider score functions\(n,m,q,k\)↦q⊤R−nGRmk\(n,m,q,k\)\\mapsto q^\{\\top\}R\_\{\-n\}GR\_\{m\}kwith the argumentsq,kq,kranging over all ofℝdk\\mathbb\{R\}^\{d\_\{k\}\}\(*unrestricted\-reachability assumption*\) and the RoPE parameterization fixed\. Then the mapG↦\[\(n,m,q,k\)↦q⊤R−nGRmk\]G\\mapsto\\bigl\[\(n,m,q,k\)\\mapsto q^\{\\top\}R\_\{\-n\}GR\_\{m\}k\\bigr\]from operators to score functions is injective, and the frozen\-GGPLGA head realizes an SDPA\-realizable score function \(for any linear reparametrization of the projections applied before rotation\) if and only ifGGlies in the RoPE commutant\. Consequently,*inside*thedk2d\_\{k\}^\{2\}\-dimensional frozen\-GGPLGA family, the score functions also realizable by pre\-RoPE\-linearly\-reparameterized SDPA form exactly the commutant subfamily, of real dimension2⋅dk/2=dk2\\cdot d\_\{k\}/2=d\_\{k\}, hence have codimension
dk2−dkreal dimensions per headd\_\{k\}^\{2\}\-d\_\{k\}\\quad\\text\{real dimensions per head\}within it, all complementary directions being absolute\-position\-sensitive\. This is a codimension statement about the SDPA\-realizable*intersection*inside the frozen\-GGfamily \(no nesting of the unrestricted global PLGA and SDPA families is asserted\) and an*operator\-dimension count for a single head under the stated assumptions*, not a function\-class separation between complete LLM families: learned projections restrict the reachable\(q,k\)\(q,k\), prior layers and alternative positional encodings can simulate score functions, and no full\-model expressivity claim is made\.
###### Proof\.
Injectivity: ifq⊤R−n\(G−G′\)Rmk=0q^\{\\top\}R\_\{\-n\}\(G\-G^\{\\prime\}\)R\_\{m\}k=0for allq,kq,kand alln,mn,m, thenR−n\(G−G′\)Rm=0R\_\{\-n\}\(G\-G^\{\\prime\}\)R\_\{m\}=0for some \(hence all\)n,mn,m, soG=G′G=G^\{\\prime\}\. Realizability: suppose linear pre\-rotation reparameterizationsA,BA,Bof the query and key projections realize the frozen\-GGscore family, i\.e\.
\(RnAq\)⊤\(RmBk\)=q⊤R−nGRmkfor alln,mand allq,k∈ℝdk\.\(R\_\{n\}Aq\)^\{\\top\}\(R\_\{m\}Bk\)\\;=\\;q^\{\\top\}R\_\{\-n\}\\,G\\,R\_\{m\}\\,k\\qquad\\text\{for all \}n,m\\text\{ and all \}q,k\\in\\mathbb\{R\}^\{d\_\{k\}\}\.Equality of the bilinear forms givesA⊤Rm−nB=R−nGRmA^\{\\top\}R\_\{m\-n\}B=R\_\{\-n\}GR\_\{m\}for alln,mn,m\. Atn=m=0n=m=0this readsA⊤B=GA^\{\\top\}B=G; atm=nm=nthe left\-hand side is againA⊤BA^\{\\top\}B, soR−nGRn=GR\_\{\-n\}GR\_\{n\}=Gfor everynn:GGlies in the commutant\. Conversely, ifGGcommutes with every rotation, the choiceA⊤=GA^\{\\top\}=G,B=IB=I\(or any factorization ofGGacross the two projections\) realizes the family, sinceA⊤Rm−nB=GRm−n=R−nGRmA^\{\\top\}R\_\{m\-n\}B=GR\_\{m\-n\}=R\_\{\-n\}GR\_\{m\}\. No invertibility ofAAorBBis assumed\. The commutant is⨁j\{cjI2\+sjJ2\}\\bigoplus\_\{j\}\\\{c\_\{j\}I\_\{2\}\+s\_\{j\}J\_\{2\}\\\}\(Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\), of real dimensiondkd\_\{k\}, insideℝdk×dk\\mathbb\{R\}^\{d\_\{k\}\\times d\_\{k\}\}of dimensiondk2d\_\{k\}^\{2\}; position sensitivity of the complement is Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)’s converse direction\. ∎
For the reference head widthdk=64d\_\{k\}=64this is40324032additional head\-level score\-function dimensions \(under the corollary’s assumptions\) that an SDPA head with RoPE and linear reparametrization before rotation cannot express\. The count itself is an arithmetic consequence of the rotation\-commutant principle discussed in Remark[4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2)\(established in the RoPE literature,\[[44](https://arxiv.org/html/2608.10288#bib.bib43),[40](https://arxiv.org/html/2608.10288#bib.bib44)\]\); what is specific to PLGA is the object being counted: the inserted head\-space operator sector that the architecture generates, exposes, and can cache\. Whether trained models actually exploit them is measurable \(projectG∗G^\{\\ast\}onto the commutant and report the residual\), which is precisely the diagnostic of Open Problem 3 in Section[9](https://arxiv.org/html/2608.10288#S9)\. SDPA is not operator\-free: in pre\-projection coordinates its scores carry the learned constant bilinear formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}, and operator\-level regularization or diagnostics*can*be posed for that object; PLGA’s distinction is that the head\-space operator is input\-conditioned, directly exposed, and generated anew from the current input’sQ~⊤Q~\\widetilde\{Q\}^\{\\top\}\\widetilde\{Q\}rather than factored into fixed learned projections, which makes the corresponding questions direct rather than reconstructive\.
### 7\.2\.An inspectable law representation
In an SDPA\-LLM the dataset\-level structure of attention is diffused through the projection weights; the operator sector is, in the language of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\], a hidden variable pinned toIIat all times\. In a PLDR\-LLM the same structural role is played by exposed tensors with proved properties: strictly positiveALMA\_\{\\mathrm\{LM\}\}with Perron–Frobenius spectral structure \(Theorem[3\.8](https://arxiv.org/html/2608.10288#S3.Thmtheorem8)\), a potential tensor generated by the unique scale\-equivariant elementwise interaction \(Corollary[6\.2](https://arxiv.org/html/2608.10288#S6.Thmtheorem2)\), a rank\-one collapse with an explicit algebraic profile \(Proposition[3\.9](https://arxiv.org/html/2608.10288#S3.Thmtheorem9)\), and an \(empirically\) invariantG∗G^\{\\ast\}with a quantified caching perturbation \(Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)\)\. Exposure is not cosmetic; it is what makes the following*operations*available in direct form \(for SDPA, operator\-level analogues can at best be posed for the fixed pre\-projection formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}; there is no input\-conditioned, exposed operator to regularize, read, or transplant\):
- •Regularization and metrics on the operator sector\.The DAG loss \([4\.11](https://arxiv.org/html/2608.10288#S4.E11)\) monitors and shapes the cycle content of the learned operators; applied as a regularizer it improved benchmark scores \(one\-pass block protocol\) over the unregularized base model without scaling model or data\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\], and as a metric it separates models whose loss curves are indistinguishable\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\]\.
- •Exponent readout\.The exponentsPPand couplingsaaare directly readable from, and monitorable in, the trained model; reading them as scaling laws of the training domain is the analogy of Remark[6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3), testable but not established\.
- •Operator transplantation\.G∗G^\{\\ast\}is a portable artifact: it can be extracted from one model and installed in another\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]; this is the experimental substrate of the operator\-transfer program \(Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3)\)\.
### 7\.3\.An intrinsic evaluation diagnostic \(prospective\)
The order parameterm\(θ\)m\(\\theta\)\(Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1)\) separates near\-critical from sub\-critical PLDR\-LLMs using nothing but the model’s own deductive fluctuations, in agreement with curated\-benchmark rankings at the phase level in the published sample\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]; whether it resolves finer differences than the phase is open \(the within\-phase reversals noted after Remark[5\.3](https://arxiv.org/html/2608.10288#S5.Thmtheorem3)are unresolved without uncertainty estimates\), so as an evaluation tool it is a diagnostic that separates the labeled conditions*in the published sample*, and a*prospective*tool beyond it: validation outside that sample \(multiple seeds, uncertainty estimates, held\-out thresholds, an independent phase criterion\) has not been carried out\. The diagnostic has no direct SDPA counterpart, for a reason worth stating precisely:
The practical weight of this advantage, if the diagnostic survives the validation protocol of Section[9](https://arxiv.org/html/2608.10288#S9), is largest exactly where benchmarks are weakest: small models, non\-language domains, and data\-limited settings, where holding out a benchmark suite is unaffordable and an intrinsic indicator of generalization is the difference between a validatable and an unvalidatable model\.
### 7\.4\.Phase\-aware, self\-instrumented training
The SOC formalization \(Definition[6\.5](https://arxiv.org/html/2608.10288#S6.Thmtheorem5)\) gives PLDR\-LLM training an internal phase diagnostic that SDPA training lacks\. The empirical record shows why this matters: sub\-critical PLDR\-LLMs achieve*lower*training loss while generating token salad, and models with different tokenizers or warm\-up schedules can be indistinguishable on loss/accuracy while differing sharply in DAG loss and deductive behavior\[[12](https://arxiv.org/html/2608.10288#bib.bib3),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The loss is the standard intrinsic scalar a pipeline monitors, and it fails to discriminate these phases; an SDPA\-LLM has other internal signals \(attention entropy, activation and gradient statistics, calibration, the pre\-projection formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}\), but none of them is a fluctuation statistic of an exposed operator sector; a PLDR pipeline monitorsm\(θ\)m\(\\theta\), the DAG losses, and dragon\-king events \(Definition[6\.5](https://arxiv.org/html/2608.10288#S6.Thmtheorem5)\) and can reject bad runs on model\-internal evidence\. Proposition[6\.4](https://arxiv.org/html/2608.10288#S6.Thmtheorem4)together with Conjecture[8\.2](https://arxiv.org/html/2608.10288#S8.Thmtheorem2)supplies a candidate operator\-level meaning of the phases \(gap closing versus gapped spectra\); until that conjecture is tested, the diagnostics are predictive rather than interpreted\.
### 7\.5\.Inference efficiency and a deployment asymmetry
With KV\-cache and G\-cache enabled, the deep PLGA subnetwork is executed once per prompt and then removed from the loop \(§[4\.4](https://arxiv.org/html/2608.10288#S4.SS4)\); empirically this yields a∼3×\\sim 3\\timesspeedup over the uncached model, with aggregate deductive\-output statistics stable to1515printed decimal digits\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]; the published cached\-versus\-uncached benchmark evaluations are unchanged \(scored under the one\-pass block protocol; see §[4\.4](https://arxiv.org/html/2608.10288#S4.SS4)and Appendix[B](https://arxiv.org/html/2608.10288#A2)\)\. A fixed\-inference deployment package*can*additionally omit the PLGA weights, since the cached path never evaluates them; enabling caching by itself does not shrink the stored parameter set, and the smaller artifact is a separate packaging step\. The reported2727–39%39\\%speed advantage over a comparably sized SDPA reference\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]is a cross\-stack comparison \(custom PLDR code versus a Hugging Face pipeline, different kernels and generation wrappers\) and is not attributable to the attention architecture alone without a matched\-implementation profile\. Two structural points frame the caching result precisely\. First, G\-cache is*exact relative to KV\-cache by construction*\(§[4\.4](https://arxiv.org/html/2608.10288#S4.SS4)\); the only approximation in either cache is freezingAAat the prompt, whose effect is measured \(empirically indistinguishable under tested decoding\) and bounded by the explicit budgets of Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)and Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13), which quantify the mechanism but, as evaluated in Appendix[D](https://arxiv.org/html/2608.10288#A4), are far too loose to certify bit\-identical decoding by themselves\. Second, the training/inference asymmetry \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\(iii\)\) separates the deployable inference artifact \(G∗G^\{\\ast\}plus the linear pathways\) from the trainable asset \(the PLGA network\): the cached artifact reproduces inference but does not permit equivalent continued training of the operator sector, so the nonlinear network can be withheld at deployment, the concealment application*proposed*in\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]\. We flag its status plainly: withholding weights is not established security; a deployed artifact can be queried, distilled, or fine\-tuned, and turning the asymmetry into a security property requires a threat model and attack evaluation that have not been carried out\. An SDPA\-LLM has no such separation: its attention computation*is*its deployable form\.
### 7\.6\.Inductive bias and transfer
The power law stage of PLGA is the unique scale\-equivariant element\-wise interaction \(Corollary[6\.2](https://arxiv.org/html/2608.10288#S6.Thmtheorem2)\): the potential stage is scale\-covariant by construction \(with the architecture\-level caveats of Remark[6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3)\), matched in form to the power law statistics of natural language\[[51](https://arxiv.org/html/2608.10288#bib.bib24)\]and of the SOC domains targeted by the transfer program\[[24](https://arxiv.org/html/2608.10288#bib.bib31)\]\. SDPA’s fixed Euclidean form carries no such bias\. SDPA is not without finite learned objects \(its projections, in particular the pre\-projection formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}, are transferable weights\); what it lacks is an input\-conditioned head\-space operator and exponent tensor extractable as a single transplantable pair\(G∗,P\)\(G^\{\\ast\},P\), which is what the operator\-transfer program of Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3)needs\. Within the family picture, SDPA is the*identity\-coupling*point: no learned operator or exponents in the score stage, while the many other inductive biases of the stack \(projections, positional encoding, FFN\) remain\.
### 7\.7\.Costs, trade\-offs, and honest limits
The advantages above are purchased, and the price should be stated with the same precision:
1. \(1\)Training\-time compute and parameters\.The metric learner dominates the attention\-parameter budget during training \(parameter ratio\#ResL/\#A≈129\\\#\\mathrm\{ResL\}/\\\#A\\approx 129–149149in the reference configurations\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]\); the overhead is removed at inference by caching, but training a PLDR\-LLM costs more per step than its SDPA base point\.
2. \(2\)The criticality search\.Reaching the near\-critical phase requires finding workable\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)pairs, is sensitive to the SwiGLU:LU ratio and to initialization details, and can fail via dragon\-king events\[[13](https://arxiv.org/html/2608.10288#bib.bib4),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]; SDPA\-LLMs, training stably at lower learning rates, are, as the source papers themselves state, easier to train and quick to infer\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The mitigations are exactly the diagnostics of §[7\.4](https://arxiv.org/html/2608.10288#S7.SS4), but the search is real work that SDPA does not require\.
3. \(3\)Benchmark parity at small scale\.At the∼\\sim100M\-parameter,∼\\sim8B\-token scale of the published experiments, average benchmark scores of PLDR\-LLMs are comparable to, not dominant over, SDPA references, with the learned\-operator advantage visible in matched ablations and in the longer4141B\-token run that overtakes the SDPA reference on average\[[12](https://arxiv.org/html/2608.10288#bib.bib3),[13](https://arxiv.org/html/2608.10288#bib.bib4),[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\. The honest summary is: a slight performance edge under matched training plus a qualitatively different capability set, not benchmark dominance; behavior at billion\-parameter scale is an open empirical question\.
4. \(4\)Theory debt\.The mechanism of §[5\.4](https://arxiv.org/html/2608.10288#S5.SS4)is a hypothesis with proved ingredients, not yet derived from the training dynamics \(Open Problem 1\), and the transfer advantages of §[7\.6](https://arxiv.org/html/2608.10288#S7.SS6)rest on Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3), which is stated to be falsified or confirmed, not assumed\.
5. \(5\)Evidence base\.All empirical anchors in this section come from single training runs per condition within one research program, without uncertainty estimates or same\-stack matched baselines; the empirical program of Section[9](https://arxiv.org/html/2608.10288#S9)lists what independent replication requires\. The claims above are calibrated to this evidence and would strengthen or fall with it\.
One sentence carries the whole comparison:*SDPA fixes the operator sector a priori; PLDR\-LLM learns it, collapses it at inference \(exactly under the invariance hypothesis, approximately as observed\), and exposes it*\. The advantages flow from the learning \(distinct training dynamics, the measured edge under matched training\), from the collapse \(caching with quantified perturbation, deployment asymmetry\), and from the exposure \(diagnostics, the prospective intrinsic evaluation, exponent readout, transfer program\), while the observed costs concentrate in reaching and holding the near\-critical regime; the theory motivates, but does not prove, that concentration\.
## 8\.Conjectures
This section collects the open claims of the research program as precise, falsifiable conjectures\. None of them is used as a premise elsewhere in the paper\.
###### Conjecture 8\.1\(Rigidity of the invariant operator\)\.
In the joint limit of widthdk→∞d\_\{k\}\\to\\infty, depth fixed, and training tokens→∞\\to\\inftyat near\-criticality \(m\(θ\)→0m\(\\theta\)\\to 0\), the deductive mapx↦GLM\(x\)x\\mapsto G\_\{\\mathrm\{LM\}\}\(x\)converges to a constantG∗G^\{\\ast\}exactly, for almost every input under the data distribution, and the finite\-size fluctuation obeys a power lawε\(dk,Ntokens\)∼dk−αNtokens−β\\varepsilon\(d\_\{k\},N\_\{\\mathrm\{tokens\}\}\)\\sim d\_\{k\}^\{\-\\alpha\}N\_\{\\mathrm\{tokens\}\}^\{\-\\beta\}with universal exponents\. \(The observed progression10−6→010^\{\-6\}\\to 0at float resolution under a5×5\\timestoken increase\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\], confounded by schedule differences, is a first data point, not an exponent estimate\. TheO\(S−1/2\)O\(S^\{\-1/2\}\)andO\(1/S\)O\(1/S\)source terms of Proposition[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10)concern the*context length*SS, notNtokensN\_\{\\mathrm\{tokens\}\}; they constrain the within\-pass fluctuation floor, whileβ\\betamust come from the training dynamics\.\)
###### Conjecture 8\.2\(Spectral form of the order parameter\)\.
There is a symmetrization of the trained attention operators \(with respect to their stationary measures, or a reversibilization\) and a family limit \(in context length and width\) under which near\-critical models develop power law spectral accumulation at11with exponentϑ\\varthetain the sense of Proposition[6\.4](https://arxiv.org/html/2608.10288#S6.Thmtheorem4)\(ii\.a\), including its uniform lower\-edge gap \(the stronger exact\-density form \(ii\.b\), with its constant, is a sharper version of the same conjecture\), while sub\-critical models have a uniform spectral gap; and benchmark\-relevant reasoning ability is a monotone function ofϑ\\vartheta\. \(The conjecture’s burden includes constructing the reversible family to which Proposition[6\.4](https://arxiv.org/html/2608.10288#S6.Thmtheorem4)applies; per\-layer, per\-token operator variation means no single\-matrix formulation is adequate\.\) This would upgrade the order parameter \([5\.1](https://arxiv.org/html/2608.10288#S5.E1)\) from a fluctuation diagnostic to a spectral one, computable from a single forward pass\.
The third conjecture formalizes the transplantation experiments of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\]: aGLMG\_\{\\mathrm\{LM\}\}learned on one dataset interval can be transplanted as the constant operator of a fresh model trained on a different interval, and performs comparably to \(and differently from\) identity or random constant operators\.
###### Conjecture 8\.3\(Operator transfer across domains: congruence alignment\)\.
Let𝒟1,𝒟2\\mathcal\{D\}\_\{1\},\\mathcal\{D\}\_\{2\}be data distributions whose generative processes lie in the same universality class \(equal critical exponents of their long\-range statistics; making this precise is part of the conjecture’s burden\)\. LetG1∗,G2∗G\_\{1\}^\{\\ast\},G\_\{2\}^\{\\ast\}be the invariant operators of PLDR\-LLMs pretrained to near\-criticality \(m→0m\\to 0\) on each\. Then there exist invertible linear mapsS\(ℓ,i\)S^\{\(\\ell,i\)\}, acting as a*common*change of frame on queries and keys and constrained to a proper subgroup fixed by the conjecture \(e\.g\. orthogonal maps, or maps of bounded condition number; the constraint is essential, since with unconstrained independent query and key frames all full\-rank matrices are left–right equivalent and the statement would be near\-vacuous\), such that
G2∗\(ℓ,i\)=\(S\(ℓ,i\)\)−⊤G1∗\(ℓ,i\)\(S\(ℓ,i\)\)−1\+o\(1\)G\_\{2\}^\{\\ast\(\\ell,i\)\}\\;=\\;\\bigl\(S^\{\(\\ell,i\)\}\\bigr\)^\{\-\\top\}G\_\{1\}^\{\\ast\(\\ell,i\)\}\\,\\bigl\(S^\{\(\\ell,i\)\}\\bigr\)^\{\-1\}\+o\(1\)in the joint limit of model width and data\. The transformation law is the*congruence*action, sinceGLMG\_\{\\mathrm\{LM\}\}is a bilinear form on query–key pairs: ifq′=Sqq^\{\\prime\}=Sqandk′=Skk^\{\\prime\}=Sk, preservation ofq⊤Gkq^\{\\top\}GkrequiresG′=S−⊤GS−1G^\{\\prime\}=S^\{\-\\top\}GS^\{\-1\}\.121212The similarity actionSGS−1SGS^\{\-1\}would be the wrong transformation law here, and eigenvalues the wrong invariants: eigenvalues of a nonsymmetric bilinear\-form matrix are not invariants of the congruence action\. The correct invariants are those of the constrained congruence classes, e\.g\. the signature of the symmetric part and the rank data of the pairing; congruence of nonsymmetric forms has subtleties beyond the symmetric\-part signature, part of the conjecture’s burden\.Equivalently, the congruence\-invariants of the operators \(not their eigenvalues\) are universality\-class data\.
Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3)concerns the operators only; it does not by itself move a model between domains\. The stronger, separate claim is:
###### Conjecture 8\.4\(Whole\-model domain transfer\)\.
Under the hypotheses of Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3), there is an explicit transport procedure, freezing the extracted pair\(G∗,P\)\(G^\{\\ast\},P\)of the source model and re\-fitting only a stated list of components \(tokenizer/embedding maps, linear projections, and readout; the nonlinear PLGA network stays frozen\), such that the transported model reaches benchmark parity on the target domain, within a stated tolerance, with a from\-scratch control trained at matched parameter, token, and FLOP budget\. \(Operator congruence alone, Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3), does not imply this: tokenizer, embeddings, projections, FFN, normalization, and readout are not transported by a congruence ofG∗G^\{\\ast\}; which of them must be re\-fit, and at what cost, is exactly what this conjecture asserts to be small\.\)
Both conjectures are falsifiable with existing tooling: train on two same\-class synthetic sources \(e\.g\. two SOC sandpile family simulators\), extractG∗G^\{\\ast\}, and test alignment of the constrained congruence\-invariants \(Conjecture[8\.3](https://arxiv.org/html/2608.10288#S8.Thmtheorem3)\); then run the transport procedure against its matched from\-scratch control \(Conjecture[8\.4](https://arxiv.org/html/2608.10288#S8.Thmtheorem4)\)\.
## 9\.Discussion and Open Problems
We collected the analytical skeleton of the PLDR\-LLM program, with each piece at its own epistemic level: a five\-operator attention mechanism whose power law stage is the unique scale\-covariant elementwise interaction \(Corollary[6\.2](https://arxiv.org/html/2608.10288#S6.Thmtheorem2), with the scope limits of Remark[6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3)\); Perron–Frobenius structure of the positive interaction tensor and rank\-one algebra of the collapsed generatorAA\(Theorem[3\.8](https://arxiv.org/html/2608.10288#S3.Thmtheorem8), Proposition[3\.9](https://arxiv.org/html/2608.10288#S3.Thmtheorem9)\); an exact algebraic collapse of inference onto a constant\-operator model*under the invariance hypothesis*, with SDPA as the identity point of the family \(Theorem[5\.4](https://arxiv.org/html/2608.10288#S5.Thmtheorem4)\) and quantified perturbation bounds \(Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7), Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13), evaluated in Appendix[D](https://arxiv.org/html/2608.10288#A4)\); a conditional three\-stage analysis \(rotary twirl, statistical concentration, and row\-map contraction\) of the origin of the observed deductive\-output invariance, with its hypotheses stated and its measurable diagnostics computed directly on a released checkpoint \(Section[5\.4](https://arxiv.org/html/2608.10288#S5.SS4), Hypothesis[5\.14](https://arxiv.org/html/2608.10288#S5.Thmtheorem14)\); walk\-counting and commutant identities connecting the regularizer and positional geometry to spectral theory \(Theorem[4\.4](https://arxiv.org/html/2608.10288#S4.Thmtheorem4)with Remarks[4\.5](https://arxiv.org/html/2608.10288#S4.Thmtheorem5)and[4\.6](https://arxiv.org/html/2608.10288#S4.Thmtheorem6), Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\); a phenomenological SOC framework for training with an intrinsic order parameter \(Section[6](https://arxiv.org/html/2608.10288#S6)\); and a consolidated account of the advantages over SDPA\-LLMs at their actual strength, including a conditional head\-level positional operator codimension \(Section[7](https://arxiv.org/html/2608.10288#S7)\)\.
Open problems, beyond the conjectures of Section[8](https://arxiv.org/html/2608.10288#S8):
1. \(1\)Dynamics of the collapse\.Derive, from the gradient flow of \([4\.11](https://arxiv.org/html/2608.10288#S4.E11)\)\-augmented cross\-entropy, the convergence of the shared row mapφ\\varphi\(Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\) to the constant\-map manifold at critical\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\), turning the three\-part mechanism of Section[5\.4](https://arxiv.org/html/2608.10288#S5.SS4)into a theorem about the training dynamics, ideally identifying the constant\-map manifold as an attracting invariant manifold whose stability changes at the critical point\.
2. \(2\)An honest RG map\.Construct an explicit coarse\-graining on token sequences under which the layer map of Definition[4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3)is \(approximately\) covariant, upgrading the RG reading of Section[6\.3](https://arxiv.org/html/2608.10288#S6.SS3)from correspondence to theorem\.
3. \(3\)Commutant diagnostics\.Measure the distance of trainedG∗G^\{\\ast\}from the RoPE commutant of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)across layers; this quantifies how much absolute\-position geometry natural language demands, a question with no direct analogue for SDPA’s fixed identity operator \(its pre\-projection formWQWK⊤W\_\{Q\}W\_\{K\}^\{\\top\}is constant, not input\-conditioned\)\. A first such measurement now exists: on the audited checkpoint the commutant residual‖GLM−ΠcommGLM‖F/‖GLM‖F\\left\\lVert G\_\{\\mathrm\{LM\}\}\-\\Pi\_\{\\mathrm\{comm\}\}G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}/\\left\\lVert G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}is reported per layer and head in Appendix[D](https://arxiv.org/html/2608.10288#A4); the open problem is its behavior across scales, checkpoints, and training trajectories\. Relatedly, measure the empirical size of the off\-commutant part ofD~/S\\widetilde\{D\}/Sagainst theO\(1/S\)O\(1/S\)twirl bound of Lemma[5\.9](https://arxiv.org/html/2608.10288#S5.Thmtheorem9)\.
4. \(4\)The empirical program\.The empirical claims of this paper rest on single training runs per condition from one research program\. What would settle them: multiple pretraining seeds per\(ηmax,Tw\)\(\\eta\_\{\\max\},T\_\{w\}\)cell over a dense grid, at several model scales and on independent corpora, with uncertainty on bothm\(θ\)m\(\\theta\)\(in the stabilized normalization of Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1)\) and benchmark scores; an*independent*phase criterion fixed in advance of measuring the order parameter and evaluated blind against behavior, to break the labeling circularity noted in §[6\.3](https://arxiv.org/html/2608.10288#S6.SS3); parameter\-, token\-, and FLOP\-matched SDPA, constant\-GG, no\-power, and randomized\-PPbaselines run in the same software stack \(with a profiler breakdown for any speed claims\), together with frozen\-, transferred\-, and random\-GGinference controls; the commutant\-projection intervention \(project the trainedGLMG\_\{\\mathrm\{LM\}\}onto the RoPE commutant and measure the behavioral change, the causal counterpart of the occupancy measurement in Appendix[D](https://arxiv.org/html/2608.10288#A4)\); element\-wise cached\-versus\-recomputed comparisons at score, softmax, logit, KL/TV, and token\-decision levels across prompts, layers, heads, and decoding steps, extended to held\-out prompt sets, longer contexts, adversarial distribution shifts, and multiple checkpoints; the historical\-row context\-interaction probe of Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)run at a sequence of*training*checkpoints, locating when prefix consistency \(Definition[4\.8](https://arxiv.org/html/2608.10288#S4.Thmtheorem8)\) emerges as the operator collapses, together with the sequential\-versus\-block score gap along the same trajectory \(extending the two\-released\-checkpoint contrast of Appendix[D](https://arxiv.org/html/2608.10288#A4)\); the blockwise\-CE\-versus\-sequential\-NLL gap along the same training trajectory: the target\-exposure channel of §[4\.3](https://arxiv.org/html/2608.10288#S4.SS3)is present precisely while the metric learner is input\-sensitive, so a collapse measured only at the end state does not settle its role during training; sequential rescoring of the actual benchmark items across*every*checkpoint used for causal claims about operator, DAG, or phase effects, the stated precondition for any future claim of benchmark superiority or causal phase effects \(the two released checkpoints are covered in Appendix[D](https://arxiv.org/html/2608.10288#A4)\); singular\-value spectra and numerical ranks in place of determinants; measured row\-map Jacobians, frequency\-resolved twirl residuals, commutant residuals, and end\-to\-end Lipschitz factors along training \(extending Appendix[D](https://arxiv.org/html/2608.10288#A4)from one checkpoint to trajectories\); and, for the SOC reading specifically, finite\-size scaling and data collapse, susceptibility, avalanche or1/f1/fstatistics, and an identified self\-tuning feedback mechanism\. Independent replication outside this research program would materially change the strength of every empirical claim above\.
The broader claim of the program is that PLDR\-LLM is not “a transformer variant” but a family of models parameterized by an operatorGG, containing SDPA atG=IG=I, whose training in the critical\-like regime produces, and exposes, approximately invariant operators\. The exact inclusion and the exposure are theorems and architecture; the invariance is a measured hypothesis with a proposed mechanism; and the reading of the invariant operators as laws of the training domain is the program’s open ambition, to be earned by the empirical program above rather than assumed from the mathematics\.
## Acknowledgments
I am grateful to my parents for their support and patience\. This research was conducted independently without support from a grant or corporation\.
## Disclosure of the use of AI tools
The author discloses that \(Claude, model Claude Fable 5, Anthropic\) was used in the preparation of this article: in organizing the material of the source papers\[[10](https://arxiv.org/html/2608.10288#bib.bib1),[11](https://arxiv.org/html/2608.10288#bib.bib2),[12](https://arxiv.org/html/2608.10288#bib.bib3),[13](https://arxiv.org/html/2608.10288#bib.bib4),[14](https://arxiv.org/html/2608.10288#bib.bib5)\], in drafting the text and the proofs of the stated results, in cross\-checking the architectural equations of Sections[3](https://arxiv.org/html/2608.10288#S3)–[4](https://arxiv.org/html/2608.10288#S4)against the reference implementations listed in Appendix[B](https://arxiv.org/html/2608.10288#A2), in drafting the Lean 4 formalization of selected proofs released with this article \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\), which was verified mechanically by the Lean proof checker, and in implementing and running the numerical audit of Appendix[D](https://arxiv.org/html/2608.10288#A4)on the released checkpoint\. \(Codex, model GPT\-5\.6\-Sol, OpenAI\) was used as a non\-authorial review and verification tool during revision of this manuscript\. Its involvement included checking mathematical definitions, dimensions, proofs, and claim scope; comparing the manuscript with the cited PLDR\-LLM and PLGA implementations; building and testing the Lean 4 formalization; running the supplied numerical\-audit tests; checking citations and cross\-references; conducting targeted literature searches; and suggesting technical and expository revisions\. All definitions, theorems, proofs, and claims were reviewed and verified by the author, who takes full responsibility for the content of this article\.
## Appendix ANotation
### Result dependency map
Direct proof inputs of the main numbered results, as a navigation aid through the cross\-references\. Each result’s statement carries its own hypotheses; classical inputs are cited at the point of use in the body text\.
## Appendix BCode, Models, and Verification Resources
Reference implementations and released models accompanying the source papers\. The architectural equations of Sections[3](https://arxiv.org/html/2608.10288#S3)and[4](https://arxiv.org/html/2608.10288#S4)were verified for this article against the reference implementations at the exact snapshots recorded in the verification manifest below; “every equation verified” throughout this paper means verified against those snapshots\.
##### Verification manifest\.
The verified snapshots are the following; each identifier is set in fixed1616\-character groups with explicit line breaks, so no engine’s paragraph breaker can carry it past the text block\.
The sequential\-validation audit of Appendix[D](https://arxiv.org/html/2608.10288#A4)additionally pins its held\-out corpus and benchmark items as parquet files at the following Hugging Face dataset revisions \(same grouping convention; the SHA\-256 of each individual file is recorded in the audit’s results JSON, and the release checklist verifies that every identifier in this manifest resolves against its public remote\):
The spaces inside each identifier are grouping only; the identifier is the concatenation of its groups\. The numerical audit of Appendix[D](https://arxiv.org/html/2608.10288#A4)loads the model at the pinned model revision above and records the same hash in its results file\.
##### Model implementations\.
- •
- •PLDR\-LLM with KV\-cache and G\-cache \(PyTorch v510; identical exceptWVW\_\{V\}keeps framework\-default initialization; includes v510G \(predefinedGLMG\_\{\\mathrm\{LM\}\}\) and v510Gi \(transplantedGLMG\_\{\\mathrm\{LM\}\}\) ablation models\):[https://github\.com/burcgokden/PLDR\-LLM\-with\-KVG\-cache](https://github.com/burcgokden/PLDR-LLM-with-KVG-cache)
- •
- •
- •
##### Machine\-checked proofs and audit code \(Lean 4 \+ Python\)\.
The Lean 4/mathlib formalization of the elementary proof cores listed in the introduction, building with no unproved obligations \(lake buildre\-verifies every proof with the Lean kernel\), together with the numerical\-audit code, its Python dependency specification, and the full results files and raw arrays of Appendix[D](https://arxiv.org/html/2608.10288#A4)\(in the repository’saudit/directory\):[https://github\.com/burcgokden/PLDR\-LLM\-Math\-Foundations](https://github.com/burcgokden/PLDR-LLM-Math-Foundations)
##### Released models \(Hugging Face\)\.
The organization[https://huggingface\.co/fromthesky](https://huggingface.co/fromthesky)hosts the pretrained model families of the source papers, including thepldrllmv5/v9series of\[[12](https://arxiv.org/html/2608.10288#bib.bib3)\], thePLDR\-LLM\-v51andPLDR\-LLM\-v51Gseries of\[[13](https://arxiv.org/html/2608.10288#bib.bib4)\], thePLDR\-LLM\-v51\-SOCseries of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\], andv52fine\-tuned variants\. ThePLDR\-LLM\-v51\-SOCrepositories ship a custom Hugging Face Transformers port \(modeling\_pldrllm\.py,PldrllmForCausalLM\) whose configuration records the reference hyperparameters used throughout this paper:Adff=170A\_\{d\\\!f\\\!f\}\{=\}170,Nres=8N\_\{\\mathrm\{res\}\}\{=\}8,nA=2n\_\{A\}\{=\}2,dk=64d\_\{k\}\{=\}64,εLN=10−6\\varepsilon\_\{\\operatorname\{LN\}\}\{=\}10^\{\-6\}, RoPE base10410^\{4\}, untied embeddings, biases on all projections\.
##### Evaluation wrappers and scoring protocol\.
The benchmark evaluations of the source papers run through forks of the EleutherAI evaluation harness with a PLDR\-LLM model wrapper:
- •
- •
The wrapper’s log\-likelihood path scores each answer candidate by*one\-pass block scoring*\(Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\): context and candidate are concatenated, all tokens but the last are evaluated in one call, and the candidate log\-probabilities are summed\. This is a whole\-candidate block score; it coincides with sequential final\-row scoring exactly for one\-token candidates and, on the audited checkpoint, to within the measured gaps of Appendix[D](https://arxiv.org/html/2608.10288#A4)for multi\-token candidates\. The log\-likelihood path performs a single uncached pass by construction, so no generation\-time cache participates in the published scores; the wrapper’s*generation*path is not used by any published score cited here or by any audit in this paper\.
## Appendix CLean Formalization: Exact Coverage
The Lean 4/mathlib repository \([https://github\.com/burcgokden/PLDR\-LLM\-Math\-Foundations](https://github.com/burcgokden/PLDR-LLM-Math-Foundations)\) builds with nosorry,admit,axiom, orunsafedeclarations;lake buildre\-verifies every proof with the Lean kernel\. The precise trust statement is: kernel\-checked*relative to mathlib’s standard classical principles*, not axiom\-free in a foundational sense, and not constructive\. An axiom audit in the repository’s continuous integration \(scripts/check\_axioms\.py\) sweeps\#print axiomsover every exported theorem and fails if anything appears beyondpropext,Classical\.choice, andQuot\.sound\. This appendix states*exactly*which part of each numbered result is kernel\-checked and which is not\. Formalization of a proof core is not evidence for the adjacent unformalized claims, and nothing in this paper cites the formalization as such; in particular, kernel\-checking the algebraic cores could not have detected a defect in the numerical audit’s own bookkeeping, which is why the audit carries its separate semantic test oracles \(Appendix[D](https://arxiv.org/html/2608.10288#A4)\)\.
Outside the scope of the formalization entirely: Perron–Frobenius theory \(Thm\.[3\.8](https://arxiv.org/html/2608.10288#S3.Thmtheorem8); not yet in mathlib\), matrix Bernstein concentration \(Prop\.[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10)\), the reversible\-family spectral dictionary \(Prop\.[6\.4](https://arxiv.org/html/2608.10288#S6.Thmtheorem4)\), the perturbation bounds of Prop\.[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)and Cor\.[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13), and everything stated as a hypothesis, analogy, or conjecture\.
## Appendix DNumerical Audit on a Released Checkpoint
This appendix evaluates, on a released checkpoint, the measurable quantities that the results of this paper depend on: singular\-value spectra in place of determinants \(trained and at random initialization\), the LayerNorm scale error of Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)computed directly, the twirl energy of Lemma[5\.9](https://arxiv.org/html/2608.10288#S5.Thmtheorem9)resolved against*pre\-rotation*data, the commutant residual of the trained operators against Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1)’s direction count, the contraction diagnostics of Proposition[5\.12](https://arxiv.org/html/2608.10288#S5.Thmtheorem12)measured on the full composition, the budget of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)as a sample\-extrema proxy against measured decoding margins, per\-step cached\-operator deviations, DAG\-loss values, and the order parameter of Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1)in both normalizations, together with the online\-contract measurements of Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\(historical\-row movement under suffix changes, padding at the Gram boundary, and sequential\-versus\-block candidate scoring\), run on two released checkpoints \(Sections[D\.10](https://arxiv.org/html/2608.10288#A4.SS10)–[D\.11](https://arxiv.org/html/2608.10288#A4.SS11)\)\. The design rule of the audit is that every reported quantity is either the named quantity computed directly, or is explicitly labeled a bound or proxy with its formula shown, and where a measurement contradicts a convenient assumption, the contradiction is reported\. Semantically sensitive constructions are guarded in code rather than by convention: the power stage is built by a shape\-checked helper with a sentinel test that fails under any head\- or row\-broadcast indexing of the exponent parameter, and the assembled perturbation chain is checked against a miniature\-decoder finite\-difference oracle that fails under any assembly omitting the injection layer’s own post\-attention factors\.
### D\.1\.Setup and provenance
Model:fromthesky/PLDR\-LLM\-v51\-SOC\-110M\-5\(the4141B\-token near\-critical model of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\];L=5L=5layers,h=14h=14heads,dk=64d\_\{k\}=64, float32\), loaded through its released Hugging Face port, unmodified,*at the pinned model revision of the verification manifest*\(Appendix[B](https://arxiv.org/html/2608.10288#A2)\), undertransformers4\.56\.14\.56\.1\(the version the port targets\) and run with eager attention on a single consumer GPU\.131313The audit code, its Python dependency specification, and the results are published with the Lean formalization \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\)\. The results file records the model revision, the SHA\-256 of the fetchedmodeling\_pldrllm\.py, library versions, and the random seed; the run pins the cuBLAS workspace and forces deterministic CUDA algorithms, making repeated launches bitwise\-reproducible \(see Scope\); the raw per\-instance arrays behind every summary below ship in a compressed archive whose SHA\-256 the results file records, and the repository’s model\-free unit tests recompute the summaries from those raw arrays\. Undertransformers5\.x the released port needs two load\-compatibility fixes, documented there; the pinned environment avoids them\.Inputs: eight fixed English prompts of3232–4343tokens \(150150–250250characters\) spanning technical, narrative, review, and instructional registers\. Aggregation \(min/median/max\) is over all layers, heads, and prompts unless stated; spectra are computed in float64\.
### D\.2\.Singular values and numerical rank \(in place of determinants\), trained versus initialized
Over all5×14×8=5605\\times 14\\times 8=560\(layer, head, prompt\) instances: the generatorAAhas numerical rank11in*every*instance \(tolerance64εf32σ164\\,\\varepsilon\_\{\\mathrm\{f32\}\}\\sigma\_\{1\}\), withσ2/σ1≤1\.4×10−8\\sigma\_\{2\}/\\sigma\_\{1\}\\leq 1\.4\\times 10^\{\-8\}\(median8\.7×10−178\.7\\times 10^\{\-17\}\); its rows agree within each head to a row\-variance/entry\-variance ratio of≈4×10−13\\approx 4\\times 10^\{\-13\}, and across heads to relative RMS≤1\.6×10−8\\leq 1\.6\\times 10^\{\-8\}\(median0at float resolution\)\. This is a direct, spectrum\-based confirmation of the rank\-one, cross\-head\-identical singularity condition \(Proposition[3\.9](https://arxiv.org/html/2608.10288#S3.Thmtheorem9), Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\(iii\)\) on this checkpoint\. By contrast,ALMA\_\{\\mathrm\{LM\}\}is*not*numerically low\-rank:σ2/σ1\\sigma\_\{2\}/\\sigma\_\{1\}has median0\.270\.27, and the numerical rank has median62\.562\.5of6464\(range11–6464\), even though its float determinant is0: withσ1\\sigma\_\{1\}ranging over6\.4×10−86\.4\\times 10^\{\-8\}–6\.2×1036\.2\\times 10^\{3\}and many small singular values, the determinant underflows while the matrix is far from rank one\. The measurement thus confirms both halves of Remark[3\.10](https://arxiv.org/html/2608.10288#S3.Thmtheorem10):AA’s collapse is real and sharp, and the entrywise nonlinear imageALMA\_\{\\mathrm\{LM\}\}is generically of \(near\-\)full numerical rank, so float\-zero determinants ofALMA\_\{\\mathrm\{LM\}\}carry no rank information\.
The trained\-versus\-initialized control \(Proposition[3\.6](https://arxiv.org/html/2608.10288#S3.Thmtheorem6)\) runs the same battery on the same architecture at random initialization, under*two*initialization laws, each labeled by its provenance: the Hugging Face port’s own initializer, which drawsWW,PP,aaand dense weights Xavier\-uniform, and the native training law of the released code, which drawsWW,PP,aaXavier\-normal with dense weights Xavier\-uniform \(the law the audited checkpoint was actually trained from; the two laws differ only in the distribution of the metric\-tensor parameters\)\. Under*both*laws the results agree to the displayed precision:AAhas numerical rank6363on every instance withσ2/σ1\\sigma\_\{2\}/\\sigma\_\{1\}of median0\.690\.69,ALMA\_\{\\mathrm\{LM\}\}has numerical rank of median6464, and the cross\-head relative RMS ofAAis1\.11\.1–1\.31\.3\(heads disagree at order one\)\. The value6363itself is architecturally expected, not a signature: at initialization the row map’s final LayerNorm hasγ=𝟏\\gamma=\\mathbf\{1\},β=0\\beta=0, so every row ofAAis exactly centered,A𝟏=0A\\mathbf\{1\}=0, and numerical rank≤dk−1=63\\leq d\_\{k\}\-1=63is guaranteed; the generic value is then6363\. The informative contrast with the trained model is the collapse ofσ2/σ1\\sigma\_\{2\}/\\sigma\_\{1\}\(median0\.690\.69at initialization versus≤1\.4×10−8\\leq 1\.4\\times 10^\{\-8\}trained\) and of the cross\-head disagreement \(order one versus≤1\.6×10−8\\leq 1\.6\\times 10^\{\-8\}\)\. The rank\-one collapse and the cross\-head identity of the trained model are therefore training outcomes, not artifacts of the architecture or of either initialization law\.
### D\.3\.The LayerNorm scale error, measured directly
The row variancesviv\_\{i\}ofD~/S\\widetilde\{D\}/S\(all layers/heads, head rows pooled;35,84035\{,\}840rows\) have median2\.6×10−52\.6\\times 10^\{\-5\}and range0to0\.290\.29\. The discrepancy∥LN\(D~\)i,:−LN\(D~/S\)i,:∥2\\left\\lVert\\operatorname\{LN\}\(\\widetilde\{D\}\)\_\{i,:\}\-\\operatorname\{LN\}\(\\widetilde\{D\}/S\)\_\{i,:\}\\right\\rVert\_\{2\}is computed*directly*for every captured row, in float64, with the checkpoint’sγ\\gamma,β\\beta,εLN\\varepsilon\_\{\\operatorname\{LN\}\}, and each prompt’s actualSS; the relative error divides by∥LN\(D~/S\)i,:∥2\\left\\lVert\\operatorname\{LN\}\(\\widetilde\{D\}/S\)\_\{i,:\}\\right\\rVert\_\{2\}, with the convention0when the two outputs coincide\. Measured: median relative error4\.6×10−44\.6\\times 10^\{\-4\}, with a long tail \(9090th percentile4\.3×10−24\.3\\times 10^\{\-2\},9999th percentile0\.440\.44, maximum4\.34\.3onε\\varepsilon\-dominated rows\); median absolute error5\.4×10−105\.4\\times 10^\{\-10\}, maximum1\.1×10−51\.1\\times 10^\{\-5\}\. The upper proxyεLN/\(2vi\)\\varepsilon\_\{\\operatorname\{LN\}\}/\(2v\_\{i\}\)of Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(ii\) \(whose exact first\-order coefficient is\(1−S−2\)εLN/\(2vi\)\(1\-S^\{\-2\}\)\\varepsilon\_\{\\operatorname\{LN\}\}/\(2v\_\{i\}\)\) is reported alongside*as a proxy*: its median is1\.9×10−21\.9\\times 10^\{\-2\}, an overestimate of the measured relative error by a factor≈35\\approx 35at matched rows, and it diverges on the1414exactly\-constant rows, where the true discrepancy is exactly0\(both centered rows vanish\)\. Conclusions: the idealizationLN\(D~\)=LN\(D~/S\)\\operatorname\{LN\}\(\\widetilde\{D\}\)=\\operatorname\{LN\}\(\\widetilde\{D\}/S\)is approximately valid for the typical row \(∼0\.05%\\sim 0\.05\\%error\) but not uniformly, since16\.4%16\.4\\%of rows \(5865/358405865/35840\) exceed2%2\\%relative error and the worst rows deviate at order one; any quantitative use of the Stage\-2 concentration argument must carry the measured distribution, not the proxy; and the divergence of an upper bound is not evidence about the quantity it bounds\.
### D\.4\.Twirl energy against pre\-rotation data
Analytically \(dk=64d\_\{k\}=64, base10410^\{4\}\):CΘ=4\.4968×104C\_\{\\Theta\}=4\.4968\\times 10^\{4\}, soCΘ/S≈43\.9C\_\{\\Theta\}/S\\approx 43\.9atS=1024S=1024\(the uniform bound of Lemma[5\.9](https://arxiv.org/html/2608.10288#S5.Thmtheorem9)is vacuous there\)\. The per\-frequency multipliers\|sin\(Sω/2\)\|/\(S\|sin\(ω/2\)\|\)\\left\\lvert\\sin\(S\\omega/2\)\\right\\rvert/\(S\\left\\lvert\\sin\(\\omega/2\)\\right\\rvert\)have median0\.27/0\.054/0\.0130\.27/0\.054/0\.013atS=64/256/1024S=64/256/1024, with worst value≈1\\approx 1at everySS\. Empirically, the audit captures the*pre\-rotation*queries position by position \(head0of each layer, all prompts\) and forms both aggregates1S∑nqnqn⊤\\frac\{1\}\{S\}\\sum\_\{n\}q\_\{n\}q\_\{n\}^\{\\top\}and1S∑nRnqnqn⊤Rn⊤\\frac\{1\}\{S\}\\sum\_\{n\}R\_\{n\}q\_\{n\}q\_\{n\}^\{\\top\}R\_\{n\}^\{\\top\}, so what the actual finite\-SSrotation removed is identified rather than inferred from the post\-rotation aggregate alone \(the rotated aggregate reproduces the capturedD~/S\\widetilde\{D\}/Sto relative error≤1\.2×10−7\\leq 1\.2\\times 10^\{\-7\}, validating the decomposition pipeline\)\. Two aggregations are reported, each under its exact label, since ratios do not commute with pooling\.*Per \(prompt, layer\) instance*\(all8×5=408\\times 5=40instances\): in the rotary eigenbasis the commutant \(zero\-frequency\) component carries a median6\.9%6\.9\\%\(3\.93\.9–15\.7%15\.7\\%\) of the pre\-rotation energy and10\.9%10\.9\\%\(4\.24\.2–18\.0%18\.0\\%\) of the post\-rotation energy; the off\-commutant energy removed by the actual rotation at the prompts’ lengths \(S=32S=32–4343\) has median9\.5%9\.5\\%, range−1\.4%\-1\.4\\%to46%46\\%: on one instance the rotated off\-commutant energy slightly*exceeds*the unrotated, so the finite\-SSrotation does not even monotonically suppress instance by instance\.*Prompt\-pooled per layer*\(energy\-weighted over prompts; range over the five layers only\): commutant fractions median7\.2%7\.2\\%\(4\.54\.5–12\.6%12\.6\\%\) pre\-rotation and11\.1%11\.1\\%\(4\.54\.5–13\.8%13\.8\\%\) post\-rotation; pooled off\-commutant suppression median10\.6%10\.6\\%\(range0\.009%0\.009\\%–43%43\\%\)\. The stationarity idealization of Proposition[5\.10](https://arxiv.org/html/2608.10288#S5.Thmtheorem10)is also directly checked, per \(prompt, layer\), and is far from holding: first\-half and second\-half aggregates ofqnqn⊤q\_\{n\}q\_\{n\}^\{\\top\}differ by a median relative33%33\\%\(up to90%90\\%\)\. Conclusion: at the audited context lengths the actual twirl removes only a modest fraction \(∼10%\\sim 10\\%median under either aggregation\) of the off\-commutant, instance\-specific energy, so Stage 1 cannot by itself account for the observed invariance at these lengths; attribution to Stage 3 rests on the direct Stage\-3 measurements below, not on subtraction\.
### D\.5\.Commutant residual of the trained operators
Answering the commutant\-diagnostics item of Section[9](https://arxiv.org/html/2608.10288#S9)on the released checkpoint: for every \(prompt, layer, head\) instance, the relative off\-commutant energy‖GLM−ΠcommGLM‖F/‖GLM‖F\\left\\lVert G\_\{\\mathrm\{LM\}\}\-\\Pi\_\{\\mathrm\{comm\}\}G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}/\\left\\lVert G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}, whereΠcomm\\Pi\_\{\\mathrm\{comm\}\}is the orthogonal projection onto the RoPE commutant of Proposition[4\.1](https://arxiv.org/html/2608.10288#S4.Thmtheorem1)\(computed exactly in the rotation eigenbasis; under nonresonance the commutant is the equal\-frequency entries, of real dimensiondkd\_\{k\}\)\. Measured over all560560instances: minimum0\.420\.42, median0\.810\.81, maximum0\.990\.99; per\-layer medians0\.98,0\.99,0\.93,0\.71,0\.660\.98,\\ 0\.99,\\ 0\.93,\\ 0\.71,\\ 0\.66for layers11–55\. The trained operators thus place the*majority*of their Frobenius energy in thedk2−dkd\_\{k\}^\{2\}\-d\_\{k\}absolute\-position\-sensitive directions counted by Corollary[7\.1](https://arxiv.org/html/2608.10288#S7.Thmtheorem1): on this checkpoint the learnedGLMG\_\{\\mathrm\{LM\}\}is far from the commutant subfamily in every layer and head, most strongly in the early layers\. This is a checkpoint\-scoped occupancy measurement, not a claim that those directions are causally used by decoding; the ablation that would test causal use \(projectingGLMG\_\{\\mathrm\{LM\}\}onto the commutant and measuring the behavioral change\) is part of the empirical program of Section[9](https://arxiv.org/html/2608.10288#S9)\.
### D\.6\.Row\-map contraction, measured on the composition
The quantity that controls diameters is the Jacobian of the full compositionφ\\varphi, and the audit computes it directly: at128128rows per layer of realLN\(D~\)\\operatorname\{LN\}\(\\widetilde\{D\}\)data, drawn from*every*prompt \(1616per prompt\), the largest singular value ofJφJ\\varphihas per\-layer medians3\.2×10−11,1\.4×10−16,1\.0×10−19,4\.9×10−12,3\.5×10−73\.2\\times 10^\{\-11\},\\ 1\.4\\times 10^\{\-16\},\\ 1\.0\\times 10^\{\-19\},\\ 4\.9\\times 10^\{\-12\},\\ 3\.5\\times 10^\{\-7\}for layers11–55\(pooled range1\.0×10−191\.0\\times 10^\{\-19\}–3\.6×10−73\.6\\times 10^\{\-7\}, nearly constant across rows and prompts within each layer\)\. Empirical pairwise contraction ratios‖φ\(r\)−φ\(r′\)‖/‖r−r′‖\\left\\lVert\\varphi\(r\)\-\\varphi\(r^\{\\prime\}\)\\right\\rVert/\\left\\lVert r\-r^\{\\prime\}\\right\\rVert, sampled both within prompts and across prompts \(10001000draws attempted each way;765765within\-prompt and694694cross\-prompt pairs retained after skipping coincident prompt indices and denominators below10−1210^\{\-12\}\), are*exactly zero at float resolution for≈95%\\approx 95\\%of retained pairs*and at most4\.6×10−64\.6\\times 10^\{\-6\}otherwise: sampled visited rows are mapped to outputs indistinguishable at float32\. For continuity with the per\-unit view: individual units are*not*contractive \(per\-unit Jacobian norms along the same trajectories have median0\.470\.47and maximum20\.820\.8\), so the strong hypothesisLj≤κ<1L\_\{j\}\\leq\\kappa<1of Proposition[5\.12](https://arxiv.org/html/2608.10288#S5.Thmtheorem12)isfalseper unit on this checkpoint, while the measured composition is contractive by seven or more orders of magnitude at every sampled row\. These are sampled pointwise statistics on visited data, not a tube\-uniform Lipschitz certificate \(Proposition[5\.12](https://arxiv.org/html/2608.10288#S5.Thmtheorem12)’s hypothesis line\); within that scope, they directly support the locally\-constant reading of the trained row map that Proposition[3\.11](https://arxiv.org/html/2608.10288#S3.Thmtheorem11)\(iii\) makes diagnostic\.
### D\.7\.The invariance budget as a sample\-extrema proxy, against measured margins
What follows is a*sample\-extrema worst\-case proxy*for the chain of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13), not a certificate, for five reasons stated up front: \(1\) the activation, operator, and preactivation extrema are maxima/minima over the eight prompt\-only forward passes; the capture hooks are removed before the decoding experiment, so the extrema do not cover the grown\-context states whose margins the final coefficient is compared against; \(2\) the LayerNorm factors use minimum*endpoint*row variances, which do not lower\-bound variances along an interpolation tube; \(3\) the measured minimum entry ofALMA\_\{\\mathrm\{LM\}\}is not a lower bound along a perturbation path through aniSwiGLU\\operatorname\{iSwiGLU\}zero crossing, where only the architectural floor10−910^\{\-9\}applies \(the constant is therefore evaluated at both endpoints below\); \(4\) the factors’ extrema are attained at unrelated sample points, and no common perturbation set is defined; \(5\) only the derivative suprema are certified, by closed\-form envelopes on the sampled interval \(sup\|u\|≤U\|iSwiGLU′\|≤2Uς\(U\)\+U2/4\\sup\_\{\\left\\lvert u\\right\\rvert\\leq U\}\\left\\lvert\\operatorname\{iSwiGLU\}^\{\\prime\}\\right\\rvert\\leq 2U\\varsigma\(U\)\+U^\{2\}/4andsup\|u\|≤U\|σs′\|≤ς\(U\)\+U/4\\sup\_\{\\left\\lvert u\\right\\rvert\\leq U\}\\left\\lvert\\sigma\_\{s\}^\{\\prime\}\\right\\rvert\\leq\\varsigma\(U\)\+U/4, withς\\varsigmathe logistic sigmoid; these dominate the true suprema, which a sampled grid maximum, being a lower estimate, does not\)\. A genuine certificate would require interval or otherwise verified propagation along every compared decoding state\.
With that scope: per\-head constants of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)are‖W‖2\\left\\lVert W\\right\\rVert\_\{2\}median0\.0220\.022\(max2\.32\.3\);‖a‖2\\left\\lVert a\\right\\rVert\_\{2\}median0\.920\.92; sampled preactivation boundUUmedian0\.0180\.018\(max78\.778\.7\), giving the certified envelopesup\|u\|≤U\|iSwiGLU′\|\\sup\_\{\\left\\lvert u\\right\\rvert\\leq U\}\\left\\lvert\\operatorname\{iSwiGLU\}^\{\\prime\}\\right\\rvertmedian0\.0180\.018\(max1\.7×1031\.7\\times 10^\{3\}\); the power\-stage factor on the*sampled*entry range ofALMA\_\{\\mathrm\{LM\}\}has median1\.0×10101\.0\\times 10^\{10\}\(max5\.8×10115\.8\\times 10^\{11\}\), because sampled entries ofALMA\_\{\\mathrm\{LM\}\}reach theϵ=10−9\\epsilon=10^\{\-9\}floor and exponentsPij<1P\_\{ij\}<1makemAPij−1m\_\{A\}^\{P\_\{ij\}\-1\}enormous\. The resultingCGsampC\_\{G\}^\{\\mathrm\{samp\}\}\(measuredmAm\_\{A\}; valid for pairs of visited states, Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)\) has median7\.5×1067\.5\\times 10^\{6\}and max1\.2×10151\.2\\times 10^\{15\}per head; the path\-validCGtubeC\_\{G\}^\{\\mathrm\{tube\}\}\(mA=10−9m\_\{A\}=10^\{\-9\}, the architectural floor\) has median9\.2×1079\.2\\times 10^\{7\}while the max is essentially unchanged \(1\.3×10151\.3\\times 10^\{15\}\), because the worst heads’ sampled minima already sit at the floor\. TheseCGC\_\{G\}values are Frobenius\-to\-Frobenius constants; converting a uniform per\-row boundεA\\varepsilon\_\{A\}onΔA\\Delta Ainto the spectral hypothesis of Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)costs the additional factordk=8\\sqrt\{d\_\{k\}\}=8of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)\(εG=CGdkεA\\varepsilon\_\{G\}=C\_\{G\}\\sqrt\{d\_\{k\}\}\\,\\varepsilon\_\{A\}\); the assembled chain below does not use this conversion, entering instead at the directly hypothesized spectral radiusεspec\\varepsilon\_\{\\mathrm\{spec\}\}\. The head\-level factor of Proposition[5\.7](https://arxiv.org/html/2608.10288#S5.Thmtheorem7)has median8\.58\.5\(max7474\)\. Per decoder layer: the two LayerNorm factors use the exact Jacobian bound‖γ‖∞/vmin\+ε\\left\\lVert\\gamma\\right\\rVert\_\{\\infty\}/\\sqrt\{v\_\{\\min\}\+\\varepsilon\}of Lemma[5\.11](https://arxiv.org/html/2608.10288#S5.Thmtheorem11)\(iv\) with the sampled minimum row variance of each LayerNorm’s inputs \(values8\.98\.9–23\.023\.0for the post\-attention LayerNorm and0\.510\.51–15\.215\.2for the post\-FFN one\); the gated\-FFN factor uses the product rule‖W3‖\(max\|x2\|⋅max\|σs′\|⋅‖W1‖\+max\|σs\(x1\)\|⋅‖W2‖\)\\left\\lVert W\_\{3\}\\right\\rVert\(\\max\\left\\lvert x\_\{2\}\\right\\rvert\\cdot\\max\\left\\lvert\\sigma\_\{s\}^\{\\prime\}\\right\\rvert\\cdot\\left\\lVert W\_\{1\}\\right\\rVert\+\\max\\left\\lvert\\sigma\_\{s\}\(x\_\{1\}\)\\right\\rvert\\cdot\\left\\lVert W\_\{2\}\\right\\rVert\)with activation extrema sampled on the prompt passes and the certifiedσs′\\sigma\_\{s\}^\{\\prime\}envelope \(values727727–39663966\); the attention factor is the coarse assembly from sampled operator norms \(values4\.7×1044\.7\\times 10^\{4\}–1\.1×1051\.1\\times 10^\{5\}\)\. The assembly respects the injection point: a perturbationΔGLM\\Delta G\_\{\\mathrm\{LM\}\}enters at the attention output of its layer, so its term carries that layer’s*own*post\-attention remainder \(post\-attention LayerNorm, FFN residual, post\-FFN LayerNorm,LLN1\(1\+LFFN\)LLN2L\_\{\\mathrm\{LN1\}\}\(1\+L\_\{\\mathrm\{FFN\}\}\)L\_\{\\mathrm\{LN2\}\}, with values6\.6×1036\.6\\times 10^\{3\},6\.9×1046\.9\\times 10^\{4\},1\.8×1041\.8\\times 10^\{4\},1\.8×1041\.8\\times 10^\{4\},1\.1×1051\.1\\times 10^\{5\}for layers11–55\) before the product of the downstream layers’ full factors; no\(1\+Lattn\)\(1\+L\_\{\\mathrm\{attn\}\}\)enters at the injection layer itself, since the perturbation arrives at the attention output, not the layer input\. The full per\-layer factors are3\.1×1083\.1\\times 10^\{8\}–1\.2×10101\.2\\times 10^\{10\},‖Wvocab‖2=175\\left\\lVert W\_\{\\mathrm\{vocab\}\}\\right\\rVert\_\{2\}=175, and the assembled end\-to\-end coefficient multiplying‖ΔGLM‖2\\left\\lVert\\Delta G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{2\}in the logit bound is≈7\.3×1046\\approx 7\.3\\times 10^\{46\}\. The audit’s test suite checks this assembly semantically, against finite differences of a miniature decoder: the coefficient must dominate the measured logit sensitivity of an actual two\-layer post\-attention path, and any assembly that omits the same\-layer remainder fails that oracle on a small\-variance input\. Two perturbation scales must not be conflated: the*measured*cached\-versus\-recomputed operator deviation is exactly0\(bitwise, next paragraph\), and the proxy is therefore stress\-tested instead at the*hypothetical single\-rounding radius*εspec=2−24max‖GLM‖F=3\.4×10−5\\varepsilon\_\{\\mathrm\{spec\}\}=2^\{\-24\}\\max\\left\\lVert G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}=3\.4\\times 10^\{\-5\}\(measuredmax‖GLM‖F=563\\max\\left\\lVert G\_\{\\mathrm\{LM\}\}\\right\\rVert\_\{F\}=563\): what one final elementwise float32 rounding of an exact\-real operator could contribute under a relative\-error model\. It is not a measurement, and it excludes accumulated arithmetic roundoff\.
Measured fidelity, greedy decoding of4848tokens on44prompts, cached \(frozen promptAA/ALMA\_\{\\mathrm\{LM\}\}/GLMG\_\{\\mathrm\{LM\}\}plus KV\-cache\) versus full uncached recomputation at every step: the recomputedGLMG\_\{\\mathrm\{LM\}\}on the grown context is compared with the prompt\-frozen value at*all4848steps of all44prompts on every layer*\(960960per\-layer comparisons\) and is*bitwise equal at every one*\(εG=0\\varepsilon\_\{G\}=0in both relative\-RMS and entrywise\-maximum senses; the per\-step records ship in the raw archive\); the maximum logit deviation is3\.6×10−53\.6\\times 10^\{\-5\}\(medians≈10−5\\approx 10^\{\-5\}\), attributable to floating\-point path differences of the two execution orders rather than to the operator; the minimum realized top\-two margin is6\.0×10−36\.0\\times 10^\{\-3\}\(per\-prompt medians1\.21\.2–3\.53\.5\); the greedy token choice agrees at every step of every prompt; and the corrected sufficient criterion of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)holds with room to spare,2B=7\.2×10−5<6\.0×10−3=Δ2B=7\.2\\times 10^\{\-5\}<6\.0\\times 10^\{\-3\}=\\Delta\. The decoded continuations are stored in the results file as a neutral, human\-inspectable record of the audited decoding, with repetition statistics: greedy decoding of this110110M\-parameter checkpoint produces repetition loops \(duplicate44\-gram fractions0\.490\.49–0\.800\.80across the four prompts\), while the sampled continuations of the order\-parameter section do not \(0\.00\.0\)\. No claim about generation quality is based on these texts; the checkpoint’s language capabilities are documented by the benchmark evaluations of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\], and the tensor measurements above are indifferent to text quality\.
Verdict on the proxy\.At the hypothetical single\-rounding radius the assembled proxy evaluates to7\.3×1046×3\.4×10−5≈2\.4×10427\.3\\times 10^\{46\}\\times 3\.4\\times 10^\{\-5\}\\approx 2\.4\\times 10^\{42\}, which exceeds the halved minimum marginΔ/2=3\.0×10−3\\Delta/2=3\.0\\times 10^\{\-3\}of the corrected criterion by≈45\\approx 45orders of magnitude\. Any certificate obtained by replacing each sampled factor in this*same*factorwise norm\-product assembly by a supremum over a containing tube is no smaller, so this assembly cannot certify the margin; a different direct or margin\-aware bound on the composite map \(exploiting alignment of singular directions, cancellation along residual paths, or decision\-relevant directions only\) could in principle be smaller, and none is constructed here\. This closes the factorwise norm\-product route: the bound\-form chain of Corollary[5\.13](https://arxiv.org/html/2608.10288#S5.Thmtheorem13)does*not*certify bit\-identical decoding on this checkpoint by this route, exactly as the corollary itself anticipates; what the measurements support is the empirical statement: on the tested workloads the cached model is indistinguishable from the uncached model at the10−510^\{\-5\}logit level, two orders below the smallest realized decision margin, with the cached operator itself exactly invariant at float resolution at every compared step\. A certificate would require margin\-aware, non\-worst\-case propagation \(e\.g\. interval or randomized smoothing analysis\) along every compared decoding state, which we leave as future work\.
### D\.8\.DAG\-loss values
The implemented per\-instance DAG loss\|log\(treM⊙M/dk\)\|\\left\\lvert\\log\(\\operatorname\{tr\}e^\{M\\odot M\}/d\_\{k\}\)\\right\\rvert, computed overflow\-safely over the560560instances of each tensor: forALMA\_\{\\mathrm\{LM\}\}, minimum0\(floating\-point underflow readings, exactly as Remark[4\.5](https://arxiv.org/html/2608.10288#S4.Thmtheorem5)predicts: the analytic floor of this normalized logarithmic loss islog\(1\+ϵ2\)≈10−18\\log\(1\+\\epsilon^\{2\}\)\\approx 10^\{\-18\}, and the floordkϵ2≈6\.4×10−17d\_\{k\}\\epsilon^\{2\}\\approx 6\.4\\times 10^\{\-17\}of the un\-normalized NOTEARS quantityhhis likewise sub\-resolution; both lie below float64 resolution of a naive log evaluation\), median3\.6×10−103\.6\\times 10^\{\-10\}, maximum6\.1×10−56\.1\\times 10^\{\-5\}; forAPA\_\{P\}, minimum6363, median1\.1×1031\.1\\times 10^\{3\}, maximum4\.1×1044\.1\\times 10^\{4\}; forGLMG\_\{\\mathrm\{LM\}\}\(no floor asserted\), minimum4\.94\.9, median1\.3×1031\.3\\times 10^\{3\}, maximum2\.8×1042\.8\\times 10^\{4\}\. On this checkpoint all three regularized tensors thus carry strictly positive measured cycle content except for the underflow readings ofALMA\_\{\\mathrm\{LM\}\}\.141414The exponent parameterPPis a per\-head tensor of shape\[h,dk,dk\]\[h,d\_\{k\},d\_\{k\}\]with*no*batch axis, unlike the batched activation tensors returned alongside it; this is an indexing hazard NumPy does not flag, since broadcast accepts several wrong pairings silently\. The audit buildsAPA\_\{P\}through a shape\-checked helper with a sentinel semantic test that fails under any head\- or row\-broadcast indexing ofPP, and the chain constants read the parameter directly under the same guard\.
### D\.9\.Order parameter, two normalizations
Two independent stochastic continuations \(6464tokens, temperature11\) of55prompts, whose decoded texts are likewise stored in the results file \(as neutral records with repetition statistics\)\. Two normalizations are computed and reported separately for every tensor: the RMS normalization of Definition[5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1), and a*symmetrized signed\-mean*normalization with denominator12\(\|μ1\|\+\|μ2\|\)\\tfrac\{1\}\{2\}\(\\left\\lvert\\mu\_\{1\}\\right\\rvert\+\\left\\lvert\\mu\_\{2\}\\right\\rvert\), a symmetric adaptation of the source papers’ convention rather than the convention itself \(there, the denominator is the first run’s\|μ1\|\\left\\lvert\\mu\_\{1\}\\right\\rvertalone, resp\.\|μC\|\\left\\lvert\\mu\_\{C\}\\right\\rvertagainst a cached run\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]\)\. Measured: forGLMG\_\{\\mathrm\{LM\}\}the order parameter is exactly0at float resolution on every layer under both normalizations; forAPA\_\{P\}, computed from the per\-head exponent tensor \(see the DAG footnote above\), the maximum is8\.6×10−138\.6\\times 10^\{\-13\}RMS\-normalized \(3\.3×10−123\.3\\times 10^\{\-12\}signed\), i\.e\. zero to twelve digits but not bitwise; forALMA\_\{\\mathrm\{LM\}\}, maximum3\.4×10−93\.4\\times 10^\{\-9\}RMS\-normalized and5\.6×10−95\.6\\times 10^\{\-9\}signed\-normalized \(medians0; the two normalizations’ maxima differ and are quoted separately\); for the most sensitive tensorAA, the RMS\-normalized value has median0and maximum4\.9×10−94\.9\\times 10^\{\-9\}\(signed\-mean maximum3\.1×10−83\.1\\times 10^\{\-8\}\)\. On this checkpoint the two normalizations agree in order of magnitude wherever they are nonzero\. Of the piecewise definition’s branches, the zero\-numerator branch*is*exercised \(every bitwise\-equal pair, in particular everyGLMG\_\{\\mathrm\{LM\}\}comparison, returns0through it\), while the degenerate positive\-numerator/zero\-signed\-denominator case, the case the stabilized definition exists to guard, is not observed; the stabilized definition matters for models whose deductive outputs are centered near zero, and is retained for that reason\. These values reproduce, with an independent implementation and the stabilized statistic, them≈0m\\approx 0\(float\) report of\[[14](https://arxiv.org/html/2608.10288#bib.bib5)\]for this model\.
### D\.10\.The online contract, historical\-row movement, and padding
The properties separated in Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)are measured by a dedicated online\-contract script published alongside the main audit \(same repository; raw arrays and result\-file SHA\-256 recorded as for the main audit\)\. The script runs in float32 eager mode on the same single consumer GPU and under the same determinism configuration as the main audit \(pinned cuBLAS workspace, forced deterministic algorithms; a CPU fallback is provided\); two complete back\-to\-back runs reproduce both the results file and the raw archive byte\-for\-byte\. It loads*two*released checkpoints at pinned revisions: the auditedPLDR\-LLM\-v51\-SOC\-110M\-5model above and, as a contrast within the same architecture and training family,PLDR\-LLM\-v51\-SOC\-110M\-1\(the manifest of Appendix[B](https://arxiv.org/html/2608.10288#A2)pins both\)\.
*Historical\-row movement\.*Same\-length inputs sharing a prefix and differing only in the suffix are compared at the shared\-prefix rows \(a same\-length design, so shape\-dependent floating\-point execution cannot masquerade as context sensitivity\)\. A fixed three\-token\-prefix pair is complemented by a seeded randomized protocol \(1616pairs, length1212, prefix length44; suffixes also compared with every suffix position attention\-masked\)\. On the audited checkpoint every comparison is*bitwise equal*: zero changed logit entries in all2,048,0002\{,\}048\{,\}000randomized comparisons and at every fixed\-pair position, withGLMG\_\{\\mathrm\{LM\}\}bitwise equal at every layer \(the only movement anywhere is in the layer\-44precursors:270270entries ofAAat magnitude≤1\.17×10−10\\leq 1\.17\\times 10^\{\-10\}and one entry ofALMA\_\{\\mathrm\{LM\}\}at2\.7×10−122\.7\\times 10^\{\-12\}\)\. On the contrast checkpoint the movement is systematic:74\.5%74\.5\\%of the randomized\-comparison entries change \(maximum9\.4×10−49\.4\\times 10^\{\-4\}; masked variant74\.3%74\.3\\%, maximum7\.4×10−47\.4\\times 10^\{\-4\}\), the fixed pair moves31,97831\{,\}978and31,57331\{,\}573of32,00032\{,\}000vocabulary entries at its two multi\-key prefix positions, and the per\-layerGLMG\_\{\\mathrm\{LM\}\}deviations reach3\.0×10−23\.0\\times 10^\{\-2\}: the historical rows move exactly through the deductive\-tensor path of \([4\.15](https://arxiv.org/html/2608.10288#S4.E15)\)\. Prefix position0is structurally insensitive on both checkpoints \(a single allowed key makes the softmax row\(1\)\(1\)regardless of the score\)\. The pair of measurements is the checkpoint\-level content of Remark[4\.9](https://arxiv.org/html/2608.10288#S4.Thmtheorem9)\(iii\): historical\-row prefix consistency fails at generic weights of this architecture and holds bitwise on the tested pairs for the collapsed checkpoint \(a finite measurement, not the universal decoder property\), tying its emergence to the collapse phenomenon\.
*Online determinism\.*Repeated identical prefix calls are bitwise equal on both checkpoints, so the step\-ttconditional of \([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\) is a deterministic function of the prefix in this configuration\.
*Padding at the Gram boundary\.*Appending four attention\-masked padding tokens to a prompt and comparing the final*real*\-token row against the unpadded call: on the audited checkpoint the deviation is≤8\.6×10−6\\leq 8\.6\\times 10^\{\-6\}\(exactly zero on one of the two prompts\) withGLMG\_\{\\mathrm\{LM\}\}*bitwise unchanged*and the deviation independent of the padding content, i\.e\. pure shape\-level float sensitivity of the longer call, not operator contamination; on the contrast checkpointGLMG\_\{\\mathrm\{LM\}\}itself moves \(maxima4\.5×10−34\.5\\times 10^\{\-3\}–1\.2×10−21\.2\\times 10^\{\-2\}, roughly50,00050\{,\}000changed entries\) and the final\-row deviation \(3\.53\.5–7\.4×10−57\.4\\times 10^\{\-5\}\)*depends on the padding content*, demonstrating that masked rows enter the Gram\. Stripping the padding restores the unpadded call bitwise on both checkpoints\. This is the measured basis for the unpadded\-S=tS\{=\}tcontract of Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)\.
### D\.11\.Sequential versus one\-pass block scores
The same script compares the two candidate\-scoring semantics of Section[4\.5](https://arxiv.org/html/2608.10288#S4.SS5)on a fixed suite: for each of the eight prompts, four multi\-token candidates \(the model’s own greedy continuations of lengths44and88, a second\-best\-first\-token continuation of length44, and a seeded random candidate of length44\) and four single\-token control candidates\. For each candidate the*sequential*score∑ilogpθ\(yi∣x,y1:i−1\)\\sum\_\{i\}\\log p\_\{\\theta\}\(y\_\{i\}\\mid x,y\_\{1:i\-1\}\)is computed by one final\-row prefix call per token, and the*block*score by one call on\(x,y\)\(x,y\)minus its last token with candidate log\-probabilities read at the aligned rows \(log\-softmax in float64 on the float32 logits; the exact index algebra is unit\-tested against a prefix\-consistent toy decoder, for which the two protocols must coincide identically, and a global toy decoder, for which they must not\)\.
Single\-token candidates coincide*exactly*on both checkpoints, as they must \(the two protocols invoke the model on the identical tensor\)\. For multi\-token candidates, on the audited checkpoint the per\-token gap is at most3\.3×10−63\.3\\times 10^\{\-6\}with median exactly0\(most per\-token gaps vanish\), and the per\-candidate total\-score gap is at most3\.3×10−63\.3\\times 10^\{\-6\}; on the contrast checkpoint the gaps are roughly fifty\-fold larger \(per\-token maximum1\.5×10−41\.5\\times 10^\{\-4\}, per\-candidate maximum2\.2×10−42\.2\\times 10^\{\-4\}\), a genuinely semantic difference, consistent with the historical\-row movement above\. On*neither*checkpoint does any of the3232candidate rankings change: zero argmax flips and zero discordant pairs out of4848ordered comparisons\. On the audited checkpoint the residual gap is at the shape\-level float\-sensitivity scale \(its historical rows are bitwise invariant in same\-length comparisons\), so one\-pass block scores and sequential scores are interchangeable there at far below score\-decision scales; published benchmark scores remain labeled as block scores \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\), and the contrast checkpoint quantifies the regime where the two semantics genuinely differ\. This constructed suite is a controlled*diagnostic*, not a validation of benchmark evaluations; the corresponding measurement on held\-out text and on real benchmark items is Section[D\.12](https://arxiv.org/html/2608.10288#A4.SS12)\.
### D\.12\.Held\-out sequential NLL and real benchmark items under both scoring protocols
The same comparison is run on held\-out text and on*real*benchmark items \(scriptaudit\_seq\.py; both released checkpoints; the same GPU float32 eager configuration, cuBLAS workspace and deterministic\-algorithm pins, and two byte\-identical back\-to\-back official runs as the preceding subsections\)\. Every dataset enters as a parquet file pinned by full revision hash \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\), with per\-file SHA\-256 recorded in the results JSON; the benchmark request templates and the published TruthfulQA metric are transcriptions of the pinned evaluation wrapper’s task configurations, unit tested against hand\-built documents and hand\-computed metric values, and the wrapper’s encode\-pair convention \(trailing\-whitespace shift, joint\-encoding split\) is reproduced: on all7,1727\{,\}172scored requests the independently encoded context was verifiably a token prefix of the joint encoding\. Every summary below is re\-derived offline from the raw arrays by the model\-free test suite\.
*Held\-out blockwise CE versus sequential NLL\.*On4848seeded disjoint257257\-token windows of the pinned WikiText\-2 \(raw\) validation split\[[25](https://arxiv.org/html/2608.10288#bib.bib46)\]\(12,28812\{,\}288scored positions per checkpoint\), each position is scored twice: by its row of one full\-block call \(the training objective’s per\-row quantity, \([4\.10](https://arxiv.org/html/2608.10288#S4.E10)\)\) and by one final\-rowS=tS=tcall \(\([4\.8](https://arxiv.org/html/2608.10288#S4.E8)\), the deployed chain\-rule quantity\)\. On the audited checkpoint the two aggregates coincide to nine decimal digits: blockwise CE3\.4794463\.479446nats/token against sequential NLL3\.4794463\.479446\(difference1\.8×10−91\.8\\times 10^\{\-9\}\), with per\-token\|gap\|\\left\\lvert\\text\{gap\}\\right\\rvertmedian9\.6×10−79\.6\\times 10^\{\-7\}, maximum2\.4×10−52\.4\\times 10^\{\-5\}, and signed median exactly0\. On the contrast checkpoint the aggregates still agree to4\.3×10−64\.3\\times 10^\{\-6\}nats/token \(3\.6848833\.684883versus3\.6848793\.684879\), but the per\-token gaps are two orders larger \(median1\.8×10−51\.8\\times 10^\{\-5\}, maximum2\.8×10−32\.8\\times 10^\{\-3\}\) and nearly sign\-balanced \(49\.3%49\.3\\%of positions score worse sequentially\); their magnitude*decreases*with position within the window \(per\-position\-bucket medians2\.52\.5–3\.8×10−53\.8\\times 10^\{\-5\}early,5\.0×10−65\.0\\times 10^\{\-6\}in the final bucket\), which is the suffix\-length dependence the first\-order mechanism \([4\.15](https://arxiv.org/html/2608.10288#S4.E15)\) predicts: a later historical row leaves less suffix to move the Gram\. On this sample, then, the blockwise objective’s value*is*the autoregressive NLL to float resolution on the collapsed checkpoint, and a close, sign\-balanced surrogate for it on the noncollapsed one; neither fact extends beyond the tested sample, and neither removes the target\-exposure structure of §[4\.3](https://arxiv.org/html/2608.10288#S4.SS3), whose training\-time role is a separate, unmeasured question \(Section[9](https://arxiv.org/html/2608.10288#S9)\)\.
*Real benchmark items\.*Seeded samples of100100actual items from each of the eight published zero\-shot tasks \(ARC\-Easy and ARC\-Challenge test splits, HellaSwag, PIQA, Social\-IQa, WinoGrande, and TruthfulQA validation splits, OpenBookQA test split;800800items,3,0523\{,\}052candidates,32,24632\{,\}246scored candidate tokens per checkpoint\) are scored under the one\-pass block protocol and under the sequential chain\-rule protocol\. TruthfulQA enters under its publishedtruthfulqa\_mc2protocol\[[23](https://arxiv.org/html/2608.10288#bib.bib47)\]: each question carries several true and several false reference answers, and the published score is not an argmax but the normalized probability mass on the true answers,∑i∈trueeℓi/∑jeℓj\\sum\_\{i\\in\\mathrm\{true\}\}e^\{\\ell\_\{i\}\}\\big/\\sum\_\{j\}e^\{\\ell\_\{j\}\}over the per\-candidate total log\-likelihoodsℓ\\ell, transcribed from the pinned harness task configuration, computed by an overflow\-safe softmax equivalent, and evaluated once from block and once from sequential candidate scores\. On the seven argmax tasks the outcome\-level result is uniform:*zero*raw argmax changes,*zero*length\-normalized argmax changes, and*zero*discordant candidate pairs \(of2,9012\{,\}901\) on*both*checkpoints, so every per\-task accuracy, raw and normalized, is identical under the two protocols on this sample; single\-token candidates coincide exactly, as they must\. On TruthfulQA the published metric agrees between the two protocols to8\.5×10−108\.5\\times 10^\{\-10\}in the100100\-item mean on the audited checkpoint \(0\.3976620\.397662under both; per\-item metric gap median0, maximum3\.7×10−73\.7\\times 10^\{\-7\}\) and to5\.7×10−75\.7\\times 10^\{\-7\}on the contrast checkpoint \(block0\.3997480\.399748, sequential0\.3997470\.399747; per\-item maximum4\.7×10−54\.7\\times 10^\{\-5\}\); a candidate\-level argmax \(a diagnostic, not the published metric\) does not change on either checkpoint\. The score\-level gaps mirror the constructed suite: per\-candidate absolute gaps reach2\.8×10−52\.8\\times 10^\{\-5\}on the audited checkpoint \(median exactly0on five of the eight tasks\) and2\.3×10−32\.3\\times 10^\{\-3\}on the contrast checkpoint, growing monotonically with candidate token length there \(bucket maxima3\.1×10−43\.1\\times 10^\{\-4\}at22–44tokens up to2\.3×10−32\.3\\times 10^\{\-3\}beyond2020tokens; single\-token bucket exactly0\)\. The decision margins explain the argmax stability: per\-task median block top\-two margins are0\.510\.51–12\.912\.9nats, and on every item except one the item’s own maximal candidate gap stays below half of that item’s realized margin\. The single flagged item \(the same on both checkpoints\) is a Social\-IQa instance whose top two*candidate strings are identical in the source data*, an exact tie of margin0that both protocols score bitwise identically: a dataset artifact, not a protocol discrepancy\. An auxiliary single\-gold TruthfulQA probe \(the mc1 variant of the same dataset, on the same seeded rows;534534candidates per checkpoint\) supplies the argmax structure the published metric lacks and is reported separately in the audit artifact, outside the published\-task count: zero argmax changes and zero discordant pairs \(of1,3661\{,\}366\) on both checkpoints\.
These are sampled measurements:100100items per task under one seed, two checkpoints, one software stack\. They support the statement that on these samples the published one\-pass block protocol and sequential chain\-rule scoring select the same answers on the argmax tasks and agree on the published TruthfulQA probability\-mass metric to within5×10−55\\times 10^\{\-5\}per item \(even on the noncollapsed contrast checkpoint, whose score*values*differ measurably between protocols\), and they do not revalidate the complete published benchmark tables, other checkpoints, or the training trajectory \(Section[9](https://arxiv.org/html/2608.10288#S9), empirical program\)\.
### D\.13\.Scope
This audit is one checkpoint, eight prompts, and one software stack \(the online\-contract and sequential\-validation sections additionally probe a second released checkpoint, seeded held\-out text windows, and seeded fixed\-size samples of real benchmark items, still far from the full validation sets, checkpoint fleets, and training trajectories of the source papers, which remain listed in Section[9](https://arxiv.org/html/2608.10288#S9)’s program\); its Jacobian and pairwise statistics are sampled at visited points; it establishes the measured facts above for this model and does not by itself support generalization across seeds, scales, or datasets\. Individual measured values carry floating\-point path sensitivity: they depend on the device and library versions, and quantities downstream of the elementwise power stage \(whose local amplification reaches∼1010\\sim 10^\{10\}\) can vary by more than low\-order digits across configurations\. This sensitivity extends to repeated launches in a*fixed*environment: CUDA matrix\-multiply algorithm selection can differ between process launches, and preparing this audit we observed the smallest order\-parameter statistics shift by several orders of magnitude \(within≲10−3\\lesssim 10^\{\-3\}relative\) across launches of the identical seeded script\. The shipped run therefore pins the cuBLAS workspace and forces deterministic algorithms, after which repeated launches reproduce every reported number bitwise \(verified by back\-to\-back complete reruns\); order\-parameter readings below that stability scale should in general be quoted as upper bounds\. The qualitative findings above are stable under all of this\. The audit code, its Python dependency specification, the results files, and the raw per\-instance arrays are published with the Lean formalization \(Appendix[B](https://arxiv.org/html/2608.10288#A2)\)\.
## References
- \[1\]J\. Aczél\(1966\)Lectures on functional equations and their applications\.Academic Press,New York\.Cited by:[§6\.1](https://arxiv.org/html/2608.10288#S6.SS1.1.p1.9)\.
- \[2\]J\. L\. Ba, J\. R\. Kiros, and G\. E\. Hinton\(2016\)Layer normalization\.Note:arXiv:1607\.06450Cited by:[§2\.2](https://arxiv.org/html/2608.10288#S2.SS2.p1.28)\.
- \[3\]P\. Bak, C\. Tang, and K\. Wiesenfeld\(1988\)Self\-organized criticality\.Physical Review A38\(1\),pp\. 364–374\.Cited by:[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p1.1)\.
- \[4\]J\. M\. Beggs and D\. Plenz\(2003\)Neuronal avalanches in neocortical circuits\.Journal of Neuroscience23\(35\),pp\. 11167–11177\.Cited by:[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p2.4)\.
- \[5\]Y\. N\. Dauphin, A\. Fan, M\. Auli, and D\. Grangier\(2017\)Language modeling with gated convolutional networks\.InProceedings of the 34th International Conference on Machine Learning \(ICML\),pp\. 933–941\.Cited by:[Definition 3\.1](https://arxiv.org/html/2608.10288#S3.Thmtheorem1.p1.2)\.
- \[6\]R\. Dickman, M\. A\. Muñoz, A\. Vespignani, and S\. Zapperi\(2000\)Paths to self\-organized criticality\.Brazilian Journal of Physics30\(1\),pp\. 27–41\.Cited by:[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p1.1)\.
- \[7\]Y\. Dong, J\. Cordonnier, and A\. Loukas\(2021\)Attention is not all you need: pure attention loses rank doubly exponentially with depth\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),pp\. 2793–2803\.Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[8\]A\. El\-Nouby, H\. Touvron, M\. Caron, P\. Bojanowski, M\. Douze, A\. Joulin, I\. Laptev, N\. Neverova, G\. Synnaeve, J\. Verbeek, and H\. Jégou\(2021\)XCiT: cross\-covariance image transformers\.InAdvances in Neural Information Processing Systems 34,Note:arXiv:2106\.09681Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[9\]B\. Gao and L\. Pavel\(2017\)On the properties of the softmax function with application in game theory and reinforcement learning\.Note:arXiv:1704\.00805Cited by:[Lemma 5\.6](https://arxiv.org/html/2608.10288#S5.Thmtheorem6)\.
- \[10\]B\. Gokden\(2019\)CoulGAT: an experiment on interpretability of graph attention networks\.Note:arXiv:1912\.08409Cited by:[5th item](https://arxiv.org/html/2608.10288#A2.I1.i5.p1.1),[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[Remark 3\.4](https://arxiv.org/html/2608.10288#S3.Thmtheorem4.p1.6),[Remark 3\.5](https://arxiv.org/html/2608.10288#S3.Thmtheorem5.p1.1),[Disclosure of the use of AI tools](https://arxiv.org/html/2608.10288#Sx2.p1.1)\.
- \[11\]B\. Gokden\(2021\)Power law graph transformer for machine translation and representation learning\.Note:arXiv:2107\.02039Cited by:[4th item](https://arxiv.org/html/2608.10288#A2.I1.i4.p1.1),[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p2.2),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p2.8),[§2\.2](https://arxiv.org/html/2608.10288#S2.SS2.p1.28),[§3\.4](https://arxiv.org/html/2608.10288#S3.SS4.p1.7),[Definition 3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2),[Definition 3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2.p1.22),[Remark 3\.3](https://arxiv.org/html/2608.10288#S3.Thmtheorem3.p1.2),[Remark 3\.5](https://arxiv.org/html/2608.10288#S3.Thmtheorem5.p1.1),[Disclosure of the use of AI tools](https://arxiv.org/html/2608.10288#Sx2.p1.1)\.
- \[12\]B\. Gokden\(2024\)PLDR\-LLM: large language model from power law decoder representations\.Note:arXiv:2410\.16703Cited by:[3rd item](https://arxiv.org/html/2608.10288#A2.I1.i3.p1.1),[Appendix B](https://arxiv.org/html/2608.10288#A2.SS0.SSS0.Px4.p1.6),[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p1.6),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p2.2),[Definition 3\.1](https://arxiv.org/html/2608.10288#S3.Thmtheorem1.p1.2),[Definition 3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2),[Remark 3\.3](https://arxiv.org/html/2608.10288#S3.Thmtheorem3.p1.2),[Remark 3\.5](https://arxiv.org/html/2608.10288#S3.Thmtheorem5.p1.1),[§4\.3](https://arxiv.org/html/2608.10288#S4.SS3.p2.9),[Definition 4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3),[Remark 4\.6](https://arxiv.org/html/2608.10288#S4.Thmtheorem6.p1.4),[§5\.4\.1](https://arxiv.org/html/2608.10288#S5.SS4.SSS1.p2.21),[Remark 6\.6](https://arxiv.org/html/2608.10288#S6.Thmtheorem6.p1.2),[§7](https://arxiv.org/html/2608.10288#S7.4.4.7.2.3.1.1),[1st item](https://arxiv.org/html/2608.10288#S7.I1.i1.p1.1),[item 3](https://arxiv.org/html/2608.10288#S7.I2.i3.p1.3),[§7\.4](https://arxiv.org/html/2608.10288#S7.SS4.p1.2),[Disclosure of the use of AI tools](https://arxiv.org/html/2608.10288#Sx2.p1.1)\.
- \[13\]B\. Gokden\(2025\)PLDR\-LLMs learn a generalizable tensor operator that can replace its own deep neural net at inference\.Note:arXiv:2502\.13502Cited by:[Appendix B](https://arxiv.org/html/2608.10288#A2.SS0.SSS0.Px4.p1.6),[item 1](https://arxiv.org/html/2608.10288#S1.I1.i1.p1.4),[item 2](https://arxiv.org/html/2608.10288#S1.I1.i2.p1.4),[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px4.p1.3.3.6.2.2.1.1),[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p1.6),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p2.8),[§3\.2](https://arxiv.org/html/2608.10288#S3.SS2.p1.5),[§3\.2](https://arxiv.org/html/2608.10288#S3.SS2.p3.4),[Remark 3\.10](https://arxiv.org/html/2608.10288#S3.Thmtheorem10.p1.23),[Remark 3\.3](https://arxiv.org/html/2608.10288#S3.Thmtheorem3.p1.2),[§4\.4](https://arxiv.org/html/2608.10288#S4.SS4.p1.2),[§4\.4](https://arxiv.org/html/2608.10288#S4.SS4.p1.3),[Definition 4\.3](https://arxiv.org/html/2608.10288#S4.Thmtheorem3),[item \(iii\)](https://arxiv.org/html/2608.10288#S5.I2.i3.p1.4),[item \(ii\)](https://arxiv.org/html/2608.10288#S5.I4.i2.p1.12),[§5\.3](https://arxiv.org/html/2608.10288#S5.SS3.p1.3),[§5\.4\.4](https://arxiv.org/html/2608.10288#S5.SS4.SSS4.p2.4),[Remark 5\.8](https://arxiv.org/html/2608.10288#S5.Thmtheorem8.p1.1),[§7](https://arxiv.org/html/2608.10288#S7.1.1.1.1.1.1),[§7](https://arxiv.org/html/2608.10288#S7.4.4.4.1.1.1),[§7](https://arxiv.org/html/2608.10288#S7.4.4.8.3.3.1.1),[3rd item](https://arxiv.org/html/2608.10288#S7.I1.i3.p1.1),[item 1](https://arxiv.org/html/2608.10288#S7.I2.i1.p1.2),[item 2](https://arxiv.org/html/2608.10288#S7.I2.i2.p1.1),[item 3](https://arxiv.org/html/2608.10288#S7.I2.i3.p1.3),[§7\.1](https://arxiv.org/html/2608.10288#S7.SS1.p1.3),[§7\.5](https://arxiv.org/html/2608.10288#S7.SS5.p1.6),[§8](https://arxiv.org/html/2608.10288#S8.p2.1),[Disclosure of the use of AI tools](https://arxiv.org/html/2608.10288#Sx2.p1.1),[footnote 4](https://arxiv.org/html/2608.10288#footnote4)\.
- \[14\]B\. Gokden\(2026\)PLDR\-LLMs reason at self\-organized criticality\.Note:arXiv:2603\.23539Cited by:[Appendix B](https://arxiv.org/html/2608.10288#A2.SS0.SSS0.Px4.p1.6),[§D\.1](https://arxiv.org/html/2608.10288#A4.SS1.p1.9),[§D\.7](https://arxiv.org/html/2608.10288#A4.SS7.p3.21),[§D\.9](https://arxiv.org/html/2608.10288#A4.SS9.p1.22),[item 3](https://arxiv.org/html/2608.10288#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px4.p1.3.3.6.2.2.1.1),[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[§2\.1](https://arxiv.org/html/2608.10288#S2.SS1.p1.6),[Definition 3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2),[Definition 3\.2](https://arxiv.org/html/2608.10288#S3.Thmtheorem2.p1.22),[Remark 3\.3](https://arxiv.org/html/2608.10288#S3.Thmtheorem3.p1.2),[Remark 3\.4](https://arxiv.org/html/2608.10288#S3.Thmtheorem4.p1.6),[item \(iii\)](https://arxiv.org/html/2608.10288#S5.I4.i3.p1.4),[item 3](https://arxiv.org/html/2608.10288#S5.I5.i3.p1.1),[§5\.1](https://arxiv.org/html/2608.10288#S5.SS1.p1.8),[§5\.4\.4](https://arxiv.org/html/2608.10288#S5.SS4.SSS4.p2.4),[Definition 5\.1](https://arxiv.org/html/2608.10288#S5.Thmtheorem1.p1.12),[2nd item](https://arxiv.org/html/2608.10288#S6.I2.i2.p1.1),[3rd item](https://arxiv.org/html/2608.10288#S6.I2.i3.p1.1),[§6\.2](https://arxiv.org/html/2608.10288#S6.SS2.p1.4),[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p1.1),[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p2.4),[§7](https://arxiv.org/html/2608.10288#S7.2.2.2.1.1.1),[§7](https://arxiv.org/html/2608.10288#S7.3.3.3.1.1.1),[item 2](https://arxiv.org/html/2608.10288#S7.I2.i2.p1.1),[item 3](https://arxiv.org/html/2608.10288#S7.I2.i3.p1.3),[§7\.2](https://arxiv.org/html/2608.10288#S7.SS2.p1.4),[§7\.3](https://arxiv.org/html/2608.10288#S7.SS3.p1.1),[§7\.4](https://arxiv.org/html/2608.10288#S7.SS4.p1.2),[Conjecture 8\.1](https://arxiv.org/html/2608.10288#S8.Thmtheorem1.p1.13.13),[Disclosure of the use of AI tools](https://arxiv.org/html/2608.10288#Sx2.p1.1),[footnote 4](https://arxiv.org/html/2608.10288#footnote4),[footnote 5](https://arxiv.org/html/2608.10288#footnote5)\.
- \[15\]N\. Goldenfeld\(1992\)Lectures on phase transitions and the renormalization group\.Addison\-Wesley\.Cited by:[Remark 6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3.p1.8),[Remark 6\.6](https://arxiv.org/html/2608.10288#S6.Thmtheorem6.p1.2)\.
- \[16\]D\. Ha, A\. Dai, and Q\. V\. Le\(2016\)HyperNetworks\.Note:arXiv:1609\.09106Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[17\]J\. Hesse and T\. Gross\(2014\)Self\-organized criticality as a fundamental property of neural systems\.Frontiers in Systems Neuroscience8,pp\. 166\.Cited by:[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p2.4)\.
- \[18\]R\. A\. Horn and C\. R\. Johnson\(2013\)Matrix analysis\.2nd edition,Cambridge University Press\.Cited by:[§3\.2](https://arxiv.org/html/2608.10288#S3.SS2.3.p1.1),[§3\.2](https://arxiv.org/html/2608.10288#S3.SS2.4.p1.21)\.
- \[19\]K\. Irie, I\. Schlag, R\. Csordás, and J\. Schmidhuber\(2021\)Going beyond linear transformers with recurrent fast weight programmers\.InAdvances in Neural Information Processing Systems 34,Note:arXiv:2106\.06295Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[20\]A\. Katharopoulos, A\. Vyas, N\. Pappas, and F\. Fleuret\(2020\)Transformers are RNNs: fast autoregressive transformers with linear attention\.InProceedings of the 37th International Conference on Machine Learning,Note:arXiv:2006\.16236Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[21\]J\. Kim, J\. Jun, and B\. Zhang\(2018\)Bilinear attention networks\.InAdvances in Neural Information Processing Systems 31,Note:arXiv:1805\.07932Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[22\]D\. Le, T\. Nguyen, C\. Nguyen, and A\. T\. Luu\(2026\)Don’t read everything: a curvature\-conditioned query for linear attention\.Note:arXiv:2606\.01294Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[23\]S\. Lin, J\. Hilton, and O\. Evans\(2021\)TruthfulQA: measuring how models mimic human falsehoods\.Note:arXiv:2109\.07958; published at ACL 2022Cited by:[§D\.12](https://arxiv.org/html/2608.10288#A4.SS12.p3.30)\.
- \[24\]D\. Marković and C\. Gros\(2014\)Power laws and self\-organized criticality in theory and nature\.Physics Reports536\(2\),pp\. 41–74\.Cited by:[§6\.3](https://arxiv.org/html/2608.10288#S6.SS3.p2.4),[§7\.6](https://arxiv.org/html/2608.10288#S7.SS6.p1.2)\.
- \[25\]S\. Merity, C\. Xiong, J\. Bradbury, and R\. Socher\(2016\)Pointer sentinel mixture models\.Note:arXiv:1609\.07843; the WikiText corporaCited by:[§D\.12](https://arxiv.org/html/2608.10288#A4.SS12.p2.20)\.
- \[26\]C\. W\. Misner, K\. S\. Thorne, and J\. A\. Wheeler\(1973\)Gravitation\.W\. H\. Freeman,San Francisco\.Note:Ch\. 21\.12Cited by:[Remark 3\.4](https://arxiv.org/html/2608.10288#S3.Thmtheorem4.p1.6),[Remark 5\.8](https://arxiv.org/html/2608.10288#S5.Thmtheorem8.p1.1)\.
- \[27\]M\. E\. J\. Newman\(2005\)Power laws, Pareto distributions and Zipf’s law\.Contemporary Physics46\(5\),pp\. 323–351\.Cited by:[Remark 6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3.p1.8)\.
- \[28\]S\. Ostmeier, B\. Axelrod, M\. Varma, M\. E\. Moseley, A\. Chaudhari, and C\. Langlotz\(2024\)LieRE: Lie rotational positional encodings\.Note:arXiv:2406\.10322Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[Remark 4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2.p1.4)\.
- \[29\]B\. Qin, J\. Li, S\. Tang, and Y\. Zhuang\(2022\)DBA: efficient transformer with dynamic bilinear low\-rank attention\.Note:arXiv:2211\.16368Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[30\]M\. E\. Sander, P\. Ablin, M\. Blondel, and G\. Peyré\(2021\)Sinkformers: transformers with doubly stochastic attention\.Note:arXiv:2110\.11773Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[31\]I\. Schlag, K\. Irie, and J\. Schmidhuber\(2021\)Linear transformers are secretly fast weight programmers\.InProceedings of the 38th International Conference on Machine Learning,Note:arXiv:2102\.11174Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[32\]S\. Schug, S\. Kobayashi, Y\. Akram, J\. Sacramento, and R\. Pascanu\(2025\)Attention as a hypernetwork\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2406\.05816Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[33\]E\. Seneta\(2006\)Non\-negative matrices and Markov chains\.2nd edition,Springer\.Cited by:[§3\.2](https://arxiv.org/html/2608.10288#S3.SS2.3.p1.1)\.
- \[34\]N\. Shazeer, Z\. Lan, Y\. Cheng, N\. Ding, and L\. Hou\(2020\)Talking\-heads attention\.Note:arXiv:2003\.02436Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[35\]N\. Shazeer\(2020\)GLU variants improve transformer\.Note:arXiv:2002\.05202Cited by:[Definition 3\.1](https://arxiv.org/html/2608.10288#S3.Thmtheorem1.p1.2)\.
- \[36\]D\. Sornette and G\. Ouillon\(2012\)Dragon\-kings: mechanisms, statistical methods and empirical evidence\.European Physical Journal Special Topics205,pp\. 1–26\.Cited by:[item 3](https://arxiv.org/html/2608.10288#S5.I5.i3.p1.1),[3rd item](https://arxiv.org/html/2608.10288#S6.I2.i3.p1.1)\.
- \[37\]H\. E\. Stanley\(1999\)Scaling, universality, and renormalization: three pillars of modern critical phenomena\.Reviews of Modern Physics71\(2\),pp\. S358–S366\.Cited by:[Remark 6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3.p1.8)\.
- \[38\]J\. Su, Y\. Lu, S\. Pan, B\. Wen, and Y\. Liu\(2021\)RoFormer: enhanced transformer with rotary position embedding\.Note:arXiv:2104\.09864Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[§4\.1](https://arxiv.org/html/2608.10288#S4.SS1.p1.1),[Remark 4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2.p1.4)\.
- \[39\]Y\. Tay, D\. Bahri, D\. Metzler, D\. Juan, Z\. Zhao, and C\. Zheng\(2020\)Synthesizer: rethinking self\-attention in transformer models\.Note:arXiv:2005\.00743Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[40\]V\. Tran, V\. K\. Bui, V\. Trinh, T\. L\. Ngoc, and T\. M\. Nguyen\(2026\)Functional equivalence in attention: a comprehensive study with applications to linear mode connectivity\.Note:arXiv:2606\.17830Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[Remark 4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2.p1.4),[§7\.1](https://arxiv.org/html/2608.10288#S7.SS1.p3.5)\.
- \[41\]J\. A\. Tropp\(2015\)An introduction to matrix concentration inequalities\.Foundations and Trends in Machine Learning8\(1–2\),pp\. 1–230\.Cited by:[§5\.4\.2](https://arxiv.org/html/2608.10288#S5.SS4.SSS2.1.p1.8)\.
- \[42\]R\. Vashisht and H\. G\. Ramaswamy\(2026\)Faster query\-key learning sharpens attention in self\-attention models\.InProceedings of the International Conference on Machine Learning \(ICML\),Note:arXiv:2608\.06776Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[Remark 5\.5](https://arxiv.org/html/2608.10288#S5.Thmtheorem5.p1.3)\.
- \[43\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems 30 \(NIPS\),pp\. 6000–6010\.Cited by:[§1](https://arxiv.org/html/2608.10288#S1.p1.4),[§5\.2](https://arxiv.org/html/2608.10288#S5.SS2.1.p1.7)\.
- \[44\]H\. Wang and K\. Wang\(2025\)Complete characterization of gauge symmetries in transformer architectures\.InNeurIPS Workshop on Symmetry and Geometry in Neural Representations \(NeurReps\),Note:[https://openreview\.net/forum?id=KrkbYbK0cH](https://openreview.net/forum?id=KrkbYbK0cH)Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[Remark 4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2.p1.4),[§7\.1](https://arxiv.org/html/2608.10288#S7.SS1.p3.5)\.
- \[45\]K\. G\. Wilson and J\. Kogut\(1974\)The renormalization group and theϵ\\epsilonexpansion\.Physics Reports12\(2\),pp\. 75–199\.Cited by:[Remark 6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3.p1.8),[Remark 6\.6](https://arxiv.org/html/2608.10288#S6.Thmtheorem6.p1.2)\.
- \[46\]S\. Yang, Y\. Shen, K\. Wen, S\. Tan, M\. Mishra, L\. Ren, R\. Panda, and Y\. Kim\(2025\)PaTH attention: position encoding via accumulating Householder transformations\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2505\.16381Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[47\]Y\. Yang, J\. Shang, Y\. Li, G\. Zhao, S\. Wang, and D\. Yu\(2026\)Autonomy\-of\-heads: data\-free sparse attention from frozen query\-key geometry\.Note:arXiv:2608\.06849Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17)\.
- \[48\]H\. Yu, T\. Jiang, S\. Jia, S\. Yan, S\. Liu, H\. Qian, G\. Li, S\. Dong, H\. Zhang, and C\. Yuan\(2025\)ComRoPE: scalable and robust rotary position embedding parameterized by trainable commuting angle matrices\.Note:arXiv:2506\.03737Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[Remark 4\.2](https://arxiv.org/html/2608.10288#S4.Thmtheorem2.p1.4)\.
- \[49\]H\. Yukawa\(1935\)On the interaction of elementary particles\. I\.Proceedings of the Physico\-Mathematical Society of Japan17,pp\. 48–57\.Cited by:[Remark 3\.4](https://arxiv.org/html/2608.10288#S3.Thmtheorem4.p1.6)\.
- \[50\]X\. Zheng, B\. Aragam, P\. Ravikumar, and E\. P\. Xing\(2018\)DAGs with NO TEARS: continuous optimization for structure learning\.InAdvances in Neural Information Processing Systems 31 \(NeurIPS\),pp\. 9492–9503\.Cited by:[§1](https://arxiv.org/html/2608.10288#S1.SS0.SSS0.Px2.p1.17),[§4\.3](https://arxiv.org/html/2608.10288#S4.SS3.p2.9),[Theorem 4\.4](https://arxiv.org/html/2608.10288#S4.Thmtheorem4)\.
- \[51\]G\. K\. Zipf\(1949\)Human behavior and the principle of least effort\.Addison\-Wesley\.Cited by:[Remark 6\.3](https://arxiv.org/html/2608.10288#S6.Thmtheorem3.p1.8),[§7\.6](https://arxiv.org/html/2608.10288#S7.SS6.p1.2)\.相似文章
图注意力何时应稀疏?学习逐边的 Tsallis 指数
本文提出了 LTGA,一种图注意力层,学习逐边的 Tsallis 熵指数,以在重尾、softmax 和紧支撑注意力之间插值,提供可解释的稀疏注意力,并在图基准上取得有竞争力的性能。
Parallax: 参数化局部线性注意力机制用于语言建模
介绍Parallax,一种参数化局部线性注意力机制,结合硬件感知优化,提升LLM预训练效率和性能,在0.6B和1.7B规模实现帕累托改进。
Dynamic Linear Attention
DLA引入了自适应状态合并和容量受限的内存建模,用于多状态线性注意力,提升了长上下文LLM的性能。
Exact Linear Attention
本文介绍了一种名为Exact Linear Attention (ELA) 的机制,该机制通过利用核函数分解,在不引入近似误差的情况下实现了Transformer注意力的线性计算复杂度,并通过约束核函数解决了梯度爆炸和词元稀释问题。文中还提出了包括超链接(Hyper Link)、记忆叶(Memory Lobe)以及面向混合专家模型的路由偏置在内的工程创新。
混合线性注意力大语言模型中的大规模激活:注意力前尖峰与尖峰间平台期
本文首次系统研究了混合线性注意力大语言模型中的大规模激活现象,揭示了由抵消时机控制的注意力前尖峰和尖峰间平台期,并展示了其形态如何在完全注意力极限下恢复。