Symmetry without a manifold: intrinsic dimension on orbits
Summary
The paper demonstrates that standard neural scaling law derivations fail when data forms group orbits, as intrinsic dimension is undefined, leading to exponential rather than power law scaling in model performance.
View Cached Full Text
Cached at: 09/17/26, 09:03 AM
# Symmetry without a manifold: intrinsic dimension on orbits Source: [https://arxiv.org/html/2609.17926](https://arxiv.org/html/2609.17926) Chon\-Fai Kam††thanks:Corresponding author\.Email:[dubussygauss@gmail\.com](mailto:[email protected])Affiliation:Affiliation:Dipartimento di Fisica e Chimica, Università degli Studi di Palermo, via Archirafi 36, I\-90123 Palermo, Italy Université Paris Cité and Université de La Réunion, BIGR, INSERM UMR\_S1134, F\-75014 Paris, France and EnergyLab, Université de La Réunion, F\-97715 Saint\-Denis, France and Université Paris Cité and Université de La Réunion, BIGR, INSERM UMR\_S1134, F\-75014 Paris, France PEACCEL, AI for Biologics, F\-75013 Paris, FranceFrédéric CadetEmail:[frederic\.cadet\.run@gmail\.com](mailto:[email protected])Affiliation:Affiliation: ###### Abstract The standard geometric derivation of neural scaling exponents takes the intrinsic dimension of a data manifold as its input\. On modular addition inℤp\\mathbb\{Z\}\_\{p\}that derivation has no input\. The exact algebraic solution is an orbit ofℤp\\mathbb\{Z\}\_\{p\}acting by isometries\. Transitivity alone makes the ratio statistic underlying the standard dimension estimator a point mass, so the estimator is undefined, and here the two nearest neighbour distances coincide exactly\. Breaking the symmetry at scaleϵ\\epsilonreturns a number, but one that tracks1/ϵ1/\\epsilonwith no scale free plateau\. We show that the failure is general, since on any finite orbit of a group acting by isometries the estimator reports the resolution at which the set is probed rather than a dimension\. What replaces the power law is exponential in hidden width,L\(h\)=L∞\+Aexp\(−chα\)L\(h\)=L\_\{\\infty\}\+A\\exp\(\-c\\,h^\{\\alpha\}\), withR2R^\{2\}between0\.9820\.982and0\.9950\.995against0\.8570\.857to0\.9060\.906for a power law admitting the same floor and fitted under the same protocol\. Where the data supply is sufficient the rate belongs to the regulariser rather than to the group, since weight decay movesccby a factor of4747while group order moves it by1\.101\.10, a residual below seed to seed resolution, for every fixedα\\alphabetween0\.750\.75and22\. The critical width falls with group order rather than rising, against capacity counting that assigns a fixed number of neurons to each irreducible representation\. ††proceedings:: Preprint\. Under review at NeurReps 2026###### keywords representational geometry, neural scaling laws, group representations, modular arithmetic, intrinsic dimension, grokking, weight decay ## 1Introduction A power law is a strong claim about a system\. It says that no scale is preferred, that the same relative improvement follows from the same relative investment at every size\. Neural scaling laws are reported as power laws across seven orders of magnitude in model size and across architectures, modalities and languages\([Kaplan et al\., 2020](https://arxiv.org/html/2609.17926#bib.bib13);[Hoffmann et al\., 2022](https://arxiv.org/html/2609.17926#bib.bib11);[Bahri et al\., 2024](https://arxiv.org/html/2609.17926#bib.bib2)\), and the derivations that explain why the form should be a power law all pass through an assumption about the data\. This paper reports a task on which that assumption fails in the strongest sense available: the quantity the standard derivation takes as input does not exist\. The failure is structural rather than numerical, and it does not stay confined to the task, because the object responsible is a group orbit\. Wherever a representation is a finite orbit of a group acting by isometries, the two nearest neighbour estimator is degenerate, and any dimension it reports under perturbation is a property of the probe scale rather than of the set\. The condition is finiteness, not symmetry\. A continuous ring attractor or toroidal population code\([Gardner et al\., 2022](https://arxiv.org/html/2609.17926#bib.bib7)\)is not covered, and a continuous torus given the same treatment returns a stable value near two\. What is covered is the sampled case, which is the usual experimental one: a population recorded atNNequally spaced stimulus conditions is itself a finite orbit, and an intrinsic dimension estimated from it reports the probe scale rather than the geometry of the underlying attractor\. Two derivations dominate\. The first routes the exponent through geometry, obtainingL∼N−4/dL\\sim N^\{\-4/d\}from the cell size adddimensional manifold partitioned byNNparameters admits\([Sharma and Kaplan, 2022](https://arxiv.org/html/2609.17926#bib.bib25)\)\. The second routes it through counting: if a task decomposes into subtasks whose use frequencies follow a Zipf distribution and each learned subtask reduces the loss by a fixed amount, the cumulative loss is again a power law\([Michaud et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib19);[Brill, 2024](https://arxiv.org/html/2609.17926#bib.bib3)\)\. The two accounts differ in almost everything except the ingredient that produces the exponent, which in both cases is a scale free distribution over structure in the data\. Algebraic tasks supply neither ingredient, and the literature has already noticed that something goes wrong\. These tasks entered the field through grokking\([Power et al\., 2022](https://arxiv.org/html/2609.17926#bib.bib24);[Liu et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib17);[Varma et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib30)\), and loss curves on algorithmic problems show pronounced phase transitions that depart from the established power law trend\([Naidu et al\., 2026](https://arxiv.org/html/2609.17926#bib.bib21)\)\. The geometric relation has been contradicted directly on a one dimensional regression problem where the measured exponent was11against a prediction of44\([Liu and Tegmark, 2023](https://arxiv.org/html/2609.17926#bib.bib16)\)\. What has not been asked is what replaces the power law when it fails, and whether the quantity the geometric derivation requires as input exists at all\. For modular addition it does not, and the reason can be stated before any measurement: the exact solution is known in closed form\([Nanda et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib22);[Gromov, 2023](https://arxiv.org/html/2609.17926#bib.bib8);[Chughtai et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib4)\)and its image is an orbit ofℤp\\mathbb\{Z\}\_\{p\}acting by isometries, not a sample from a density\. The governing hypothesis is that the rate of decay factorises aslogc=f\(λ\)\+g\(p\)\\log c=f\(\\lambda\)\+g\(p\)withggflat, so that the regularisation budget and not the order of the group sets the rate, and the design of Sec\.[2\.4](https://arxiv.org/html/2609.17926#S2.SS4)is built so that the two sweeps can disagree\. The choice of architecture is what makes the separation readable, since a two layer network with quadratic activation admits an exact solution whose Fourier content is known term by term\([Gromov, 2023](https://arxiv.org/html/2609.17926#bib.bib8);[Doshi et al\., 2024](https://arxiv.org/html/2609.17926#bib.bib5)\)\. We contribute the following\. Intrinsic dimension is undefined on group orbit representations, not merely mismeasured: the two nearest neighbour ratio is a point mass on any orbit of a finite group acting by isometries \(Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)\), equals exactly11for the representations at issue \(Lemma[6](https://arxiv.org/html/2609.17926#Thmtheorem6)\), and under perturbation of scaleϵ\\epsilonthe estimate tracks1/ϵ1/\\epsilonwith no plateau\. The test loss onℤp\\mathbb\{Z\}\_\{p\}is exponential in hidden width rather than polynomial in parameter count, over four and a half decades, and a power law admitting the same floor, fitted under the same protocol, is rejected at every group order and at every fixed exponent between0\.750\.75and22\. Given sufficient data the rate factorises: weight decay moves it by a factor of4747and group order by1\.101\.10, a residual below seed to seed resolution, at every fixed exponent in that range\. And the width at which generalisation appears falls as the group grows, contradicting in sign as well as magnitude any capacity argument that assigns a fixed number of neurons to each irreducible representation\. On this task the geometric route has no input, the power law is replaced by an exponential in hidden width, and the rate of that exponential belongs to the training regime and not to the group\. ## 2Problem formulation ### 2\.1Task and model The task is addition in the cyclic groupℤp\\mathbb\{Z\}\_\{p\}forppprime\. Inputs are pairs\(a,b\)∈ℤp×ℤp\(a,b\)\\in\\mathbb\{Z\}\_\{p\}\\times\\mathbb\{Z\}\_\{p\}encoded as two concatenated one hot vectors, so the input dimension is2p2p, and the target is\(a\+b\)modp\(a\+b\)\\bmod ptreated asppway classification\. Half of thep2p^\{2\}pairs are held out, the partition drawn once from a fixed seed so that it is identical across every width, weight decay and initialisation\. Table[1](https://arxiv.org/html/2609.17926#A1.T1)collects the notation\. The network follows[Gromov \(2023\)](https://arxiv.org/html/2609.17926#bib.bib8)\. Writingxxfor the input, f\(x\)=W2σ\(W1x\),σ\(z\)=z2,f\(x\)\\;=\\;W\_\{2\}\\,\\sigma\\\!\\left\(W\_\{1\}x\\right\),\\qquad\\sigma\(z\)=z^\{2\},\(1\)withW1∈ℝh×2pW\_\{1\}\\in\\mathbb\{R\}^\{h\\times 2p\}andW2∈ℝp×hW\_\{2\}\\in\\mathbb\{R\}^\{p\\times h\}and no biases\. The quadratic activation is what puts the Fourier content of the solution in closed form, since the cross term at frequencykkin the square of a sum of cosines producescosωk\(a\+b\)\\cos\\omega\_\{k\}\(a\+b\)directly\. The parameter count isN=3phN=3ph, so width and parameter count are proportional at fixedpp\. Training is full batch AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.17926#bib.bib18)\), with weight decay the principal control variable, swept over\{0,0\.1,0\.25,0\.5,1,2,4\}\\\{0,\\,0\.1,\\,0\.25,\\,0\.5,\\,1,\\,2,\\,4\\\}and defaulting to11where a single value is needed\. Three seeds are run at every point\. Reported quantities are the cross entropy and the accuracy on the held out half, together with the two spectral ones of Sec\.[2\.3](https://arxiv.org/html/2609.17926#S2.SS3)\. Appendix[E](https://arxiv.org/html/2609.17926#A5)gives the optimiser settings, the initialisation and the step budget\. ### 2\.2The exact solution as a group orbit The geometry the paper measures is fixed by representation theory before any network is trained\. Writingωk=2πk/p\\omega\_\{k\}=2\\pi k/p, the real irreducible representations ofℤp\\mathbb\{Z\}\_\{p\}are the trivial one together withKmax=\(p−1\)/2K\_\{\\mathrm\{max\}\}=\(p\-1\)/2two dimensional ones, in whichρk\(m\)\\rho\_\{k\}\(m\)is the rotation byωkm\\omega\_\{k\}mandkkandp−kp\-kare equivalent\.*Frequency*and*irreducible representation*are therefore the same object counted in two languages, andKmaxK\_\{\\mathrm\{max\}\}is the number of either available at group orderpp\. For a nonemptyS⊆\{1,…,Kmax\}S\\subseteq\\\{1,\\dots,K\_\{\\mathrm\{max\}\}\\\}define the embedding ΦS:ℤp→ℝ2\|S\|,ΦS\(n\)=\(cosωkn,sinωkn\)k∈S,\\Phi\_\{S\}:\\mathbb\{Z\}\_\{p\}\\to\\mathbb\{R\}^\{2\|S\|\},\\qquad\\Phi\_\{S\}\(n\)\\;=\\;\\big\(\\cos\\omega\_\{k\}n,\\ \\sin\\omega\_\{k\}n\\big\)\_\{k\\in S\},\(2\)and writeXS=ΦS\(ℤp\)X\_\{S\}=\\Phi\_\{S\}\(\\mathbb\{Z\}\_\{p\}\)for its image\. The exact algebraic solutions of the task, known in closed form\([Nanda et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib22);[Gromov, 2023](https://arxiv.org/html/2609.17926#bib.bib8);[Chughtai et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib4)\), place theppinput tokens at the points ofXSX\_\{S\}for someSS\. The group acts onXSX\_\{S\}by m⋅ΦS\(n\)=ΦS\(n\+m\)=\(⨁k∈Sρk\(m\)\)ΦS\(n\),m\\cdot\\Phi\_\{S\}\(n\)\\;=\\;\\Phi\_\{S\}\(n\+m\)\\;=\\;\\Big\(\\bigoplus\_\{k\\in S\}\\rho\_\{k\}\(m\)\\Big\)\\,\\Phi\_\{S\}\(n\),\(3\)which is block diagonal in rotations and hence an isometry of the ambient space\. The action is transitive by construction, and it is free becauseppis prime andSSis nonempty, soΦS\\Phi\_\{S\}is injective and\|XS\|=p\|X\_\{S\}\|=p\. The object of study is thus a single orbit of a finite group acting by isometries, not a sample from a density on a manifold, and Sec\.[4](https://arxiv.org/html/2609.17926#S4)shows that this is what breaks the geometric derivation\. ### 2\.3Candidate controlling variables The literature advances three quantities as what sets performance on a task of this kind, and the study measures all three across problem sizes rather than at one\. The first is the intrinsic dimensiondd, which the geometric derivation\([Sharma and Kaplan, 2022](https://arxiv.org/html/2609.17926#bib.bib25)\)takes as its input, estimated by the two nearest neighbour method of[Facco et al\. \(2017\)](https://arxiv.org/html/2609.17926#bib.bib6)onXSX\_\{S\}or on the embeddings a trained network learns\. The second is the number of irreducible representations the network carries, measured by the aggregate participation ratioKeffK\_\{\\mathrm\{eff\}\}of the spectral power in the embedding block ofW1W\_\{1\}and by the per neuron ratio obtained by the same construction within a single hidden unit, both defined in Appendix[C](https://arxiv.org/html/2609.17926#A3), Eq\. \([34](https://arxiv.org/html/2609.17926#A3.E34)\)\. The third is the widthhch\_\{c\}at which held out accuracy first crosses0\.90\.9, where a capacity argument would locate the dependence on group order\. ### 2\.4Hypotheses Competing accounts of what limits performance on this task make opposite predictions, and the design is chosen so that they can disagree\. Under a*capacity*account the binding constraint is the number of irreducible representations the network can carry, so the rate at which the loss falls with width, and the widthhch\_\{c\}at which generalisation appears, should both scale withKmaxK\_\{\\mathrm\{max\}\}and hence withpp\. Under a*budget*account the binding constraint is not the availability of modes but the norm in which they are expressed, so the rate should be a function of the regularisation strengthλ\\lambdaalone\. Writingccfor that rate, the budget account predicts the factorisation logc=f\(λ\)\+g\(p\),gconstant inp,\\log c\\;=\\;f\(\\lambda\)\\;\+\\;g\(p\),\\qquad g\\ \\text\{constant in\}\\ p,\(4\)which is falsified by any resolvable dependence ofggon group order\. Sweeping width at fixed data separates capacity from optimisation, and sweeping weight decay at fixed width separates the norm budget from capacity\. We measure the test loss for eight group orders between2323and113113, across fourteen widths from88to128128and seven weight decay strengths, with three seeds at every point\. Three questions organise what follows: whetherddexists onXSX\_\{S\}at all \(Sec\.[4](https://arxiv.org/html/2609.17926#S4)\), which functional family describesL\(h\)L\(h\)\(Sec\.[3](https://arxiv.org/html/2609.17926#S3)\), and which sweep moves the rate \(Sec\.[6](https://arxiv.org/html/2609.17926#S6)\)\. Figure 1:Test loss against width at weight decay11with fits to Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\) atα=1\\alpha=1\(a\), and the fitted floor against the number of available frequencies \(b\)\. BeyondKmax=26K\_\{\\mathrm\{max\}\}=26the floor falls into the plateau noise documented in Appendix[E](https://arxiv.org/html/2609.17926#A5)and the flattening there is not fitted\. ## 3The loss is exponential in width The first question is whether the received functional form describes the data at all\. FittinglogL\\log LagainstlogN\\log Npooled over all112112configurations at weight decay11givesR2=0\.726R^\{2\}=0\.726, while fittinglogL\\log LagainsthhgivesR2=0\.924R^\{2\}=0\.924, and per group order the best power law reachesR2R^\{2\}between0\.8570\.857and0\.9060\.906\. The residuals are not scattered but arc shaped, giving a Wald–Wolfowitz runs statistic ofz=−2\.78z=\-2\.78at every one of the eight group orders \(Appendix[E](https://arxiv.org/html/2609.17926#A5)\)\. The local slope varies in magnitude from0\.040\.04to8\.28\.2and turns negative nearh=12h=12to1616, and a power law cannot accommodate a sign change in its own exponent\. A plain exponential improves matters but introduces a new instability, since the curve saturates at the top of the width range and a form without a floor compensates by tilting, drifting by up to105%105\\%under choices of cutoff\. Admitting a floor removes it, as Fig\.[1](https://arxiv.org/html/2609.17926#S2.F1)\(a\) shows\. Writing L\(h\)=L∞\+Ae−chα,L\(h\)\\;=\\;L\_\{\\infty\}\\;\+\\;A\\,e^\{\-c\\,h^\{\\alpha\}\},\(5\)and fixingα=1\\alpha=1for the moment, the fit reachesR2R^\{2\}between0\.9820\.982and0\.9950\.995across the eight group orders, the rate becomesc=0\.1180±0\.0065c=0\.1180\\pm 0\.0065, and the drift under truncation falls to0\.3%0\.3\\%\(Appendix[E](https://arxiv.org/html/2609.17926#A5)\)\. The comparison is not one of parameter count\. A power law admitting the same floor, L\(h\)=L∞\+Ah−γ,L\(h\)\\;=\\;L\_\{\\infty\}\\;\+\\;A\\,h^\{\-\\gamma\},\(6\)fitted tologL\\log Lover the same width range under the same rejection rules and the same seed aggregation, reachesR2R^\{2\}between0\.8570\.857and0\.9060\.906against0\.9820\.982to0\.9950\.995for Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\), withΔAIC\\Delta\\mathrm\{AIC\}between25\.525\.5and40\.940\.9in favour of the exponential at every group order and at every fixedα\\alphabetween0\.750\.75and22\. The third parameter is not what decides it: the floor recovered by Eq\. \([6](https://arxiv.org/html/2609.17926#S3.E6)\) is not identified by the data, since a power law already approaches zero at a polynomial rate and has nothing for a floor to absorb \(Appendix[E](https://arxiv.org/html/2609.17926#A5)\)\. The extra parameter absorbs a systematic feature of the data and thereby stabilises the parameter of interest\. The exponentα\\alphais not determined by these measurements and the paper does not claim it\. Left free it converges to1\.583±0\.2541\.583\\pm 0\.254, and over\[0\.75,2\]\[0\.75,2\]the fittedccranges by a factor of214214whileR2R^\{2\}moves only from0\.9780\.978to0\.9940\.994\(Appendix[E](https://arxiv.org/html/2609.17926#A5)\)\. Section[6](https://arxiv.org/html/2609.17926#S6)reports what survives the degeneracy, and it survives because the claim made there concerns howccresponds to a change in conditions, not its value\. The floorL∞L\_\{\\infty\}falls with group order, from3\.7×10−33\.7\\times 10^\{\-3\}atp=23p=23to3\.1×10−53\.1\\times 10^\{\-5\}atp=113p=113as Fig\.[1](https://arxiv.org/html/2609.17926#S2.F1)\(b\) shows, its logarithm correlating withlogKmax\\log K\_\{\\mathrm\{max\}\}at−0\.94\-0\.94\. AboveKmax=26K\_\{\\mathrm\{max\}\}=26the apparent flattening is an artefact of measurement rather than a property of the task, since the loss there executes a random walk within a flat basin whose amplitude exceeds the differences between adjacent floors \(Appendix[E](https://arxiv.org/html/2609.17926#A5)\), so the slope is fitted only forKmax≤26K\_\{\\mathrm\{max\}\}\\leq 26\. ## 4The geometric route has no input Having established that the loss is not a power law, we turn to the derivation that predicts one\. The relationαN=4/d\\alpha\_\{N\}=4/drequires a manifold dimensiondd, and onℤp×ℤp\\mathbb\{Z\}\_\{p\}\\times\\mathbb\{Z\}\_\{p\}the naive answer is22\. Fitted over the width range of Sec\.[3](https://arxiv.org/html/2609.17926#S3), the best power law returns an exponent of3\.873\.87atp=47p=47and between2\.872\.87and4\.204\.20across the eight group orders, implying dimensions between0\.950\.95and1\.391\.39\. The implied dimension is not close to22, and it is not stable, since the exponent rises by half again fromp=23p=23top=113p=113, where the derivation supplies no mechanism for it to move at all\. What the geometric route delivers here is not a number to be checked against4/d4/d\. The deeper problem is thatdddoes not exist here\. The standard estimator of[Facco et al\. \(2017\)](https://arxiv.org/html/2609.17926#bib.bib6)recoversddas the shape of the Pareto law followed by the ratioμi=r2/r1\\mu\_\{i\}=r\_\{2\}/r\_\{1\}of each point’s two nearest neighbour distances, and is accurate on manifolds of known dimension even at the sample sizes available here \(Appendix[B](https://arxiv.org/html/2609.17926#A2)\)\. On the exact algebraic solution it returns nothing at all\. The reason is structural and is stated as follows\. ###### Proposition 1\(Degeneracy on a group orbit\)\. LetGGbe a finite group acting transitively on a finite setX⊂ℝmX\\subset\\mathbb\{R\}^\{m\}by isometries\. Thenμi=r2\(i\)/r1\(i\)\\mu\_\{i\}=r\_\{2\}\(i\)/r\_\{1\}\(i\)takes the same value for everyi∈Xi\\in X, the empirical distribution ofμ\\muis a point mass, and the two nearest neighbour regression is undefined\. Transitivity alone suffices, and the proof is one line: an isometry takingxxtoyycarries the distances out ofxxonto those out ofyy, sor1r\_\{1\},r2r\_\{2\}and their ratio agree at every point and the regression has no variation in its design variable\. The action of Eq\. \([3](https://arxiv.org/html/2609.17926#S2.E3)\) is in addition free, which fixes\|XS\|=p\|X\_\{S\}\|=p\. Appendix[B](https://arxiv.org/html/2609.17926#A2)proves the proposition and shows that at full frequency content theppembedding vectors form a regular simplex with pairwise cosine similarity−1/\(p−1\)\-1/\(p\-1\)identically, which atp=47p=47is−0\.0217\-0\.0217with a sample standard deviation of zero to machine precision\. Breaking the symmetry with additive noise of scaleϵ\\epsilonproduces a number, but not a stable one, tracking1/ϵ1/\\epsilonover a factor of thirty in the probe scale with no plateau, where a genuine torus given the same treatment returns a stable value near22\(Appendix[B](https://arxiv.org/html/2609.17926#A2)\)\. The dimension of an algebraic solution is not a property of the set but of the resolution at which it is probed, and4/d4/dhas no input to take\. What is not immediate is the scope\. This is not small sample noise, since it holds at everyNNand every embedding; not special toℤp\\mathbb\{Z\}\_\{p\}, since transitivity is the only hypothesis; and not repaired by perturbation, since the number that then appears tracks the probe rather than the set\. Trained networks reproduce this: atp=47p=47and width6464, at unit test accuracy, the estimator returns6868on the learned embeddings against an ambient dimension of6464, and grows with width\. The simplex describes full frequency content, not what a trained network shows, and[Tan et al\. \(2026\)](https://arxiv.org/html/2609.17926#bib.bib28)report a low rank cyclic geometry instead; both are orbits, so Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)covers either \(Appendix[B](https://arxiv.org/html/2609.17926#A2)\)\. Figure 2:The rateccagainst weight decay for four group orders \(a\)\. The curves forp≥47p\\geq 47coincide to within measurement error while the rate itself moves by a factor of4747\. Per neuron spectral participation ratio against held out accuracy over the full grid \(b\), showing that rank one structure holds only where the task is solved\. ## 5The irreducible representation count does not control the loss If the geometric variable is unavailable, the mechanistic one seems obvious\. Each neuron converges to a single irreducible representation under gradient flow\([He et al\., 2026b](https://arxiv.org/html/2609.17926#bib.bib10);[He et al\., 2026a](https://arxiv.org/html/2609.17926#bib.bib9)\), which makes the count of distinct represented modes well defined, and it is the quantity we expected to organise the data\. It does not\. Within a single group order the correlation is convincing, returningR2=0\.90R^\{2\}=0\.90atp=47p=47\. Across group orders it collapses, the loss ranging over a factor of25002500atKeff≈20K\_\{\\mathrm\{eff\}\}\\approx 20andKeffK\_\{\\mathrm\{eff\}\}returningR2=0\.590R^\{2\}=0\.590pooled againstR2=0\.917R^\{2\}=0\.917for width alone\. The within group correlation is an artefact of an occupancy process rather than a measure of capacity: a zero parameter occupancy model accounts forR2=0\.907R^\{2\}=0\.907of the variance inKeffK\_\{\\mathrm\{eff\}\}\. It is visible only across problem sizes, since within a single size capacity and sampling are confounded by construction\. The claim delimits[He et al\. \(2026a\)](https://arxiv.org/html/2609.17926#bib.bib9)without disagreeing with them\. Appendix[D](https://arxiv.org/html/2609.17926#A4)gives the measurements and a second finding, that regularisation imposes the rank one structure of the neurons rather than the architecture supplying it \(Fig\.[2](https://arxiv.org/html/2609.17926#S4.F2)\(b\)\)\. ## 6The rate factorises What remains is the rate\. Weight decay is the natural candidate, since Appendix[A](https://arxiv.org/html/2609.17926#A1)shows the unconstrained frequency restricted solution already achieves zero loss with four frequencies, so the constraint that binds cannot be the availability of modes\. Sweeping weight decay at fixed group order moves the rate by a factor of4747, from0\.1710\.171at weight decay0\.10\.1to0\.003640\.00364at weight decay44, monotone decreasing in the mean over group orders once runs that fail to solve the task are excluded\. Retaining them produces a spurious maximum at intermediate weight decay, because the three parameter form fitted to an unsolved run places the recovered floor above the bulk of the data\. Appendix[E](https://arxiv.org/html/2609.17926#A5)states the validity criterion, which rejects a fit when the recovered floor exceeds three times the smallest observed loss\. Against that factor of4747the group order does almost nothing\. This is the test of the budget hypothesis of Sec\.[2\.4](https://arxiv.org/html/2609.17926#S2.SS4), Eq\. \([4](https://arxiv.org/html/2609.17926#S2.E4)\)\. Over four group orders and five weight decay strengths, tabulated in Table[6](https://arxiv.org/html/2609.17926#A5.T6), the decomposition returns a spread inggof a factor of1\.0971\.097amongp≥47p\\geq 47, against the factor of4747contributed byff, and Fig\.[2](https://arxiv.org/html/2609.17926#S4.F2)\(a\) shows the four curves\. The residual spread is not merely small but unresolvable: in the regime where the floor is identified the standard error of a three seed mean is4\.3%4\.3\\%and the observed spread across group orders is also4\.3%4\.3\\%, so group order dependence, if present, is below what this experiment can detect \(Appendix[E](https://arxiv.org/html/2609.17926#A5)\)\. Because the exponentα\\alphais not identified, the factorisation must also be checked against the choice of fixing it at unity\. Repeating the fit atα∈\{0\.75,…,2\}\\alpha\\in\\\{0\.75,\\dots,2\\\}moves the rate by a factor of214214while the coefficient of variation across group order stays between4\.0%4\.0\\%and6\.8%6\.8\\%\(Table[10](https://arxiv.org/html/2609.17926#A5.T10)\), so the claim concerns the structure of the dependence, not a number\. The single exception is instructive\. Atp=23p=23the rate sits a factor of0\.7930\.793below the others, a real deviation of4\.84\.8standard errors, and that group order also fails outright at low regularisation\. Its training set contains264264pairs against63846384atp=113p=113\. The factorisation is a statement about the regime in which data are sufficient, andp=23p=23marks one edge of it\. There is a second edge in the same direction: raising the training fraction from a half to four fifths moves the rate at both activations, and under the quadratic one it carries the residual past zero \(Appendix[F](https://arxiv.org/html/2609.17926#A6)\)\. Whether the factorisation belongs to the task or to the quadratic activation is settled in part by repeating the sweep with ReLU\. The functional form transfers, the exponential family reachingR2=0\.963R^\{2\}=0\.963against0\.8690\.869for the best power law\. The factorisation does not, since the rate differs between group orders by61%61\\%on average, and truncation does not account for it\. The critical width tracks the discrepancy, in the pattern of Sec\.[7](https://arxiv.org/html/2609.17926#S7), where the quadratic activation itself departs atp=23p=23with264264training pairs\. The factorisation holds wherever data are plentiful relative to what the architecture needs per mode, and ReLU reaches that boundary at a larger group order\. Supplying more data closes the gap \(Appendix[F](https://arxiv.org/html/2609.17926#A6)\)\. Appendix[E\.10](https://arxiv.org/html/2609.17926#A5.SS10)gives the protocol and the full comparison\. Figure 3:Held out accuracy against width at weight decay11for eight group orders \(a\), and the resulting critical width against group order \(b\), with the least squares line\. The requirement falls as the group grows, where capacity counting that assigns a fixed number of neurons per irreducible representation requires it to rise\. ## 7Critical width falls as the group grows The factorisation says that group order does not enter the rate\. Whether it enters anywhere is best asked ofhch\_\{c\}, where a capacity argument would put it\. Defininghch\_\{c\}as the width at which held out accuracy crosses0\.90\.9by linear interpolation, it falls from34\.134\.1atp=23p=23to26\.626\.6atp=113p=113, a linear fit againstppgiving a slope of−0\.0738\-0\.0738withR2=0\.934R^\{2\}=0\.934as Fig\.[3](https://arxiv.org/html/2609.17926#S6.F3)\(b\) shows\. The requirement falls as the group grows, by twenty two percent over a range of4\.94\.9in the order of the group\. This is the opposite of what capacity counting predicts\. Any account assigning a fixed number of neurons to each irreducible representation makeshch\_\{c\}grow withpp, and our own Proposition[2](https://arxiv.org/html/2609.17926#Thmtheorem2)is of that form, needing4\(p−1\)4\(p\-1\)neurons at full frequency content, whereas the expressible class of a monomial activation is fixed by its degree alone and does not grow withpp\([Kam et al\., 2026](https://arxiv.org/html/2609.17926#bib.bib12)\)\. Such bounds give a width at which the solution can be represented and not the width at which training finds it, but the disagreement here is in sign as well as magnitude and cannot be absorbed by a constant\. The training set explains the direction, not the representation theory\. Pairs grow asp2/2p^\{2\}/2while frequencies grow asp/2p/2, so larger groups are better supplied with data per mode and overfit less, and[Tian \(2025\)](https://arxiv.org/html/2609.17926#bib.bib29)prove thatO\(MlogM\)O\(M\\log M\)samples suffice on a group arithmetic task of orderMM, which grows more slowly than the pairs the task supplies\. This also accounts for the failure ofp=23p=23at low regularisation reported in Sec\.[6](https://arxiv.org/html/2609.17926#S6)\. An independent argument reaches the same conclusion from the representation side, since[Tan et al\. \(2026\)](https://arxiv.org/html/2609.17926#bib.bib28)derive a critical regularisation strength scaling as the inverse of the number of classes: larger groups should tolerate weaker weight decay, which is what we observe\. ## 8Discussion The power law form of neural scaling laws is a property of the data, not of learning as such, and modular addition supplies neither ingredient the derivations require\. Its exact solution is a group orbit and not a sampled manifold, its irreducible representations are equivalent under the symmetry of the group and carry no preferred scale, and they combine coherently rather than additively, since the logit is built by constructive interference of modes into a peak\. Exponential scaling is not without precedent\([Sorscher et al\., 2022](https://arxiv.org/html/2609.17926#bib.bib26);[Liu et al\., 2025](https://arxiv.org/html/2609.17926#bib.bib15)\), though there the mechanism is a change in the effective distribution over structure rather than a property of the task\. Whether the form follows from the coherence is left open: the stretched exponential of Appendix[A](https://arxiv.org/html/2609.17926#A1)depends on group order where the measured rate does not\. Both departures from the factorisation, the quadratic activation atp=23p=23and ReLU atp=47p=47, occur where the training set is smallest relative to the width the architecture needs, and in both the fitted rate falls below the common value\. That predicts more data closes the deficit, and it does, by an amount that tracks the size of the deficit at both activations \(Appendix[F](https://arxiv.org/html/2609.17926#A6)\)\. Data supply bounds the factorisation, not the activation\. The most portable result is methodological\. Three quantities that behave like controlling variables within a single problem size turn out not to be controlling variables at all: the intrinsic dimension, which is undefined on group orbits yet returns a plausible number under perturbation; the aggregate spectral occupancy, which tracks an occupancy process rather than capacity; and the fitted power law exponent itself, which is stable enough within one group order to be reported but moves by half again across them, where the derivation that produces it gives no reason for movement\. Each would have survived a study at a single problem size\. Wherever a representation carries a group action, the estimators that summarise its geometry must be validated against that action before their output is read as a dimension\. The validation is two checks, both available beforehand: the empirical distribution of the ratio statistic, and the stability of the estimate under a change of probe scale\. The torus control of Appendix[B](https://arxiv.org/html/2609.17926#A2)passes both and the orbit fails both\. The non abelian case is the natural continuation, withKmaxK\_\{\\mathrm\{max\}\}replaced by∑ρdρ\\sum\_\{\\rho\}d\_\{\\rho\}: whether the factorisation survives that replacement would say whether the dimensions of the representations enter the rate\([Chughtai et al\., 2023](https://arxiv.org/html/2609.17926#bib.bib4);[Stander et al\., 2024](https://arxiv.org/html/2609.17926#bib.bib27)\)\. ## References - Ansuini et al\. \(2019\)Alessio Ansuini, Alessandro Laio, Jakob H\. Macke, and Davide Zoccolan\.Intrinsic dimension of data representations in deep neural networks\.In*Advances in Neural Information Processing Systems*, volume 32, 2019\. - Bahri et al\. \(2024\)Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma\.Explaining neural scaling laws\.*Proceedings of the National Academy of Sciences*, 121\(27\):e2311878121, 2024\. - Brill \(2024\)Ari Brill\.Neural scaling laws rooted in the data distribution\.*arXiv preprint arXiv:2412\.07942*, 2024\. - Chughtai et al\. \(2023\)Bilal Chughtai, Lawrence Chan, and Neel Nanda\.A toy model of universality: Reverse engineering how networks learn group operations\.In*International Conference on Machine Learning*, 2023\. - Doshi et al\. \(2024\)Darshil Doshi, Tianyu He, Aritra Das, and Andrey Gromov\.Grokking modular polynomials\.*arXiv preprint arXiv:2406\.03495*, 2024\. - Facco et al\. \(2017\)Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio\.Estimating the intrinsic dimension of datasets by a minimal neighborhood information\.*Scientific Reports*, 7:12140, 2017\. - Gardner et al\. \(2022\)Richard J\. Gardner, Erik Hermansen, Marius Pachitariu, Yoram Burak, Nils A\. Baas, Benjamin A\. Dunn, May\-Britt Moser, and Edvard I\. Moser\.Toroidal topology of population activity in grid cells\.*Nature*, 602:123–128, 2022\. - Gromov \(2023\)Andrey Gromov\.Grokking modular arithmetic\.*arXiv preprint arXiv:2301\.02679*, 2023\. - He et al\. \(2026a\)Jianliang He, Leda Wang, Siyu Chen, and Zhuoran Yang\.On the mechanism and dynamics of modular addition: Fourier features, lottery ticket, and grokking\.*arXiv preprint arXiv:2602\.16849*, 2026a\. - He et al\. \(2026b\)Jianliang He, Leda Wang, Fengzhuo Zhang, Siyu Chen, and Zhuoran Yang\.Neural networks provably learn spectral representations for group composition\.*arXiv preprint arXiv:2606\.02993*, 2026b\. - Hoffmann et al\. \(2022\)Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al\.Training compute\-optimal large language models\.*arXiv preprint arXiv:2203\.15556*, 2022\. - Kam et al\. \(2026\)Chon\-Fai Kam, Xavier Cadet, Miloud Bessafi, and Frederic Cadet\.Algebraic representability as the limiting regime of grokking: An exactly solvable model with holomorphic activations\.*arXiv preprint arXiv:2607\.13749*, 2026\. - Kaplan et al\. \(2020\)Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B\. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei\.Scaling laws for neural language models\.*arXiv preprint arXiv:2001\.08361*, 2020\. - Levina and Bickel \(2004\)Elizaveta Levina and Peter J\. Bickel\.Maximum likelihood estimation of intrinsic dimension\.In*Advances in Neural Information Processing Systems*, volume 17, 2004\. - Liu et al\. \(2025\)Yizhou Liu, Ziming Liu, and Jeff Gore\.Superposition yields robust neural scaling\.*arXiv preprint arXiv:2505\.10465*, 2025\. - Liu and Tegmark \(2023\)Ziming Liu and Max Tegmark\.A neural scaling law from lottery ticket ensembling\.*arXiv preprint arXiv:2310\.02258*, 2023\. - Liu et al\. \(2023\)Ziming Liu, Eric J\. Michaud, and Max Tegmark\.Omnigrok: Grokking beyond algorithmic data\.In*International Conference on Learning Representations*, 2023\. - Loshchilov and Hutter \(2019\)Ilya Loshchilov and Frank Hutter\.Decoupled weight decay regularization\.In*International Conference on Learning Representations*, 2019\. - Michaud et al\. \(2023\)Eric J\. Michaud, Ziming Liu, Uzay Girit, and Max Tegmark\.The quantization model of neural scaling\.In*Advances in Neural Information Processing Systems*, volume 36, 2023\. - Morwani et al\. \(2024\)Depen Morwani, Benjamin L\. Edelman, Costin\-Andrei Oncescu, Rosie Zhao, and Sham Kakade\.Feature emergence via margin maximization: Case studies in algebraic tasks\.In*International Conference on Learning Representations*, 2024\. - Naidu et al\. \(2026\)Prudhviraj Naidu, Zixian Wang, Leon Bergen, and Ramamohan Paturi\.Quiet feature learning in algorithmic tasks\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 40, pages 37756–37764, 2026\.[10\.1609/aaai\.v40i44\.41111](https://doi.org/10.1609/aaai.v40i44.41111)\. - Nanda et al\. \(2023\)Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt\.Progress measures for grokking via mechanistic interpretability\.In*International Conference on Learning Representations*, 2023\. - Papyan et al\. \(2020\)Vardan Papyan, X\. Y\. Han, and David L\. Donoho\.Prevalence of neural collapse during the terminal phase of deep learning training\.*Proceedings of the National Academy of Sciences*, 117\(40\):24652–24663, 2020\. - Power et al\. \(2022\)Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra\.Grokking: Generalization beyond overfitting on small algorithmic datasets\.*arXiv preprint arXiv:2201\.02177*, 2022\. - Sharma and Kaplan \(2022\)Utkarsh Sharma and Jared Kaplan\.Scaling laws from the data manifold dimension\.*Journal of Machine Learning Research*, 23\(9\):1–34, 2022\. - Sorscher et al\. \(2022\)Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S\. Morcos\.Beyond neural scaling laws: Beating power law scaling via data pruning\.In*Advances in Neural Information Processing Systems*, volume 35, 2022\. - Stander et al\. \(2024\)Dashiell Stander, Qinan Yu, Honglu Fan, and Stella Biderman\.Grokking group multiplication with cosets\.*arXiv preprint arXiv:2312\.06581*, 2024\. - Tan et al\. \(2026\)Hu Tan, Kuo Gai, and Shihua Zhang\.Beyond neural collapse: Task\-intrinsic geometry governs neural representations in modular arithmetic\.*arXiv preprint arXiv:2606\.08985*, 2026\. - Tian \(2025\)Yuandong Tian\.Provable scaling laws of feature emergence from learning dynamics of grokking\.*arXiv preprint arXiv:2509\.21519*, 2025\. - Varma et al\. \(2023\)Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar\.Explaining grokking through circuit efficiency\.*arXiv preprint arXiv:2309\.02390*, 2023\. ## Appendix AThe frequency restricted solution and its loss Table 1:Notation\.This appendix works out what the exact algebraic solution costs when it is restricted to a subset of the available irreducible representations and its logits are held to a fixed norm\. Four things are established\. The architecture of Eq\. \([1](https://arxiv.org/html/2609.17926#S2.E1)\) represents the target logit exactly, and the explicit construction that does so is far more expensive in width than what the trained network finds\. The resulting logit is a Dirichlet kernel whose behaviour at full frequency content is exactly a discrete delta\. Without a norm constraint the loss reaches zero at very small mode counts, so mode availability is not what limits the trained network\. Under a norm constraint the loss becomes a stretched exponential in the mode count, but with a range of validity narrower than a naive expansion suggests and with an optimal amplitude assignment that is not the uniform one\. ### A\.1From the architecture to the logit WriteW1=\[U;V\]W\_\{1\}=\[U;V\]withU,V∈ℝp×hU,V\\in\\mathbb\{R\}^\{p\\times h\}, so that the preactivation of hidden unitjjon the input pair\(a,b\)\(a,b\)isUaj\+VbjU\_\{aj\}\+V\_\{bj\}and the logit assigned to classccis zc\(a,b\)=∑j=1h\(Uaj\+Vbj\)2W2,jc\.z\_\{c\}\(a,b\)\\;=\\;\\sum\_\{j=1\}^\{h\}\\big\(U\_\{aj\}\+V\_\{bj\}\\big\)^\{2\}\\,W\_\{2,jc\}\.\(7\)The quadratic activation in Eq\. \([7](https://arxiv.org/html/2609.17926#A1.E7)\) matters because the square of a sum of two cosines contains a cross term at the sum frequency\. Writingφa=ωka\\varphi\_\{a\}=\\omega\_\{k\}aandφb=ωkb\\varphi\_\{b\}=\\omega\_\{k\}b, expanding\(cosφa\+cosφb\)2\(\\cos\\varphi\_\{a\}\+\\cos\\varphi\_\{b\}\)^\{2\}produces2cosφacosφb=cos\(φa−φb\)\+cos\(φa\+φb\)2\\cos\\varphi\_\{a\}\\cos\\varphi\_\{b\}=\\cos\(\\varphi\_\{a\}\-\\varphi\_\{b\}\)\+\\cos\(\\varphi\_\{a\}\+\\varphi\_\{b\}\), so the dependence ona\+ba\+bthat the task requires is already present after one nonlinearity\. The difficulty is that the same expansion also producescos\(φa−φb\)\\cos\(\\varphi\_\{a\}\-\\varphi\_\{b\}\)and the two squared terms, none of which depend ona\+ba\+b, and these must be removed\. They can be removed exactly\. The following construction uses eight hidden units per frequency and second layer weights taking values in\{0,±14\}\\\{0,\\pm\\tfrac\{1\}\{4\}\\\}\. ###### Proposition 2\(Exact representability\)\. FixS⊆\{1,…,Kmax\}S\\subseteq\\\{1,\\dots,K\_\{\\mathrm\{max\}\}\\\}and amplitudesβk\\beta\_\{k\}\. There is a choice ofW1W\_\{1\}andW2W\_\{2\}withh=8\|S\|h=8\|S\|for which zc\(a,b\)=∑k∈Sβkcos\(ωk\(a\+b−c\)\)z\_\{c\}\(a,b\)\\;=\\;\\sum\_\{k\\in S\}\\beta\_\{k\}\\cos\\\!\\big\(\\omega\_\{k\}\(a\+b\-c\)\\big\)\(8\)holds identically onℤp×ℤp\\mathbb\{Z\}\_\{p\}\\times\\mathbb\{Z\}\_\{p\}\. ###### Proof\. Fixk∈Sk\\in Sand writeφa=ωka\\varphi\_\{a\}=\\omega\_\{k\}aandφb=ωkb\\varphi\_\{b\}=\\omega\_\{k\}bas above\. Allocate four units carrying the preactivationscosφa±cosφb\\cos\\varphi\_\{a\}\\pm\\cos\\varphi\_\{b\}andsinφa±sinφb\\sin\\varphi\_\{a\}\\pm\\sin\\varphi\_\{b\}\. Squaring and combining with signs\+,−,−,\+\+,\-,\-,\+gives \(cosφa\+cosφb\)2−\(cosφa−cosφb\)2\\displaystyle\(\\cos\\varphi\_\{a\}\+\\cos\\varphi\_\{b\}\)^\{2\}\-\(\\cos\\varphi\_\{a\}\-\\cos\\varphi\_\{b\}\)^\{2\}=4cosφacosφb,\\displaystyle=4\\cos\\varphi\_\{a\}\\cos\\varphi\_\{b\},\(9\)\(sinφa\+sinφb\)2−\(sinφa−sinφb\)2\\displaystyle\(\\sin\\varphi\_\{a\}\+\\sin\\varphi\_\{b\}\)^\{2\}\-\(\\sin\\varphi\_\{a\}\-\\sin\\varphi\_\{b\}\)^\{2\}=4sinφasinφb,\\displaystyle=4\\sin\\varphi\_\{a\}\\sin\\varphi\_\{b\},\(10\)and subtracting the second line from the first leaves4cos\(φa\+φb\)4\\cos\(\\varphi\_\{a\}\+\\varphi\_\{b\}\)\. Every term that does not depend ona\+ba\+bhas cancelled\. Allocate four further units carryingcosφa±sinφb\\cos\\varphi\_\{a\}\\pm\\sin\\varphi\_\{b\}andsinφa±cosφb\\sin\\varphi\_\{a\}\\pm\\cos\\varphi\_\{b\}, whose squared differences give4cosφasinφb4\\cos\\varphi\_\{a\}\\sin\\varphi\_\{b\}and4sinφacosφb4\\sin\\varphi\_\{a\}\\cos\\varphi\_\{b\}, and whose sum is4sin\(φa\+φb\)4\\sin\(\\varphi\_\{a\}\+\\varphi\_\{b\}\)\. Assigning second layer weights14βkcosωkc\\tfrac\{1\}\{4\}\\beta\_\{k\}\\cos\\omega\_\{k\}cto the first group and14βksinωkc\\tfrac\{1\}\{4\}\\beta\_\{k\}\\sin\\omega\_\{k\}cto the second and summing overkkproduces ∑k∈Sβk\[cos\(ωk\(a\+b\)\)cosωkc\+sin\(ωk\(a\+b\)\)sinωkc\],\\sum\_\{k\\in S\}\\beta\_\{k\}\\Big\[\\cos\\big\(\\omega\_\{k\}\(a\+b\)\\big\)\\cos\\omega\_\{k\}c\+\\sin\\big\(\\omega\_\{k\}\(a\+b\)\\big\)\\sin\\omega\_\{k\}c\\Big\],\(11\)which is Eq\. \([8](https://arxiv.org/html/2609.17926#A1.E8)\) by the cosine subtraction formula\. ∎ We have verified Eq\. \([8](https://arxiv.org/html/2609.17926#A1.E8)\) numerically against this construction atp=47p=47for\|S\|=1\|S\|=1,44and2323, with agreement to10−1410^\{\-14\}\. The width Proposition[2](https://arxiv.org/html/2609.17926#Thmtheorem2)consumes is worth comparing against measurement\. Atp=47p=47with all frequencies retained it needs8Kmax=1848K\_\{\\mathrm\{max\}\}=184hidden units, whereas the critical width reported in Sec\.[7](https://arxiv.org/html/2609.17926#S7)is30\.930\.9\. The trained network is therefore about six times more economical than the most obvious exact solution, which is consistent with the finding of Refs\.\([He et al\., 2026b](https://arxiv.org/html/2609.17926#bib.bib10);[He et al\., 2026a](https://arxiv.org/html/2609.17926#bib.bib9)\)that a single unit carries a single irreducible representation and that the unwanted terms are suppressed by incoherent cancellation across units rather than removed exactly\. The construction above should be read as a proof of representability rather than as a model of what gradient descent finds\. That the solutions selected by training are the max margin ones, and that these use Fourier features on algebraic tasks, is established by[Morwani et al\. \(2024\)](https://arxiv.org/html/2609.17926#bib.bib20)\. Which targets a quadratic network can reach at all is a prior question, separate from the width reaching one of them consumes\.[Kam et al\. \(2026\)](https://arxiv.org/html/2609.17926#bib.bib12)settle it for the monomial familyσ\(z\)=zk\\sigma\(z\)=z^\{k\}on roots of unity inputs, where the reachable functions span thek\+1k\+1characters whose frequency pair sums tokk, and a target outside that span cannot be fitted even on the training set\. Proposition[2](https://arxiv.org/html/2609.17926#Thmtheorem2)sits on the far side of that question\. The target of Eq\. \([8](https://arxiv.org/html/2609.17926#A1.E8)\) is reachable, and what the width buys is the removal of the terms the square produces alongside it\. ### A\.2The logit as a Dirichlet kernel Everything downstream depends on the shape of Eq\. \([8](https://arxiv.org/html/2609.17926#A1.E8)\) as a function of the class index, so we record it in closed form\. TakingS=\{1,…,K\}S=\\\{1,\\dots,K\\\}with unit amplitudes and writings=a\+b−cmodps=a\+b\-c\\bmod p, the logit is the partial sum DK\(s\)=∑k=1Kcos\(ωks\)\.D\_\{K\}\(s\)\\;=\\;\\sum\_\{k=1\}^\{K\}\\cos\(\\omega\_\{k\}s\)\.\(12\) ###### Lemma 3\(Closed form\)\. Fors≢0modps\\not\\equiv 0\\bmod p, DK\(s\)=−12\+sin\(\(2K\+1\)πs/p\)2sin\(πs/p\),D\_\{K\}\(s\)\\;=\\;\-\\frac\{1\}\{2\}\\;\+\\;\\frac\{\\sin\\\!\\big\(\(2K\+1\)\\pi s/p\\big\)\}\{2\\sin\(\\pi s/p\)\},\(13\)whileDK\(0\)=KD\_\{K\}\(0\)=K\. ###### Proof\. Writeθ=2πs/p\\theta=2\\pi s/p, which is not a multiple of2π2\\pi, sosin\(θ/2\)≠0\\sin\(\\theta/2\)\\neq 0\. The product to sum identity gives2sin\(θ/2\)cos\(kθ\)=sin\(\(k\+12\)θ\)−sin\(\(k−12\)θ\)2\\sin\(\\theta/2\)\\cos\(k\\theta\)=\\sin\\big\(\(k\+\\tfrac\{1\}\{2\}\)\\theta\\big\)\-\\sin\\big\(\(k\-\\tfrac\{1\}\{2\}\)\\theta\\big\)\. Summing overkkfrom11toKKmakes the right hand side telescope tosin\(\(K\+12\)θ\)−sin\(θ/2\)\\sin\\big\(\(K\+\\tfrac\{1\}\{2\}\)\\theta\\big\)\-\\sin\(\\theta/2\)\. Dividing by2sin\(θ/2\)2\\sin\(\\theta/2\)and substitutingθ=2πs/p\\theta=2\\pi s/pgives Eq\. \([13](https://arxiv.org/html/2609.17926#A1.E13)\)\. The value ats=0s=0is immediate from Eq\. \([12](https://arxiv.org/html/2609.17926#A1.E12)\)\. ∎ The case of full frequency content is special and exact\. ###### Corollary 4\(Delta kernel\)\. Letppbe odd andKmax=\(p−1\)/2K\_\{\\mathrm\{max\}\}=\(p\-1\)/2\. ThenDKmax\(0\)=KmaxD\_\{K\_\{\\mathrm\{max\}\}\}\(0\)=K\_\{\\mathrm\{max\}\}, andDKmax\(s\)=−12D\_\{K\_\{\\mathrm\{max\}\}\}\(s\)=\-\\tfrac\{1\}\{2\}for everys≢0s\\not\\equiv 0\. ###### Proof\. AtK=KmaxK=K\_\{\\mathrm\{max\}\}the numerator of Eq\. \([13](https://arxiv.org/html/2609.17926#A1.E13)\) issin\(\(2Kmax\+1\)πs/p\)=sin\(πs\)\\sin\\big\(\(2K\_\{\\mathrm\{max\}\}\+1\)\\pi s/p\\big\)=\\sin\(\\pi s\), which vanishes for every integerss\. ∎ The kernel at full frequency content is thus a discrete delta sitting on a constant background, and the separation between the correct class and every competitor is exactlyKmax\+12K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}per unit amplitude\. Numerical evaluation atp=47p=47agrees to2×10−142\\times 10^\{\-14\}\. This is the algebraic statement behind the simplex geometry established in Appendix[B](https://arxiv.org/html/2609.17926#A2), since a kernel that is constant off the diagonal is precisely what makes all pairwise inner products equal\. ### A\.3Cross entropy Because the group acts transitively on the input pairs, the loss does not depend on which pair is presented, and the cross entropy for the solution \([8](https://arxiv.org/html/2609.17926#A1.E8)\) with uniform amplitudeβ\\betais L\(K,β\)=log\(1\+∑s≠0e−β\(K−DK\(s\)\)\)\.L\(K,\\beta\)\\;=\\;\\log\\\!\\Big\(1\+\\sum\_\{s\\neq 0\}e^\{\-\\beta\\left\(K\-D\_\{K\}\(s\)\\right\)\}\\Big\)\.\(14\)Two regimes follow according to whetherβ\\betais free\. #### A\.3\.1Unconstrained amplitude Supposeβ\\betamay grow without bound\. Every exponent in Eq\. \([14](https://arxiv.org/html/2609.17926#A1.E14)\) is then driven to−∞\-\\inftyprovidedK−DK\(s\)\>0K\-D\_\{K\}\(s\)\>0for alls≠0s\\neq 0, that is provided the kernel has a strict maximum at the origin\. By Lemma[3](https://arxiv.org/html/2609.17926#Thmtheorem3)the second term ofDK\(s\)D\_\{K\}\(s\)is bounded by1/\(2sin\(π/p\)\)1/\(2\\sin\(\\pi/p\)\)in absolute value, so the condition holds for everyK≥1K\\geq 1onceppis large enough, and it holds for everyK≥1K\\geq 1at the values ofppstudied here\. HenceL→0L\\to 0for any nonempty frequency set\. The rate at which this happens is fast\. Optimisingβ\\betaat eachKKforp=47p=47drives the loss below10−1210^\{\-12\}fromK=4K=4onwards, whereas trained networks whose measured mode count is comparable sit nearL≈3L\\approx 3\. The gap between these two numbers is the reason Sec\.[6](https://arxiv.org/html/2609.17926#S6)looks to the norm budget rather than to mode availability for the binding constraint\. #### A\.3\.2Constrained amplitude Weight decay does not permit unbounded logits, so the relevant question is how well a fixed budget can be spent\. Allow the amplitudes to differ and write𝜷=\(βk\)k∈S\\boldsymbol\{\\beta\}=\(\\beta\_\{k\}\)\_\{k\\in S\}subject to‖𝜷‖2=B\\\|\\boldsymbol\{\\beta\}\\\|\_\{2\}=B\. The logit separation between the correct class and the competitor at offsetssis∑k∈Sβk\(1−cos\(ωks\)\)\\sum\_\{k\\in S\}\\beta\_\{k\}\\big\(1\-\\cos\(\\omega\_\{k\}s\)\\big\), so the quantity that controls the loss is the worst case separation ΔS\(𝜷\)=min∑k∈Ss≠0βk\(1−cos\(ωks\)\),\\Delta\_\{S\}\(\\boldsymbol\{\\beta\}\)\\;=\\;\\min\_\{s\\neq 0\}\\ \\sum\_\{k\\in S\}\\beta\_\{k\}\\big\(1\-\\cos\(\\omega\_\{k\}s\)\\big\),\(15\)and Eq\. \([14](https://arxiv.org/html/2609.17926#A1.E14)\) is bounded byL≤log\(1\+\(p−1\)e−ΔS\(𝜷\)\)L\\leq\\log\\big\(1\+\(p\-1\)e^\{\-\\Delta\_\{S\}\(\\boldsymbol\{\\beta\}\)\}\\big\)with equality when a single competitor dominates\. Choosing the best allocation of a fixed budget is therefore the max\-min problemΔS∗\(B\)=max‖𝜷‖=BΔS\(𝜷\)\\Delta^\{\*\}\_\{S\}\(B\)=\\max\_\{\\\|\\boldsymbol\{\\beta\}\\\|=B\}\\Delta\_\{S\}\(\\boldsymbol\{\\beta\}\)\. At full frequency content the answer is uniform allocation, and Corollary[4](https://arxiv.org/html/2609.17926#Thmtheorem4)makes it exact\. ###### Proposition 5\(Optimal allocation\)\. LetS=\{1,…,Kmax\}S=\\\{1,\\dots,K\_\{\\mathrm\{max\}\}\\\}\. The uniform allocationβk=B/Kmax\\beta\_\{k\}=B/\\sqrt\{K\_\{\\mathrm\{max\}\}\}is optimal, attaining ΔS∗\(B\)=B\(Kmax\+12\)Kmax=Bp2\(p−1\),\\Delta^\{\*\}\_\{S\}\(B\)\\;=\\;\\frac\{B\\,\(K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}\)\}\{\\sqrt\{K\_\{\\mathrm\{max\}\}\}\}\\;=\\;\\frac\{B\\,p\}\{\\sqrt\{2\(p\-1\)\}\},\(16\)and the loss satisfies L=log\(1\+\(p−1\)exp\[−BKmax\+12Kmax\]\)≃\(p−1\)e−BKmaxL\\;=\\;\\log\\\!\\Big\(1\+\(p\-1\)\\,\\exp\\\!\\Big\[\-B\\,\\frac\{K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}\}\{\\sqrt\{K\_\{\\mathrm\{max\}\}\}\}\\Big\]\\Big\)\\;\\simeq\\;\(p\-1\)\\,e^\{\-B\\sqrt\{K\_\{\\mathrm\{max\}\}\}\}\(17\)for largeBB\. ###### Proof\. WriteMsk=1−cos\(ωks\)M\_\{sk\}=1\-\\cos\(\\omega\_\{k\}s\)fors≠0s\\neq 0andk∈Sk\\in S, so thatΔS\(𝜷\)=mins≠0\(M𝜷\)s\\Delta\_\{S\}\(\\boldsymbol\{\\beta\}\)=\\min\_\{s\\neq 0\}\(M\\boldsymbol\{\\beta\}\)\_\{s\}\. Uniform allocation attains the stated value\. Settingβk=β\\beta\_\{k\}=\\betafor allkkgives\(M𝜷\)s=β\(Kmax−DKmax\(s\)\)\(M\\boldsymbol\{\\beta\}\)\_\{s\}=\\beta\\big\(K\_\{\\mathrm\{max\}\}\-D\_\{K\_\{\\mathrm\{max\}\}\}\(s\)\\big\), which by Corollary[4](https://arxiv.org/html/2609.17926#Thmtheorem4)equalsβ\(Kmax\+12\)\\beta\(K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}\)for everys≠0s\\neq 0\. The minimum is therefore attained simultaneously at allp−1p\-1competitors, and substitutingβ=B/Kmax\\beta=B/\\sqrt\{K\_\{\\mathrm\{max\}\}\}withKmax\+12=p/2K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}=p/2gives Eq\. \([16](https://arxiv.org/html/2609.17926#A1.E16)\)\. No allocation does better\. Since∑s=0p−1cos\(ωks\)=0\\sum\_\{s=0\}^\{p\-1\}\\cos\(\\omega\_\{k\}s\)=0fork≢0k\\not\\equiv 0, we have∑s≠0cos\(ωks\)=−1\\sum\_\{s\\neq 0\}\\cos\(\\omega\_\{k\}s\)=\-1and hence∑s≠0Msk=\(p−1\)\+1=p\\sum\_\{s\\neq 0\}M\_\{sk\}=\(p\-1\)\+1=pfor everykk\. Summing the components ofM𝜷M\\boldsymbol\{\\beta\}therefore givesp∑kβkp\\sum\_\{k\}\\beta\_\{k\}, so the minimum component obeys ΔS\(𝜷\)≤pp−1∑k∈Sβk≤pp−1Kmax‖𝜷‖2=Bp2\(p−1\),\\Delta\_\{S\}\(\\boldsymbol\{\\beta\}\)\\;\\leq\\;\\frac\{p\}\{p\-1\}\\sum\_\{k\\in S\}\\beta\_\{k\}\\;\\leq\\;\\frac\{p\}\{p\-1\}\\sqrt\{K\_\{\\mathrm\{max\}\}\}\\,\\\|\\boldsymbol\{\\beta\}\\\|\_\{2\}\\;=\\;\\frac\{B\\,p\}\{\\sqrt\{2\(p\-1\)\}\},\(18\)the second step by Cauchy–Schwarz and the last byKmax=\(p−1\)/2K\_\{\\mathrm\{max\}\}=\(p\-1\)/2\. The bound coincides with the value attained above, and Cauchy–Schwarz is tight only for uniform𝜷\\boldsymbol\{\\beta\}\. SubstitutingΔ∗\\Delta^\{\*\}into Eq\. \([14](https://arxiv.org/html/2609.17926#A1.E14)\), with allp−1p\-1competitors contributing equally, gives Eq\. \([17](https://arxiv.org/html/2609.17926#A1.E17)\)\. ∎ SinceKmax=\(p−1\)/2K\_\{\\mathrm\{max\}\}=\(p\-1\)/2, Eq\. \([17](https://arxiv.org/html/2609.17926#A1.E17)\) says that at full frequency content the loss falls asexp\(−Bp/2\)\\exp\(\-B\\sqrt\{p/2\}\)up to the prefactor\. Holding a small target lossLLfixed as the group grows therefore requiresB≃2log\(\(p−1\)/L\)/pB\\simeq\\sqrt\{2\}\\,\\log\\\!\\big\(\(p\-1\)/L\\big\)/\\sqrt\{p\}, which decreases withpp\. The measured Frobenius norm ofW1W\_\{1\}does the opposite, growing from about4646atp=47p=47to about7777atp=113p=113\. The two are not the same quantity, sinceBBis the norm of the coefficient vector in the logit expansion while‖W1‖F\\\|W\_\{1\}\\\|\_\{F\}also carries the width and the parametrisation of the construction, but the comparison gives no support to the budget account and we do not claim it does\. For proper subsets the uniform allocation is not optimal, and the discrepancy is not small\. Solving Eq\. \([15](https://arxiv.org/html/2609.17926#A1.E15)\) numerically atp=47p=47by Nelder–Mead from twelve random starts gives worst case separations exceeding the uniform value by25%25\\%atK=4K=4and by28%28\\%atK=10K=10, while atK=KmaxK=K\_\{\\mathrm\{max\}\}the optimiser recovers the uniform allocation to within1\.6%1\.6\\%, as Proposition[5](https://arxiv.org/html/2609.17926#Thmtheorem5)requires\. The reason is that the constraint set loses its symmetry once frequencies are removed, so the coefficients1−cos\(ωks\)1\-\\cos\(\\omega\_\{k\}s\)no longer take a common value and the budget is better spent on the frequencies that separate the nearest competitor\. With that caveat recorded, the uniform allocation remains the natural reference because it is what a network with no preference among its retained modes would produce\. Under it,β=B/K\\beta=B/\\sqrt\{K\}and the separation is governed by the dominant competitor\. Numerically the dominant competitor iss=1s=1for everyKKbelowKmaxK\_\{\\mathrm\{max\}\}atp=47p=47, and expandingcos\(2πk/p\)=1−2π2k2/p2\+O\(k4/p4\)\\cos\(2\\pi k/p\)=1\-2\\pi^\{2\}k^\{2\}/p^\{2\}\+O\(k^\{4\}/p^\{4\}\)gives K−DK\(1\)=2π2p2∑k=1Kk2\+O\(K5p4\)=2π23K3p2\+O\(K2p2\)\+O\(K5p4\),K\-D\_\{K\}\(1\)\\;=\\;\\frac\{2\\pi^\{2\}\}\{p^\{2\}\}\\sum\_\{k=1\}^\{K\}k^\{2\}\+O\\\!\\left\(\\frac\{K^\{5\}\}\{p^\{4\}\}\\right\)\\;=\\;\\frac\{2\\pi^\{2\}\}\{3\}\\,\\frac\{K^\{3\}\}\{p^\{2\}\}\+O\\\!\\left\(\\frac\{K^\{2\}\}\{p^\{2\}\}\\right\)\+O\\\!\\left\(\\frac\{K^\{5\}\}\{p^\{4\}\}\\right\),\(19\)so that ΔK\(B\)≃2π23BK5/2p2\.\\Delta\_\{K\}\(B\)\\;\\simeq\\;\\frac\{2\\pi^\{2\}\}\{3\}\\,\\frac\{B\\,K^\{5/2\}\}\{p^\{2\}\}\.\(20\)The loss is then a stretched exponential in the mode count with exponent5/25/2and an explicitp−2p^\{\-2\}prefactor\. The range over which Eq\. \([20](https://arxiv.org/html/2609.17926#A1.E20)\) is usable is narrower than the derivation suggests, and it is worth being explicit about where it fails\. Table[2](https://arxiv.org/html/2609.17926#A1.T2)compares the expansion against the exactK−DK\(1\)K\-D\_\{K\}\(1\)atp=47p=47\. Agreement is good belowK=12K=12, which is roughly half ofKmaxK\_\{\\mathrm\{max\}\}, and deteriorates rapidly above it\. ByK=20K=20the expansion overstates the separation by35%35\\%, and atK=KmaxK=K\_\{\\mathrm\{max\}\}it gives36\.236\.2against the exact valueKmax\+12=23\.5K\_\{\\mathrm\{max\}\}\+\\tfrac\{1\}\{2\}=23\.5supplied by Corollary[4](https://arxiv.org/html/2609.17926#Thmtheorem4)\. The cubic growth cannot continue because the separation saturates once the kernel becomes a delta\. Table 2:Exact separation from the dominant competitor against its smallKKexpansion, atp=47p=47whereKmax=23K\_\{\\mathrm\{max\}\}=23\.Evaluating Eq\. \([14](https://arxiv.org/html/2609.17926#A1.E14)\) numerically without the expansion atp=47p=47and fittinglogL\\log LagainstKKgivesR2R^\{2\}between0\.9810\.981and0\.9950\.995, against0\.8260\.826to0\.8740\.874for a power law inKK, with decay constants0\.2200\.220,0\.3470\.347and0\.6060\.606atB=1\.5B=1\.5,2\.02\.0and3\.03\.0\. The functional form that the exact solution produces under a budget is therefore exponential rather than polynomial in the mode count, which is the same qualitative statement that Sec\.[3](https://arxiv.org/html/2609.17926#S3)makes about width\. ### A\.4What the calculation does not explain The agreement in functional form should not be mistaken for a derivation of the measurement\. Equation \([20](https://arxiv.org/html/2609.17926#A1.E20)\) carries an explicitp−2p^\{\-2\}and Eq\. \([17](https://arxiv.org/html/2609.17926#A1.E17)\) carries an explicitKmax\\sqrt\{K\_\{\\mathrm\{max\}\}\}, so both predict a rate that varies strongly with group order\. The measured rate does not, as Sec\.[6](https://arxiv.org/html/2609.17926#S6)establishes over a range of4\.94\.9inpp\. Three assumptions stand between the calculation and the experiment, and each fails in a way that is documented elsewhere in this paper\. The calculation treats the mode countKKas the independent variable, whereas the quantity that is swept is width, and Appendix[C](https://arxiv.org/html/2609.17926#A3)shows that the relation between them is an occupancy law rather than a proportionality\. The calculation holds the budget fixed whileKKvaries, whereas the measured Frobenius norm saturates with width, so that the amplitude per mode falls as width rises rather than staying constant\. The calculation assumes a single allocation of the budget across modes, whereas the numerical max\-min solution above shows that the optimal allocation depends onKKand departs from uniform by a quarter at small mode counts\. A derivation that respects all three would need to describe how gradient descent under weight decay distributes a saturating norm over a mode set that it is simultaneously enlarging\. We do not have that description, and we record its absence rather than presenting the calculation of this appendix as an explanation of Sec\.[6](https://arxiv.org/html/2609.17926#S6)\. ## Appendix BDegeneracy of the two nearest neighbour estimator on a group orbit Section[4](https://arxiv.org/html/2609.17926#S4)claims that the intrinsic dimension is not merely mismeasured on the algebraic solution but undefined\. This appendix supplies the argument in four parts\. The estimator is stated together with the assumption it rests on and validated on manifolds of known dimension\. The assumption is then shown to fail on any group orbit, in a way that is structural rather than statistical\. The geometry at full frequency content is identified as a regular simplex, and the sense in which partial frequency sets fall short of that is made precise\. Finally the behaviour under a symmetry breaking perturbation is characterised, and the absence of a scale free plateau is established numerically over a factor of thirty in the probe scale\. ### B\.1The estimator and its assumption LetX=\{x1,…,xN\}⊂ℝmX=\\\{x\_\{1\},\\dots,x\_\{N\}\\\}\\subset\\mathbb\{R\}^\{m\}and writer1\(i\)r\_\{1\}\(i\)andr2\(i\)r\_\{2\}\(i\)for the distances fromxix\_\{i\}to its first and second nearest neighbours\. The estimator of[Facco et al\. \(2017\)](https://arxiv.org/html/2609.17926#bib.bib6), which refines the maximum likelihood construction of[Levina and Bickel \(2004\)](https://arxiv.org/html/2609.17926#bib.bib14)and has been applied to the internal representations of trained networks by[Ansuini et al\. \(2019\)](https://arxiv.org/html/2609.17926#bib.bib1), works with the ratio μi=r2\(i\)r1\(i\)≥1\.\\mu\_\{i\}\\;=\\;\\frac\{r\_\{2\}\(i\)\}\{r\_\{1\}\(i\)\}\\;\\geq\\;1\.\(21\)The ratio \([21](https://arxiv.org/html/2609.17926#A2.E21)\) is useful because of the following observation\. Suppose the points are drawn independently from a density that is constant on the scale ofr2r\_\{2\}, on a manifold of dimensiondd\. The probability that the ball of radiusrraboutxix\_\{i\}contains no other point falls asexp\(−ρvdrd\)\\exp\(\-\\rho\\,v\_\{d\}r^\{d\}\)withvdv\_\{d\}the volume of the unitddball andρ\\rhothe local density, so the first two neighbour distances have the joint density of the first two order statistics of a Poisson process with intensityρvddrd−1\\rho\\,v\_\{d\}\\,d\\,r^\{d\-1\}\. Changing variables toμ\\muremoves bothρ\\rhoandvdv\_\{d\}and leaves Pr\(μi≤μ\)=1−μ−d,μ≥1,\\Pr\(\\mu\_\{i\}\\leq\\mu\)\\;=\\;1\-\\mu^\{\-d\},\\qquad\\mu\\geq 1,\(22\)a Pareto law whose shape parameter is the dimension\. The cancellation of the density is what makes the estimator local and free of a bandwidth choice\. Taking logarithms of the survival function gives a line through the origin, −log\(1−F\(μi\)\)=dlogμi,\-\\log\\big\(1\-F\(\\mu\_\{i\}\)\\big\)\\;=\\;d\\,\\log\\mu\_\{i\},\(23\)andddis recovered as the slope, in practice after discarding the largest ten percent of theμi\\mu\_\{i\}where the assumption of constant density is worst\. The estimator behaves as advertised on manifolds whose dimension is known\. Table[3](https://arxiv.org/html/2609.17926#A2.T3)reports our own calibration\. Accuracy is good up tod=5d=5and degrades slowly above it, which is the usual behaviour of nearest neighbour estimators as the dimension grows\. Small sample size is not the obstacle at the sizes relevant here\. A two torus sampled at only4747points, which is the number of token embeddings available atp=47p=47, returns2\.01±0\.342\.01\\pm 0\.34over forty draws, against2\.01±0\.052\.01\\pm 0\.05at20002000points\. Table 3:Calibration of the estimator\. Linear manifolds aredddimensional subspaces embedded inℝ40\\mathbb\{R\}^\{40\}, sampled at20002000points\. Torus figures are means over independent draws with the standard deviation across draws\.Equation \([22](https://arxiv.org/html/2609.17926#A2.E22)\) requires that the points be a sample from a density\. A group orbit is not\. ### B\.2Degeneracy on an orbit ###### of Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)\. Letx,y∈Xx,y\\in X\. Transitivity givesg∈Gg\\in Gwithy=gxy=gx, and becauseggacts by isometries the mapz↦gzz\\mapsto gzis a bijection ofXXpreserving all pairwise distances\. The multiset of distances fromyytoX∖\{y\}X\\setminus\\\{y\\\}is therefore the image underggof the multiset of distances fromxxtoX∖\{x\}X\\setminus\\\{x\\\}, and the two multisets coincide\. Their ordered elements agree in particular, sor1\(x\)=r1\(y\)r\_\{1\}\(x\)=r\_\{1\}\(y\)andr2\(x\)=r2\(y\)r\_\{2\}\(x\)=r\_\{2\}\(y\), whenceμx=μy\\mu\_\{x\}=\\mu\_\{y\}\. Sincexxandyywere arbitrary the empirical distribution ofμ\\muis a point mass at a single valueμ0\\mu\_\{0\}\. The regression \([23](https://arxiv.org/html/2609.17926#A2.E23)\) has sloped^=∑ixiyi/∑ixi2\\hat\{d\}=\\sum\_\{i\}x\_\{i\}y\_\{i\}/\\sum\_\{i\}x\_\{i\}^\{2\}withxi=logμix\_\{i\}=\\log\\mu\_\{i\}\. Everyxix\_\{i\}equalslogμ0\\log\\mu\_\{0\}, so the design carries no variation\. Ifμ0=1\\mu\_\{0\}=1thenxi=0x\_\{i\}=0for alliiand the ratio is0/00/0\. Ifμ0\>1\\mu\_\{0\}\>1then the empirical distribution function takes a different value at each index while the abscissa does not, so the fitted slope is set by the plotting convention used to defineFFrather than by the geometry\. ∎ For the solutions of interest the degenerate case is the first one, and this is stronger than the proposition requires\. ###### Lemma 6\(Coincident neighbour shells\)\. LetS⊆\{1,…,Kmax\}S\\subseteq\\\{1,\\dots,K\_\{\\mathrm\{max\}\}\\\}be nonempty and letΦS\\Phi\_\{S\}be the embedding of Eq\. \([2](https://arxiv.org/html/2609.17926#S2.E2)\)\. Thenr1=r2r\_\{1\}=r\_\{2\}at every point ofXSX\_\{S\}, soμ0=1\\mu\_\{0\}=1\. ###### Proof\. Squared distances depend only on the difference of the arguments, ‖ΦS\(n\)−ΦS\(n′\)‖2=2∑k∈S\(1−cosωk\(n−n′\)\)≡ψ\(n−n′\),\\\|\\Phi\_\{S\}\(n\)\-\\Phi\_\{S\}\(n^\{\\prime\}\)\\\|^\{2\}\\;=\\;2\\sum\_\{k\\in S\}\\big\(1\-\\cos\\omega\_\{k\}\(n\-n^\{\\prime\}\)\\big\)\\;\\equiv\\;\\psi\(n\-n^\{\\prime\}\),\(24\)andψ\\psiis even in its argument moduloppbecause the cosine is\. Henceψ\(t\)=ψ\(−t\)\\psi\(t\)=\\psi\(\-t\)for everyt≠0t\\neq 0, so the neighbours at offsets\+t\+tand−t\-tlie at equal distance and every distance from a given point occurs with even multiplicity\. The two smallest are therefore equal\. ∎ Numerically atp=47p=47the ratior2/r1r\_\{2\}/r\_\{1\}equals11to within10−1410^\{\-14\}for\|S\|=1\|S\|=1,\|S\|=5\|S\|=5and\|S\|=23\|S\|=23, with the common nearest neighbour distance taking the values0\.13360\.1336,0\.58600\.5860and1\.42951\.4295respectively on the normalised embedding\. That the hypotheses of Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)hold for the map \([2](https://arxiv.org/html/2609.17926#S2.E2)\) was established in Sec\.[2\.2](https://arxiv.org/html/2609.17926#S2.SS2)\. ### B\.3The simplex at full frequency content The inner product structure follows from character orthogonality, and it separates the complete frequency set from every proper subset\. ###### Lemma 7\(Inner products\)\. SetΦ^S=ΦS/\|S\|\\hat\{\\Phi\}\_\{S\}=\\Phi\_\{S\}/\\sqrt\{\|S\|\}\. For every nonemptySS, 1p−1∑n′≠n⟨Φ^S\(n\),Φ^S\(n′\)⟩=−1p−1,\\frac\{1\}\{p\-1\}\\sum\_\{n^\{\\prime\}\\neq n\}\\big\\langle\\hat\{\\Phi\}\_\{S\}\(n\),\\hat\{\\Phi\}\_\{S\}\(n^\{\\prime\}\)\\big\\rangle\\;=\\;\-\\frac\{1\}\{p\-1\},\(25\)independently ofnnand ofSS\. IfS=\{1,…,Kmax\}S=\\\{1,\\dots,K\_\{\\mathrm\{max\}\}\\\}the individual inner products all equal−1/\(p−1\)\-1/\(p\-1\), so the image is a regular simplex inℝp−1\\mathbb\{R\}^\{p\-1\}\. For proper subsets they do not\. ###### Proof\. Every point of the image has the same norm, since‖ΦS\(n\)‖2=\|S\|\\\|\\Phi\_\{S\}\(n\)\\\|^\{2\}=\|S\|for allnn\. The normalised inner product is therefore\|S\|−1∑k∈Scosωk\(n−n′\)\|S\|^\{\-1\}\\sum\_\{k\\in S\}\\cos\\omega\_\{k\}\(n\-n^\{\\prime\}\)\. Character orthogonality gives ∑t=0p−1cosωkt=0fork≢0,\\sum\_\{t=0\}^\{p\-1\}\\cos\\omega\_\{k\}t\\;=\\;0\\quad\\text\{for \}k\\not\\equiv 0,\(26\)and hence∑t≠0cosωkt=−1\\sum\_\{t\\neq 0\}\\cos\\omega\_\{k\}t=\-1\. Summing overn′≠nn^\{\\prime\}\\neq ntherefore yields−\|S\|∑k∈S−11=−1\-\|S\|^\{\-1\}\\sum\_\{k\\in S\}1=\-1, and Eq\. \([25](https://arxiv.org/html/2609.17926#A2.E25)\) follows on dividing byp−1p\-1\. For the complete set the sum∑k=1Kmaxcosωkt\\sum\_\{k=1\}^\{K\_\{\\mathrm\{max\}\}\}\\cos\\omega\_\{k\}tisDKmax\(t\)D\_\{K\_\{\\mathrm\{max\}\}\}\(t\), which Corollary[4](https://arxiv.org/html/2609.17926#Thmtheorem4)evaluates as−12\-\\tfrac\{1\}\{2\}for everyt≠0t\\neq 0\. Dividing byKmax=\(p−1\)/2K\_\{\\mathrm\{max\}\}=\(p\-1\)/2gives−1/\(p−1\)\-1/\(p\-1\)for every pair\. A set ofppunit vectors with all pairwise inner products equal to−1/\(p−1\)\-1/\(p\-1\)is a regular simplex\. ∎ Atp=47p=47the common value is−0\.021739\-0\.021739and the sample standard deviation across all10811081pairs is5×10−165\\times 10^\{\-16\}\. For\|S\|=1\|S\|=1and\|S\|=3\|S\|=3the mean is the same to six decimals, as Eq\. \([25](https://arxiv.org/html/2609.17926#A2.E25)\) requires, while the standard deviations are0\.6990\.699and0\.3850\.385\. The mean is thus uninformative about the geometry and only the vanishing of the spread identifies the simplex\. The distinction matters for the argument of Sec\.[4](https://arxiv.org/html/2609.17926#S4)\. A regular simplex spansp−1p\-1dimensions with all points mutually equidistant, so any notion of local dimension either returns the ambient value or fails to be defined\. Proper subsets are not simplices, but Lemma[6](https://arxiv.org/html/2609.17926#Thmtheorem6)shows they are still orbits withμ0=1\\mu\_\{0\}=1, so the estimator fails on them for the same reason\. The configuration identified by Lemma[7](https://arxiv.org/html/2609.17926#Thmtheorem7)is the simplex equiangular tight frame that appears as the terminal geometry in neural collapse\([Papyan et al\., 2020](https://arxiv.org/html/2609.17926#bib.bib23)\), where the class means of a trained classifier acquire pairwise cosine−1/\(p−1\)\-1/\(p\-1\)acrossppclasses\. The coincidence is worth flagging because it is not an instance of that phenomenon\. Neural collapse describes what optimisation drives the last layer towards, whereas Lemma[7](https://arxiv.org/html/2609.17926#Thmtheorem7)is a statement about the exact algebraic solution and holds whether or not any network finds it\. The two come apart on this task in a way recently made precise, since[Tan et al\. \(2026\)](https://arxiv.org/html/2609.17926#bib.bib28)argue that modular addition does not reach the simplex but settles into a low rank cyclic geometry, the simplex gaining only anO\(1\)O\(1\)advantage in cross entropy against aΘ\(p\)\\Theta\(p\)advantage for the cyclic solution under a weight decay surrogate\. Our own measurements sit between these descriptions, sinceKeffK\_\{\\mathrm\{eff\}\}approachesKmaxK\_\{\\mathrm\{max\}\}at large width, which is nearer the simplex than the rank two picture, plausibly because the quadratic activation used here makes many frequencies cheap to carry\. What matters for the present argument is that both configurations are orbits, so Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)applies to either and the intrinsic dimension is undefined in both cases\. ### B\.4Behaviour under a symmetry breaking perturbation An orbit returns no number, so the natural next question is what happens when the symmetry is broken slightly\. Adding independent Gaussian noise of scaleϵ\\epsilonto each coordinate produces a finite estimate\. It does not produce a stable one\. ###### Lemma 8\(No scale free plateau\)\. LetX=\{x1,…,xN\}⊂ℝmX=\\\{x\_\{1\},\\dots,x\_\{N\}\\\}\\subset\\mathbb\{R\}^\{m\}be a group orbit in the sense of Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1), with common nearest neighbour distanced1d\_\{1\}and first shell multiplicityn1=\|\{k:‖xi−xk‖=d1\}\|≥2n\_\{1\}=\\left\|\\\{k:\\\|x\_\{i\}\-x\_\{k\}\\\|=d\_\{1\}\\\}\\right\|\\geq 2, the same for everyii\. LetXϵ=\{xi\+ϵξi\}X\_\{\\epsilon\}=\\\{x\_\{i\}\+\\epsilon\\xi\_\{i\}\\\}withξi\\xi\_\{i\}independent standard Gaussian vectors inℝm\\mathbb\{R\}^\{m\}\. Then forϵ≪d1/m\\epsilon\\ll d\_\{1\}/\\sqrt\{m\}the two nearest neighbour estimate satisfies d^\(ϵ\)=Cϵ\(1\+O\(ϵ\)\),C=d12∑iGiyi∑iGi2,\\hat\{d\}\(\\epsilon\)\\;=\\;\\frac\{C\}\{\\epsilon\}\\,\\big\(1\+O\(\\epsilon\)\\big\),\\qquad C\\;=\\;\\frac\{d\_\{1\}\}\{\\sqrt\{2\}\}\\,\\frac\{\\sum\_\{i\}G\_\{i\}\\,y\_\{i\}\}\{\\sum\_\{i\}G\_\{i\}^\{2\}\},\(27\)whereyi=−log\(1−Fi\)y\_\{i\}=\-\\log\(1\-F\_\{i\}\)are the plotting positions used by the regression andGiG\_\{i\}are the first shell gap variables defined in the proof\. NeitherCCnorGiG\_\{i\}depends onϵ\\epsilon\. ###### Proof\. Writeuik=\(xi−xk\)/‖xi−xk‖u\_\{ik\}=\(x\_\{i\}\-x\_\{k\}\)/\\\|x\_\{i\}\-x\_\{k\}\\\|andδik=ξi−ξk\\delta\_\{ik\}=\\xi\_\{i\}\-\\xi\_\{k\}, so thatδik\\delta\_\{ik\}is Gaussian with covariance2Im2I\_\{m\}\. The perturbed squared distance is ‖xi−xk\+ϵδik‖2=d2\+2ϵd⟨uik,δik⟩\+ϵ2‖δik‖2,\\\|x\_\{i\}\-x\_\{k\}\+\\epsilon\\delta\_\{ik\}\\\|^\{2\}=d^\{2\}\+2\\epsilon d\\,\\langle u\_\{ik\},\\delta\_\{ik\}\\rangle\+\\epsilon^\{2\}\\\|\\delta\_\{ik\}\\\|^\{2\},\(28\)withd=‖xi−xk‖d=\\\|x\_\{i\}\-x\_\{k\}\\\|\. Taking the square root and expanding, rik\(ϵ\)=d\+ϵ2Zik\+ϵ22d\(‖δik‖2−⟨uik,δik⟩2\)\+O\(ϵ3\),r\_\{ik\}\(\\epsilon\)\\;=\\;d\+\\epsilon\\sqrt\{2\}\\,Z\_\{ik\}\+\\frac\{\\epsilon^\{2\}\}\{2d\}\\Big\(\\\|\\delta\_\{ik\}\\\|^\{2\}\-\\langle u\_\{ik\},\\delta\_\{ik\}\\rangle^\{2\}\\Big\)\+O\(\\epsilon^\{3\}\),\(29\)whereZik=⟨uik,δik⟩/2Z\_\{ik\}=\\langle u\_\{ik\},\\delta\_\{ik\}\\rangle/\\sqrt\{2\}is standard Gaussian\. The second order term has meanϵ2\(m−1\)/d\\epsilon^\{2\}\(m\-1\)/d, so the first order term dominates wheneverϵm≪d\\epsilon\\sqrt\{m\}\\ll d, which is the stated hypothesis\. We now identify which pairs supplyr1r\_\{1\}andr2r\_\{2\}\. Ifn1<N−1n\_\{1\}<N\-1letd2\>d1d\_\{2\}\>d\_\{1\}be the next distance in the common multiset, which exists and is the same for everyiiby Proposition[1](https://arxiv.org/html/2609.17926#Thmtheorem1)\. The fluctuations in Eq\. \([29](https://arxiv.org/html/2609.17926#A2.E29)\) areO\(ϵ\)O\(\\epsilon\)whiled2−d1d\_\{2\}\-d\_\{1\}is fixed, so the eventEϵE\_\{\\epsilon\}that every perturbed first shell distance is smaller than every perturbed second shell distance has probability at least1−N2exp\(−\(d2−d1\)2/8ϵ2\)1\-N^\{2\}\\exp\\\!\\big\(\-\(d\_\{2\}\-d\_\{1\}\)^\{2\}/8\\epsilon^\{2\}\\big\)by a Gaussian tail bound and a union over pairs\. Since the noise is unbounded no deterministic threshold inϵ\\epsiloncan forceEϵE\_\{\\epsilon\}, butℙ\(Eϵ\)→1\\mathbb\{P\}\(E\_\{\\epsilon\}\)\\to 1asϵ→0\\epsilon\\to 0at a rate fixed byXX, and the statement below is made onEϵE\_\{\\epsilon\}\. Ifn1=N−1n\_\{1\}=N\-1, which is the case of a regular simplex, there is no second shell and the statement is vacuous\. In both casesr1\(i\)r\_\{1\}\(i\)andr2\(i\)r\_\{2\}\(i\)are the two smallest among then1n\_\{1\}first shell distances\. LetZi\(1\)≤Zi\(2\)Z\_\{i\(1\)\}\\leq Z\_\{i\(2\)\}denote the two smallest of\{Zik:‖xi−xk‖=d1\}\\\{Z\_\{ik\}:\\\|x\_\{i\}\-x\_\{k\}\\\|=d\_\{1\}\\\}and set Gi=Zi\(2\)−Zi\(1\)≥0\.G\_\{i\}\\;=\\;Z\_\{i\(2\)\}\-Z\_\{i\(1\)\}\\;\\geq\\;0\.\(30\)The law ofGiG\_\{i\}is determined byn1n\_\{1\}and by the Gram matrix of the first shell directionsuiku\_\{ik\}, which fixes the correlationsCov\(Zik,Zil\)=⟨uik,uil⟩/2\\mathrm\{Cov\}\(Z\_\{ik\},Z\_\{il\}\)=\\langle u\_\{ik\},u\_\{il\}\\rangle/2\. It does not involveϵ\\epsilon\. Substituting into Eq\. \([29](https://arxiv.org/html/2609.17926#A2.E29)\), μi=r2\(i\)r1\(i\)=d1\+ϵ2Zi\(2\)\+O\(ϵ2\)d1\+ϵ2Zi\(1\)\+O\(ϵ2\)=1\+2d1ϵGi\+O\(ϵ2\),\\mu\_\{i\}\\;=\\;\\frac\{r\_\{2\}\(i\)\}\{r\_\{1\}\(i\)\}\\;=\\;\\frac\{d\_\{1\}\+\\epsilon\\sqrt\{2\}\\,Z\_\{i\(2\)\}\+O\(\\epsilon^\{2\}\)\}\{d\_\{1\}\+\\epsilon\\sqrt\{2\}\\,Z\_\{i\(1\)\}\+O\(\\epsilon^\{2\}\)\}\\;=\\;1\+\\frac\{\\sqrt\{2\}\}\{d\_\{1\}\}\\,\\epsilon\\,G\_\{i\}\+O\(\\epsilon^\{2\}\),\(31\)so withκ=2/d1\\kappa=\\sqrt\{2\}/d\_\{1\}we haveμi−1=ϵκGi\\mu\_\{i\}\-1=\\epsilon\\kappa G\_\{i\}to leading order and xi≡logμi=ϵκGi\(1\+O\(ϵ\)\)\.x\_\{i\}\\;\\equiv\\;\\log\\mu\_\{i\}\\;=\\;\\epsilon\\kappa G\_\{i\}\\,\\big\(1\+O\(\\epsilon\)\\big\)\.\(32\) The response variable does not move\. The estimator sorts theμi\\mu\_\{i\}and assignsFi=i/\(N\+1\)F\_\{i\}=i/\(N\+1\)by rank, soyi=−log\(1−Fi\)y\_\{i\}=\-\\log\(1\-F\_\{i\}\)depends on the data only through the ordering of theμi\\mu\_\{i\}\. By Eq\. \([32](https://arxiv.org/html/2609.17926#A2.E32)\) that ordering coincides with the ordering of theGiG\_\{i\}, which carries noϵ\\epsilon\. The same applies to the discarding of the largest ten percent of theμi\\mu\_\{i\}, since that too is a rank operation\. The regression \([23](https://arxiv.org/html/2609.17926#A2.E23)\) has no intercept, so its slope is d^\(ϵ\)=∑ixiyi∑ixi2=ϵκ∑iGiyiϵ2κ2∑iGi2\(1\+O\(ϵ\)\)=1ϵκ∑iGiyi∑iGi2\(1\+O\(ϵ\)\),\\hat\{d\}\(\\epsilon\)\\;=\\;\\frac\{\\sum\_\{i\}x\_\{i\}y\_\{i\}\}\{\\sum\_\{i\}x\_\{i\}^\{2\}\}\\;=\\;\\frac\{\\epsilon\\kappa\\sum\_\{i\}G\_\{i\}y\_\{i\}\}\{\\epsilon^\{2\}\\kappa^\{2\}\\sum\_\{i\}G\_\{i\}^\{2\}\}\\big\(1\+O\(\\epsilon\)\\big\)\\;=\\;\\frac\{1\}\{\\epsilon\\kappa\}\\,\\frac\{\\sum\_\{i\}G\_\{i\}y\_\{i\}\}\{\\sum\_\{i\}G\_\{i\}^\{2\}\}\\big\(1\+O\(\\epsilon\)\\big\),\(33\)and substitutingκ=2/d1\\kappa=\\sqrt\{2\}/d\_\{1\}gives Eq\. \([27](https://arxiv.org/html/2609.17926#A2.E27)\)\. Both factors on the right are independent ofϵ\\epsilon, the first by construction and the second becauseGGandyyare\. ∎ Two consequences of Lemma[8](https://arxiv.org/html/2609.17926#Thmtheorem8)are testable separately\. The constant is proportional tod1d\_\{1\}at fixed first shell structure, so rescaling the embedding by a factorssshould rescaleCCby the same factor\. Multiplying the single frequency solution atp=47p=47bys∈\{0\.5,1,2,4\}s\\in\\\{0\.5,1,2,4\\\}givesC/sC/sequal to0\.064940\.06494in all four cases, and doing the same to the full frequency solution gives2\.91382\.9138,2\.91762\.9176,2\.91952\.9195and2\.92052\.9205at fixed noise realisation, a drift of two parts in a thousand across a factor of eight in scale\. The remaining factor depends only on the first shell, whose multiplicity isn1=2n\_\{1\}=2for a single frequency, where the neighbours sit at offsets±1\\pm 1, andn1=p−1=46n\_\{1\}=p\-1=46for the simplex, where every other point is a nearest neighbour\. The measured ratiosC/d1C/d\_\{1\}are0\.4860\.486and2\.0412\.041accordingly\. Theϵ\\epsilondependence of Eq\. \([27](https://arxiv.org/html/2609.17926#A2.E27)\) is exact to the precision of the measurement\. Table[4](https://arxiv.org/html/2609.17926#A2.T4)listsd^ϵ\\hat\{d\}\\,\\epsilonacross a factor of thirty inϵ\\epsilonand it is constant to about one percent, at0\.0650\.065for a single frequency and3\.023\.02for the complete set\. Table 4:Estimated dimension times probe scale, atp=47p=47\. Constancy of the product is the absence of a plateau\. Entries are means over six independent noise draws, four for the final row, which is a genuine two torus sampled at15001500points subjected to the same treatment\.In raw terms the estimate on the single frequency solution falls from6565atϵ=10−3\\epsilon=10^\{\-3\}to6\.56\.5at10−210^\{\-2\}and1\.761\.76at10−110^\{\-1\}, and on the full twenty three frequency solution it reaches about30003000atϵ=10−3\\epsilon=10^\{\-3\}, while a genuine continuous torus given the same treatment returns2\.072\.07,2\.322\.32and3\.753\.75on a single draw at those three scales\. The contrast with the torus is the point of the table\. There the estimate itself is stable, moving from1\.981\.98to2\.202\.20across the same two decades before the noise begins to fill the ambient space and inflate it\. An object with a dimension reports the same dimension at every scale below its own curvature\. An orbit reports whatever the probe scale dictates\. Trained networks reproduce the pathology rather than escaping it\. Atp=47p=47and width6464, with held out accuracy exactly unity, the estimator applied to the learned token embeddings returns6868against an ambient dimension of6464, and the estimate rises with width instead of converging\. ### B\.5Consequence for the geometric derivation The relationαN=4/d\\alpha\_\{N\}=4/dtakes a manifold dimension as input and returns a scaling exponent\. On this task that input does not exist\. The failure is not that the manifold is of high dimension, nor that the estimator is imprecise at the available sample size, both of which would leave the derivation meaningful\. It is that the object whose dimension is sought is a group orbit, on which the statistic the estimator uses is constant by symmetry and the regression that defines the estimate has no design variation\. This is a stronger statement than a failed prediction\. A theory that predicts the wrong exponent can be corrected\. A theory whose input is undefined on a class of problems has to be restricted to the complement of that class, and the results of Sec\.[3](https://arxiv.org/html/2609.17926#S3)indicate that algebraic tasks lie outside it\. ## Appendix CAggregate spectral occupancy as an occupancy process Section[5](https://arxiv.org/html/2609.17926#S5)reports that the aggregate participation ratioKeffK\_\{\\mathrm\{eff\}\}correlates with the loss within a group order and fails to organise it across group orders\. This appendix explains why by deriving whatKeffK\_\{\\mathrm\{eff\}\}would be if the network carried no information at all beyond the number of units and the number of available frequencies\. The prediction has no free parameters and accounts for ninety percent of the variance in the measurement, which is the sense in whichKeffK\_\{\\mathrm\{eff\}\}is a sampling statistic rather than a capacity variable\. ### C\.1Definition of the participation ratios WritingE∈ℝp×hE\\in\\mathbb\{R\}^\{p\\times h\}for the block ofW1W\_\{1\}embedding the first argument andPk=∑j\|E^kj\|2P\_\{k\}=\\sum\_\{j\}\|\\hat\{E\}\_\{kj\}\|^\{2\}for the power at frequencykksummed over hidden units, with the constant mode removed and the rest normalised, the aggregate participation ratio is Keff=\(∑kPk2\)−1\.K\_\{\\mathrm\{eff\}\}\\;=\\;\\Big\(\\textstyle\\sum\_\{k\}P\_\{k\}^\{2\}\\Big\)^\{\-1\}\.\(34\)The same construction applied within a single hidden unit gives a per neuron participation ratio, whose energy weighted mean over units is what is reported alongsideKeffK\_\{\\mathrm\{eff\}\}throughout\. ### C\.2The model Take as given that each hidden unit carries a single irreducible representation\. This is proved for gradient flow on this architecture by[He et al\. \(2026b\)](https://arxiv.org/html/2609.17926#bib.bib10), who also establish uniform diversification across the nontrivial representations in the abelian case, and it is confirmed here by the per neuron participation ratios reported in Sec\.[5](https://arxiv.org/html/2609.17926#S5)for configurations that solve the task\. Model the assignment ashhunits drawing independently and uniformly from theKmaxK\_\{\\mathrm\{max\}\}available frequencies, and writemkm\_\{k\}for the number of units landing on frequencykk, so that∑kmk=h\\sum\_\{k\}m\_\{k\}=hand the vector\(mk\)\(m\_\{k\}\)is multinomial with equal cell probabilities\. Two quantities follow\. The number of frequencies represented at all is the number of occupied cells, and the participation ratio is a softer count that weights each frequency by the power it carries\. ### C\.3Distinct frequencies ###### Lemma 9\(Occupied cells\)\. Under the model, 𝔼\[Kdistinct\]=Kmax\[1−\(1−1Kmax\)h\]≃Kmax\(1−e−h/Kmax\),\\mathbb\{E\}\\big\[K\_\{\\mathrm\{distinct\}\}\\big\]\\;=\\;K\_\{\\mathrm\{max\}\}\\left\[1\-\\left\(1\-\\frac\{1\}\{K\_\{\\mathrm\{max\}\}\}\\right\)^\{h\}\\right\]\\;\\simeq\\;K\_\{\\mathrm\{max\}\}\\left\(1\-e^\{\-h/K\_\{\\mathrm\{max\}\}\}\\right\),\(35\)the approximation holding forKmax≫1K\_\{\\mathrm\{max\}\}\\gg 1\. ###### Proof\. Frequencykkis unoccupied exactly when allhhdraws avoid it, which happens with probability\(1−1/Kmax\)h\(1\-1/K\_\{\\mathrm\{max\}\}\)^\{h\}\. Linearity of expectation over theKmaxK\_\{\\mathrm\{max\}\}indicator variables gives the first equality, and\(1−1/Kmax\)h=exp\[hlog\(1−1/Kmax\)\]→e−h/Kmax\(1\-1/K\_\{\\mathrm\{max\}\}\)^\{h\}=\\exp\\\!\\big\[h\\log\(1\-1/K\_\{\\mathrm\{max\}\}\)\\big\]\\to e^\{\-h/K\_\{\\mathrm\{max\}\}\}gives the second\. ∎ Equation \([35](https://arxiv.org/html/2609.17926#A3.E35)\) of Lemma[9](https://arxiv.org/html/2609.17926#Thmtheorem9)is the classical occupancy count, and it saturates atKmaxK\_\{\\mathrm\{max\}\}oncehhexceeds a few multiples ofKmaxK\_\{\\mathrm\{max\}\}\. It is not, however, the quantity we measure\. ### C\.4The participation ratio If every unit carries comparable spectral energy then the aggregate power at frequencykkis proportional tomkm\_\{k\}, and the participation ratio \([34](https://arxiv.org/html/2609.17926#A3.E34)\) becomes Keff=\(∑kmk\)2∑kmk2=h2∑kmk2\.K\_\{\\mathrm\{eff\}\}\\;=\\;\\frac\{\\big\(\\sum\_\{k\}m\_\{k\}\\big\)^\{2\}\}\{\\sum\_\{k\}m\_\{k\}^\{2\}\}\\;=\\;\\frac\{h^\{2\}\}\{\\sum\_\{k\}m\_\{k\}^\{2\}\}\.\(36\)The denominator is the only random quantity, and its expectation is available in closed form\. ###### Proposition 10\(Expected participation ratio\)\. Under the model, 𝔼\[∑kmk2\]=h\(1−1Kmax\)\+h2Kmax,\\mathbb\{E\}\\Big\[\\sum\_\{k\}m\_\{k\}^\{2\}\\Big\]\\;=\\;h\\left\(1\-\\frac\{1\}\{K\_\{\\mathrm\{max\}\}\}\\right\)\+\\frac\{h^\{2\}\}\{K\_\{\\mathrm\{max\}\}\},\(37\)so that to first order Keff≃hKmaxh\+Kmax−1\.K\_\{\\mathrm\{eff\}\}\\;\\simeq\\;\\frac\{h\\,K\_\{\\mathrm\{max\}\}\}\{h\+K\_\{\\mathrm\{max\}\}\-1\}\.\(38\)The estimate is a lower bound on𝔼\[Keff\]\\mathbb\{E\}\[K\_\{\\mathrm\{eff\}\}\]\. ###### Proof\. Eachmkm\_\{k\}is binomial withhhtrials and success probability1/Kmax1/K\_\{\\mathrm\{max\}\}, so𝔼\[mk\]=h/Kmax\\mathbb\{E\}\[m\_\{k\}\]=h/K\_\{\\mathrm\{max\}\}andVar\(mk\)=hKmax−1\(1−Kmax−1\)\\mathrm\{Var\}\(m\_\{k\}\)=hK\_\{\\mathrm\{max\}\}^\{\-1\}\(1\-K\_\{\\mathrm\{max\}\}^\{\-1\}\)\. Hence𝔼\[mk2\]=hKmax−1\(1−Kmax−1\)\+h2Kmax−2\\mathbb\{E\}\[m\_\{k\}^\{2\}\]=hK\_\{\\mathrm\{max\}\}^\{\-1\}\(1\-K\_\{\\mathrm\{max\}\}^\{\-1\}\)\+h^\{2\}K\_\{\\mathrm\{max\}\}^\{\-2\}, and summing over theKmaxK\_\{\\mathrm\{max\}\}frequencies gives the stated expectation\. Substituting into Eq\. \([36](https://arxiv.org/html/2609.17926#A3.E36)\) and replacing the denominator by its mean gives Eq\. \([38](https://arxiv.org/html/2609.17926#A3.E38)\)\. Sincex↦1/xx\\mapsto 1/xis convex, Jensen’s inequality gives𝔼\[h2/∑kmk2\]≥h2/𝔼\[∑kmk2\]\\mathbb\{E\}\[h^\{2\}/\\sum\_\{k\}m\_\{k\}^\{2\}\]\\geq h^\{2\}/\\mathbb\{E\}\[\\sum\_\{k\}m\_\{k\}^\{2\}\], so the substitution understates the expectation\. ∎ Equation \([38](https://arxiv.org/html/2609.17926#A3.E38)\) is a harmonic combination of the two counts that bound the answer\. It reduces tohhwhenh≪Kmaxh\\ll K\_\{\\mathrm\{max\}\}, which is the regime in which every unit lands on its own frequency, and toKmaxK\_\{\\mathrm\{max\}\}whenh≫Kmaxh\\gg K\_\{\\mathrm\{max\}\}, which is the regime in which every frequency is occupied\. Nothing in it refers to the task\. ### C\.5Comparison with measurement Table[5](https://arxiv.org/html/2609.17926#A3.T5)sets Eq\. \([38](https://arxiv.org/html/2609.17926#A3.E38)\) against the measuredKeffK\_\{\\mathrm\{eff\}\}at weight decay11\. Across the full grid of eight group orders and fourteen widths the model accounts forR2=0\.907R^\{2\}=0\.907of the variance inKeffK\_\{\\mathrm\{eff\}\}, with a mean ratio of measurement to prediction of1\.131\.13and a standard deviation of0\.140\.14\. Table 5:Measured aggregate participation ratio against Eq\. \([38](https://arxiv.org/html/2609.17926#A3.E38)\), at weight decay11\.The residual is systematic, the prediction sitting about eleven percent below the measurement, and its sign is predicted by Proposition[10](https://arxiv.org/html/2609.17926#Thmtheorem10)\. Two further effects work in the same direction\. Units do not carry exactly equal spectral energy, which spreads the power across occupied frequencies more evenly than the multiplicities alone would, and weight decay removes power from frequencies that contribute little, which trims the tail of the multiplicity distribution during training\([He et al\., 2026a](https://arxiv.org/html/2609.17926#bib.bib9)\)\. Neither is large enough to change the character of the agreement\. The agreement is also not exact in the sense that would be required to call the model correct\. It is close enough to establish the negative claim, which is that the growth ofKeffK\_\{\\mathrm\{eff\}\}with width needs no explanation in terms of what the task demands\. ### C\.6WhyKeffK\_\{\\mathrm\{eff\}\}cannot collapse the data The consequence for interpretation follows from Eq\. \([38](https://arxiv.org/html/2609.17926#A3.E38)\) directly\. SinceKeffK\_\{\\mathrm\{eff\}\}is a deterministic function ofhhandKmaxK\_\{\\mathrm\{max\}\}up to the residual above, it carries essentially the information already present in those two variables and no more\. Inverting the relation gives h≃Keff\(Kmax−1\)Kmax−Keff,h\\;\\simeq\\;\\frac\{K\_\{\\mathrm\{eff\}\}\(K\_\{\\mathrm\{max\}\}\-1\)\}\{K\_\{\\mathrm\{max\}\}\-K\_\{\\mathrm\{eff\}\}\},\(39\)so two configurations at different group orders that share a value ofKeffK\_\{\\mathrm\{eff\}\}have different widths, by a factor that grows as the group orders diverge\. The loss depends on width, so those two configurations have different losses\. This is what produces the spread of a factor of25002500atKeff≈20K\_\{\\mathrm\{eff\}\}\\approx 20reported in Sec\.[5](https://arxiv.org/html/2609.17926#S5), and it would occur under the model even if the network’s mode content were irrelevant to its performance\. A regression of the loss onKeffK\_\{\\mathrm\{eff\}\}within a single group order therefore cannot distinguish a capacity mechanism from a sampling one, since the two predict the same monotone relation\. Only comparison across group orders separates them, and there the sampling account is what survives\. ## Appendix DMeasurements behind the mode count This appendix gives the measurements summarised in Sec\.[5](https://arxiv.org/html/2609.17926#S5)\. The network solves the task by carrying Fourier modes, so the number of modes it carries should play the role that manifold dimension plays elsewhere\. Recent work proves that each neuron converges to a single irreducible representation under gradient flow\([He et al\., 2026b](https://arxiv.org/html/2609.17926#bib.bib10);[He et al\., 2026a](https://arxiv.org/html/2609.17926#bib.bib9)\), which makes the count of distinct represented modes a well defined quantity\. It is also the quantity we expected to organise the data, and reporting that it does not is the purpose of this appendix\. Within a single group order the correlation is convincing\. Atp=47p=47the loss falls monotonically asKeffK\_\{\\mathrm\{eff\}\}rises from6\.36\.3to22\.422\.4, and a fit oflogL\\log Lagainstlog\(Kmax−Keff\)\\log\(K\_\{\\mathrm\{max\}\}\-K\_\{\\mathrm\{eff\}\}\)returnsR2=0\.90R^\{2\}=0\.90\. Across group orders it collapses\. AtKeff≈20K\_\{\\mathrm\{eff\}\}\\approx 20the loss ranges from1\.9×10−41\.9\\times 10^\{\-4\}atp=41p=41to0\.4880\.488atp=113p=113, a spread of a factor of25002500at fixed value of the supposed controlling variable\. Pooled over all configurations,logL\\log Lregressed onKeffK\_\{\\mathrm\{eff\}\}returnsR2=0\.590R^\{2\}=0\.590whilelogL\\log Lregressed on width alone returnsR2=0\.917R^\{2\}=0\.917\. The normalised variableKeff/KmaxK\_\{\\mathrm\{eff\}\}/K\_\{\\mathrm\{max\}\}is worse still atR2=0\.363R^\{2\}=0\.363\. The within group correlation was therefore an artefact of the fact thatKeffK\_\{\\mathrm\{eff\}\}andhhincrease together at fixedpp, for a reason that has nothing to do with capacity\. Among the251251configurations that reach test accuracy above0\.990\.99the energy weighted per neuron participation ratio lies between1\.041\.04and1\.761\.76with median1\.1391\.139, confirming that neurons are rank one in frequency\. An occupancy model in which each neuron selects one ofKmaxK\_\{\\mathrm\{max\}\}frequencies uniformly then predictsKeff=hKmax/\(h\+Kmax−1\)K\_\{\\mathrm\{eff\}\}=hK\_\{\\mathrm\{max\}\}/\(h\+K\_\{\\mathrm\{max\}\}\-1\)with no free parameters, accounting forR2=0\.907R^\{2\}=0\.907of the measured variance; Appendix[C](https://arxiv.org/html/2609.17926#A3)derives the prediction and gives the comparison in full\. Aggregate spectral occupancy grows with width because more draws cover more bins, and a quantity that grows for that reason cannot be read as a measure of what the network needs\. The rank one structure itself turns out to be conditional, which is worth recording because it is usually stated as a property of the architecture\. Across the full grid the per neuron participation ratio reaches6\.616\.61, and every one of the130130configurations above33occurs at weight decay0\.250\.25or below and at test accuracy at most0\.9030\.903\. Within a fixed weight decay the rank correlation between accuracy and participation ratio is between−0\.40\-0\.40and−0\.64\-0\.64for weight decay at most0\.50\.5, and is indistinguishable from zero at weight decay22and above\. Figure[2](https://arxiv.org/html/2609.17926#S4.F2)\(b\) shows the scatter\. Convergence of each neuron to a single irreducible representation is proved for gradient flow in Ref\.\([He et al\., 2026b](https://arxiv.org/html/2609.17926#bib.bib10)\), and what these measurements add is that the property is something regularisation imposes rather than something the architecture supplies\. A memorising network carries several frequencies per unit\. The result is a delimitation rather than a disagreement with the mechanistic literature\.[He et al\. \(2026a\)](https://arxiv.org/html/2609.17926#bib.bib9)characterise grokking on this task as a three stage process driven by the competition between loss minimisation and weight decay, with a diversification condition on the frequencies a solution carries\. What is measured here is not whether such conditions hold but whether the number of frequencies, once diversified, predicts the loss across problem sizes\. It does not, and Appendix[C](https://arxiv.org/html/2609.17926#A3)identifies the occupancy process that produces the appearance that it does\. ## Appendix EFitting protocol and identifiability This appendix records the rules used to fit Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\), the triage that removes degenerate fits, and the checks that establish which conclusions survive the choices involved\. Three of those choices could plausibly have been made differently, namely the space in which the fit is performed, the threshold at which a fit is rejected, and the value at which the exponent is held\. Each is examined below, and the exponent turns out to be the only one that matters\. ### E\.1Protocol Full batch AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.17926#bib.bib18)\)is used throughout, with learning rate3×10−33\\times 10^\{\-3\}, moments\(0\.9,0\.98\)\(0\.9,0\.98\), numerical stabiliser10−810^\{\-8\}and80008000steps\. Weights are initialised from a centred normal with variance set by fan in\. The three seeds are trained simultaneously by carrying the seed index as a leading tensor dimension, which leaves the models mathematically independent because AdamW acts elementwise\. For each pair\(p,λ\)\(p,\\lambda\)the mean test loss over three seeds is computed at each of the fourteen widths, and Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\) is fitted tologL\\log Lby nonlinear least squares withL∞L\_\{\\infty\},logA\\log Aandccfree andα\\alphafixed unless stated otherwise\. Fitting in the logarithm rather than in the loss is not cosmetic\. The measured range spans four and a half decades, so least squares onLLitself would be determined almost entirely by the two or three smallest widths and would carry no information about the floor, which is precisely the parameter that stabilises the rate\. Working inlogL\\log Lweights each decade equally, which is the appropriate choice when the quantity of interest is a rate of decay rather than an absolute level\. Two independent criteria then decide whether a fit enters the analysis\. The first concerns the floor\. A fit is rejected when the recoveredL∞L\_\{\\infty\}exceeds three times the smallest observed loss\. The rationale is that a resolved floor places the largest widths on the asymptote, soL∞≈minhL\(h\)L\_\{\\infty\}\\approx\\min\_\{h\}L\(h\)is the signature of success rather than of failure\. A criterion demandingL∞<minhL\(h\)L\_\{\\infty\}<\\min\_\{h\}L\(h\)strictly rejects ten of the twenty cells that solve the task, including several withR2R^\{2\}above0\.980\.98, and is therefore too aggressive\. The failures the criterion is meant to catch place the floor one to five orders of magnitude above the data, which occurs when the run never solves the task and the fit interprets the memorisation plateau as an upper asymptote\. When the recovered floor instead falls below10−310^\{\-3\}of the smallest observed loss it is not identified by the data at all, and the two parameter fit is reported in its place\. The second concerns learning\. A cell is excluded when held out accuracy at the largest width is below0\.990\.99\. This is not implied by the floor test and does not imply it\. A run at weight decay44attainsR2=0\.996R^\{2\}=0\.996with a well behaved two parameter fit while barely learning, since a curve that is flat and straight is easy to fit and says nothing\. Goodness of fit alone is therefore not evidence that a rate is meaningful\. In practice the second criterion does almost all of the work\. Of the twenty eight cells in the grid, eight are excluded, and seven of those fail on accuracy alone\. Only one cell is caught by the floor test, and that cell fails the accuracy test as well\. The floor criterion is thus close to inert on this data set, which is worth stating because it is the more arbitrary of the two\. Of the twenty cells that remain, Table[6](https://arxiv.org/html/2609.17926#A5.T6)lists eighteen; weight decay0\.10\.1solves the task only atp=71p=71andp=113p=113, and is omitted\. Table 6:Fitted rateccatα=1\\alpha=1for four group orders and five weight decay strengths, with the coefficient of variation ofccacross group order taken overp≥47p\\geq 47alone\. Entries are omitted where the run failed to reach test accuracy0\.990\.99at the largest width\. ### E\.2Runs excluded by the validity criterion The spurious maximum reported in Sec\.[6](https://arxiv.org/html/2609.17926#S6)came from two runs whose final accuracies were0\.180\.18and0\.850\.85\. For these the three parameter fit places the recovered floor above the bulk of the data rather than below it\. The same criterion excludesp=23p=23at weight decay0\.10\.1, which reaches an accuracy of0\.0890\.089wherep=113p=113already reaches unity, and it is the smallest training set in the study,264264pairs against11041104atp=47p=47and63846384atp=113p=113, that places it outside the regime in which the factorisation is claimed\. ### E\.3Sensitivity to the rejection threshold Because the threshold of three is a judgement, we vary it\. Table[7](https://arxiv.org/html/2609.17926#A5.T7)reports the number of cells retained and the resulting factorisation statistics\. Nothing changes above three, since no cell has a recovered floor between three and ten times the smallest observed loss\. At two the threshold removes three further cells and moves the group order spread by half a percent\. Table 7:Effect of the floor rejection threshold on the factorisation of Sec\.[6](https://arxiv.org/html/2609.17926#S6)\. The spread inggis taken overp≥47p\\geq 47\. ### E\.4The matched three parameter comparison Equation \([5](https://arxiv.org/html/2609.17926#S3.E5)\) carries three parameters and the power law of Sec\.[3](https://arxiv.org/html/2609.17926#S3)carries two, so the two are compared here on equal terms\. The alternative is Eq\. \([6](https://arxiv.org/html/2609.17926#S3.E6)\), fitted tologL\\log Lby nonlinear least squares withL∞L\_\{\\infty\},logA\\log Aandγ\\gammafree, over the same width range, with the same rejection rules and the same three seed aggregation\. SinceN=3phN=3ph, at fixed group order a power law inNNand a power law inhhdiffer only by a multiplicative constant absorbed intoAA; the fittedγ\\gammais the same either way, so nothing here depends on which of the two is taken as the abscissa\. Both families have three parameters, so the penalty term cancels andΔAIC\\Delta\\mathrm\{AIC\}reduces tonlog\(RSSpow/RSSexp\)n\\log\(\\mathrm\{RSS\}\_\{\\mathrm\{pow\}\}/\\mathrm\{RSS\}\_\{\\mathrm\{exp\}\}\)\. Table[8](https://arxiv.org/html/2609.17926#A5.T8)gives the outcome atα=1\\alpha=1\. The exponential is preferred at every group order, by between25\.525\.5and40\.940\.9units of AIC, and the same holds at every fixedα\\alphain\{0\.75,1,1\.25,1\.5,1\.75,2\}\\\{0\.75,1,1\.25,1\.5,1\.75,2\\\}, where the smallest margin over the whole set is18\.818\.8atα=0\.75\\alpha=0\.75\. The recovered floor of Eq\. \([6](https://arxiv.org/html/2609.17926#S3.E6)\) lies between10−2310^\{\-23\}and10−1910^\{\-19\}, against smallest observed losses of order10−510^\{\-5\}to10−310^\{\-3\}, so it is not identified by the data and the fit coincides with the two parameter form to the precision reported here\. The power law therefore does not use its third parameter, and scoring it with two rather than three would moveΔAIC\\Delta\\mathrm\{AIC\}by22, which changes no comparison in the table\. A power law has nothing for a floor to absorb, because it already approaches zero at a polynomial rate\. The asymmetry that motivated this appendix does not, in the event, favour the exponential: the floor is a feature the exponential needs and the power law cannot use\. Table 8:Matched three parameter comparison atα=1\\alpha=1\. Both families are fitted tologL\\log Lover the same width range, with the same rejection rules and the same seed aggregation\.zzis the Wald–Wolfowitz runs statistic on residuals ordered by width; drift is the largest change in the rate parameter on dropping the one and two largest widths\.Δ\\DeltaAIC is positive where the exponential is preferred\.The exponent returned by the matched fit is reported in Sec\.[4](https://arxiv.org/html/2609.17926#S4)and is not the quantity the geometric derivation expects\. It rises from2\.872\.87atp=23p=23to4\.204\.20atp=113p=113, implying dimensions from1\.391\.39down to0\.950\.95where the naive answer is22\. ### E\.5Residual structure of the power law Section[3](https://arxiv.org/html/2609.17926#S3)rejects the power law on the strength of its residuals rather than its coefficient of determination, and the test deserves to be stated\. FittinglogL\\log LagainstlogN\\log Nat weight decay11and recording the signs of the residuals in order of increasing width, the Wald–Wolfowitz runs statistic counts how many times the sign changes\. Independent errors would give a number of runs near2n1n2/\(n1\+n2\)\+12n\_\{1\}n\_\{2\}/\(n\_\{1\}\+n\_\{2\}\)\+1, which is8\.08\.0for these samples with a standard deviation of1\.81\.8\. Every one of the eight group orders returns exactly three runs, soz=−2\.78z=\-2\.78in each case\. The local slope over the width range varies from0\.040\.04to8\.28\.2and turns negative nearh=12h=12to1616, where training accuracy has reached unity while test accuracy is stalled near0\.600\.60\. The residuals form a single negative stretch, then a positive one, then a negative one, which is the signature of fitting a curve with a straight line\. That the count is identical across group orders spanning a factor of4\.94\.9indicates a systematic feature of the functional form rather than a coincidence of one data set\. The statistic must be computed against the family actually being rejected, and it is\. Adding the floor of Eq\. \([6](https://arxiv.org/html/2609.17926#S3.E6)\) does not disturb the residual structure: the runs statistic is−2\.78\-2\.78at seven of the eight group orders and−2\.76\-2\.76at the eighth, so the arc survives the matched fit of Appendix[E\.4](https://arxiv.org/html/2609.17926#A5.SS4)\. This is the expected consequence of a floor that the data do not identify, and it is what allows Sec\.[3](https://arxiv.org/html/2609.17926#S3)to reject the power law on residual structure rather than onR2R^\{2\}\. The exponential is not free of structure either, and the comparison is reported in both directions\. Its runs statistic ranges from−1\.48\-1\.48to−2\.23\-2\.23and exceeds the conventional threshold of1\.961\.96in absolute value at three of the eight group orders, against eight of eight for the matched power law, with the statistic more negative under the power law at every group order\. ### E\.6Truncation The instability that motivates the floor is visible in the local slopes, which flatten at the top of the width range: atp=47p=47the slope betweenh=64h=64andh=96h=96is0\.0610\.061and betweenh=96h=96andh=128h=128it falls to0\.0260\.026\. A plain exponential fit tologL\\log Lagainsthhaccordingly drifts as Sec\.[3](https://arxiv.org/html/2609.17926#S3)reports\. Admitting the floor of Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\) removes this\. A rate extracted from a saturating curve can be an artefact of where the curve is cut\. Refitting after removing the largest widths givesc=0\.1180c=0\.1180,0\.11830\.1183and0\.11790\.1179on dropping none, one and two, so the drift is0\.3%0\.3\\%\. Dropping three moves it to0\.10450\.1045, a fall of eleven percent, because the fit then has no points on the asymptote andL∞L\_\{\\infty\}ceases to be identified\. The reported values therefore rest on the presence of at least two widths in the saturated region, and the protocol keeps the full range for that reason\. ### E\.7Critical width by group order The values ofhch\_\{c\}plotted in Fig\.[3](https://arxiv.org/html/2609.17926#S6.F3)\(b\), obtained by linear interpolation of held out accuracy through0\.90\.9at weight decay11, are given in Table[9](https://arxiv.org/html/2609.17926#A5.T9)\. Table 9:Critical width against group order\. ### E\.8Seed error by regime The uncertainty onccis not a single number, and pooling across regimes would misstate it in both directions\. Fitting each seed separately gives a relative standard deviation of7\.5%7\.5\\%where the floor is identified, at weight decay between0\.250\.25and11, and0\.9%0\.9\\%where it is not, at weight decay22and above\. The difference is a property of the estimator rather than of the network, since a three parameter fit performed on a single noisy curve is intrinsically less stable than a two parameter one\. The relevant comparison for Sec\.[6](https://arxiv.org/html/2609.17926#S6)is therefore within the floor resolved regime, where the standard error of a three seed mean is4\.3%4\.3\\%\. The observed spread ofccacross group orders in the same regime is also4\.3%4\.3\\%\. Group order dependence, if it exists, is below the resolution of this experiment\. ### E\.9Identifiability of the rate and the exponent Fitting Eq\. \([5](https://arxiv.org/html/2609.17926#S3.E5)\) withα\\alphafree returns1\.583±0\.2541\.583\\pm 0\.254across the eight group orders, with values from1\.091\.09to1\.911\.91\. The corresponding rates span0\.00270\.0027to0\.07380\.0738, a factor of2828, so the pair is strongly anticorrelated and neither is determined on its own\. Fixingα\\alphaand refitting gives Table[10](https://arxiv.org/html/2609.17926#A5.T10), where each entry is the mean over the eight group orders\. Table 10:Rate, its coefficient of variation across group order, and the mean and minimum quality of fit, at fixedα\\alpha\.The quality of fit varies by less than two percent across a range over which the rate varies by a factor of214214\. One decade of width cannot separate the two parameters, and no claim in this paper depends on their separation\. What the table also shows is that the group order independence of the rate holds at every value ofα\\alpha, with a coefficient of variation between4\.04\.0and6\.86\.8percent throughout, and that it tightens slightly asα\\alpharises\. The conclusion of Sec\.[6](https://arxiv.org/html/2609.17926#S6)is therefore about howccresponds to a change in regularisation rather than about the value ofcc, and that distinction is what makes it robust to a parameter the data cannot fix\. Determiningα\\alphawould require widths spanning several decades rather than one\. Atp=113p=113the largest width used here is128128, and the memorisation plateau occupies everything below about1616, so the usable range is under one decade\. Reaching two would require widths near20002000at the same group order, which is feasible and which we have not done\. ### E\.10The ReLU runs The transfer experiment of Sec\.[6](https://arxiv.org/html/2609.17926#S6)uses the same data, split, optimiser and width grid as the main sweep, with three changes forced by the activation\. Initialisation follows He rather than the fan in rule, with standard deviation2/fanin\\sqrt\{2/\\mathrm\{fan\\,in\}\}in both layers\. The learning rate is10−210^\{\-2\}rather than3×10−33\\times 10^\{\-3\}\. The budget is40 00040\\,000steps rather than80008000, with the loss also recorded at20 00020\\,000so that convergence is checked per cell rather than assumed\. The learning rate was chosen by a preliminary scan over\{10−3,3×10−3,10−2\}\\\{10^\{\-3\},3\\times 10^\{\-3\},10^\{\-2\}\\\}crossed with weight decay in\{0\.05,0\.1,0\.25,0\.5\}\\\{0\.05,0\.1,0\.25,0\.5\\\}atp=47p=47andh=128h=128, run to150 000150\\,000steps with the accuracy recorded at eight logarithmically spaced checkpoints\. Nine of the twelve settings exceed accuracy0\.990\.99\. The step at which they cross it depends strongly on the learning rate, being50005000at10−210^\{\-2\}with weight decay0\.50\.5,40 00040\\,000at3×10−33\\times 10^\{\-3\}with the same weight decay, and beyond150 000150\\,000at10−310^\{\-3\}\. Three settings were still climbing at150 000150\\,000steps\. An earlier attempt at3×10−33\\times 10^\{\-3\}and20 00020\\,000steps reached only0\.650\.65and would have been reported as an architecture failure had the scan not been run\. Because activation and learning rate changed together, the quadratic sweep was repeated at10−210^\{\-2\}forp∈\{47,113\}p\\in\\\{47,113\\\}and weight decay in\{0\.1,0\.25,0\.5,1\}\\\{0\.1,0\.25,0\.5,1\\\}\. Over the eight cells the ratio of rates isc\(10−2\)/c\(3×10−3\)=0\.98±0\.04c\(10^\{\-2\}\)/c\(3\\times 10^\{\-3\}\)=0\.98\\pm 0\.04, with seven of eight between0\.890\.89and1\.031\.03\. The single outlier isp=47p=47at weight decay0\.10\.1, which attains accuracy0\.9440\.944and does not pass the learning criterion\. Within the quadratic activation the group order gap at10−210^\{\-2\}is6\.0%6\.0\\%over the cells that solve the task, against33to5%5\\%at3×10−33\\times 10^\{\-3\}\. No dead units were observed at any width or weight decay, with the fraction of hidden units never active on the training set equal to zero throughout\. Across the cells that solve the task the exponential family reachesR2=0\.963R^\{2\}=0\.963against0\.8690\.869for the best power law, against0\.9780\.978and0\.8990\.899for the quadratic activation on the same grid, so the departure from a power law survives the change of activation almost intact\. The rate does not\. It differs between group orders by61%61\\%on average, withp=113p=113abovep=47p=47at every weight decay tested, by a factor of1\.81\.8on average\. The critical width under ReLU is7979to8888atp=47p=47and2323to5858atp=113p=113, against2626to3131and2424to2727for the quadratic activation\. The two architectures agree at the larger group order and diverge threefold at the smaller one, where the training set holds11041104pairs against63846384\. ### E\.11Truncation under ReLU The ReLU gap between group orders is not a truncation artefact\. Extendingp=47p=47fromh=128h=128toh=384h=384resolves the floor in all four cells and moves the gap only from61%61\\%to57%57\\%\. ### E\.12Rescalings that do not rescue the factorisation Since ReLU needs a larger width, the rate may be expressed in the wrong units, and a dimensionless product might collapse the two activations\. Six candidates were examined, formed fromcctogether with the critical widthhch\_\{c\}, the widthh∗=\(logA−logL∞\)/ch^\{\*\}=\(\\log A\-\\log L\_\{\\infty\}\)/cat which the exponential term meets the floor, andKmaxK\_\{\\mathrm\{max\}\}\. The best of them ischcc\\,h\_\{c\}, which reduces the spread across activations at matched weight decay from38\.3%38\.3\\%to9\.4%9\.4\\%, and across group order from14\.4%14\.4\\%to11\.1%11\.1\\%\. We do not adopt it\. The ratio between activations is0\.852±0\.1160\.852\\pm 0\.116rather than unity, so a systematic offset of fifteen percent remains\. The improvement across group order is an artefact of pooling the two activations: within the quadratic activation alone, where the factorisation is claimed, the same rescaling raises the mean spread across group order at matched weight decay from3\.7%3\.7\\%to7\.3%7\.3\\%\. And the ranking is not stable across criteria, sincech∗c\\,h^\{\*\}gives the smaller spread across group order at6\.4%6\.4\\%while giving a larger one across activations at19\.3%19\.3\\%\. Selecting the best of six candidates on the quantity one wishes to collapse is a procedure that will usually succeed on noise, and the outcome here is consistent with that rather than with a change of units\. ### E\.13Noise floor at large group order Long runs at50 00050\\,000steps were used to determine whether the fitted floor at largeppis an asymptote or a stopping point\. Held out accuracy remains at unity from step40004000onwards in every case, and the weight norm changes by between0\.3%0\.3\\%and2\.6%2\.6\\%over the remaining46 00046\\,000steps, so neither forgetting nor norm starvation is occurring\. The loss nonetheless wanders\. Its increments change sign one to three times over the plateau and the ratio of maximum to minimum lies between1\.111\.11and1\.801\.80\. The step at which the minimum occurs is80008000,40004000,32 00032\\,000,16 00016\\,000,80008000,50 00050\\,000,32 00032\\,000and16 00016\\,000across the eight configurations examined, which is consistent with a random walk and not with a drift\. Reduced precision arithmetic is not the cause\. Disabling it raises the mean ratio from1\.3651\.365to1\.9691\.969rather than lowering it, which locates the effect in the optimiser rather than in the arithmetic\. The picture is of a basin flat enough that AdamW with decoupled weight decay does not settle, and the wandering amplitude is the width of that basin as seen through the cross entropy\. The consequence is visible in the scaling ofL∞L\_\{\\infty\}withKmaxK\_\{\\mathrm\{max\}\}: up toKmax=26K\_\{\\mathrm\{max\}\}=26the log log slope is−4\.5\-4\.5, and fromKmax=26K\_\{\\mathrm\{max\}\}=26onwards it flattens to−1\.1\-1\.1, the flattening being the noise rather than the task\. Since adjacent large group orders differ inL∞L\_\{\\infty\}by factors between1\.21\.2and1\.51\.5, and the wandering amplitude is comparable, the floor is not resolvable there\. The slope reported in Sec\.[3](https://arxiv.org/html/2609.17926#S3)is accordingly restricted toKmax≤26K\_\{\\mathrm\{max\}\}\\leq 26, where the fitted floors exceed7×10−57\\times 10^\{\-5\}and sit above the noise\. ## Appendix FThe training fraction experiment Section[6](https://arxiv.org/html/2609.17926#S6)reports that the factorisation fails under ReLU, and the Discussion attributes the failure to data supply rather than to the activation\. That attribution is a prediction and this appendix tests it\. ### F\.1Design Raising the training fraction shrinks the held out set, so a rate fitted at one fraction is not directly comparable with a rate fitted at another\. The splits are nested: the partition is drawn once from a fixed seed, so the training set at a fraction of one half is contained in the training set at four fifths and the held out set at four fifths is contained in the held out set at one half\. Every rate reported here is therefore fitted on the set held out at four fifths, which is442442pairs atp=47p=47and25542554atp=113p=113and is disjoint from the training set at both fractions\. The items are identical across the comparison and only the quantity of training data changes\. The quadratic activation is run as a control\. If the training fraction moved the rate where the factorisation already holds, the ReLU arm would be uninformative\. The control needs24 00024\\,000steps rather than the80008000of the main sweep: at four fifths there are sixty percent more training pairs, and at80008000steps the loss is still falling between the last two checkpoints by18%18\\%to30%30\\%at weight decay0\.50\.5and below\. Fitting a rate to a curve that has not levelled off biases it, and biases it differently at the two group orders, which is exactly the comparison at issue\. Weight decay0\.250\.25is excluded from the comparison\. Its cells are the least stable in the sweep: of the five whose loss rises rather than falls between the last two checkpoints, which is the plateau wandering of Appendix[E](https://arxiv.org/html/2609.17926#A5)and not a failure to converge, three are at this weight decay, and the largest excursion anywhere in the study is a near doubling atp=47p=47and four fifths\. Weight decay11and0\.50\.5are stable at both fractions and both group orders\. ### F\.2Result Table[11](https://arxiv.org/html/2609.17926#A6.T11)gives the fitted rates\. Writing the discrepancy with its sign, as\(c47−c113\)\(c\_\{47\}\-c\_\{113\}\)over their mean, so that a negative value means the smaller group order lags\. This is not the same quantity as the61%61\\%quoted in Sec\.[6](https://arxiv.org/html/2609.17926#S6), which is an unsigned mean over four weight decays on the native holdout in the main ReLU sweep; here it is signed, restricted to the two weight decays resolved at all four combinations of group order and fraction, and measured on the common holdout\. At the half fraction the two agree in magnitude to within five points\. The movement is\+17\.0\+17\.0points under the quadratic activation and\+52\.7\+52\.7under ReLU\. It has the same sign in both: additional data raises the rate at the smaller group order relative to the larger one\. What differs is the size of the deficit available to be closed\. Under ReLU the deficit was large and closes to within the seed to seed resolution of the experiment; under the quadratic activation it was already small and the same intervention carries it past zero\. The critical width says the same thing without any fit\. Atλ=1\\lambda=1the ratiohc\(47\)/hc\(113\)h\_\{c\}\(47\)/h\_\{c\}\(113\)falls from1\.531\.53to0\.950\.95under ReLU as the fraction rises, so the two group orders come to agree and the anomaly that motivated the experiment disappears\. Under the quadratic activation, which never showed that anomaly, the ratio goes from1\.141\.14to1\.301\.30\. These values are measured on the common holdout, at24 00024\\,000steps under the quadratic activation and40 00040\\,000under ReLU, and so differ slightly from Table[9](https://arxiv.org/html/2609.17926#A5.T9), which is the main sweep at80008000steps on the native one\. Table 11:Fitted rateccon the common holdout at two training fractions\. Quadratic runs use24 00024\\,000steps, ReLU runs40 00040\\,000\. ### F\.3What this does and does not establish The prediction the Discussion makes is confirmed for the case it was made about\. It is not established that the training fraction is inert elsewhere: under the quadratic activation it moves the rate too, and at weight decay0\.50\.5it moves it enough to reverse the sign of the residual\. The factorisation reported in Sec\.[6](https://arxiv.org/html/2609.17926#S6)is therefore a statement about a regime in the data supply as well as in the architecture, and the boundary of that regime is what the two activations locate at different places\.
Similar Articles
How much of the weight-space perception gap is actually symmetry? Evidence from ~1.8M fitted SIRENs [R]
This paper investigates the role of parameter symmetry in weight-space perception, showing that symmetry scatter alone accounts for nearly the entire degradation in performance between shared and independently-fitted neural networks. The study uses SIRENs and large-scale experiments to argue that computational advantages may justify weight-space methods over informational equivalence to function access.
Observable- and Positional-Encoding-Dependent Symmetry Readout from Neural Network Weights
This paper shows that the geometric symmetry visible from neural network weights depends on the positional encoding and readout observable, and validates this using MLPs trained on 2D signed distance functions with multiple symmetry groups.
[R] Measuring the Symmetry--Data Exchange Rate
This paper empirically measures the symmetry–data exchange rate predicted by equivariance theory, finding that wrong-group symmetry constraints are actively harmful, augmentation with test-time orbit averaging matches equivariant architectures, and the theoretical |G|-fold sample complexity reduction is only weakly confirmed with wide confidence intervals. The study is explicitly exploratory and not pre-registered.
Statistically Meaningful Geometry and Gauge Symmetry Breaking: A Geometric Foundation for Scientific Discovery and Intelligence Emergence
This paper introduces Statistically Meaningful Geometry (SMG), a geometric framework for modeling over-parameterized learning systems as infinite-dimensional non-parametric Orlicz fiber bundles. It proposes that under out-of-distribution stimuli, the system undergoes a gauge symmetry break, leading to the emergence of new causal axes that can distinguish genuine scientific discovery from hallucinations.
Group Invariant Spectral Embedding
This paper proposes incorporating symmetries into affinity kernels for spectral embedding, proving convergence of invariant graph Laplacians on quotient manifolds with improved sample complexity.