Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws
Summary
This paper introduces Coupled Scaling, a task-conditioned framework that explains how neural scaling laws vary based on the relationship between task structure and the geometric representations accessible by architecture-optimization systems.
View Cached Full Text
Cached at: 09/04/26, 06:30 AM
# 1. Introduction
Source: [https://arxiv.org/html/2609.03533](https://arxiv.org/html/2609.03533)
Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws
Jie Wang\*\*\*LLM assistance disclosure\. OpenAI’s ChatGPT and Anthropic’s Claude assisted with literature discovery, code drafting and debugging, statistical cross\-checks, formal presentation, and language revision\. The author originated the research question and framework; determined the claims, derivations, and empirical design; independently verified the citations, code, and reported results; and takes full responsibility for the work\.
School of Civil and Commercial Law, Southwest University of Political Science and Law
September 2026
Abstract
Existing theories derive neural scaling from data geometry or a specified data–model spectrum, but systems trained on the same data can scale differently in practice when architecture or optimization changes the representations they can efficiently reach\. We introduce*Coupled Scaling*, a task\-conditioned framework in which finite\-budget scaling depends on the relation between task structure and the geometry accessible to an architecture–optimization system\. In a solvable mode\-truncation model, loss separates exactly into target energy outside architectural support and an unresolved supported tail\. For an arbitrary priority order, the residual lies between the best\-NNsupported tail and the tail beyond the largest completed high\-value prefix\. If the cumulative\-tail and coverage log\-rates areγA,T\\gamma\_\{A,T\}andρA,O,T\\rho\_\{A,O,T\}, the residual exponent lies in\[ρA,O,TγA,T,γA,T\]\[\\rho\_\{A,O,T\}\\gamma\_\{A,T\},\\gamma\_\{A,T\}\]\. The interval separates prefix\-limited scaling from dispersed acquisition\. Under bounded off\-prefix gain, the completed prefix is rate\-determining and the lower endpoint is exact,αA,O,T=ρA,O,TγA,T\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\\gamma\_\{A,T\}; the power\-law caseaA,T,j≍j−bA,Ta\_\{A,T,j\}\\asymp j^\{\-b\_\{A,T\}\}recoversαA,O,T=ρA,O,T\(bA,T−1\)\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\(b\_\{A,T\}\-1\)\. A fixed\-kernel specialization derives the training\-time exponent from the near\-zero tail of a task\-weighted spectral measure defined independently of the loss fit\. These results sharpen the separation between architectural support and finite\-budget acquisition and motivate two empirical tests: static task\-relevant geometry should track loss at a common budget, while multiscale geometry should track coupling\-specific exponent ordering, including its preregistered reversal across contrasting tasks\. An audit of released emergence trajectories identifies the controls needed for a direct test and organizes them in a factorial design that measures geometry separately from the scaling fit\.
Neural scaling laws describe how loss changes with model size, data, and compute, and they guide both resource allocation and forecasts of future capability \(Hestness et al\., 2017; Kaplan et al\., 2020; Hoffmann et al\., 2022\)\. A central comparative question remains open: when should systems trained on the same data share a scaling law, and when should their slopes or attainable losses diverge?
Existing geometric and spectral theories address the task\-side part of this question by explaining how task structure enters scaling\. Manifold accounts relate parameter scaling to intrinsic dimension under smoothness assumptions \(Sharma and Kaplan, 2022\); kernel and random\-feature models derive learning curves from a specified feature map, data–model spectrum, target alignment, and statistical regime \(Maloney et al\., 2022; Bahri et al\., 2024\)\. These theories apply most directly when the representation is fixed or sufficiently flexible\.
The comparison changes when systems trained on the same data reach different task\-relevant geometries\. Equivariant and non\-equivariant models trained on the same neural\-force\-field data have different parameter\-, data\-, and compute\-scaling exponents \(Ngo and Ravanbakhsh, 2026\)\. Feature learning changes training\-time and compute exponents for targets poorly aligned with the initial representation \(Bordelon et al\., 2025\), preconditioning changes fitted model\-size exponents in controlled random\-feature regression \(Ramani and Jain, 2026\), and weight\-decay interventions on superposition give a direct geometric mechanism for width scaling \(Liu, Liu, and Gore, 2025\)\. Together, these studies place the learning system inside the scaling relation\. Across distinct resource axes, they reveal the same structural possibility: a model may contain directions relevant to a target yet assign them too little strength or reach them too late to matter at the observed scale\. Architecture shapes availability; training shapes when available directions become useful\. Their effects call for an object indexed by the learning system, the task, and the available budget\.
We call the resulting relation*Coupled Scaling*\.111Cheng et al\. \(2026\) use ‘coupling’ for a depth–width pathL=u\(N\)L=u\(N\), with training and test sample sizes entering the finite\-sample reliability conditions under which gains along that path remain observable\. Here coupling refers to the relation between an architecture–optimization system’s accessible geometry and task structure\.LetBBbe a finite envelope over parameters, data, compute, and training time, and let𝒫=\(𝒟train,ℓ\)\\mathcal\{P\}=\(\\mathcal\{D\}\_\{\\mathrm\{train\}\},\\ell\)fix the training problem and objective\. For architectureAAand training procedureOO, the representational accessibility setℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)contains the geometries reachable within that envelope\. Architecture can shape structural support, feature learning can reshape the realized representation, and optimization can change which represented directions are acquired first\. Coupled Scaling tracks these channels through their interaction with the task\.
Its empirical signature is a coupling\-by\-task interaction with distinct level and rate margins\. At a common finite budget, static task\-relevant geometry is paired with loss at that budget\. Across scales, geometry growth is paired with fitted exponent differences\. Shared task\-relevant access predicts convergence toward the applicable data\-side rate\. Under the common\-tail and prefix\-adequacy conditions formalized in §3\.3, reversed acquisition\-rate ordering across tasks predicts reversed exponent ordering\. Static probes measure useful task structure at a fixed budget; multiscale trajectories measure its growth\. The geometry quantities are specified independently of the loss curve\.
The interaction has an exact witness in a mode\-truncation model\. Architecture supplies a supportSAS\_\{A\}, and the architecture–optimization coupling supplies a priority orderπA,O\\pi\_\{A,O\}over supported directions\. Population loss decomposes into target energy outside support and an unresolved supported tail\. For an arbitrary priority order, the cumulative supported tail and the largest completed high\-value prefix give the universal bracket
A¯A,T\(N\)≤EA,O,T\(N\)≤A¯A,T\(JA,O,T\(N\)\)\.\\overline\{A\}\_\{A,T\}\(N\)\\leq E\_\{A,O,T\}\(N\)\\leq\\overline\{A\}\_\{A,T\}\(J\_\{A,O,T\}\(N\)\)\.If the cumulative\-tail and completed\-prefix log\-rates areγA,T\\gamma\_\{A,T\}andρA,O,T\\rho\_\{A,O,T\}, the residual exponent lies in\[ρA,O,TγA,T,γA,T\]\[\\rho\_\{A,O,T\}\\gamma\_\{A,T\},\\gamma\_\{A,T\}\]\. Bounded off\-prefix gain makes the completed prefix rate\-determining and gives
αA,O,T=ρA,O,TγA,T\.\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\\gamma\_\{A,T\}\.ForaA,T,j≍j−bA,Ta\_\{A,T,j\}\\asymp j^\{\-b\_\{A,T\}\},γA,T=bA,T−1\\gamma\_\{A,T\}=b\_\{A,T\}\-1, yielding the power\-law specialization\. The interval also captures dispersed acquisition: Appendix B\.2 constructs a fixed nested order with the same completed\-prefix rate and a strictly faster residual rate\. A fixed\-kernel specialization derives the training\-time exponent from task\-weighted spectral mass near zero\.
At the theoretical level, Coupled Scaling supplies a comparison framework and a solvable witness\. It identifies the strict floor, gives an exact tail–coverage bracket for arbitrary priority orders, states when a scalar completed\-prefix rate predicts the exponent, and anchors the construction in a fixed\-kernel spectral limit\.
The empirical contribution is an identification design that pairs static geometry with matched\-budget loss and multiscale geometry with fitted rates\. The released emergence trajectories identify the needed controls; the proposed factorial test crosses couplings with contrasting tasks\. Figure[1](https://arxiv.org/html/2609.03533#Sx1.F1)summarizes the logic\.
TargetfTf\_\{T\}Architecture supportSAS\_\{A\}Unsupported target energy∑k∉SAcT,k2\\sum\_\{k\\notin S\_\{A\}\}c\_\{T,k\}^\{2\}strict floorLA∞\(T\)L\_\{A\}^\{\\infty\}\(T\)Cumulative supported tailA¯A,T\(m\)=∑j\>maA,T,j\\overline\{A\}\_\{A,T\}\(m\)=\\sum\_\{j\>m\}a\_\{A,T,j\}log\-rateγA,T\\gamma\_\{A,T\}Canonical completed prefixJA,O,T\(N\)J\_\{A,O,T\}\(N\)log\-rateρA,O,T\\rho\_\{A,O,T\}Residual rateργ≤α≤γ\\rho\\gamma\\leq\\alpha\\leq\\gammabounded off\-prefix gain:α=ργ\\alpha=\\rho\\gamma\(A\) Minimal support–tail–coverage witnessFixed kernelKqK\_\{q\}and targetfTf\_\{T\}Task\-weighted spectralmeasureνq,T\\nu\_\{q,T\}Zero atomνq,T\(\{0\}\)\\nu\_\{q,T\}\(\\\{0\\\}\)⟶Lq,T∞\\longrightarrow L\_\{q,T\}^\{\\infty\}Positive near\-zero tailFq,T\(x\)∼cxβq,Tℓ\(1/x\)F\_\{q,T\}\(x\)\\sim cx^\{\\beta\_\{q,T\}\}\\ell\(1/x\)Lq,T\(t\)−Lq,T∞∼Ct−βq,Tℓ\(t\)L\_\{q,T\}\(t\)\-L\_\{q,T\}^\{\\infty\}\\sim Ct^\{\-\\beta\_\{q,T\}\}\\ell\(t\)training\-time exponentαt\(q,T\)=βq,T\\alpha\_\{t\}\(q,T\)=\\beta\_\{q,T\}\(B\) Fixed\-kernel specializationMeasure geometry at each preregisteredNNStatic levelgq,tlevel\(N¯\)g^\{\\mathrm\{level\}\}\_\{q,t\}\(\\bar\{N\}\)Compare with loss at the same budgetLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)Rate\-sensitive trajectorygq,t\(N\)g\_\{q,t\}\(N\)Estimateρ\(g\)\\rho^\{\(g\)\}, or use fixed\-kernelβ\\betaCompare with fitted exponentαq,t\\alpha\_\{q,t\}Test matched\-budget tracking, rate prediction, and task reversal\(C\) Direct2×22\\times 2geometry\-tracking testFigure 1:Coupled Scaling as a support–tail–coverage framework and a direct test\. \(A\) Target energy outside architectural support determines the strict floor\. Within support, the cumulative target tail and the canonical completed\-prefix coverage give a universal rate interval; bounded off\-prefix gain makes the lower endpoint exact\. ForaA,T,j≍j−bA,Ta\_\{A,T,j\}\\asymp j^\{\-b\_\{A,T\}\}, the cumulative\-tail rate isγA,T=bA,T−1\\gamma\_\{A,T\}=b\_\{A,T\}\-1\. \(B\) In the fixed\-kernel specialization, regular variation of task\-weighted spectral mass near zero gives the training\-time exponent, up to a slowly varying factor\. \(C\) In the factorial design, static geometry atN¯\\bar\{N\}is compared with loss at the same preregistered budget\. The multiscale trajectory suppliesρ\(g\)\\rho^\{\(g\)\}, and the fixed\-kernel alternative usesβ\\beta\. The common\-tail and prefix\-adequacy criteria define the product\-law comparison\. Within\-task contrasts are primary, and the fitted exponent ordering is tested for a preregistered reversal across tasks\.
## 2\. From Data\-Side Scaling to Model\-Side Dependence
### 2\.1 Data Geometry and Fixed\-Representation Accounts
Empirical scaling work established regular power\-law or power\-law\-plus\-constant relations across model size, data, and compute \(Hestness et al\., 2017; Kaplan et al\., 2020; Hoffmann et al\., 2022\)\. These relations became useful for forecasting and allocation; their fitted parameters remained principally descriptive\.222Empirical scaling curves can exhibit smoothly broken power laws, delayed inflections, regime changes, and nonmonotonic transitions \(Caballero et al\., 2023\)\. Accordingly, every exponent and claim in this paper is indexed to a stated axis, loss, and fitted regime\.Data\-side theory supplied mechanisms by connecting scaling to the structure of the learning problem\. Sharma and Kaplan \(2022\), for example, relate parameter scaling to intrinsic dimension under smoothness and generic\-function assumptions\. Bahri et al\. \(2024\) derive multiple statistical regimes from a data–model spectrum, target alignment, and finite\-sample effects; solvable kernel and random\-feature models express learning curves through eigenspectra, target coefficients, noise, regularization, and resource regime \(Maloney et al\., 2022; Bordelon et al\., 2020; Canatar et al\., 2021\)\.
Discrete coverage accounts provide a closely related tail perspective\. Zou et al\. \(2026\) formalize a sharp, monotone effective frontier in a ranked Zipfian pattern space and derive resource\-dependent loss from the mass beyond that frontier\. Song et al\. \(2026\) construct a corpus\-intrinsic predictive\-contribution spectrum and infer a moving data\-scale cutoff by matching observed excess loss to its residual tail\. Together, these papers establish tail–frontier composition as a direct neighboring mechanism\. Coupled Scaling extends that mechanism to systems that may differ in support and strict floor, permits acquisition paths that interleave target ranks, and obtains the empirical acquisition rate from geometry before the loss fit\.
The prefix\-exact specialization recovers Zou et al\.’s sharp\-frontier law\. Under full support andKR=\{1,…,k⋆\(R\)\}K\_\{R\}=\\\{1,\\ldots,k\_\{\\star\}\(R\)\\\}, the canonical completed prefix isJ\(R\)=k⋆\(R\)J\(R\)=k\_\{\\star\}\(R\), and Proposition 2 givesE\(R\)=A¯\(k⋆\(R\)\)E\(R\)=\\overline\{A\}\(k\_\{\\star\}\(R\)\)\. Henceaj≍j−ba\_\{j\}\\asymp j^\{\-b\}andk⋆\(R\)≍Rρk\_\{\\star\}\(R\)\\asymp R^\{\\rho\}yieldE\(R\)≍R−ρ\(b−1\)E\(R\)\\asymp R^\{\-\\rho\(b\-1\)\}\. This recovery isolates the added comparative objects: coupling\-indexed support, interleaved acquisition, the prefix\-adequacy criterion, and cross\-task reversal\. Zou et al\.’s resource\-specific frontier derivations and max\-bottleneck analysis address complementary questions\. Song et al\. infer the cutoff by matching observed excess loss to residual tail mass; Coupled Scaling estimates acquisition from a geometry trajectory specified before loss fitting and evaluates it as a predictor of exponent ordering\.
The matched\-data cross\-system question is how architecture–optimization systems make task\-relevant directions strong, available early, or reachable within a feasible budget\.
### 2\.2 Same\-Data Evidence for Model\-Side Dependence
Architecture comparisons make model\-side dependence concrete\. Tay et al\. \(2023\) report architecture\-dependent scaling behavior and rank changes under matched language\-model pretraining\. Ngo and Ravanbakhsh \(2026\) compare equivariant and non\-equivariant neural\-force\-field models on the same data and task and find different parameter\-, data\-, and compute\-scaling exponents, with task\-matched equivariance scaling more efficiently\. In a controlled hierarchical\-language generator, Cagnetta et al\. \(2025\) find that convolutional networks aligned through locality and weight sharing scale faster than Transformers; representation probes track acquisition of the latent hierarchy\. Defilippis, Krzakala, Loureiro, and Maillard \(2026a\) sharpen this structural picture for two\-layer networks on hierarchical multi\-index targets, deriving representation\-limited regimes, subspace\-recovery transitions, plateaus, and spectral structure\.
A recent resource\-matched architecture comparison studies recurrent depth while closely matching per\-token FLOPs, total non\-embedding parameters, and KV\-cache size between looped and unlooped sparse MoE Transformers\. Wang et al\. \(2026\) loop the middle half of the layers twice and fit separate Chinchilla\-style scaling surfaces for the two architectures across four model sizes\. The looped recipe has a steeper compute\-optimal frontier; its downstream advantage is largest on code and grows with sample length and the number of in\-context examples\. Because budget matching also changes width, expert count, and attention configuration, the experiment identifies a resource\-matched architecture recipe rather than recurrence in isolation\. In Coupled Scaling terms, it supplies unusually controlled evidence of model\-side rate dependence while leaving open whether the operative accessibility channel is support, completed\-prefix acquisition, or dispersed off\-prefix gain\. Distinguishing these possibilities requires an independently measured, task\-conditioned geometry trajectory\.
Feature learning and optimization supply a second group of mechanisms\. Bordelon, Atanasov, and Pehlevan \(2025\) show that feature learning improves training\-time and compute exponents for hard targets outside the initial kernel’s reproducing space, with little change for aligned targets\. Ramani and Jain \(2026\) isolate a spectral optimization channel: preconditioning changes fitted model\-size exponents in controlled random\-feature regression\. Defilippis et al\. \(2026b\) derive excess\-risk phase diagrams for diagonal and quadratic shallow networks and connect those regimes to trained\-weight spectra\. In GPT\-style models, Jha and Reagen \(2026\) hold the architecture family, data recipe, and FFN\-width schedule fixed across four optimizers and observe different effective\-rank scaling\. Their extended\-training control reaches approximately matched validation perplexity for AdamW and low\-rank Dion and preserves the rank difference\. Huang et al\. \(2026\) add a capacity\-allocation account in which larger models retain rare and complex task features by reducing resource competition and gradient interference\. Volkova et al\. \(2026\) and Bansal et al\. \(2022\) identify shared\-exponent regimes captured by coefficients or effective\-resource rescaling\.
Representational interference through superposition supplies a third, explicitly geometric mechanism\. In a controlled autoencoder, Liu, Liu, and Gore \(2025\) vary weight decay and obtain weak\- and strong\-superposition regimes\. The weak regime inherits its exponent from the feature\-frequency tail; the strong regime produces a robust inverse\-width contribution through geometric interference among representation vectors\. Four open language\-model families exhibit the overlap signature associated with the strong regime\.
Across these studies, interventions on architecture, feature learning, optimization, and superposition alter fitted scaling, learned geometry, or both under fixed or closely controlled data\. These results motivate a common coupling\-by\-task design that tests whether independently measured geometry predicts the scaling interaction\.
### 2\.3 The Missing Relational Variable
The missing variable is relational: model\-side interventions matter through the task structure they make easy or difficult to acquire\. Connecting data\-side theory to those interventions requires an object indexed by task and budget that records structural support, realized representation, and acquisition order, and reduces to spectrum plus target alignment in a fixed\-representation limit\. The object describes what a learning system can reach at feasible cost and preserves feasible\-budget distinctions even when unbounded expressivity is shared\. We call it*representational accessibility*\.
Representational accessibility also identifies a shared\-access regime\. When compared systems expose effectively the same task\-relevant geometry, coupling\-specific exponent differences should disappear\. Neural scaling universality offers a hypothesis for this regime\. Liu and Gore \(2026\) argue that time, width, and depth exponents for current dense Transformers are fixed within a universality class, with architecture and data changing coefficients\. Volkova et al\. \(2026\) obtain more stable language\-model extrapolation by fixing Chinchilla exponents to an AdamW reference and fitting optimizer\-specific resource rescalings\. Coefficient\-only descriptions are natural under shared task\-relevant access\. Coupled Scaling becomes discriminating when a controlled intervention changes that access differently across tasks\.
The resulting design crosses couplings with tasks\. Protocol\-asymptotic support sets strict floors, finite\-budget acquisition sets residual rates, and the completed\-prefix bracket determines when a scalar coverage rate predicts the exponent\. Equivalent access defines the null; reversed acquisition\-rate and exponent orderings across tasks supply the confirmatory interaction\.
## 3\. The Coupled Scaling Framework
### 3\.1 Budget\-Relative Representational Accessibility
A neural\-network architecture defines an*information\-flow structure*: a pattern induced by connectivity, attention, normalization, activation, parameter sharing, and other inductive biases\. Convolution makes spatial locality inexpensive; attention permits content\-dependent interaction across positions; equivariant architectures restrict representations to respect specified symmetries\. These inductive biases change the resources required to realize task\-relevant representations, even among architectures with comparable unrestricted expressivity\. Feasible\-scale comparison therefore requires a budget\-relative notion of representational reach\.
LetBBdenote a finite budget envelope over parameters, data, compute, and optimization time\. Fix a training problem𝒫=\(𝒟train,ℓ\)\\mathcal\{P\}=\(\\mathcal\{D\}\_\{\\mathrm\{train\}\},\\ell\), whereℓ\\ellincludes the supervision or self\-supervised objective and the associated preprocessing\. For architectureAAtrained by procedureOO, define the*representational accessibility set*ℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)as the representational geometries attainable by the specified system withinBBon𝒫\\mathcal\{P\}\. HereOOincludes parameterization, optimizer, schedule, and the resulting training dynamics; different random seeds index draws from the stochastic component ofOO\. Membership records feasibility withinBB; computational optimality lies outside the definition\. An empirical instantiation must prespecify the representation operator or projection, the equivalence or tolerance criterion, and the resource schedule encoded byBB\. Universal approximation at unbounded capacity is therefore compatible with sharply different accessibility at feasible scale\.
An operational form makes the suppressed choices explicit\. Fix a representation descriptorℳ\\mathcal\{M\}and an equivalence relation∼\\simon its output space\. Let
𝒰B\(A,O;𝒫\)=\{τ:τis an admissible run of\(A,O\)on𝒫,cost\(τ\)⪯B\}\.\\mathcal\{U\}\_\{B\}\(A,O;\\mathcal\{P\}\)=\\left\\\{\\tau:\\ \\tau\\ \\text\{is an admissible run of \}\(A,O\)\\text\{ on \}\\mathcal\{P\},\\ \\operatorname\{cost\}\(\\tau\)\\preceq B\\right\\\}\.
The corresponding accessibility set is
ℛBℳ,∼\(A,O,𝒫\)=\{\[ℳ\(θτ\)\]∼:τ∈𝒰B\(A,O,𝒫\)\}\.\\mathcal\{R\}\_\{B\}^\{\\mathcal\{M\},\\sim\}\(A,O;\\mathcal\{P\}\)=\\left\\\{\[\\mathcal\{M\}\(\\theta\_\{\\tau\}\)\]\_\{\\sim\}:\\tau\\in\\mathcal\{U\}\_\{B\}\(A,O;\\mathcal\{P\}\)\\right\\\}\.
The superscripts are suppressed when the descriptor and equivalence convention are fixed\. A closure may be taken only after a topology on the descriptor space has been specified\. WhenOOis stochastic, the empirical instantiation must state whether the reported object concerns the support of the induced distribution, an existence claim, a high\-probability region, or an expectation\-level summary\. Budget monotonicity follows when every run admissible underB1B\_\{1\}remains admissible underB2⪰B1B\_\{2\}\\succeq B\_\{1\}\. Empirical tests therefore use prespecified task\-relevant projections as operational summaries of the reachable set\.
By*representation geometry*we mean the task\-relevant relational structure induced by a representation—operationally, a feature, similarity, covariance, or kernel operator—modulo coordinate changes that preserve the relevant operator\. This quotient convention makes accessibility invariant to coordinate reparameterizations that preserve the relevant operator\.
We impose four admissibility requirements onℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)\.
1. 1\.Budget monotonicity\.IfB1⪯B2B\_\{1\}\\preceq B\_\{2\}componentwise, thenℛB1\(A,O,𝒫\)⊆ℛB2\(A,O,𝒫\)\\mathcal\{R\}\_\{B\_\{1\}\}\(A,O;\\mathcal\{P\}\)\\subseteq\\mathcal\{R\}\_\{B\_\{2\}\}\(A,O;\\mathcal\{P\}\)\. Every geometry feasible under the smaller envelope remains feasible under the larger one when the smaller\-budget procedure remains available\.
2. 2\.Task conditioning\.Predictions concern a task projectionPT,𝒟ℛB\(A,O,𝒫\)P\_\{T,\\mathcal\{D\}\}\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\): the portion of accessible geometry relevant to evaluation targetTTunder distribution𝒟\\mathcal\{D\}\. An architecture can be well aligned with one target and poorly aligned with another on the same inputs\. When training and evaluation tasks coincide,TTis already encoded inℓ\\ell; the projection notation keeps the comparison explicit\.
3. 3\.Mechanism separation\.Structural support, realized representation, and learning order are distinct channels\. Architecture may alter available directions; feature learning may reweight or create realized directions; optimization may change which represented directions are reached within the budget\.
4. 4\.Kernel consistency\.In a fixed\-kernel or appropriate lazy\-training limit, the task projection must reduce to a description in terms of available eigendirections, their eigenvalues, and target alignment\.
These requirements distinguish the relevant levels of description\.ℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)is conditional on the training problem; its task projection is the comparison object for a particular evaluation problem\. An*effective spectrum*is one representation of that projection, anddaccessd\_\{\\mathrm\{access\}\}is a possible scalar summary\. Both must be measured independently of the scaling curve\.
The dependence onOOcaptures variation within a fixed architecture: lazy and rich training regimes can realize different portions of representation space under different parameterizations and dynamics \(Chizat et al\., 2019; Yang and Hu, 2021\)\. The dependence onAAcaptures structural variation under shared data, as illustrated by equivariant and non\-equivariant architectures \(Ngo and Ravanbakhsh, 2026\)\. Coupled Scaling hypothesizes that both sources of variation change the task\-relevant geometry reachable within the budget and thereby enter scaling behavior\.
### 3\.2 Task Projection and the Dual Bottleneck
Consider a taskTTunder data distribution𝒟\\mathcal\{D\}\. Only part of the input geometry is relevant to the task loss\. Coupled Scaling separates that task\-relevant structure into two budget\-relative components\.
Accessible structureis task\-relevant structure that some geometry inℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)resolves efficiently\. Increasing scale can progressively refine this component, with rates governed by the applicable geometric, spectral, and statistical regime\.
Currently inaccessible structureis task\-relevant structure not efficiently resolved by any geometry inℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)\. It may become reachable at larger budgets or under other couplings\. It contributes to residual loss at the stated budget because the current coupling supplies no sufficiently efficient pathway\. Accessibility is graded: a high\-cost direction can behave as inaccessible over the observed range\.
Two notions of a floor must be distinguished\. LetB\(N\)B\(N\)denote the experiment’s resource schedule as model size grows, including the associated data, compute, and training\-time allocation\. The*strict asymptotic floor*under that protocol is
LA,O∞\(𝒟\)=lim infN→∞L\(N,A,O,𝒟\),L^\{\\infty\}\_\{A,O\}\(\\mathcal\{D\}\)=\\liminf\_\{N\\to\\infty\}L\(N;A,O,\\mathcal\{D\}\),
where the remaining resources followB\(N\)B\(N\)\. This liminf defines an asymptotic lower envelope without assuming an ordinary limit or an observed plateau\. The task/target index is suppressed in this notation\. The additive floor\-plus\-tail form applies to regimes in whichL\(N\)→L∞L\(N\)\\to L^\{\\infty\}and the residual tail is nonnegative\. The*finite\-budget attainable loss*at a maximum feasible model sizeN¯\\bar\{N\}is
LA,Oeff\(𝒟,N¯\)=infN≤N¯L\(N,A,O,𝒟\)\.L^\{\\mathrm\{eff\}\}\_\{A,O\}\(\\mathcal\{D\};\\bar\{N\}\)=\\inf\_\{N\\leq\\bar\{N\}\}L\(N;A,O,\\mathcal\{D\}\)\.
For monotone learning curves this isL\(N¯\)L\(\\bar\{N\}\)\. In a nonnegative floor\-plus\-tail regime it contains both the strict floor and whatever task\-relevant tail remains unresolved atN¯\\bar\{N\}\. For a nonmonotone curve, an early low\-loss excursion can place the finite\-range infimum below the asymptotic lower envelope\. A plateau estimated from a bounded empirical range may therefore combineLA,O∞L^\{\\infty\}\_\{A,O\}with an unresolved tail\.
Within an active scaling regime, the standard phenomenological form becomes
L\(N,A,O,𝒟\)≈LA,O∞\(𝒟\)\+C\(A,O,𝒟\)N−α\(A,O,𝒟\)\.L\(N;A,O,\\mathcal\{D\}\)\\approx L^\{\\infty\}\_\{A,O\}\(\\mathcal\{D\}\)\+C\(A,O,\\mathcal\{D\}\)N^\{\-\\alpha\(A,O,\\mathcal\{D\}\)\}\.
The terms have distinct roles\.α\\alphadescribes the rate at which the coupling resolves the accessible task\-relevant tail in that regime, andCCcaptures its scale\.Leff\(N¯\)L^\{\\mathrm\{eff\}\}\(\\bar\{N\}\)summarizes the best loss attained over the feasible range\. A geometry measurement made at a particular budgetN¯\\bar\{N\}is state\-matched toL\(N¯\)L\(\\bar\{N\}\)\. It is also state\-matched toLeff\(N¯\)L^\{\\mathrm\{eff\}\}\(\\bar\{N\}\)when the curve is monotone over the studied range\. A strictL∞L^\{\\infty\}contains target structure outside the ultimately reachable support together with irreducible loss\. Its empirical identification requires a range that separates an asymptote from curvature or an unresolved tail\. For realistic networks, the expression is a regime\-indexed hypothesis: its exponent is tied to a stated resource axis and fitted range, allowing broken or changing laws across scales\.
This notation separates two roles of optimization\. In the general framework,OOcan affectL∞L^\{\\infty\}if the training mechanism changes the support ultimately reached under the specified protocol\. In the minimal model of §3\.3, support is assigned to architecture andOOchanges priority order\. The strict floor in that special case is therefore architecture\- and data\-dependent, whileOOchanges the finite\-budget residual, prefactor, and exponent\. Existing optimizer evidence concerns finite\-budget residuals, prefactors, and exponents; strict\-floor estimation is a next empirical target\.
The dual bottleneck yields three regimes\.
*Accessibility\-unconstrained regime\.*The task projection ofℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)covers the relevant structure over the fitted range\. Scaling is then governed by the applicable data\-side regime, and system differences may be confined to constants or vanish within a prespecified equivalence margin\.
*Accessibility\-constrained regime\.*Some relevant directions are absent, weakly represented, or reached too late\. The observed exponent can change, and finite\-budget residual loss rises\. If the omitted structure lies outside ultimate support, the strict floor rises as well\.
*Task\-aligned regime\.*A coupling covers the relevant structure with little representational waste and prioritizes high\-value directions\. It can combine a low attainable loss with a steep fitted slope for that task\. Alignment and any resulting advantage are task\-specific\.
### 3\.3 A Minimal Solvable Instance
The following construction provides an exact witness for the mechanism\. It specializes the budget\-relative framework by separating a protocol\-asymptotic support from finite\-budget acquisition\. Letf∗∈L2\(𝒟\)f^\{\*\}\\in L^\{2\}\(\\mathcal\{D\}\)be a target expanded in an orthonormal basis\{φk\}k≥1\\\{\\varphi\_\{k\}\\\}\_\{k\\geq 1\}:
f∗=∑kckφk,ck=⟨f∗,φk⟩\.f^\{\*\}=\\sum\_\{k\}c\_\{k\}\\varphi\_\{k\},\\qquad c\_\{k\}=\\langle f^\{\*\},\\varphi\_\{k\}\\rangle\.
Represent a coupling by two objects: an architecture\-family supportSA⊆ℕS\_\{A\}\\subseteq\\mathbb\{N\}, defined as the union of directions available as the modeled capacity coordinate grows under the fixed protocol, and a priority orderπA,O\\pi\_\{A,O\}enumeratingSAS\_\{A\}\. For the asymptotic construction,SAS\_\{A\}is countably infinite andπA,O:ℕ→SA\\pi\_\{A,O\}:\\mathbb\{N\}\\to S\_\{A\}is a bijection\. A finite support exhausts after\|SA\|\|S\_\{A\}\|and belongs to a finite\-tail regime\. We stipulate joint realizability of successive prefixes by a nested capacity family\. ThusSAS\_\{A\}denotes the protocol\-asymptotic support in this idealized instance\. At budgetNN, the learner’s finite\-budget access consists of the firstNNprioritized modes,
f^N=∑r=1NcπA,O\(r\)φπA,O\(r\)\.\\hat\{f\}\_\{N\}=\\sum\_\{r=1\}^\{N\}c\_\{\\pi\_\{A,O\}\(r\)\}\\varphi\_\{\\pi\_\{A,O\}\(r\)\}\.
Orthonormality gives the exact decomposition
L\(f^N\)=∑k∉SAck2⏟LA∞\(T\)\+∑r\>NcπA,O\(r\)2⏟EA,O,T\(N\)\.L\(\\hat\{f\}\_\{N\}\)=\\underbrace\{\\sum\_\{k\\notin S\_\{A\}\}c\_\{k\}^\{2\}\}\_\{L\_\{A\}^\{\\infty\}\(T\)\}\+\\underbrace\{\\sum\_\{r\>N\}c\_\{\\pi\_\{A,O\}\(r\)\}^\{2\}\}\_\{E\_\{A,O,T\}\(N\)\}\.
Proposition 1 \(architecture\-dependent strict floor\)\.LetVA=span¯\{φk:k∈SA\}V\_\{A\}=\\overline\{\\mathrm\{span\}\}\\\{\\varphi\_\{k\}:k\\in S\_\{A\}\\\}\. Then
LA∞\(T\)=‖ΠVA⟂f∗‖2\.L\_\{A\}^\{\\infty\}\(T\)=\\\|\\Pi\_\{V\_\{A\}^\{\\perp\}\}f^\{\*\}\\\|^\{2\}\.
The strict floor is the target energy outside architectural support\. It vanishes if and only if the task\-relevant support off∗f^\{\*\}is contained inSAS\_\{A\}\. Within this construction it is invariant to priority order and finite budget\. The proof is given in Appendix B\.2\.
At maximum feasible budgetN¯\\bar\{N\}, however,
LA,O,Teff\(N¯\)=LA∞\(T\)\+EA,O,T\(N¯\),L^\{\\mathrm\{eff\}\}\_\{A,O,T\}\(\\bar\{N\}\)=L\_\{A\}^\{\\infty\}\(T\)\+E\_\{A,O,T\}\(\\bar\{N\}\),
so the full coupling governs attainable loss, while architectural support alone governs the strict floor in this construction\.
For the residual rate, fix a taskTTand letIA,O\(N\)=\{πA,O\(r\):1≤r≤N\}I\_\{A,O\}\(N\)=\\\{\\pi\_\{A,O\}\(r\):1\\leq r\\leq N\\\}be the acquired set\. List the modes inSAS\_\{A\}with nonzero target power asi1,i2,…i\_\{1\},i\_\{2\},\\ldotsin nonincreasing target power and writeaA,T,j=cij2a\_\{A,T,j\}=c\_\{i\_\{j\}\}^\{2\}\. Ties are resolved by a deterministic rule fixed ex ante and independently of the coupling\. Let
KA,O,T\(N\)=\{j:ij∈IA,O\(N\)\}K\_\{A,O,T\}\(N\)=\\\{j:i\_\{j\}\\in I\_\{A,O\}\(N\)\\\}
be the acquired rank set and define the*canonical completed\-prefix coverage*
JA,O,T\(N\)=max\{m≥0:\{1,…,m\}⊆KA,O,T\(N\)\},J\_\{A,O,T\}\(N\)=\\max\\bigl\\\{m\\geq 0:\\\{1,\\ldots,m\\\}\\subseteq K\_\{A,O,T\}\(N\)\\bigr\\\},
with the maximum of the empty prefix equal to zero\. Finally define the cumulative supported target tail
A¯A,T\(m\)=∑j\>maA,T,j\.\\overline\{A\}\_\{A,T\}\(m\)=\\sum\_\{j\>m\}a\_\{A,T,j\}\.
Proposition 2 \(tail–coverage bracket and product law\)\.SupposeJA,O,T\(N\)→∞J\_\{A,O,T\}\(N\)\\to\\inftyand the cumulative\-tail and coverage log\-rates exist:
γA,T=limm→∞−logA¯A,T\(m\)logm∈\(0,∞\),ρA,O,T=limN→∞logJA,O,T\(N\)logN∈\(0,1\]\.\\gamma\_\{A,T\}=\\lim\_\{m\\to\\infty\}\\frac\{\-\\log\\overline\{A\}\_\{A,T\}\(m\)\}\{\\log m\}\\in\(0,\\infty\),\\qquad\\rho\_\{A,O,T\}=\\lim\_\{N\\to\\infty\}\\frac\{\\log J\_\{A,O,T\}\(N\)\}\{\\log N\}\\in\(0,1\]\.
Then, for everyNN,
A¯A,T\(N\)≤EA,O,T\(N\)≤A¯A,T\(JA,O,T\(N\)\),\\overline\{A\}\_\{A,T\}\(N\)\\;\\leq\\;E\_\{A,O,T\}\(N\)\\;\\leq\\;\\overline\{A\}\_\{A,T\}\\\!\\left\(J\_\{A,O,T\}\(N\)\\right\),
and hence
ρA,O,TγA,T≤lim infN→∞−logEA,O,T\(N\)logN≤lim supN→∞−logEA,O,T\(N\)logN≤γA,T\.\\rho\_\{A,O,T\}\\gamma\_\{A,T\}\\;\\leq\\;\\liminf\_\{N\\to\\infty\}\\frac\{\-\\log E\_\{A,O,T\}\(N\)\}\{\\log N\}\\;\\leq\\;\\limsup\_\{N\\to\\infty\}\\frac\{\-\\log E\_\{A,O,T\}\(N\)\}\{\\log N\}\\;\\leq\\;\\gamma\_\{A,T\}\.
If bounded off\-prefix gain holds—that is, for someη∈\(0,1\]\\eta\\in\(0,1\]and all sufficiently largeNN,
EA,O,T\(N\)≥ηA¯A,T\(JA,O,T\(N\)\),E\_\{A,O,T\}\(N\)\\geq\\eta\\,\\overline\{A\}\_\{A,T\}\\\!\\left\(J\_\{A,O,T\}\(N\)\\right\),
then the residual exponent exists and satisfies
αA,O,T=ρA,O,TγA,T\.\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\\gamma\_\{A,T\}\.
The same product law holds under the weaker rate condition
A¯A,T\(JA,O,T\(N\)\)EA,O,T\(N\)=No\(1\)\.\\frac\{\\overline\{A\}\_\{A,T\}\(J\_\{A,O,T\}\(N\)\)\}\{E\_\{A,O,T\}\(N\)\}=N^\{o\(1\)\}\.
If, in addition,
aA,T,j≍j−bA,T,bA,T\>1,JA,O,T\(N\)≍NρA,O,T,a\_\{A,T,j\}\\asymp j^\{\-b\_\{A,T\}\},\\qquad b\_\{A,T\}\>1,\\qquad J\_\{A,O,T\}\(N\)\\asymp N^\{\\rho\_\{A,O,T\}\},
thenγA,T=bA,T−1\\gamma\_\{A,T\}=b\_\{A,T\}\-1and bounded off\-prefix gain yields the stronger asymptotic\-order result
EA,O,T\(N\)≍N−ρA,O,T\(bA,T−1\)\.E\_\{A,O,T\}\(N\)\\asymp N^\{\-\\rho\_\{A,O,T\}\(b\_\{A,T\}\-1\)\}\.
The lower bound is a best\-NN\-mode bound: budgetNNcan resolve at mostNNnonzero\-target modes, and the firstNNranks capture the greatest possible target power\. The upper bound uses the completed prefix\. The bounded\-gain assumption makes these bounds rate\-matched; it can be weakened from a constant\-factor condition to the displayed subpolynomial ratio\. WhenρA,O,T=1\\rho\_\{A,O,T\}=1, the universal interval collapses andαA,O,T=γA,T\\alpha\_\{A,O,T\}=\\gamma\_\{A,T\}even without a separate off\-prefix assumption\. Appendix B\.2 proves the result, quantifies the gap between the two bounds, and gives a fixed nested order in which dispersed acquisition adds a second rate component\. Figure[2](https://arxiv.org/html/2609.03533#Sx7.F2)evaluates the equality cases and the interleaved counterexample at finite budgets\.
Because the supported sequence is defined after restriction toSAS\_\{A\},γA,T\\gamma\_\{A,T\}, likebA,Tb\_\{A,T\}, is generally architecture–task dependent\. Common full support in a basis fixed ex ante independently of the coupling is one sufficient condition under which compared systems share a coupling\-invariantγT\\gamma\_\{T\}\. With full support andρ=1\\rho=1, the representational floor vanishes and the residual rate reduces to the common task\-side tail rate\.
Corollary 1 \(coupling\-by\-task reversal\)\.Consider two couplingsq1=\(A1,O1\)q\_\{1\}=\(A\_\{1\},O\_\{1\}\)andq2=\(A2,O2\)q\_\{2\}=\(A\_\{2\},O\_\{2\}\)and two tasksT1,T2T\_\{1\},T\_\{2\}\. Suppose that, within each taskTtT\_\{t\}, the two couplings have full support in the same ex ante task\-side basis and induce a common cumulative supported\-tail log\-rate
γA1,Tt=γA2,Tt=:γt\>0\.\\gamma\_\{A\_\{1\},T\_\{t\}\}=\\gamma\_\{A\_\{2\},T\_\{t\}\}=:\\gamma\_\{t\}\>0\.
Suppose also that both couplings satisfy Proposition 2’s bounded or subpolynomial off\-prefix condition\. If
ρq1,T1\>ρq2,T1andρq1,T2<ρq2,T2,\\rho\_\{q\_\{1\},T\_\{1\}\}\>\\rho\_\{q\_\{2\},T\_\{1\}\}\\qquad\\text\{and\}\\qquad\\rho\_\{q\_\{1\},T\_\{2\}\}<\\rho\_\{q\_\{2\},T\_\{2\}\},
then their residual\-exponent ordering reverses across tasks\. The corresponding residual\-loss ordering reverses for all sufficiently largeNN; equal within\-task floors extend the conclusion to total loss\. Equal support, common tail behavior, and equal coverage rates imply no exponent difference in this instance\.
The construction supplies an existence witness for the support–acquisition mechanism\. Under stipulated support, acquisition order, and an omniscient within\-order allocation, it generates coupling\-dependent exponents and cross\-task reversals\. Its abstraction from sample noise and optimization error isolates the quantities that a fuller theory must derive from architecture and training dynamics: support, cumulative supported\-tail behavior, acquisition order, and the completed\-prefix coverage rate\.
Formal\-to\-empirical bridge\.The solvable quantities identify the targets for realistic measurement\.SAS\_\{A\}corresponds to protocol\-asymptotic task\-relevant support in this instance, andIA,O\(N\)I\_\{A,O\}\(N\)to the finite\-budget acquired set\. The cumulative\-tail rateγA,T\\gamma\_\{A,T\}summarizes the remaining target power after a high\-value prefix;bA,T−1b\_\{A,T\}\-1is its power\-law special case\.ρA,O,T\\rho\_\{A,O,T\}measures how quickly the coupling completes that prefix\. The tail is basis\-relative; its data\-side reading requires a basis fixed ex ante independently of the coupling or supplied by a task/data\-side operator\. Empirical interpretation ofLA∞L\_\{A\}^\{\\infty\}uses a scale range that separates the asymptote from slow acquisition\. Section 3\.5 separates static geometry levels from rate\-sensitive coverage trajectories\. A confirmatory product\-law test preregisters a prefix\-adequacy route: either a rank\-resolved off\-prefix diagnostic or a design\- or theory\-based argument that makes off\-prefix gain bounded or exponent\-neutral\. All quantities are fixed independently of the fitted loss curve\. The empirical test asks whether they predict matched\-budget loss and the direction of exponent differences under the stated tail and acquisition conditions\.
### 3\.4 Effective Spectrum and a Fixed\-Kernel Specialization
For a fixed task and data distribution, an effective spectrum operationalizes one projection of representational accessibility: which directions are represented, with what strength, and at what learning rate\. Kernel and random\-feature theories make this description precise through an architecture\- and data\-dependent eigenspectrum together with the target’s decomposition over its eigenfunctions\. Target power in large\-eigenvalue modes is typically learned earlier, whereas power in small\- or zero\-eigenvalue modes is learned slowly or remains unexpressed \(Bordelon et al\., 2020; Canatar et al\., 2021; Bahri et al\., 2024\)\.
The three model\-side channels enter this projection differently\.
- •Structural accessibility\.Architecture shapes support, basis directions, information flow, and inductive bias\. Wiring\-level results on copying and associative recall illustrate that architectural constraints can create large task\-specific efficiency gaps \(Jelassi et al\., 2024; Arora et al\., 2024\)\.
- •Dynamical accessibility\.Feature learning changes the representation realized during training and can reweight task\-relevant directions relative to the initial kernel \(Yang and Hu, 2021; Bordelon et al\., 2025\)\.
- •Optimization\-conditioned accessibility\.Preconditioning and changes in the training metric can alter the order or rate at which already represented directions are resolved while preserving support in the controlled setting of Ramani and Jain \(2026\)\.
Superposition supplies one concrete operational instance\. In the controlled model of Liu, Liu, and Gore \(2025\), overlap among feature vectors contributes an interference term to loss, and a weight\-decay intervention changes the superposition regime\. Within Coupled Scaling, this can be interpreted as changing the effective strength and interference of accessible directions\. More generally, the common object across the three channels is budget\-relative learnability: architecture can move a support frontier, feature learning can reshape realized geometry, and optimization can change the rate at which the system approaches that frontier\. The support/order split in §3\.3 is a minimal formal separation of these roles; richer models may allow feature learning and optimization to change both support and order\.
Kernel correspondence\.In kernel regression and appropriate infinite\-width lazy\-training limits, the task projection of accessibility has a precise analogue\. Learning curves depend on sample size, kernel eigendirections and eigenvalues, target alignment, and—where applicable—noise and regularization \(Bordelon et al\., 2020; Canatar et al\., 2021; Caponnetto and De Vito, 2007\)\. Kernel support corresponds to available directions, while eigenvalue\-weighted target alignment supplies an acquisition hierarchy\. The lazy\-limit analogue ofℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)is therefore a fixed spectrum\-and\-alignment description conditional on the remaining statistical assumptions\. This anchors the framework in a known limit\. When representation is fixed, Coupled Scaling recovers the static spectrum\-and\-alignment description\. Beyond that limit, independently measured evolving geometry supplies the empirical bridge to feature learning and cross\-family variation\.
Restricted operational anchor\.Letqqindex a bounded, positive semidefinite, self\-adjoint fixed\-kernel operatorKqK\_\{q\}onL2\(𝒟\)L^\{2\}\(\\mathcal\{D\}\), while the optimizer and noiseless population gradient\-flow dynamics are held fixed\. Assume a complete countable orthonormal eigenbasis\{ϕq,k\}k\\\{\\phi\_\{q,k\}\\\}\_\{k\}, including the zero\-eigenvalue subspace, with eigenvaluesλq,k≥0\\lambda\_\{q,k\}\\geq 0, and writecq,T,k=⟨fT,ϕq,k⟩c\_\{q,T,k\}=\\langle f\_\{T\},\\phi\_\{q,k\}\\rangle\. For a nonzero targetfT∈L2\(𝒟\)f\_\{T\}\\in L^\{2\}\(\\mathcal\{D\}\), define the normalized task\-weighted spectral measure
νq,T=∑kcq,T,k2∑jcq,T,j2δλq,k\.\\nu\_\{q,T\}=\\sum\_\{k\}\\frac\{c\_\{q,T,k\}^\{\\,2\}\}\{\\sum\_\{j\}c\_\{q,T,j\}^\{\\,2\}\}\\,\\delta\_\{\\lambda\_\{q,k\}\}\.
This object is determined by the fixed kernel spectrum and target alignment before any power\-law fit\. For squared\-loss gradient flow from zero initialization, using residual dynamicsr˙=−Kqr\\dot\{r\}=\-K\_\{q\}r, equivalentlyr\(t\)=e−tKqr\(0\)r\(t\)=e^\{\-tK\_\{q\}\}r\(0\),
Lq,T\(t\)Lq,T\(0\)=∫e−2λtdνq,T\(λ\),Hq,T\(t\)=1−∫e−2λtdνq,T\(λ\)\.\\frac\{L\_\{q,T\}\(t\)\}\{L\_\{q,T\}\(0\)\}=\\int e^\{\-2\\lambda t\}\\,d\\nu\_\{q,T\}\(\\lambda\),\\qquad H\_\{q,T\}\(t\)=1\-\\int e^\{\-2\\lambda t\}\\,d\\nu\_\{q,T\}\(\\lambda\)\.
Spectral mass on\(0,∞\)\(0,\\infty\)represents target energy learnable under the fixed\-kernel dynamics, whileνq,T\(\{0\}\)\\nu\_\{q,T\}\(\\\{0\\\}\)is the normalized strict\-residual fraction\. The spectral filtere−2λte^\{\-2\\lambda t\}supplies a dynamics\-derived soft acquisition profile, andHq,T\(t\)H\_\{q,T\}\(t\)is the corresponding resolved\-energy summary\. Independent geometry measurements enter the empirical tests separately\.
Proposition 3 \(classical fixed\-kernel spectral\-tail bridge\)\.Letmq,T=νq,T\(\{0\}\)m\_\{q,T\}=\\nu\_\{q,T\}\(\\\{0\\\}\)andFq,T\(x\)=νq,T\(\(0,x\]\)F\_\{q,T\}\(x\)=\\nu\_\{q,T\}\(\(0,x\]\)\. If, for someβq,T\>0\\beta\_\{q,T\}\>0,Fq,TF\_\{q,T\}is regularly varying at zero,
limx↓0Fq,T\(ax\)Fq,T\(x\)=aβq,T\(a\>0\),\\lim\_\{x\\downarrow 0\}\\frac\{F\_\{q,T\}\(ax\)\}\{F\_\{q,T\}\(x\)\}=a^\{\\beta\_\{q,T\}\}\\qquad\(a\>0\),
then, ast→∞t\\to\\infty,
Lq,T\(t\)−Lq,T∞Lq,T\(0\)∼Γ\(βq,T\+1\)Fq,T\(\(2t\)−1\),Lq,T∞=mq,TLq,T\(0\)\.\\frac\{L\_\{q,T\}\(t\)\-L\_\{q,T\}^\{\\infty\}\}\{L\_\{q,T\}\(0\)\}\\sim\\Gamma\(\\beta\_\{q,T\}\+1\)F\_\{q,T\}\(\(2t\)^\{\-1\}\),\\qquad L\_\{q,T\}^\{\\infty\}=m\_\{q,T\}L\_\{q,T\}\(0\)\.
HereΓ\(⋅\)\\Gamma\(\\cdot\)is the Euler gamma function\.
In particular, ifFq,T\(x\)∼cq,Txβq,Tℓq,T\(1/x\)F\_\{q,T\}\(x\)\\sim c\_\{q,T\}x^\{\\beta\_\{q,T\}\}\\ell\_\{q,T\}\(1/x\), whereℓq,T\\ell\_\{q,T\}is slowly varying, then the residual loss is asymptotic to
Lq,T\(0\)cq,TΓ\(βq,T\+1\)\(2t\)−βq,Tℓq,T\(2t\)\.L\_\{q,T\}\(0\)c\_\{q,T\}\\Gamma\(\\beta\_\{q,T\}\+1\)\(2t\)^\{\-\\beta\_\{q,T\}\}\\ell\_\{q,T\}\(2t\)\.
Thus independently specified spectral mass near zero determines a fixed\-kernel training\-time exponentαt\(q,T\)=βq,T\\alpha\_\{t\}\(q,T\)=\\beta\_\{q,T\}, up to a slowly varying correction\. To see the result, suppressq,Tq,T, separate the zero atom, sets=2ts=2t, and use Stieltjes integration by parts:
∫\(0,∞\)e−sλ𝑑F\(λ\)=s∫0∞e−sλF\(λ\)𝑑λ=∫0∞e−uF\(u/s\)𝑑u\.\\int\_\{\(0,\\infty\)\}e^\{\-s\\lambda\}\\,dF\(\\lambda\)=s\\int\_\{0\}^\{\\infty\}e^\{\-s\\lambda\}F\(\\lambda\)\\,d\\lambda=\\int\_\{0\}^\{\\infty\}e^\{\-u\}F\(u/s\)\\,du\.
After division byF\(1/s\)F\(1/s\), regular variation and Potter bounds give dominated convergence to∫0∞e−uuβ𝑑u=Γ\(β\+1\)\\int\_\{0\}^\{\\infty\}e^\{\-u\}u^\{\\beta\}du=\\Gamma\(\\beta\+1\)\. Proposition 3 uses the Abelian direction of the classical Karamata Laplace–Stieltjes theorem \(Bingham, Goldie, and Teugels, 1987\) as a restricted calibration: regular variation of the task\-weighted spectral measure implies the stated loss asymptotic\. The reverse identification from an observed power\-law loss curve requires additional Tauberian conditions\. A positive spectral gap yields exponential decay and defines a different regime\.
For the positive eigenmodes indexed so thatλ1≥λ2≥⋯↓0\\lambda\_\{1\}\\geq\\lambda\_\{2\}\\geq\\cdots\\downarrow 0, letwj=cq,T,j2/∑kcq,T,k2w\_\{j\}=c\_\{q,T,j\}^\{2\}/\\sum\_\{k\}c\_\{q,T,k\}^\{2\}be their weights relative to total target energy, including any zero\-eigenvalue target mass in the denominator\. If this eigenvalue order is asymptotically aligned with the target\-power order of Proposition 2 andλj∼κj−a\\lambda\_\{j\}\\sim\\kappa j^\{\-a\},wj∼dj−bw\_\{j\}\\sim dj^\{\-b\}, whereκ,d\>0\\kappa,d\>0,a\>0a\>0, andb\>1b\>1,
F\(x\)∼db−1κ−\(b−1\)/ax\(b−1\)/a\.F\(x\)\\sim\\frac\{d\}\{b\-1\}\\kappa^\{\-\(b\-1\)/a\}x^\{\(b\-1\)/a\}\.
Because the filter effectively resolves modes withtλj≳1t\\lambda\_\{j\}\\gtrsim 1, its soft cutoff satisfiesJeff\(t\)≍t1/aJ\_\{\\mathrm\{eff\}\}\(t\)\\asymp t^\{1/a\}\. Henceβ=\(b−1\)/a\\beta=\(b\-1\)/amatches the rate product with the analoguesγ=b−1\\gamma=b\-1andρeff=1/a\\rho\_\{\\mathrm\{eff\}\}=1/a\. This soft\-filter correspondence is distinct from the hard acquired\-set bounds in Proposition 2\. If one setsN=tN=tand retains the normalizationJ\(N\)≤NJ\(N\)\\leq N, thena≥1a\\geq 1is required\. For asymptotically misaligned eigenvalue and target\-power orders, the joint spectral measureFFis the appropriate object\. Oppositeβ\\beta\-orderings across two tasks give the same eventual residual\-order reversal as Corollary 1; equal within\-task floors extend the reversal to total loss\.
The construction provides a mathematically closed fixed\-kernel instantiation of the framework; feature\-learning networks require an independently specified accessibility measure\.
### 3\.5 Operationalization and Measurement
The empirical design maps each theoretical object to a prespecified proxy and comparison\. Table 1 summarizes the resulting measurement and identification structure\.
Table 1\.Theoretical objects, empirical roles, candidate measurements, and identification status\.
ObjectEmpirical roleCandidate measurementIdentification statusℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)Budget\-relative reachable set conditional on the training problemPrespecified projections of a fixed representation descriptorTheoretical object; full\-set recovery is unspecifiedPT,𝒟ℛB\(A,O,𝒫\)P\_\{T,\\mathcal\{D\}\}\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\)Task\-relevant projectionCross\-system probes, task\-conditioned representation comparisonPartly measurableStatic task\-conditioned geometry levelgq,tlevel\(N¯\)g^\{\\mathrm\{level\}\}\_\{q,t\}\(\\bar\{N\}\)Task\-relevant breadth or alignment at a common finite budget; primary outcomeLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)Held\-out probes, effective rank, alignment, or task\-conditioned spectral summariesPartly measurable; proxy\-dependentCumulative\-tail rateγA,T\\gamma\_\{A,T\}Task\-power mass remaining after a high\-value prefixEx ante task basis, supported target spectrum, or controlled synthetic constructionPrecise in the witness; basis\- and support\-dependent in generalCompleted\-prefix rateρA,O,T\\rho\_\{A,O,T\}Rate at which high\-value task structure becomes accessible; primary outcomeαq,t\\alpha\_\{q,t\}under the off\-prefix conditionJq,t\(g\)\(N\)J^\{\(g\)\}\_\{q,t\}\(N\), a preregistered geometry trajectoryPrecise in the witness; proxy\-based in generalFixed\-kernel spectral tailβq,T\\beta\_\{q,T\}Rate\-sensitive fixed\-kernel quantity; primary outcomeαt\(q,T\)\\alpha\_\{t\}\(q,T\)Near\-zero tail ofFq,TF\_\{q,T\}Precise under Proposition 3 conditionsScaling saturationConsequence of limited accessibilityRegime transition, fitted curvature, or plateauDiagnostic proxy requiring independent validationThe static quantitygq,tlevel\(N¯\)g^\{\\mathrm\{level\}\}\_\{q,t\}\(\\bar\{N\}\)summarizes task\-relevant breadth, alignment, or effective rank at the selected comparison budget, typically the largest feasible point on the resource grid\. A directional contrast uses a scalar proxy or scalarization fixed in advance\. Candidate measurements include the subspace used for effective adaptation \(Aghajanyan et al\., 2021\), activation\-space intrinsic dimension \(Ansuini et al\., 2019\), anddalignd\_\{\\mathrm\{align\}\}, which tests whether a local residual direction transfers from training to held\-out data \(Zhang et al\., 2026\)\. In the fixed\-kernel setting, the full task\-weighted spectral measureνq,T\\nu\_\{q,T\}supplies the corresponding geometry\.
The level is paired withLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)under a common resource convention\. Parameter\-matched and compute\-matched comparisons form separate analyses because they may rank the same systems differently\. The couplings also share the probe population, task projection, representation layer, and scalar orientation\. Repetition across seeds identifies stable geometric separation\. The finite\-range summaryLq,teff\(N¯\)L^\{\\mathrm\{eff\}\}\_\{q,t\}\(\\bar\{N\}\)may be reported alongside the primary state\-matched loss and coincides with it when the curve is monotone over the studied range\.
Exponent prediction uses a trajectory\. Letgq,t\(N\)g\_\{q,t\}\(N\)be the held\-out geometry measurement at each scale on the fixed grid\. When the measure has a completed high\-value\-prefix interpretation, defineJq,t\(g\)\(N\)J^\{\(g\)\}\_\{q,t\}\(N\)and estimate
Jq,t\(g\)\(N\)≍Nρq,t\(g\)\.J^\{\(g\)\}\_\{q,t\}\(N\)\\asymp N^\{\\rho^\{\(g\)\}\_\{q,t\}\}\.The empirical rateρq,t\(g\)\\rho^\{\(g\)\}\_\{q,t\}is the analogue of the canonical completed\-prefix rate in Proposition 2\. Confirmatory cells follow one of two prefix\-adequacy routes\. A rank\-resolved route compares unresolved supported energy with the cumulative tail beyond the completed prefix, using geometry quantities specified before loss fitting\. A design\- or theory\-based route establishes bounded or exponent\-neutral off\-prefix gain through construction or independent theory\. Polynomial off\-prefix acceleration selects the universal exponent interval in place of the product\-law endpoint\.
In the fixed\-kernel specialization, the near\-zero spectral\-tail indexβq,T\\beta\_\{q,T\}supplies the rate\-sensitive quantity directly; Appendix B\.2 gives the corresponding diagnostic, finite\-budget illustration, and counterexample\. Level and rate can move independently: systems can differ atN¯\\bar\{N\}yet grow at similar rates, or cross within the measured range despite different rates\. Estimatingρ\(g\)\\rho^\{\(g\)\}therefore uses the full trajectory and its uncertainty\. Nikolaou et al\. \(2026\) track the eNTK eigenvalue scale currently driving loss reduction through a loss\-gradient\-weighted spectral position, while Liu, Paquette, and Sous \(2026\) use activation\-covariance and per\-sample gradient spectra as early diagnostics of optimization and token efficiency\. A rank\-and\-coverage interpretation fixed before loss fitting connects these evolving spectra to the completed\-prefix quantities\. Saturation diagnostics can then locate candidate bottlenecks, cross\-system representation comparisons measure task\-relevant geometry, and direct inductive\-bias interventions test a manipulable structural mechanism such as locality or equivariance\. The predictions below pair static geometry with matched\-budget loss and completed\-prefix growth with exponent ordering under the stated tail and prefix\-adequacy criteria\.
### 3\.6 Testable Predictions
Coupled Scaling yields two discriminating predictions\. At a common comparison budget, greater task\-relevant geometric access should predict lower loss on the corresponding task\. Across scales, a completed\-prefix rate measured apart from performance should predict the fitted exponent when the within\-task cumulative\-tail rate is common and off\-prefix gain is exponent\-neutral\. Tasks favoring the opposite coupling should reverse both orderings\. Equivalent task\-relevant access defines the null, with coupling\-specific exponent differences expected to fall within a prespecified margin\.
The exact decomposition in §3\.3 motivates the level prediction by expressing fixed\-budget loss as target energy outside architectural support plus an unresolved within\-support tail\. Its empirical use requires a preoriented scalar geometry measure and a common resource convention\. Proposition 2 supplies the formal route for the rate prediction: the geometry trajectory carries the completed\-prefix meaning, the compared couplings share the cumulative supported\-tail rate within a task, and off\-prefix gain is bounded or exponent\-neutral\. Proposition 3 supplies the fixed\-kernel comparison between an independently measured spectral\-tail index and the training\-time exponent\. Cross\-task geometry interactions additionally require common standardization\.
Three outcome patterns challenge the accessibility hypothesis under the stated conditions: verified geometry separation paired with equivalent matched\-budget loss or exponent; a scaling shift opposite to the measured rate; or an interaction accounted for by task\-side variables after the coupling intervention and geometry separation are established\. In a mode\-resolved test, a residual exponent outside the universal interval\[ργ,γ\]\[\\rho\\gamma,\\gamma\]also violates the solvable witness\.
## 4\. Evidence, Identification, and a Direct Test
### 4\.1 What Controlled Studies Establish
Same\-data comparisons provide the clearest evidence that scaling can depend on the learning system\. Ngo and Ravanbakhsh \(2026\) compare equivariant and non\-equivariant neural force fields on the same task and data, finding architecture\-dependent parameter\-, data\-, and compute\-scaling exponents\. Tay et al\. \(2023\) observe scaling differences and rank changes across ten language\-model architectures under a matched pretraining and evaluation protocol\. Structured shallow\-network theories add representation\-limited regimes, phase transitions, and spectral structure, and feature\-learning analyses connect risk exponents to the learned representation\. These results show that architecture and structural alignment can shape scaling under fixed or closely controlled data\. Each effect remains indexed to the resource axis measured: parameter\-, data\-, compute\-, and training\-time exponents are distinct empirical quantities\.
Feature learning and optimization supply complementary mechanisms\. Bordelon et al\. \(2025\) obtain different training\-time and compute exponents for hard tasks under feature learning\. Ramani and Jain \(2026\) obtain optimizer\-dependent fitted model\-size exponents in a random\-feature setting\. Jha and Reagen \(2026\) measure optimizer\-dependent FFN effective\-rank scaling in a common GPT\-style family and include an extended\-training comparison at matched perplexity\. Activation and gradient spectra offer multiscale diagnostics of optimization\.
Superposition and capacity\-based mechanisms broaden this picture\. Liu, Liu, and Gore \(2025\) show that superposition\-induced geometric interference can generate robust model\-width scaling, while capacity, interference, and rare\-task retention provide another route through which finite resources shape access to task features\. Shared\-exponent cases identify the corresponding null regime\. In Bansal et al\. \(2022\), the tested architecture, task setup, filtering, and iid\-noise interventions have little effect on data\-scaling exponents; back\-translated data is the reported exception and degrades the fitted exponent\. Volkova et al\. \(2026\) find ill\-conditioned separate optimizer fits and improved stability and extrapolation under a constrained shared\-exponent rescaling\. Current evidence supports both difference and equivalence hypotheses at the level of finite\-range exponent or geometry effects\. Architecture\-specific strict floors remain an empirical target\. A direct Coupled Scaling test therefore brings four elements into one design: an explicit coupling intervention, contrasting tasks, multiscale geometry measurement, and a separately fitted scaling interaction\. Appendix C\.1 maps the existing evidence to these elements; §§4\.2–4\.3 derive the corresponding controls and test\.
### 4\.2 An Identification Audit of Released Emergence Trajectories
We use the released emergence trajectories to ask what a direct test of Coupled Scaling must control\. Letxxbe a monotone scaling coordinate and suppose that, within a fitted regime,
Lt\(x,A,O\)=Lt∞\(A,O\)\+Ct\(A,O\)x−αt\(A,O\)\.L\_\{t\}\(x;A,O\)=L\_\{t\}^\{\\infty\}\(A,O\)\+C\_\{t\}\(A,O\)x^\{\-\\alpha\_\{t\}\(A,O\)\}\.
For a task\-specific thresholdεt\\varepsilon\_\{t\}, the corresponding continuous crossing scale is
xt\(A,O\)=\(Ct\(A,O\)εt−Lt∞\(A,O\)\)1/αt\(A,O\)\.x\_\{t\}\(A,O\)=\\left\(\\frac\{C\_\{t\}\(A,O\)\}\{\\varepsilon\_\{t\}\-L\_\{t\}^\{\\infty\}\(A,O\)\}\\right\)^\{1/\\alpha\_\{t\}\(A,O\)\}\.
Thus the exponent, prefactor, and attainable floor can each change an emergence scale\. For unrelated tasks, crossing order can reverse across couplings because each task has its own curve parameters\. A valid prerequisite ordering has an additional task\-side interpretation: success on a composite must entail competence on its component at commensurate thresholds\. Appendix A gives the continuous and discrete validity conditions\.
Liu et al\. \(2026\) find a stable capability\-emergence order across the released model families and describe it as an implicit curriculum\. The Coupled Scaling question is whether, after valid task dependencies are separated out, the remaining order has a component that changes with the learning system and follows representation geometry measured on its own\. The release helps specify the required controls because it spans several families and scales while combining those changes with training and evaluation differences\.
The released data contain four accuracy trajectories from three open\-weight model families\. We evaluate five threshold\-based measures and two threshold\-free measures—maximum\-improvement timing \(max\-slope\) and normalized trajectory area under the curve \(AUC\)\. The primary universe contains 29 tasks and 46 prerequisite edges, and a registry\-native universe provides a robustness check\. The matched estimand compares prerequisite\-linked elemental–composite pairs \(category A\) with cross\-level but non\-prerequisite pairs \(C1C\_\{1\}\) within the same composite:
Δ\(c\)=S¯A\(c\)−S¯C1\(c\),Δ¯=\|𝒞\|−1∑c∈𝒞Δ\(c\)\.\\Delta\(c\)=\\bar\{S\}\_\{A\(c\)\}\-\\bar\{S\}\_\{C\_\{1\}\(c\)\},\\qquad\\bar\{\\Delta\}=\|\\mathcal\{C\}\|^\{\-1\}\\sum\_\{c\\in\\mathcal\{C\}\}\\Delta\(c\)\.
Matching fixes the composite; elemental identity, difficulty, format, and exposure remain uncontrolled, so the estimand is descriptive\. Task\-universe construction, validity rules, resampling, and degeneracy checks appear in Appendix A\.
Table 2\.Within\-composite matched contrasts for the two threshold\-free measures\.Δ¯\\bar\{\\Delta\}is the composite\-equal\-weight mean ofΔ\(c\)\\Delta\(c\);rpermr\_\{\\mathrm\{perm\}\}is the tail area under the within\-composite label\-permutation reference distribution in theΔ¯\>0\\bar\{\\Delta\}\>0direction specified before the matched\-contrast implementation and serves as an uncalibrated descriptive reference quantity; ranges are 2\.5–97\.5% composite\-resampling percentiles\. All quantities are descriptive under the observational design \(Appendix A\.5\)\.
MeasureTask setncn\_\{c\}Δ¯\\bar\{\\Delta\}rpermr\_\{\\mathrm\{perm\}\}2\.5–97\.5% percentile rangemax\-slopePrimary17−\-0\.0680\.925\(−\-0\.175, \+0\.033\)max\-slopeRegistry12−\-0\.1031\.000\(−\-0\.228,−\-0\.011\)AUCPrimary20\+0\.0290\.176\(−\-0\.035, \+0\.093\)AUCRegistry17\+0\.0840\.020\(−\-0\.031, \+0\.209\)Table 2 changes the interpretation of the unmatched benchmark\. In the primary universe, the unmatched A–B comparison is positive for max\-slope \(0\.278; permutation\-reference tail area 0\.007\), AUC \(0\.128; 0\.018\), and all five thresholds\. That comparison also mixes prerequisite structure with task level, difficulty, format, and valid\-comparison structure\. The within\-composite A–C1C\_\{1\}contrast holds the composite fixed and sharply attenuates the pattern\. Max\-slope becomes negative in both universes \(−\-0\.068 and−\-0\.103\), whereas AUC remains small and positive \(\+0\.029 and \+0\.084\), with composite\-resampling ranges spanning zero\. The stable conclusion from the matched analysis is measure dependence and substantial attenuation\.
Sensitivity analyses identify AUC as the more stable threshold\-free summary\. Max\-slope changes materially with checkpoint structure, valid\-comparison weighting, and model deletion and is retained as exploratory\. AUC remains positive under every model deletion and after exclusion of six operationally degenerate composites \(\+0\.040 primary; \+0\.044 registry\)\. Appendix A reports the full seven\-measure battery, the alternative integration axis, and the remaining sensitivity analyses\. The audit’s descriptive interpretation therefore rests on AUC\.
The audit therefore serves as an identification analysis\. In the released suite, coupling, corpus, scale, checkpoint structure, and stochastic realization vary jointly\. A direct test requires coupling to vary independently of the training recipe, task structure to be matched or deliberately crossed, geometry to be measured separately from performance, and stochastic variation to be estimated with repeated runs\. Together, these controls define the factorial design in §4\.3\.
### 4\.3 A Discriminating Multiscale Test
The direct test crosses at least two architecture–optimization couplings with at least two task families selected for a predicted alignment reversal\. Within each task instance, the data\-generating process, examples, objective, evaluation, schedule, and seed distribution are held fixed across couplings\. An optimizer claim is tested by crossing optimization explicitly with architecture and task\. Every coupling–task cell spans multiple scales and seeds\. Parameter\-matched and compute\-matched analyses form distinct resource comparisons, and task families are matched on format, exposure, and nominal difficulty when those features enter the claim\. Before training, the design assigns the expected alignment direction for each coupling\. Its primary comparison is the resulting cross\-task reversal\.
Geometry is measured on held\-out data at every scale on the fixed grid\. Letgq,t\(N\)g\_\{q,t\}\(N\)denote the task\-conditioned trajectory\. A preoriented scalar level, or a scalarization fixed before any loss comparison, definesgq,tlevel\(N¯\)g^\{\\mathrm\{level\}\}\_\{q,t\}\(\\bar\{N\}\)and is paired withLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)at the same budget\. The multiscale trajectory supplies the rate comparison;Lq,teff\(N¯\)L^\{\\mathrm\{eff\}\}\_\{q,t\}\(\\bar\{N\}\)remains an optional finite\-range summary\. For a measure with a completed high\-value\-prefix interpretation, defineJq,t\(g\)\(N\)J^\{\(g\)\}\_\{q,t\}\(N\)and estimate
Jq,t\(g\)\(N\)≍Nρq,t\(g\)\.J^\{\(g\)\}\_\{q,t\}\(N\)\\asymp N^\{\\rho^\{\(g\)\}\_\{q,t\}\}\.
The Proposition 2 comparison uses a common cumulative supported\-target\-tail rate within each task,
γA1,Tt=γA2,Tt=:γt\>0\.\\gamma\_\{A\_\{1\},T\_\{t\}\}=\\gamma\_\{A\_\{2\},T\_\{t\}\}=:\\gamma\_\{t\}\>0\.Common full support in a basis fixed ex ante independently of the coupling supplies one experimental route to this condition; in the power\-law case,γt=bt−1\\gamma\_\{t\}=b\_\{t\}\-1\. When the common\-tail condition, the completed\-prefix interpretation, and bounded or exponent\-neutral off\-prefix gain hold, the predicted exponent is
αq,tpred=ρq,t\(g\)γt\.\\alpha^\{\\mathrm\{pred\}\}\_\{q,t\}=\\rho^\{\(g\)\}\_\{q,t\}\\gamma\_\{t\}\.The confirmatory protocol records how each product\-law cell satisfies the common\-tail and prefix\-adequacy requirements\. The first may follow from common full support or an independent task\-side tail analysis\. The second may follow from a rank\-resolved diagnostic or a design\- or theory\-based guarantee\. A cell with polynomial off\-prefix acceleration is evaluated against the universal exponent interval\.
The contrasting task families are chosen so that different couplings acquire high\-value directions more quickly on different tasks, producing opposite within\-task rate orderings\. For fixed\-kernel cells,βq,t\\beta\_\{q,t\}serves as the geometry\-side rate variable\. Across all cells, geometry is measured on the same held\-out population and repeated across seeds and task instances\. Probe targets, normalization, scale sets, and the primary level and rate measures are fixed before the loss curves are fit\. Additional diagnostics remain secondary\. Appendix C\.2 gives the remaining held\-out, naturalistic\-task, curve\-fitting, and inferential specifications\.
For two couplingsq1,q2q\_\{1\},q\_\{2\}, define the static contrasts at a common resource axis, matching convention, and budgetN¯\\bar\{N\}:
dglevel\(t\)=gq1,tlevel\(N¯\)−gq2,tlevel\(N¯\),dloss\(t\)=Lq2,t\(N¯\)−Lq1,t\(N¯\)\.d\_\{g\}^\{\\mathrm\{level\}\}\(t\)=g^\{\\mathrm\{level\}\}\_\{q\_\{1\},t\}\(\\bar\{N\}\)\-g^\{\\mathrm\{level\}\}\_\{q\_\{2\},t\}\(\\bar\{N\}\),\\qquad d\_\{\\mathrm\{loss\}\}\(t\)=L\_\{q\_\{2\},t\}\(\\bar\{N\}\)\-L\_\{q\_\{1\},t\}\(\\bar\{N\}\)\.HereLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)is the specified cell\-level loss estimand, aggregated across seeds and task instances through the stated hierarchical or repeated\-run analysis\. With the chosen sign convention,dloss\(t\)\>0d\_\{\\mathrm\{loss\}\}\(t\)\>0means thatq1q\_\{1\}has lower loss\. The level hypothesis is
signdloss\(t\)=signdglevel\(t\)\.\\operatorname\{sign\}d\_\{\\mathrm\{loss\}\}\(t\)=\\operatorname\{sign\}d\_\{g\}^\{\\mathrm\{level\}\}\(t\)\.
The rate contrasts are
dρ\(t\)=ρq1,t\(g\)−ρq2,t\(g\),dα\(t\)=αq1,t−αq2,t\.d\_\{\\rho\}\(t\)=\\rho^\{\(g\)\}\_\{q\_\{1\},t\}\-\\rho^\{\(g\)\}\_\{q\_\{2\},t\},\\qquad d\_\{\\alpha\}\(t\)=\\alpha\_\{q\_\{1\},t\}\-\\alpha\_\{q\_\{2\},t\}\.Under the stated tail and off\-prefix conditions, the directional prediction is
signdα\(t\)=signdρ\(t\),\\operatorname\{sign\}d\_\{\\alpha\}\(t\)=\\operatorname\{sign\}d\_\{\\rho\}\(t\),and the contrasting tasks satisfy
dα\(t1\)dα\(t2\)<0\.d\_\{\\alpha\}\(t\_\{1\}\)d\_\{\\alpha\}\(t\_\{2\}\)<0\.The static and rate margins answer different questions\. The first tests whether the system with more useful geometry atN¯\\bar\{N\}has lower loss at that budget; the second measures how quickly useful structure is added\. A system can lead atN¯\\bar\{N\}and still have the shallower exponent\. Reporting both margins distinguishes task\-conditioned accessibility from a general capacity or optimization advantage\.
For the fixed\-kernel test, definedβ\(t\)=βq1,t−βq2,td\_\{\\beta\}\(t\)=\\beta\_\{q\_\{1\},t\}\-\\beta\_\{q\_\{2\},t\}and compare its sign withdα\(t\)d\_\{\\alpha\}\(t\)\. Cross\-task geometry interactions require a common standardization\. The exponent interaction
Δα=dα\(t1\)−dα\(t2\)\\Delta\_\{\\alpha\}=d\_\{\\alpha\}\(t\_\{1\}\)\-d\_\{\\alpha\}\(t\_\{2\}\)summarizes the reversal, with uncertainty propagated from the four joint curve fits\. Raw geometry differences across tasks enter only under a shared standardization; within\-task signs remain the primary geometry tests\.
The curve analysis fitsL∞L^\{\\infty\},CC, andα\\alphajointly and propagates their uncertainty to every contrast\. Candidate curve families are evaluated on held\-out sizes or out\-of\-range prediction\. Geometry\-side rates may be estimated on an early or lower\-scale subset and tested against exponent ordering at larger held\-out scales, yielding an out\-of\-sample prediction\. Seeds replicate runs within a task instance, and independently generated or sampled instances support task\-family claims\. A pilot\-informed power analysis sets both counts; a single instance from each family confines the conclusion to those tasks\. Equivalence assessment follows verified geometry\-rate separation and, for product\-law cells, the recorded prefix\-adequacy route\.
The preregistration records the expected within\-task level and rate signs, the cross\-task reversal, the applicability route for each product\-law cell, and the equivalence margins for the null\. Together, these entries define the primary joint test\. Concordant level and rate tracking supports the accessibility hypothesis\. Scaling differences with a distinct geometry pattern identify another model\-side mechanism\. Verified rate separation with exponent equivalence challenges the directional prediction\. Complete causal mediation requires a geometry intervention or an explicit causal model\.
## 5\. Implications and Research Priorities
### 5\.1 Coupling\-Specific Scaling Walls
Within this framework, a scaling wall is indexed by a coupling, task, resource axis, and budget range\. Over a bounded range, transient optimization or data bottlenecks, a large unresolved tail within ultimate support, and a positive strict floor can produce similar curves\. Comparative interventions across a sufficiently long scale range separate these possibilities\. Recovery of the prior slope under longer training of the same coupling indicates a transient wall\. On matched data, a structurally different coupling that yields a steeper slope or lowerLeff\(N¯\)L^\{\\mathrm\{eff\}\}\(\\bar\{N\}\)supports a finite\-budget accessibility constraint, while convergence of well\-chosen couplings toward an equivalentL∞L^\{\\infty\}increases the plausibility of irreducible task noise or specification limits\. Across successive model generations, these patterns motivate a staircase conjecture: structural interventions shift the accessibility frontier, after which scaling exploits the newly reachable region\.
### 5\.2 Task\-Conditional Evaluation and Model Selection
Coupled Scaling makes evaluation explicitly task\-conditional\. When two systems expose non\-nested task\-relevant geometries, aggregate rankings depend on the benchmark’s task mixture\. Changing that mixture can reverse the ranking, with each result describing its own evaluation distribution\.
Prompting and scoring choices also alter the effective task distribution\. Attributing the resulting reversal to accessibility therefore requires an independent measurement of task\-relevant geometry\.
Compositional benchmarks add a further identification problem: endpoint scores mix prerequisite structure with task level, format, exposure, and operational degeneracy\. Component trajectories, emergence gaps, and format\-aligned comparisons make those contributions visible\. For deployment, model selection should follow the target task distribution; global rankings are summaries of the distribution on which they were constructed\.
### 5\.3 Measurement Priorities and Next Tests
The factorial intervention in §4\.3 is the immediate empirical priority: multiscale coupling\-by\-task comparisons, independent representation measurement, repeated runs, and independently sampled task instances\. The fixed\-kernel specialization supplies a complete spectral object\. Feature\-learning systems need prespecified geometry proxies for support, cumulative\-tail behavior, completed\-prefix coverage, and off\-prefix acceleration\. Rank\-resolved designs estimate these quantities directly; controlled designs establish them through construction or independent theory\. Strict\-floor estimation requires enough scale to identify an asymptote after finite\-range curvature and unfinished acquisition are accounted for\. Compute\-matched recurrent MoE architectures are an especially informative direct\-test setting: cross\-visit expert routing, residual updates, and attention redistribution can be tracked across scale as candidate accessibility trajectories, while contrasting task families test whether their rate ordering reverses \(Wang et al\., 2026\)\.
These measurements also delimit the formal extension\. A network\-level theory must derive support and acquisition from architecture and stochastic optimization while reproducing the witness’s exact floor–tail decomposition, coverage bracket, and rate alternatives\. Its resource map must replace the one\-slot\-per\-mode convention and explain when off\-prefix acquisition contributes an additional exponent\.
### 5\.4 Conclusion
Coupled Scaling treats neural scaling as a relation between a task and the learning system used to reach it\. In the solvable witness, architectural support determines the strict target residual, and the cumulative target tail combines with acquisition order to determine the finite\-budget residual rate\. Proposition 2 gives
ρA,O,TγA,T≤αA,O,T≤γA,T,\\rho\_\{A,O,T\}\\gamma\_\{A,T\}\\leq\\alpha\_\{A,O,T\}\\leq\\gamma\_\{A,T\},with the lower endpoint attained under bounded or exponent\-neutral off\-prefix gain:
αA,O,T=ρA,O,TγA,T\.\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\\gamma\_\{A,T\}\.For a power\-law supported target sequence, this becomesαA,O,T=ρA,O,T\(bA,T−1\)\\alpha\_\{A,O,T\}=\\rho\_\{A,O,T\}\(b\_\{A,T\}\-1\)\. A reversal in completed\-prefix coverage ordering yields a cross\-task exponent reversal under the common\-tail and off\-prefix conditions\. The fixed\-kernel specialization expresses the same relation through the near\-zero tail of a task\-weighted spectral measure and retains the joint spectral measure when eigenvalue and target\-power orders are misaligned\. Relative to frontier accounts of progressive tail coverage, Coupled Scaling adds coupling\-indexed support, interleaved acquisition, and a coverage bracket that locates the exponent between a prefix\-limited endpoint and the best\-NNtail rate\. This structure separates strict floors from finite\-budget rates and identifies when dispersed acquisition contributes an additional exponent\.
The empirical test measures geometry before the scaling curves are fit and crosses couplings with tasks that favor different inductive biases\. It pairs static access with matched\-budget loss and completed\-prefix rates with exponent ordering and the cross\-task reversal specified in advance\. The central question is whether independent geometry measurements of support and acquisition explain why systems trained on the same data can exhibit different scaling laws\.
## Appendix A: Reanalysis Protocol
This appendix documents the methodology behind the reanalysis reported in §4\.2\. The fixed\-seed scripts and machine\-readable results described in A\.11 reproduce every reanalysis value reported in §4\.2 and this appendix from the public data\.
### A\.1 Data Source and Models
We use data from Liu et al\. \(2026\), publicly available at[https://github\.com/KaiserWhoLearns/ElementalTask](https://github.com/KaiserWhoLearns/ElementalTask)\. At the pinned commit used for the analysis \(fc08c318\), the repository contained no license file\. A project author subsequently confirmed that the project code and data are released under the MIT License\. The analysis reads the source data in place from the authors’ repository \(A\.11\)\. The released results cover four models from three open\-weight families: Pythia\-6\.9B \(EleutherAI\); Amber\-7B \(LLM360\); and OLMo2\-1B and OLMo2\-7B \(early\-training checkpoint releases\)\. The trajectories come from separate training runs spanning three model families and several scales\. Across the suite, model family, scale, training corpus or recipe, data order, checkpoint structure, and training randomness vary jointly; notably, OLMo2\-1B and OLMo2\-7B use the same named OLMo\-2 data mixtures\. Sixty\-two tasks have accuracy trajectories for all four models\. Liu et al\.’s paper describes a broader model set; the analysis here uses the trajectories in the public release\.
### A\.2 Task Typing: Two Universes
The prerequisite relation≺\\prec\(ti≺tjt\_\{i\}\\prec t\_\{j\}ifftit\_\{i\}is a compositional component oftjt\_\{j\}\) can be derived mechanically from two sources in the release, and we report both\.
Primary \(authors\-map\) universe\.Composites are all measuredcompositional\_\*tasks whose operation chain parses under the component mapping in the dataset authors’ own analysis code \(predict\_compositional\_from\_components\.py\), which maps each operation to its measured elemental task, includingreverse→\\totoken\_reversal, and we require every component to be measured for all four models\. This yields 29 tasks \(9 elemental, 20 composite\), 46 prerequisite edges, and 406 pairs\. We adopt it as primary because it uses the authors’ own operation\-to\-task identification and covers every parseable measured composite\.
Registry\-native universe\.Composites and elementals are exactly the entries of the dataset’s metadata registries \(dataset/compositional\*\.csvoperations column;dataset/simple\.csv\) that are measured for all four models: 27 tasks \(10 elemental, 17 composite\) and 351 pairs\. Thereverseoperation has no registry elemental, so reverse\-containing composites link only through their other components \(27 edges\), and the two knowledge\-retrieval elementals \(country\-to\-capital, country\-to\-currency\) appear as isolated nodes because neither is used as a registered composite operation\. This universe hews exactly to the dataset’s own metadata and is reported as a robustness check\.
The two universes have partially overlapping composite sets\. Thelastoperation has no measured elemental in the release \(last\_letterappears in the registry files but has no released trajectories\), solower\_lastandupper\_lastenter the registry universe with a single prerequisite edge each, through their case\-map component, and are absent from the primary universe, whose admission rule requires every component to be measured under the authors’ mapping\. Conversely, three\-operation chains that parse under that mapping, such aslower\_reverse\_first, may be absent from the registry files\. Degeneracy \(A\.3\) is a property of the task alone; the audit’s member lists differ across universes only because membership does\.
### A\.3 Pair Classification and Operational\-Degeneracy Audit
All pairs are classified into five categories based solely on their position in the DAG, independent of performance data, with the definitions applied in list order:
- •A\. Prerequisite\-linked:ti≺tjt\_\{i\}\\prec t\_\{j\}ortj≺tit\_\{j\}\\prec t\_\{i\}\.
- •B\. Same\-level elemental: both tasks elemental, no dependency\.
- •C\. Cross\-level incomparable: different compositional levels, no prerequisite link\.
- •D\. Shared\-prerequisite composites: same\-level composites sharing at least one component\.
- •E\. Disjoint composites: same\-level composites with no shared components\.
- •C1C\_\{1\}\. Elemental\-anchored cross\-level, non\-prerequisite\(derived subtype of C\): elemental–composite pairs with no prerequisite link\. Primary: 134; registry: 143\. For every compositecc, its elemental pairs partition exactly intoA\(c\)A\(c\)andC1\(c\)C\_\{1\}\(c\)\.
Counts for the primary universe: A 46, B 36, C 218, D 59, E 47 \(406\)\. Registry: A 27, B 45, C 173, D 48, E 58 \(351\)\. The declared matched contrast is the within\-composite comparison of §4\.2 \(A versusC1C\_\{1\}, within each composite\); the A–B contrast and the A\-versus\-all\-incomparable contrast are reported as benchmarks\.
Six composites in each universe are*operationally degenerate*: on their own item pools, at least one declared operation is extensionally trivial, meaning that deleting it from the operation chain leaves the input–output mapping unchanged on every item\. Four are single\-character case\-mapping composites\. In the primary universe these arelower\_first,lower\_reverse, andlower\_reverse\_first\(all reducing tolowercase\) andupper\_reverse\_first\(touppercase\); in the registry universe they arelower\_first,lower\_last,lower\_reverse\(tolowercase\) andupper\_last\(touppercase\)\. Their items are single characters, on which reversal and first\- and last\-letter extraction act as identities\.
Two further composites, common to both universes, reduce through a trivial trailing case\-map:plural\_lowercoincides withsingular\_to\_pluralandtranslate\_eng\_fr\_lowerwithtranslate\_eng\_fr, because irregular plurals and French glosses are already lowercase\. The correspondingtranslate\_fr\_eng\_lowerandtranslate\_sp\_eng\_lowerare*not*degenerate: their outputs contain uppercase forms on which the trailinglowercaseis non\-trivial\.
The degenerate composites account for 14 of the 46 prerequisite edges in the primary universe and 9 of the 27 in the registry universe; their declared prerequisite edges encode operational identities and lack capability\-prerequisite content\. The A\.4 validity rule retains these edges\. When the released trajectories coincide exactly, the pair is tie\-excluded; when demonstration sampling drives them apart, the pair contributes a defined, noise\-driven ordering sign\. Because composite prompts are resampled at every evaluation \(A\.4\), the rule can remove benign coincidences while retaining contaminating divergences\. Exclusion is consequently applied at the task level using the degeneracy rule above, which is computable from task definitions alone and independent of performance data; the implementation appears in the public artifact described in A\.11\. Appendix A\.6c reports the full battery with these tasks excluded; all main\-text analyses retain them unless explicitly labelled otherwise\.
### A\.4 Emergence Measures and Checkpoint Hygiene
For each modelmmand tasktt, the threshold emergence timeem\(t\)e\_\{m\}\(t\)is the first training checkpoint at which accuracy reaches or exceedsθ\\theta, evaluated atθ∈\{0\.3,0\.4,0\.5,0\.6,0\.7\}\\theta\\in\\\{0\.3,0\.4,0\.5,0\.6,0\.7\\\}\. Tasks that never cross the threshold produce undefined emergence times and are excluded pairwise; ties \(em\(ti\)=em\(tj\)e\_\{m\}\(t\_\{i\}\)=e\_\{m\}\(t\_\{j\}\)\) are excluded from the stability computation\. Two threshold\-free measures are computed on the same trajectories:*maximum\-improvement timing*\(the checkpoint at which the largest checkpoint\-to\-checkpoint improvement occurs; labelledmax\-slopein tables\) and*trajectory area under the curve \(AUC\)*\(larger normalized area read as earlier emergence\)\.
Checkpoint identifiers are parsed numerically before ordering, and release\-style rows \(main\), which duplicate the final model, are excluded from the training\-checkpoint sequence\. Lexicographic ordering of step identifiers, or ingestion of release rows as early checkpoints, would silently corrupt emergence times on this dataset\. The max\-slope measure uses the maximum checkpoint\-to\-checkpoint improvement over the released grid with checkpoint spacing left unweighted \(the OLMo2 grids are non\-uniform in training step; A\.12\)\.
The released evaluation pipeline constructs prompts differently at the two levels\. Elemental prompts are fixed: the same demonstration set appears at every checkpoint\. Compositional prompts are redrawn at each evaluation from an unseeded generator \(compositional\_task\.py,random\.shufflewith no seed set\), whereas elemental prompts remain fixed\. The resulting demonstration\-sampling component is aligned with compositional level\. The within\-composite matched contrast of §4\.2 holds the composite fixed across both arms, making this sampling component common toA\(c\)A\(c\)andC1\(c\)C\_\{1\}\(c\)\. Degenerate composites remain exposed because at least one component pairing is an identity whose ordering is sampling noise in its entirety\. The fixed\-prompt protocol specified for the controlled design of §4\.3 removes this demonstration\-sampling component\.
### A\.5 Statistical Procedures
The declared primary comparison of §4\.2 is the within\-composite matched contrast\. The A\-versus\-B comparison is a structurally motivated benchmark, and the A\-versus\-all\-incomparable comparison is a secondary aggregate benchmark\. For these observational benchmarks, a one\-sided permutation\-reference procedure \(10,000 iterations\) randomizes pair\-type labels within the compared categories\.
Pairs sharing a task induce dependence, so we additionally report task\-resampling intervals based on 2,000 draws of the task set with replacement; a pair is retained when both of its tasks are drawn\. Each replicate is computed on the induced subgraph of a random task subset and discards the multiplicity of repeated draws\. This induced\-subgraph resampling scheme has undetermined interval conservativeness\. The pair\-level permutation\-reference tail areas serve as descriptive reference quantities under shared\-task dependence, and A\.12 adds a multiplicity\-weighted sensitivity analysis\.
The matched contrast of §4\.2 adds three procedures\. First, within\-composite label permutation reassigns whichk\(c\)k\(c\)of each qualifying composite’sk\(c\)\+m\(c\)k\(c\)\+m\(c\)pooled valid pairs carry the component label, holdingSSvalues fixed \(10,000 draws from an independent generator, seed 20260721\)\. The statistic is the composite\-equal\-weight meanΔ¯\\bar\{\\Delta\}\. The quantityrpermr\_\{\\mathrm\{perm\}\}is reported as an uncalibrated permutation\-reference tail area in theΔ¯\>0\\bar\{\\Delta\}\>0direction specified before the matched\-contrast implementation and, for the robustness analyses summarized in §4\.2, in the observed direction\. The corresponding two\-sided reference tail area is twice the observed\-direction one\-sided value, capped at 1\.
TheΔ¯\>0\\bar\{\\Delta\}\>0direction and the emergence measures were documented before the matched\-contrast implementation and its fixed seeds were added; the analysis was not externally registered\. The permutation\-reference distribution has a calibrated inferential interpretation only under conditional exchangeability\. Exact randomization calibration would require randomized component labels; in the observed suite, elemental identity, difficulty, format, and cross\-composite prevalence may covary with the label and therefore propagate to the reference distribution \(§4\.2\)\.
Second, composite\-resampling intervals use 2,000 multiplicity\-preserving draws of the qualifying composites’Δ\(c\)\\Delta\(c\)values with replacement \(seed 20260722\) to give percentile intervals forΔ¯\\bar\{\\Delta\}\. The interval model treats the per\-composite contrasts as exchangeable\. Third, the leave\-one\-elemental\-out battery probes dependence induced by elementals shared across composites by rerunning the full matched analysis with each elemental removed in turn and reporting sign changes\. The full battery is reported without selection or multiplicity correction and is interpreted through the magnitude and measure\-dependence of the residual \(§4\.2, A\.12\)\.
### A\.6a Results
Primary universe\(meanS¯\\bar\{S\}with number of valid pairs in parentheses;rpermr\_\{\\mathrm\{perm\}\}one\-sided\):
MeasureABCDEA–B gap \(rpermr\_\{\\mathrm\{perm\}\}\)A–all gap \(rpermr\_\{\\mathrm\{perm\}\}\)95% resampling interval A–Bθ=0\.3\\theta=0\.3\.941 \(17\)\.691 \(27\)\.858 \(54\)\.857 \(7\)\.556 \(3\)\.250 \(\.0079\)\.143 \(\.0446\)\(−\-\.167, \.579\)θ=0\.4\\theta=0\.4\.818 \(11\)\.583 \(24\)\.863 \(51\)1\.000 \(8\)1\.000 \(3\)\.235 \(\.0715\)\.016 \(\.4497\)\(−\-\.833, \.667\)θ=0\.5\\theta=0\.5\.833 \(12\)\.754 \(19\)\.865 \(52\)\.889 \(9\)\.333 \(3\)\.079 \(\.3402\)\.010 \(\.4290\)\(−\-\.883, \.500\)θ=0\.6\\theta=0\.6\.909 \(11\)\.867 \(20\)\.939 \(55\)1\.000 \(8\)1\.000 \(2\)\.042 \(\.3856\)−\-\.020 \(\.7275\)\(−\-\.500, \.333\)θ=0\.7\\theta=0\.7\.917 \(12\)\.841 \(21\)\.922 \(60\)\.917 \(12\)1\.000 \(3\)\.075 \(\.3613\)\.010 \(\.6100\)\(−\-\.500, \.389\)max\-slope\.871 \(31\)\.593 \(27\)\.910 \(108\)\.852 \(18\)1\.000 \(19\)\.278 \(\.0070\)\.007 \(\.4992\)\(−\-\.085, \.818\)AUC\.906 \(46\)\.778 \(36\)\.890 \(218\)\.825 \(59\)\.894 \(47\)\.128 \(\.0176\)\.037 \(\.1839\)\(−\-\.040, \.322\)Registry\-native universe\(robustness\):
MeasureABCDEA–B gap \(rpermr\_\{\\mathrm\{perm\}\}\)θ=0\.3\\theta=0\.31\.000 \(6\)\.793 \(29\)\.798 \(28\)\.857 \(7\)—\.207 \(\.2316\)θ=0\.4\\theta=0\.4\.800 \(5\)\.829 \(35\)\.871 \(31\)1\.000 \(8\)—−\-\.029 \(\.7469\)θ=0\.5\\theta=0\.5\.833 \(6\)\.790 \(35\)\.882 \(34\)1\.000 \(8\)—\.043 \(\.4578\)θ=0\.6\\theta=0\.6\.833 \(6\)\.892 \(37\)\.902 \(34\)1\.000 \(7\)—−\-\.059 \(\.7425\)θ=0\.7\\theta=0\.7\.857 \(7\)\.839 \(29\)\.868 \(38\)\.923 \(13\)1\.000 \(2\)\.018 \(\.5579\)max\-slope\.784 \(17\)\.593 \(27\)\.781 \(93\)\.611 \(12\)\.818 \(22\)\.192 \(\.1016\)AUC\.901 \(27\)\.759 \(45\)\.833 \(173\)\.792 \(48\)\.839 \(58\)\.142 \(\.0163\)In the primary universe the A–B contrast is positive at every threshold and under both threshold\-free measures, and is associated with small pair\-label permutation\-reference tail areas atθ=0\.3\\theta=0\.3and under both threshold\-free measures, although shared\-task dependence prevents calibrated inferential interpretation; the task\-resampling intervals include zero throughout\. In the registry universe, only five to seven prerequisite pairs survive at threshold level, and the threshold\-wise point estimates are correspondingly unstable \(including sign flips atθ=0\.4\\theta=0\.4and0\.60\.6\); the lowest\-threshold pattern \(1\.000 against 0\.793\) and the AUC contrast \(rperm=\.0163r\_\{\\mathrm\{perm\}\}=\.0163\) match the primary universe’s direction\. Cells withn≤3n\\leq 3\(several D and E entries\) are reported for completeness and excluded from interpretation\. Throughout these tables,nncounts task pairs contributing at least one valid model comparison\.
### A\.6b Matched\-Contrast Results
Table A\.1\.Within\-composite matched contrast, full suite, all seven emergence measures and both universes\.Δ¯\\bar\{\\Delta\}is the composite\-equal\-weight mean ofΔ\(c\)\\Delta\(c\)over qualifying composites;rpermr\_\{\\mathrm\{perm\}\}is the within\-composite permutation\-reference tail area, one\-sided in theΔ¯\>0\\bar\{\\Delta\}\>0direction specified before the matched\-contrast implementation; intervals are composite\-resampling percentiles \(A\.5\)\. Pooled columns give the pooled A\-versus\-C1C\_\{1\}means over the qualifying composites’ arm pairs and the pooled gap with its permutation\-reference tail area\. Registry threshold rows rest on four qualifying composites and are retained as background\.
MeasureUniversencn\_\{c\}Δ¯\\bar\{\\Delta\}rpermr\_\{\\mathrm\{perm\}\}95% resampling intervalS¯A\\bar\{S\}\_\{A\}\(nn\)S¯C1\\bar\{S\}\_\{C\_\{1\}\}\(nn\)pooled gap \(rpermr\_\{\\mathrm\{perm\}\}\)θ=0\.3\\theta=0\.3primary9\+0\.086\.2026\(−\-0\.111, \+0\.284\)0\.938 \(16\)0\.793 \(37\)\+0\.145 \(\.1739\)θ=0\.3\\theta=0\.3registry4\+0\.202\.2095\(\+0\.042, \+0\.411\)1\.000 \(6\)0\.754 \(23\)\+0\.246 \(\.1059\)θ=0\.4\\theta=0\.4primary9\+0\.034\.4193\(−\-0\.167, \+0\.237\)0\.818 \(11\)0\.800 \(35\)\+0\.018 \(\.6678\)θ=0\.4\\theta=0\.4registry4\+0\.035\.4402\(−\-0\.243, \+0\.278\)0\.800 \(5\)0\.840 \(25\)−\-0\.040 \(\.6424\)θ=0\.5\\theta=0\.5primary9\+0\.052\.2911\(−\-0\.086, \+0\.182\)0\.833 \(12\)0\.806 \(36\)\+0\.028 \(\.3434\)θ=0\.5\\theta=0\.5registry4\+0\.007\.5435\(−\-0\.260, \+0\.223\)0\.833 \(6\)0\.857 \(28\)−\-0\.024 \(\.6755\)θ=0\.6\\theta=0\.6primary9\+0\.018\.4531\(−\-0\.093, \+0\.110\)0\.909 \(11\)0\.905 \(35\)\+0\.004 \(\.6558\)θ=0\.6\\theta=0\.6registry4\+0\.001\.6621\(−\-0\.257, \+0\.171\)0\.833 \(6\)0\.881 \(28\)−\-0\.048 \(\.7787\)θ=0\.7\\theta=0\.7primary9\+0\.015\.4228\(−\-0\.096, \+0\.101\)0\.917 \(12\)0\.904 \(38\)\+0\.013 \(\.5636\)θ=0\.7\\theta=0\.7registry4\+0\.058\.3905\(−\-0\.221, \+0\.242\)0\.857 \(7\)0\.815 \(27\)\+0\.042 \(\.4394\)max\-slopeprimary17−\-0\.068\.9251\(−\-0\.175, \+0\.033\)0\.867 \(30\)0\.931 \(68\)−\-0\.065 \(\.9721\)max\-sloperegistry12−\-0\.1031\.000\(−\-0\.228,−\-0\.011\)0\.784 \(17\)0\.859 \(59\)−\-0\.074 \(1\.000\)AUCprimary20\+0\.029\.1762\(−\-0\.035, \+0\.093\)0\.906 \(46\)0\.863 \(134\)\+0\.043 \(\.1895\)AUCregistry17\+0\.084\.0198\(−\-0\.031, \+0\.209\)0\.901 \(27\)0\.814 \(143\)\+0\.088 \(\.0619\)The table is generated from the archived results file \(matched\_contrast\_results\.json\), which additionally contains per\-compositeΔ\(c\)\\Delta\(c\)andk\(c\)/m\(c\)k\(c\)/m\(c\), pooled comparisons, audit\-excluded scenarios, and the leave\-one\-elemental\-out battery\.
### A\.6c Degeneracy\-Excluded Battery
The following tables report the full A\.6a battery recomputed on the degeneracy\-excluded universes of A\.3 \(primary minus six: 23 tasks, 32 prerequisite edges, 253 pairs; registry minus six: 21 tasks, 18 edges, 210 pairs; independent generator, seed 20260723\)\.
Primary universe, degeneracy\-excluded:
MeasureABCDEA–B gap \(rpermr\_\{\\mathrm\{perm\}\}\)95% resampling interval A–Bθ=0\.3\\theta=0\.30\.889 \(9\)0\.691 \(27\)0\.875 \(24\)—0\.667 \(2\)\+0\.198 \(\.0750\)\(−\-0\.417, \+0\.513\)θ=0\.4\\theta=0\.40\.833 \(6\)0\.583 \(24\)0\.768 \(23\)—1\.000 \(2\)\+0\.250 \(\.0914\)\(−\-0\.593, \+0\.667\)θ=0\.5\\theta=0\.50\.857 \(7\)0\.754 \(19\)0\.747 \(25\)—0\.333 \(2\)\+0\.103 \(\.2694\)\(−\-0\.833, \+0\.467\)θ=0\.6\\theta=0\.61\.000 \(6\)0\.867 \(20\)0\.942 \(23\)—1\.000 \(2\)\+0\.133 \(\.2412\)\(\+0\.000, \+0\.333\)θ=0\.7\\theta=0\.71\.000 \(6\)0\.841 \(21\)0\.944 \(24\)1\.000 \(1\)1\.000 \(2\)\+0\.159 \(\.2996\)\(\+0\.000, \+0\.400\)max\-slope0\.861 \(24\)0\.593 \(27\)0\.910 \(74\)0\.926 \(9\)1\.000 \(13\)\+0\.269 \(\.0158\)\(−\-0\.085, \+0\.714\)AUC0\.948 \(32\)0\.778 \(36\)0\.918 \(134\)0\.910 \(26\)0\.953 \(25\)\+0\.170 \(\.0023\)\(\+0\.050, \+0\.333\)Registry universe, degeneracy\-excluded:
MeasureABCDEA–B gap \(rpermr\_\{\\mathrm\{perm\}\}\)95% resampling interval A–Bθ=0\.3\\theta=0\.31\.000 \(3\)0\.793 \(29\)0\.848 \(11\)——\+0\.207 \(\.4517\)\(\+0\.000, \+0\.429\)θ=0\.4\\theta=0\.41\.000 \(2\)0\.829 \(35\)0\.722 \(12\)——\+0\.171 \(\.6574\)\(\+0\.000, \+0\.361\)θ=0\.5\\theta=0\.51\.000 \(3\)0\.790 \(35\)0\.778 \(15\)——\+0\.210 \(\.2818\)\(\+0\.000, \+0\.424\)θ=0\.6\\theta=0\.61\.000 \(3\)0\.892 \(37\)0\.911 \(15\)——\+0\.108 \(\.5684\)\(\+0\.000, \+0\.278\)θ=0\.7\\theta=0\.71\.000 \(3\)0\.839 \(29\)0\.846 \(13\)1\.000 \(1\)—\+0\.161 \(\.4356\)\(\+0\.000, \+0\.333\)max\-slope0\.810 \(14\)0\.593 \(27\)0\.891 \(61\)0\.867 \(5\)1\.000 \(10\)\+0\.217 \(\.0840\)\(−\-0\.150, \+0\.561\)AUC1\.000 \(18\)0\.759 \(45\)0\.955 \(110\)0\.870 \(18\)1\.000 \(19\)\+0\.241 \(\.0004\)\(\+0\.100, \+0\.411\)Threshold rows thin under exclusion because the degenerate composites carried a large share of threshold\-defined prerequisite pairs \(nA≤9n\_\{A\}\\leq 9primary,≤3\\leq 3registry\); the excluded battery’s weight rests on the threshold\-free measures\. Under the AUC measure the A\-versus\-B contrast strengthens in both universes relative to the full suite, and its task\-resampling interval excludes zero: primaryS¯A=0\.948\\bar\{S\}\_\{A\}=0\.948\(32\) againstS¯B=0\.778\\bar\{S\}\_\{B\}=0\.778\(36\), gap0\.1700\.170,rperm=\.0023r\_\{\\mathrm\{perm\}\}=\.0023, interval\(0\.050,0\.333\)\(0\.050,0\.333\); registryS¯A=1\.000\\bar\{S\}\_\{A\}=1\.000\(18\) against0\.7590\.759\(45\), gap0\.2410\.241,rperm=\.0004r\_\{\\mathrm\{perm\}\}=\.0004, interval\(0\.100,0\.411\)\(0\.100,0\.411\)\.
The audit\-excluded matched contrast summarized in §4\.2 is tabulated here for reference;rpermr\_\{\\mathrm\{perm\}\}is the one\-sided permutation\-reference tail area in the observed direction \(A\.5\), and the intervals are composite\-resampling percentiles\.
MeasureUniversencn\_\{c\}\(full→\\rightarrowexcl\.\)Δ¯\\bar\{\\Delta\}\(full→\\rightarrowexcl\.\)rpermr\_\{\\mathrm\{perm\}\}\(obs\.\)95% resampling interval \(excl\.\)max\-slopeprimary17→\\rightarrow12−\-0\.068→\\rightarrow−\-0\.124\.021\(−\-0\.261, 0\.000\)max\-sloperegistry12→\\rightarrow9−\-0\.103→\\rightarrow−\-0\.130\.063\(−\-0\.278,−\-0\.018\)AUCprimary20→\\rightarrow14\+0\.029→\\rightarrow\+0\.040\.098\(−\-0\.002, \+0\.088\)AUCregistry17→\\rightarrow11\+0\.084→\\rightarrow\+0\.044\.158\(\+0\.000, \+0\.106\)
### A\.7 Prerequisite Direction and Category Heterogeneity
Two readings of the sanity check “components emerge before their composites” give different answers\. Under the natural interpretation for coarse checkpoint grids, which counts ties as satisfying the constraint, components emerge no later than their composites in 63–69% of defined model\-edge comparisons in the primary universe and 74–87% in the registry universe\. Under the threshold measures, restricting the analysis to strictly ordered comparisons lowers the primary\-universe rate to 25–28%\. Strict inversions concentrate on extraction\-style elementals whose answer format is stricter than the composite’s: for example,compositional\_first\_upperreaches accuracy 0\.81 on an OLMo2 checkpoint wheresimple\_icl\_first\_letterstands at 0\.07\. This concentration points to a task\-design artifact\. Because the prerequisite constraint on timing is soft, the near\-ceiling*ordering agreement*of category A is an empirical regularity whose strength depends on task design\.
The directional composition is also measure\-dependent\. Among strictly ordered comparisons in the primary universe, components precede their composites in 25\.3% of cases under max\-slope \(87 comparisons\) but 62\.4% under AUC \(178 comparisons\); registry figures are 45\.5% \(44\) and 77\.7% \(103\)\. The two measures yield nearly opposite directional regimes for prerequisite pairs, bearing directly on the sign reversal summarized in §4\.2\.
Selection partly determines how much weight this artifact carries\. Strict, defined comparisons are far scarcer under the timing\-sensitive measure \(87 against 178 prerequisite comparisons in the primary universe\), so pairs affected by the artifact constitute a larger share of the surviving comparisons\. At category level, this heterogeneity motivates the A–B benchmark, while the matched contrast of §4\.2 remains the main test\.
### A\.8 Universe Sensitivity
The two universes differ in whethertoken\_reversalis admitted as the elemental realization of thereverseoperation \(authors’ mapping: yes; registry: no such elemental\) and whether knowledge elementals without outgoing edges are included \(registry: yes\)\. The direction of the A–B contrast at the lowest threshold and under the AUC measure is shared across both; magnitudes and threshold\-level stability differ, dominated by the number of valid prerequisite comparisons each universe admits\.
### A\.9 Threshold\-Free Robustness
Because threshold\-crossing measures are vulnerable to metric artifacts \(Schaeffer et al\., 2023\), the full battery is replicated under two threshold\-free emergence measures \(A\.4\)\. The A–B contrast remains positive under both measures in both universes\. In the primary universe, the smallest pair\-label permutation\-reference tail areas occur under the two threshold\-free measures \(max\-slope gap0\.2780\.278,rperm=\.0070r\_\{\\mathrm\{perm\}\}=\.0070; AUC gap0\.1280\.128,rperm=\.0176r\_\{\\mathrm\{perm\}\}=\.0176\); the registry AUC gap is0\.1420\.142\(rperm=\.0163r\_\{\\mathrm\{perm\}\}=\.0163\)\. These permutation\-reference tail areas serve as descriptive reference quantities under the observational design\. The subsequent sensitivity audit finds max\-slope sensitivity to denominator weighting and completeness restriction \(A\.12\)\. AUC therefore carries the threshold\-free benchmark\.
### A\.10 Within\-Family versus Cross\-Family Agreement
The released models contain exactly one same\-family pair \(OLMo2\-1B and OLMo2\-7B\), permitting an underpowered decomposition of incomparable\-pair agreement according to whether the models share a coupling family\. Under the AUC measure, the same\-family pair agrees more often than cross\-family pairs do \(primary universe: 0\.931 over 347 pairs against 0\.855 over 357; registry: 0\.927 against 0\.794\), and the same direction holds atθ≤0\.5\\theta\\leq 0\.5\(e\.g\.,θ=0\.3\\theta=0\.3: 0\.910 against 0\.763\)\. The direction reverses atθ≥0\.6\\theta\\geq 0\.6, where surviving comparisons are few\. One within\-family pair, which also differs in scale, supplies a descriptive analogue of the controlled contrast in §4\.3 and carries no inferential weight\.
### A\.11 Code and Reproducibility
The analysis is implemented in four scripts\. The public artifact repository is[https://github\.com/quintonvina/coupled\-scaling](https://github.com/quintonvina/coupled-scaling)\(tagv1\.0\-reanalysis\)\. Itsreanalysis/directory contains the four scripts together with the machine\-readable release\-of\-record results, fresh\-environment verification logs, an environment file \(requirements\.txt\), a hash manifest \(MANIFEST\.sha256\), and a single\-command entry point \(run\_all\.sh\)\. All paths are configurable through relative defaults and environment\-variable overrides\.
reanalysis\_canonical\.py\(fixed seed 20260707; CPU\-only; runs in minutes on a laptop\) produces the canonical battery summarized in §4\.2 and reported in A\.6a, and writescanonical\_results\.json\.matched\_contrast\_canonical\.pyexecutes the canonical script unmodified and adds the matched contrast of §4\.2 using independent fixed\-seed generators \(permutation 20260721; resampling 20260722; excluded battery 20260723\)\. It writesmatched\_contrast\_results\.json, including per\-compositeΔ\(c\)\\Delta\(c\)andk\(c\)/m\(c\)k\(c\)/m\(c\), pooled A\-versus\-C1C\_\{1\}comparisons, and the leave\-one\-elemental\-out battery, as well as the degeneracy\-excluded battery of A\.6c \(a6c\_excluded\_battery\.json\)\.
degeneracy\_audit\.pyimplements the audit in A\.3\.audit\_sensitivity\.pyimplements the subsequent sensitivity battery in A\.12 using independent seed 20260727\. It writesaudit\_sensitivity\_results\.jsonand the pair\-level tablepair\_level\_audit\.csvand never overwrites the canonical outputs\. Together, the scripts ingest the public repository above and reproduce every reanalysis number in §4\.2 and this appendix for both universes\. License status and in\-place data access are documented in A\.1\.
### A\.12 Subsequent Statistical Sensitivity Audit
After the canonical, matched\-contrast, and degeneracy\-excluded batteries were fixed with their seeds and archived, an independently implemented statistical audit re\-derived the pipeline from the pinned commit and tested sensitivity to weighting and implementation choices\. A second independent implementation reproduced its findings, and the canonical battery was reproduced end to end from a fresh clone, yielding a bitwise\-identicalcanonical\_results\.json\.
The sensitivity audit leaves every reported number and seed in A\.5–A\.6c unchanged\. Its analyses are reported separately wherever they enter the interpretation \(§4\.2, §5\.3\)\. The battery is implemented inaudit\_sensitivity\.py\(independent seed 20260727\) and archived inaudit\_sensitivity\_results\.jsonandpair\_level\_audit\.csv\. For each task pair and measure, the table records the universe, tasks, category, per\-model ordering signs, valid\-model count, comparison denominator,SS, anchoring composite and component status where applicable, and degeneracy\-inclusion status\.
Valid\-comparison structure\.The declared estimand weights task pairs equally regardless of how many of the six model comparisons stand behind eachSS\. Under max\-slope this matters decisively: in the primary universe, 20 of 31 valid prerequisite pairs and 19 of 27 elemental pairs rest on a single valid comparison, and the registry universe has no four\-model prerequisite or elemental pair at all\. Three estimands of the primary\-universe A\-versus\-B gap, namely task\-pair equal weight \(the declared statistic\), comparison\-denominator weight, and four\-model complete case, give\+0\.278\+0\.278,\+0\.017\+0\.017, and−0\.133\-0\.133\. The complete\-case estimate uses five prerequisite and eight elemental pairs\. The spread across estimands makes max\-slope exploratory\. Under AUC, 43 of 46 prerequisite pairs and all 36 elemental pairs use all four models, and the same three estimands give\+0\.128\+0\.128,\+0\.142\+0\.142, and\+0\.145\+0\.145\. AUC therefore carries the battery\-level interpretation\.
Sparse support and model deletion\.Δ\(c\)\\Delta\(c\)is exactly zero for 13 of 17 qualifying composites under max\-slope and 11 of 20 under AUC in the primary universe \(8 of 12 and 9 of 17 in the registry universe\): each matched sign rests on a handful of composites, supporting a sparse\-effect interpretation of the leave\-one\-elemental\-out stability\. Leave\-one\-model\-out deletion reverses the max\-slope contrast \(removing OLMo2\-1B: primary−0\.068→\+0\.022\-0\.068\\to\+0\.022; registry→0\.000\\to 0\.000on four qualifying composites\) and leaves the AUC contrast positive under every deletion \(\+0\.012\+0\.012to\+0\.068\+0\.068primary;\+0\.068\+0\.068to\+0\.105\+0\.105registry\)\.
Resampling procedure\.The interval procedure is the induced\-subgraph resampling scheme of A\.5, with undetermined conservativeness\. A multiplicity\-weighted sensitivity \(pair weightwiwjw\_\{i\}w\_\{j\}from draw counts; 20,000 draws, seed 20260727\) changes the primary\-universe AUC A–B interval from\(−0\.040,0\.322\)\(\-0\.040,0\.322\)to\(−0\.069,0\.354\)\(\-0\.069,0\.354\)and preserves the positive lower bound of the degeneracy\-excluded AUC A\-versus\-B interval in both universes \(primary\(0\.011,0\.357\)\(0\.011,0\.357\); registry\(0\.087,0\.429\)\(0\.087,0\.429\)\)\. This weighting supplies a complementary sensitivity analysis\.
AUC axis\.The canonical AUC integrates accuracy over checkpoint index, treating adjacent checkpoints as equidistant\. The Amber and Pythia grids are \(near\-\)uniform in training step; the OLMo2 grids are non\-uniform, running from step 150 to beyond step10510^\{5\}with denser early coverage\. Integrating over parsed training step \(normalized within model\) changes 3\.0% and 1\.5% of the OLMo2\-1B and OLMo2\-7B within\-model task\-pair orderings, respectively; it changes none of the Amber or Pythia orderings\. The headline statistics become\+0\.100\+0\.100\(primary\-universe A\-versus\-B\),\+0\.061\+0\.061\(primary\-universeΔ¯\\bar\{\\Delta\},nc=20n\_\{c\}=20\),\+0\.161\+0\.161\(registry A\-versus\-B\), and\+0\.121\+0\.121\(registryΔ¯\\bar\{\\Delta\},nc=17n\_\{c\}=17\), with the degeneracy\-excluded matched contrasts at\+0\.034\+0\.034\(nc=14n\_\{c\}=14\) and\+0\.034\+0\.034\(nc=11n\_\{c\}=11\): no direction changes\. The manuscript’s AUC is the checkpoint\-index integral throughout; the step\-weighted variant is archived as sensitivity, and measuring emergence against training progress directly belongs to the controlled design of §4\.3\.
Weight of evidence\.The prespecified battery, the §4\.2 analysis declaration, and their seeds remain archived separately from the subsequent sensitivity analyses, including both degeneracy\-exclusion sets\. The audit assigns max\-slope an exploratory role\. The observational interpretation rests on AUC, whose A\-versus\-B benchmark and matched residual retain their direction under reweighting, completeness restriction, task\-level exclusion, model deletion, and the integration\-axis change\. The matched residual remains small, positive, and descriptive \(§4\.2\)\.
## Appendix B: Proofs and Extensions for the Solvable Instance
Section 3\.3 states the minimal support–order model, its two main propositions, and a coupling\-by\-task corollary\. This appendix supplies the complete construction, proofs, counterexamples, connections to adjacent solvable models, and scope conditions\. Architectural support and data determine the strict floor in this instance; the cumulative supported tail and priority order determine the finite\-budget residual\. The model is a solvable existence witness whose deep\-network counterpart requires deriving these quantities from training dynamics\.
### B\.1 Setup
Let a task be specified by a data distribution𝒟\\mathcal\{D\}and target functionf∗∈L2\(𝒟\)f^\{\*\}\\in L^\{2\}\(\\mathcal\{D\}\)\. Fix a countable orthonormal basis\{φk\}k≥1\\\{\\varphi\_\{k\}\\\}\_\{k\\geq 1\}forL2\(𝒟\)L^\{2\}\(\\mathcal\{D\}\)and writeck=⟨f∗,φk⟩c\_\{k\}=\\langle f^\{\*\},\\varphi\_\{k\}\\rangle, so thatf∗=∑kckφkf^\{\*\}=\\sum\_\{k\}c\_\{k\}\\varphi\_\{k\}and∑kck2=∥f∗∥2<∞\\sum\_\{k\}c\_\{k\}^\{2\}=\\lVert f^\{\*\}\\rVert^\{2\}<\\infty\. For a hypothesisff, define population loss asL\(f\)=∥f−f∗∥2L\(f\)=\\lVert f\-f^\{\*\}\\rVert^\{2\}\. Observation noise would add a constant term and is omitted\.
An architecture–optimization system\(A,O\)\(A,O\)is represented by two objects:
- •asupportSA⊆ℕS\_\{A\}\\subseteq\\mathbb\{N\}, the protocol\-asymptotic union of directions available to the architecture family as the modeled capacity coordinate grows, taken to be countably infinite in Proposition 2; and
- •apriority orderπA,O:ℕ→SA\\pi\_\{A,O\}:\\mathbb\{N\}\\to S\_\{A\}, a bijective enumeration that specifies the order in which additional capacity resolves those directions\.
This division of labor provides a minimal specialization ofℛB\(A,O,𝒫\)\\mathcal\{R\}\_\{B\}\(A,O;\\mathcal\{P\}\): support is assigned to the architecture alone, while optimization, parameterization, and feature learning act through the order\. The acquired set at budgetNNisIA,O\(N\)=\{πA,O\(r\):1≤r≤N\}I\_\{A,O\}\(N\)=\\\{\\pi\_\{A,O\}\(r\):1\\leq r\\leq N\\\}\. A model with budgetNNrealizes the best approximation over these prioritized modes\. The toy model assumes a nested capacity family in which successive prefixes are jointly realizable\.
f^N=∑r=1NcπA,O\(r\)φπA,O\(r\)\.\\hat\{f\}\_\{N\}\\,=\\,\\sum\_\{r=1\}^\{N\}c\_\{\\pi\_\{A,O\}\(r\)\}\\,\\varphi\_\{\\pi\_\{A,O\}\(r\)\}\.
This “omniscient within\-order” allocation removes estimation noise and optimization error to isolate the representational bottleneck\. Its motivating analogies are the coarse\-to\-fine resolution intuition of Sharma and Kaplan \(2022\), the resolution\-limited spectral picture of Bahri et al\. \(2024\), and the quanta\-truncation logic of Michaud et al\. \(2023\)\.
### B\.2 Exact decomposition and rate bounds
By orthonormality,
L\(f^N\)=∑k∉SAck2⏟LA∞\(T\)\+∑r\>NcπA,O\(r\)2⏟EA,O,T\(N\)\.L\(\\hat\{f\}\_\{N\}\)\\,=\\,\\underbrace\{\\sum\_\{k\\notin S\_\{A\}\}c\_\{k\}^\{2\}\}\_\{\\textstyle L^\{\\infty\}\_\{A\}\(T\)\}\\,\+\\,\\underbrace\{\\sum\_\{r\>N\}c^\{2\}\_\{\\pi\_\{A,O\}\(r\)\}\}\_\{\\textstyle E\_\{A,O,T\}\(N\)\}\.
Proof of Proposition 1\.LetVA=span¯\{φk:k∈SA\}V\_\{A\}=\\overline\{\\mathrm\{span\}\}\\\{\\varphi\_\{k\}:k\\in S\_\{A\}\\\}\. By orthonormality, the component off∗f^\{\*\}inVA⟂V\_\{A\}^\{\\perp\}is∑k∉SAckφk\\sum\_\{k\\notin S\_\{A\}\}c\_\{k\}\\varphi\_\{k\}, and hence
LA∞\(T\)=∑k∉SAck2=∥ΠVA⟂f∗∥2\.L^\{\\infty\}\_\{A\}\(T\)=\\sum\_\{k\\notin S\_\{A\}\}c\_\{k\}^\{2\}=\\lVert\\Pi\_\{V\_\{A\}^\{\\perp\}\}f^\{\*\}\\rVert^\{2\}\.
The quantity vanishes if and only if the task\-relevant support off∗f^\{\*\}is contained inSAS\_\{A\}\. Because priority order enumerates only directions insideSAS\_\{A\}, it leaves the orthogonal component unchanged\. BecauseπA,O\\pi\_\{A,O\}is bijective and∑kck2<∞\\sum\_\{k\}c\_\{k\}^\{2\}<\\infty,EA,O,T\(N\)=∑r\>NcπA,O\(r\)2→0E\_\{A,O,T\}\(N\)=\\sum\_\{r\>N\}c^\{2\}\_\{\\pi\_\{A,O\}\(r\)\}\\to 0, so the displayed quantity is the learner’s ordinary asymptotic limit in this construction\.■\\blacksquare
The strict floor in Proposition 1 is order\-invariant because the omniscient allocation allows any order eventually to exhaust the support\. At feasible budgets, however, the unresolved tail depends on priority order\. The following remark separates that residual from the strict asymptotic floor\.
Remark \(finite\-budget residual\)\.For a maximum feasible budgetN¯\\bar\{N\}, define
LA,O,Teff\(N¯\)=LA∞\(T\)\+∑r\>N¯cπA,O\(r\)2=L\(f^N¯\)\.L^\{\\mathrm\{eff\}\}\_\{A,O,T\}\(\\bar\{N\}\)=L\_\{A\}^\{\\infty\}\(T\)\+\\sum\_\{r\>\\bar\{N\}\}c^\{2\}\_\{\\pi\_\{A,O\}\(r\)\}=L\(\\hat\{f\}\_\{\\bar\{N\}\}\)\.
This quantity is the minimum loss attained by the monotone toy learner over budgetsN≤N¯N\\leq\\bar\{N\}\. It depends on priority order as well as support and converges toLA∞\(T\)L\_\{A\}^\{\\infty\}\(T\)asN¯→∞\\bar\{N\}\\to\\infty\. Over a bounded range, the unresolved tail can appear as part of a fitted plateau \(§3\.2\)\.
Proof of Proposition 2\.Suppress the fixed\(A,T\)\(A,T\)indices\. LetKN=\{j:ij∈IA,O\(N\)\}K\_\{N\}=\\\{j:i\_\{j\}\\in I\_\{A,O\}\(N\)\\\}, letJN=max\{m:\{1,…,m\}⊆KN\}J\_\{N\}=\\max\\\{m:\\\{1,\\ldots,m\\\}\\subseteq K\_\{N\}\\\}, and letA¯\(m\)=∑j\>maj\\overline\{A\}\(m\)=\\sum\_\{j\>m\}a\_\{j\}\. Because the acquired set contains at mostNNpositive\-target modes anda1≥a2≥⋯a\_\{1\}\\geq a\_\{2\}\\geq\\cdots,
∑j∈KNaj≤∑j=1Naj\.\\sum\_\{j\\in K\_\{N\}\}a\_\{j\}\\leq\\sum\_\{j=1\}^\{N\}a\_\{j\}\.
Subtracting both sides from∑j≥1aj\\sum\_\{j\\geq 1\}a\_\{j\}gives the best\-NN\-mode lower bound
E\(N\)≥A¯\(N\)\.E\(N\)\\geq\\overline\{A\}\(N\)\.
By definition, every rankj≤JNj\\leq J\_\{N\}has been acquired\. Hence every unresolved positive\-target mode has rank greater thanJNJ\_\{N\}, which gives
E\(N\)≤A¯\(JN\)\.E\(N\)\\leq\\overline\{A\}\(J\_\{N\}\)\.
Together,
A¯\(N\)≤E\(N\)≤A¯\(JN\)\.\\overline\{A\}\(N\)\\leq E\(N\)\\leq\\overline\{A\}\(J\_\{N\}\)\.
Taking negative logarithms, dividing bylogN\\log N, and using
−logA¯\(JN\)logN=−logA¯\(JN\)logJNlogJNlogN\\frac\{\-\\log\\overline\{A\}\(J\_\{N\}\)\}\{\\log N\}=\\frac\{\-\\log\\overline\{A\}\(J\_\{N\}\)\}\{\\log J\_\{N\}\}\\frac\{\\log J\_\{N\}\}\{\\log N\}
gives the claimed liminf–limsup interval\. IfE\(N\)≥ηA¯\(JN\)E\(N\)\\geq\\eta\\overline\{A\}\(J\_\{N\}\), then1≤A¯\(JN\)/E\(N\)≤1/η1\\leq\\overline\{A\}\(J\_\{N\}\)/E\(N\)\\leq 1/\\eta, so
−logE\(N\)logN−−logA¯\(JN\)logN⟶0,\\frac\{\-\\log E\(N\)\}\{\\log N\}\-\\frac\{\-\\log\\overline\{A\}\(J\_\{N\}\)\}\{\\log N\}\\longrightarrow 0,
andα=ργ\\alpha=\\rho\\gamma\. The same argument works wheneverA¯\(JN\)/E\(N\)=No\(1\)\\overline\{A\}\(J\_\{N\}\)/E\(N\)=N^\{o\(1\)\}\. Finally, the integral test foraj≍j−ba\_\{j\}\\asymp j^\{\-b\},b\>1b\>1, gives
A¯\(m\)=∑j\>maj≍m−\(b−1\)\.\\overline\{A\}\(m\)=\\sum\_\{j\>m\}a\_\{j\}\\asymp m^\{\-\(b\-1\)\}\.
Combined withJN≍NρJ\_\{N\}\\asymp N^\{\\rho\}and the two\-sided bounded\-gain inequality, this yieldsE\(N\)≍N−ρ\(b−1\)E\(N\)\\asymp N^\{\-\\rho\(b\-1\)\}\.■\\blacksquare
Remark \(off\-prefix acceleration decomposition\)\.ForN\>1N\>1, define
δN=log\[A¯\(JN\)/E\(N\)\]logN≥0\.\\delta\_\{N\}=\\frac\{\\log\[\\overline\{A\}\(J\_\{N\}\)/E\(N\)\]\}\{\\log N\}\\geq 0\.
The exact identity
−logE\(N\)logN=−logA¯\(JN\)logJNlogJNlogN\+δN\\frac\{\-\\log E\(N\)\}\{\\log N\}=\\frac\{\-\\log\\overline\{A\}\(J\_\{N\}\)\}\{\\log J\_\{N\}\}\\frac\{\\log J\_\{N\}\}\{\\log N\}\+\\delta\_\{N\}
shows that, ifδN→δ\\delta\_\{N\}\\to\\delta, thenα=ργ\+δ\\alpha=\\rho\\gamma\+\\delta\. The universal interval gives0≤δ≤\(1−ρ\)γ0\\leq\\delta\\leq\(1\-\\rho\)\\gamma\. This quantity is an appendix\-level way to locate the gap between completed\-prefix coverage and the actual residual; the main product law needs onlyδ=0\\delta=0\.
Remark \(non\-vacuity of the prefix\-limited assumptions\)\.For any chosenρ∈\(0,1\]\\rho\\in\(0,1\], consider an instance whose architectural support is
SA=\{i1,i2,…\}∪\{z1,z2,…\},S\_\{A\}=\\\{i\_\{1\},i\_\{2\},\\ldots\\\}\\cup\\\{z\_\{1\},z\_\{2\},\\ldots\\\},
where theiji\_\{j\}are the nonzero\-target modes and thezmz\_\{m\}enumerate all remaining orthogonal zero\-target\-power directions\. Placeiji\_\{j\}at position
pj=\{⌈j1/ρ⌉,0<ρ<1,2j,ρ=1,p\_\{j\}=\\begin\{cases\}\\lceil j^\{1/\\rho\}\\rceil,&0<\\rho<1,\\\\ 2j,&\\rho=1,\\end\{cases\}
and fill the remaining positions with thezmz\_\{m\}\. The endpoint adjustment atρ=1\\rho=1leaves room for the filler sequence while preserving linear\-order coverage\. Then
J\(N\)=max\{j:pj≤N\}≍Nρ\.J\(N\)=\\max\\\{j:p\_\{j\}\\leq N\\\}\\asymp N^\{\\rho\}\.
Because the filler modes carry zero target power,
E\(N\)=∑j\>J\(N\)aj,E\(N\)=\\sum\_\{j\>J\(N\)\}a\_\{j\},
so the bounded off\-prefix\-gain condition holds withη=1\\eta=1\. This construction provides a non\-vacuous witness for the joint assumptions\. Its role is to establish their mathematical consistency; realistic training dynamics remain the target of the derivation described in §5\.3\.
Remark \(completed\-prefix growth alone is insufficient\)\.Fix0<ρ<10<\\rho<1, letaj≍j−ba\_\{j\}\\asymp j^\{\-b\}withb\>1b\>1, reserve the sparse ranksdm=2md\_\{m\}=2^\{m\}, and write𝒬=\{dm:m≥1\}\\mathcal\{Q\}=\\\{d\_\{m\}:m\\geq 1\\\}\. Place eachdmd\_\{m\}at priority positionpm=⌈dm1/ρ⌉p\_\{m\}=\\lceil d\_\{m\}^\{1/\\rho\}\\rceil, reserving those positions before the remaining order is filled\. At every other priority position, place the smallest unused rank outside𝒬\\mathcal\{Q\}\. This defines one fixed bijection and therefore a nested acquisition sequence\.
Forpm≤N<pm\+1p\_\{m\}\\leq N<p\_\{m\+1\},pm/dm\+1→∞p\_\{m\}/d\_\{m\+1\}\\to\\infty, so for all sufficiently largemmevery nonreserved rank belowdm\+1d\_\{m\+1\}has already been placed before budgetpmp\_\{m\}, whereasdm\+1d\_\{m\+1\}remains unacquired\. Thus
J\(N\)=dm\+1−1≍Nρ\.J\(N\)=d\_\{m\+1\}\-1\\asymp N^\{\\rho\}\.
By budgetNN,N−m\+O\(1\)N\-m\+O\(1\)nonreserved ranks have been acquired, so the first unacquired nonreserved rank isΘ\(N\)\\Theta\(N\)\. The unresolved energy therefore consists of the future reserved ranks and an ordinary tail beginning at orderNN:
E\(N\)≍∑r≥m\+1dr−b\+∑j≳Nj∉𝒬j−b≍N−ρb\+N−\(b−1\)\.E\(N\)\\asymp\\sum\_\{r\\geq m\+1\}d\_\{r\}^\{\-b\}\+\\sum\_\{\\begin\{subarray\}\{c\}j\\gtrsim N\\\\ j\\notin\\mathcal\{Q\}\\end\{subarray\}\}j^\{\-b\}\\asymp N^\{\-\\rho b\}\+N^\{\-\(b\-1\)\}\.
Its exponent is
min\{ρb,b−1\}\>ρ\(b−1\)\.\\min\\\{\\rho b,b\-1\\\}\>\\rho\(b\-1\)\.
ThusJ\(N\)≍NρJ\(N\)\\asymp N^\{\\rho\}can coexist with a strictly faster residual rate when dispersed acquisitions remove polynomially more of the tail\. Equivalently,δ\>0\\delta\>0in the decomposition above\. The product\-law endpoint therefore uses the prefix\-adequacy condition\.
\(a\) Prefix\-exact,ρ=1\\rho=1\. α^=0\.600\\widehat\{\\alpha\}=0\.600\.
\(b\) Prefix\-limited,ρ=0\.6\\rho=0\.6\. α^=0\.360\\widehat\{\\alpha\}=0\.360\.
\(c\) Interleaved,ρ=0\.6\\rho=0\.6\. α^=0\.606\\widehat\{\\alpha\}=0\.606\.
Figure 2:Finite\-budget evaluation of Proposition 2 and the constructions in Appendix B\.2 withaj=j−1\.6a\_\{j\}=j^\{\-1\.6\}, soγ=0\.6\\gamma=0\.6\. Dashed lines showA¯\(N\)\\overline\{A\}\(N\), dotted lines showA¯\(JN\)\\overline\{A\}\(J\_\{N\}\), and marked solid lines showE\(N\)E\(N\)\. In panel \(a\), exact prefix acquisition collapses both bounds and givesα=γ=0\.6\\alpha=\\gamma=0\.6\. In panel \(b\), the zero\-target filler construction hasE\(N\)=A¯\(JN\)E\(N\)=\\overline\{A\}\(J\_\{N\}\)andα=ργ=0\.36\\alpha=\\rho\\gamma=0\.36\. Panel \(c\) uses the same completed\-prefix rate as panel \(b\), but interleaved acquisition moves the residual toward the best\-NNtail: the fitted slope is0\.6060\.606and approaches the asymptotic value0\.60\.6, while the product\-law endpoint is0\.360\.36\. Fitted slopes use the upper half of the displayed log\-budget range\.High\-value coverage is stronger than rank density: an order could resolve many low\-power modes while leaving high\-power modes untouched\. The completed\-prefix variable records the guaranteed head, while the appendix diagnostic records what the remaining acquisitions accomplish\. Modes outsideSAS\_\{A\}contribute to the floor in Proposition 1 and never enterEA,O,T\(N\)E\_\{A,O,T\}\(N\)\. Because the supported sequence depends onSAS\_\{A\}, its cumulative\-tail rate can depend on architecture as well as data\. Its task\-side reading in the aligned full\-support regime is relative to a basis fixed ex ante independently of the coupling\. The power\-law result supplies bounded multiplicative constants; the prefactor may vary within those bounds\. Together, Propositions 1–2 establish
LA,O,T\(N\)−LA∞\(T\)≍N−αA,O,T\(N→∞\)L\_\{A,O,T\}\(N\)\-L^\{\\infty\}\_\{A\}\(T\)\\;\\asymp\\;N^\{\-\\alpha\_\{A,O,T\}\}\\qquad\(N\\to\\infty\)
under the power\-law, completed\-prefix, and bounded\-gain conditions\. The floor–tail decomposition itself is exact\.
Proof of Corollary 1\.Under common full support in the same ex ante task\-side basis, Proposition 2 givesαq,t=ρq,tγt\\alpha\_\{q,t\}=\\rho\_\{q,t\}\\gamma\_\{t\}\. Becauseγt\>0\\gamma\_\{t\}\>0, the assumed ordering ofρ\\rhois preserved forT1T\_\{1\}and reversed forT2T\_\{2\}\. Ifα1\>α2\\alpha\_\{1\}\>\\alpha\_\{2\}, then
log\(E1\(N\)/E2\(N\)\)logN⟶−\(α1−α2\)<0,\\frac\{\\log\(E\_\{1\}\(N\)/E\_\{2\}\(N\)\)\}\{\\log N\}\\longrightarrow\-\(\\alpha\_\{1\}\-\\alpha\_\{2\}\)<0,
soE1\(N\)/E2\(N\)→0E\_\{1\}\(N\)/E\_\{2\}\(N\)\\to 0\. Thus residual\-loss ordering follows exponent ordering for all sufficiently largeNN; equal within\-task floors give the same ordering for total loss\.■\\blacksquare
### B\.3 What the toy model reproduces
*\(a\) Architecture\-dependent\(α,L∞\)\(\\alpha,L^\{\\infty\}\)on fixed data\.*Two couplings applied to the same target can differ inρA,O,T\\rho\_\{A,O,T\}and support, and therefore in exponent and floor, while the data distribution remains fixed\. The architecture dependence of the exponent parallels the same\-data results of Ngo and Ravanbakhsh \(2026\); architecture\-specific strict\-floor estimation is a next test \(§5\.3\)\.
*\(b\) Feature learning under task misalignment\.*In the toy model, feature learning is represented as reprioritization within fixed support\. When the initial order serves a task poorly \(ρinit,T<1\\rho\_\{\\text\{init\},T\}<1\), feature learning can raiseρA,O,T\\rho\_\{A,O,T\}toward11and improve the exponent; for an already aligned task \(ρinit,T=1\\rho\_\{\\text\{init\},T\}=1\), the exponent is unchanged\. This order\-side mechanism parallels the easy/hard dichotomy of Bordelon et al\. \(2025\); §4\.1 states the axis\-specificity of that evidence\. Fixed support is an explicit assumption of the construction\. Jelassi et al\.’s \(2024\) task\-specific copying bound provides a wiring\-level example consistent with a support constraint\.
*\(c\) Optimizer as an order\-side channel\.*In the toy model, preconditioning changes which directions descent resolves early: it intervenes onπ\\pi, and hence onρA,O,T\\rho\_\{A,O,T\}, without affecting support\. Ramani and Jain’s \(2026\) controlled random\-feature results support the narrower claim that optimizer choice can change fitted scaling exponents\. The construction assumes that this effect operates through order and remains bounded by support; feature\-learning networks may require a richer decomposition\.
*\(d\) Superposition as a mechanism\-specific analogue\.*Liu, Liu, and Gore \(2025\) derive an interference contribution from overlaps among representation vectors\. Their weak\-superposition regime inherits its exponent from the feature\-frequency tail, while strong superposition produces a robust inverse\-width contribution\. The two constructions operate at different levels: vector\-overlap geometry supplies the mechanism\-specific calculation, while Coupled Scaling places such interference within a task\-conditioned comparison of coupling\-dependent effective geometry and its null regimes\.
*\(e\) Recovery of the accessibility\-unconstrained regime\.*With full support in a task\-side basis fixed ex ante,ρA,O,T=1\\rho\_\{A,O,T\}=1givesαA,O,T=bT−1\\alpha\_\{A,O,T\}=b\_\{T\}\-1and a zero representational floor; restoring the observation\-noise constant omitted in §B\.1 returns the floor to its irreducible noise level\. This is the data\-side limiting case described in §3\.2 when the common basis is fixed ex ante independently of the coupling\.
*\(f\) Frequency\-ordered learning as a special case\.*Under the usage\-frequency hypothesis of Michaud et al\. \(2023\), if corpus statistics alone determineπ\\pi, the order is shared across architectures and emergence ordering is model\-independent\. This is the data\-side limiting case from which §4\.2 seeks to detect a finer\-grained departure\.
*\(g\) Finite\-budget grading\.*At a maximum feasible budgetN¯\\bar\{N\}, modes ranked beyondN¯\\bar\{N\}contribute to the order\-dependent residual in the Remark, while the strict asymptotic floor remains unchanged\.
### B\.4 Relation to solvable scaling models
Maloney, Roberts, and Sully \(2022\), Bordelon, Canatar, and Pehlevan \(2020\), Canatar, Bordelon, and Pehlevan \(2021\), and Bahri et al\. \(2024\) derive learning curves from spectra, target alignment, and statistical regime\. Bordelon, Atanasov, and Pehlevan \(2024\) add a rank\-constrained dynamics in which top\-k⋆k\_\{\\star\}truncation and target spectral tails generate multiple scaling laws\. Zou et al\. \(2026\) and Song et al\. \(2026\) establish the neighboring tail–frontier line; §2\.1 shows the exact prefix specialization that recovers Zou et al\.’s law and the identification contrast with Song et al\. Proposition 2 extends that local tail calculation to coupling\-indexed support, interleaved target ranks, and a universal bracket, while Corollary 1 supplies the cross\-task comparison\.
Liu, Liu, and Gore \(2025\) derive a complementary model\-width mechanism from geometric overlap under strong superposition\. In the lazy limit, Coupled Scaling’s support and acquisition objects have spectral analogues: available kernel modes encode support, and the task\-weighted spectral measure combines eigenvalue strength with target alignment\. Asymptotic order alignment recovers the product structure; the general case is represented by the joint spectral measure\.
### B\.5 Rate Classes and Network Extension
The solvable instance fixes an ex ante basis and target ranking, assigns one budget unit to one mode, and uses omniscient within\-order allocation\. Within this normalization it yields an exact floor–tail decomposition, an exact coverage bracket, an exponent interval, and the product\-law endpoint under bounded or exponent\-neutral off\-prefix gain\. Finite support, exponentially decaying tails, and unstable tail log\-rates form separate rate classes\.
The deep\-network extension replaces stipulated support and priority with quantities derived from architecture and stochastic optimization\. Section 5\.3 identifies the corresponding measurement program: network\-native resource maps, cumulative target tails, completed\-prefix coverage, and rank\-resolved off\-prefix acceleration\.
## Appendix C: Evidence and Test Specification
### C\.1 Detailed Evidence Matrix
Table C\.1\.Reported evidence on model\-side dependence in scaling, organized by intervention, controls, identified effect, and remaining inferential scope\.
EvidenceInterventionMain controlsIdentified effectOpen inferential targetNgo & Ravanbakhsh \(2026\)Equivariant versus non\-equivariant architectureSame neural\-force\-field data and taskArchitecture\-dependent parameter\-, data\-, and compute\-scaling exponentsPositive architecture\-specific strict floor; accessibility mechanismTay et al\. \(2023\)Ten language\-model architecturesMatched pretraining and evaluation protocolArchitecture\-dependent scaling behavior and rank changes with scaleSpecific geometric mechanism; accessibility\-consistent trackingLiu, Liu, & Gore \(2025\)Weight\-decay\-controlled superpositionSame toy architecture and feature distributions across superposition regimes; four open LLM families for external consistencyGeometric interference can generate robust model\-width scaling in a controlled representation modelGeneral task\-conditioned geometry tracking across architecture–optimization couplings; data\- or training\-time scalingBordelon et al\. \(2025\)Feature\-learning regimeControlled task/model settingHard\-task training\-time and compute exponents depend on feature learningParameter\- or data\-scaling exponent dependenceRamani & Jain \(2026\)Optimizer/preconditionerSame random\-feature model within each spectral conditionOptimizer\-dependent fitted model\-size exponentsTransfer beyond controlled lazy random features; strict\-floor effectsJha & Reagen \(2026\)AdamW, Muon, NorMuon, and rank\-constrained DionCommon GPT\-style model family, FineWeb\-Edu recipe, and FFN\-width grid; one seed per cellOptimizer\-dependent FFN effective\-rank exponents; an extended\-training AdamW–low\-rank\-Dion control separates spectral scaling at matched perplexityTask\-conditioned geometry\-to\-performance tracking; seed robustnessVolkova et al\. \(2026\)Optimizer choice in LLM pretrainingCommon model family, corpus, objective, andN,DN,Dgrid within each architecture–dataset instanceSeparate fits are ill\-conditioned; a constrained shared\-exponent rescaling improves stability and extrapolationIndependently tested exponent equivalence; task\-relevant geometry and coupling\-by\-task interactionBansal et al\. \(2022\)NMT architecture and data conditionsShared corpus and scale protocol within comparisonsData\-scaling exponents minimally affected by the tested architecture/task\-setup, filtering, and iid\-noise interventions; significant degradation with back\-translated dataInvariance beyond the tested interventions and objectivesXiao et al\. \(2025\)Historical model pipelineNo factorial isolationCapability\-density trendSeparate effects of data, architecture, optimization, and evaluation
### C\.2 Test Specification
Measurement and freezing\.Geometry definitions, probes, normalization, scalarization, rank conventions, and scales are frozen\. Held\-out probes define the trajectory independently of loss outcomes\. Where feasible, lower scales estimate its rate and larger held\-out scales evaluate both the prediction and curve family\.
Tail, coverage, and prefix adequacy\.Confirmatory Proposition 2 cells preregister the cumulative supported\-target\-tail rate, a completed high\-value\-prefix trajectory, and one of two prefix\-adequacy routes\. Common full support in a preselected task basis or independent tail estimation suppliesγt\\gamma\_\{t\}\. Rank\-resolved designs estimate off\-prefix acceleration with the geometry\-side diagnostic in §3\.5; controlled designs establish bounded or exponent\-neutral off\-prefix gain through construction or independent theory\. Zero polynomial acceleration selects the product\-law endpoint, and a stable positive slope selects the universal\-interval prediction\. Other task cells contribute to the preregistered geometry\-tracking analysis\.
Cross\-task standardization\.The exponent interactionΔα=dα\(t1\)−dα\(t2\)\\Delta\_\{\\alpha\}=d\_\{\\alpha\}\(t\_\{1\}\)\-d\_\{\\alpha\}\(t\_\{2\}\)is a summary with propagated uncertainty\. A cross\-task geometry difference\-in\-differences requires common standardization against a shared baseline, null, or ceiling\. Under task\-specific rescalinggq,t′=atgq,tg^\{\\prime\}\_\{q,t\}=a\_\{t\}g\_\{q,t\},at\>0a\_\{t\}\>0, within\-task order is unchanged, while an unstandardizedΔg\\Delta\_\{g\}can change magnitude or sign\. Primary geometry tests therefore use within\-task directions\.
Curve fitting and uncertainty\.Within each cell,L∞L^\{\\infty\},CC, andα\\alphaare fitted jointly, and their joint uncertainty propagates todα\(t\)d\_\{\\alpha\}\(t\)andΔα\\Delta\_\{\\alpha\}\. A power law plus floor, a broken power law, and other preregistered families are compared on held\-out or out\-of\-range prediction\. Because floor and exponent correlate over finite ranges, a strict\-floor reading must separate an asymptote from curvature, regime change, and an unresolved tail\.
Statistical units and calibration\.Each model size is a point on a shared scaling curve\. Stochastic replication comes from seeds within a task instance, and task\-family inference comes from independently sampled task instances\. A pilot sets both counts; one instance per family restricts the claim to the studied tasks\. A randomizationpp\-value is calibrated by the experiment’s actual randomization; §4\.2 supplies observational reference distributions\. Parameter\- and compute\-matched analyses remain separate, and each static contrast shares a resource axis, matching convention, andN¯\\bar\{N\}\. The primary static contrast is state\-matched:gq,tlevel\(N¯\)g^\{\\mathrm\{level\}\}\_\{q,t\}\(\\bar\{N\}\)is paired withLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)under the same resource axis and matching convention\.Lq,teff\(N¯\)L^\{\\mathrm\{eff\}\}\_\{q,t\}\(\\bar\{N\}\)is a secondary finite\-range summary and replacesLq,t\(N¯\)L\_\{q,t\}\(\\bar\{N\}\)only when monotonicity over the studied range has been established\. The cell\-level loss is the preregistered aggregate across seeds and task instances under the stated hierarchical or repeated\-run analysis\.
Equivalence implementation\.The directional equivalence test combines verified completed\-prefix\-rate separation, the registered prefix\-adequacy route, and a joint compatibility interval fordα\(t\)d\_\{\\alpha\}\(t\)inside the prespecified margin\. Exponent equivalence under that combination challenges the directional prediction\. Cells with unresolved separation or broad compatibility intervals remain unclassified\.
## References
Aghajanyan, A\., Gupta, S\., and Zettlemoyer, L\. \(2021\)\. Intrinsic dimensionality explains the effectiveness of language model fine\-tuning\. In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 7319–7328\.
Ansuini, A\., Laio, A\., Macke, J\. H\., and Zoccolan, D\. \(2019\)\. Intrinsic dimension of data representations in deep neural networks\. In*Advances in Neural Information Processing Systems 32 \(NeurIPS\)*, pages 6111–6122\.
Arora, S\., Eyuboglu, S\., Timalsina, A\., Johnson, I\., Poli, M\., Zou, J\., Rudra, A\., and Ré, C\. \(2024\)\. Zoology: Measuring and improving recall in efficient language models\. In*The Twelfth International Conference on Learning Representations \(ICLR\)*\.
Bahri, Y\., Dyer, E\., Kaplan, J\., Lee, J\., and Sharma, U\. \(2024\)\. Explaining neural scaling laws\.*Proceedings of the National Academy of Sciences*, 121\(27\):e2311878121\.
Bansal, Y\., Ghorbani, B\., Garg, A\., Zhang, B\., Cherry, C\., Neyshabur, B\., and Firat, O\. \(2022\)\. Data scaling laws in NMT: The effect of noise and architecture\. In*Proceedings of the 39th International Conference on Machine Learning*, PMLR 162:1466–1482\.
Bingham, N\. H\., Goldie, C\. M\., and Teugels, J\. L\. \(1987\)\.*Regular Variation*\. Cambridge University Press\.
Bordelon, B\., Atanasov, A\., and Pehlevan, C\. \(2024\)\. A dynamical model of neural scaling laws\. In*Proceedings of the 41st International Conference on Machine Learning \(ICML\)*, PMLR 235:4345–4382\.
Bordelon, B\., Atanasov, A\., and Pehlevan, C\. \(2025\)\. How feature learning can improve neural scaling laws\.*Journal of Statistical Mechanics: Theory and Experiment*, 2025\(8\):084002\.
Bordelon, B\., Canatar, A\., and Pehlevan, C\. \(2020\)\. Spectrum dependent learning curves in kernel regression and wide neural networks\. In*Proceedings of the 37th International Conference on Machine Learning \(ICML\)*, PMLR 119:1024–1034\.
Caballero, E\., Gupta, K\., Rish, I\., and Krueger, D\. \(2023\)\. Broken neural scaling laws\. In*The Eleventh International Conference on Learning Representations \(ICLR\)*\.
Cagnetta, F\., Favero, A\., Sclocchi, A\., and Wyart, M\. \(2025\)\. Scaling laws and representation learning in simple hierarchical languages: Transformers versus convolutional architectures\.*Physical Review E*, 112:065312\.
Canatar, A\., Bordelon, B\., and Pehlevan, C\. \(2021\)\. Spectral bias and task\-model alignment explain generalization in kernel regression and infinitely wide neural networks\.*Nature Communications*, 12:2914\.
Caponnetto, A\. and De Vito, E\. \(2007\)\. Optimal rates for the regularized least\-squares algorithm\.*Foundations of Computational Mathematics*, 7\(3\):331–368\.
Cheng, D\., Liu, Z\., Sun, J\., Xia, F\., Zhang, B\., Liu, D\., and Zhang, Y\. \(2026\)\. A qualitative test\-risk mechanism for scaling behavior in normalized residual networks\.*arXiv preprint*, arXiv:2605\.08297\.
Chizat, L\., Oyallon, E\., and Bach, F\. \(2019\)\. On lazy training in differentiable programming\. In*Advances in Neural Information Processing Systems 32 \(NeurIPS\)*, pages 2937–2947\.
Defilippis, L\., Krzakala, F\., Loureiro, B\., and Maillard, A\. \(2026a\)\. Optimal scaling laws in learning hierarchical multi\-index models\.*arXiv preprint*, arXiv:2602\.05846\.
Defilippis, L\., Xu, Y\., Girardin, J\., Troiani, E\., Erba, V\., Zdeborová, L\., Loureiro, B\., and Krzakala, F\. \(2026b\)\. Scaling laws and spectra of shallow neural networks in the feature learning regime\. In*The Fourteenth International Conference on Learning Representations \(ICLR\)*\.
Hestness, J\., Narang, S\., Ardalani, N\., Diamos, G\., Jun, H\., Kianinejad, H\., Patwary, M\. M\. A\., Yang, Y\., and Zhou, Y\. \(2017\)\. Deep learning scaling is predictable, empirically\.*arXiv preprint*, arXiv:1712\.00409\.
Hoffmann, J\., Borgeaud, S\., Mensch, A\., Buchatskaya, E\., Cai, T\., Rutherford, E\., de Las Casas, D\., Hendricks, L\. A\., Welbl, J\., Clark, A\., Hennigan, T\., Noland, E\., Millican, K\., van den Driessche, G\., Damoc, B\., Guy, A\., Osindero, S\., Simonyan, K\., Elsen, E\., Vinyals, O\., Rae, J\. W\., and Sifre, L\. \(2022\)\. Training compute\-optimal large language models\. In*Advances in Neural Information Processing Systems 35 \(NeurIPS\)*, pages 30016–30030\.
Huang, J\., Wurgaft, D\., Bansal, R\., Ruis, L\., Saphra, N\., Alvarez\-Melis, D\., Lampinen, A\. K\., Potts, C\., and Lubana, E\. S\. \(2026\)\. Why larger models learn more: Effects of capacity, interference, and rare\-task retention\.*arXiv preprint*, arXiv:2605\.29548\.
Jha, N\. K\. and Reagen, B\. \(2026\)\. Same architecture, different capacity: Optimizer\-induced spectral scaling laws\.*arXiv preprint*, arXiv:2605\.21803\.
Jelassi, S\., Brandfonbrener, D\., Kakade, S\. M\., and Malach, E\. \(2024\)\. Repeat after me: Transformers are better than state space models at copying\. In*Proceedings of the 41st International Conference on Machine Learning \(ICML\)*, PMLR 235:21502–21521\.
Kaplan, J\., McCandlish, S\., Henighan, T\., Brown, T\. B\., Chess, B\., Child, R\., Gray, S\., Radford, A\., Wu, J\., and Amodei, D\. \(2020\)\. Scaling laws for neural language models\.*arXiv preprint*, arXiv:2001\.08361\.
Liu, A\. Z\., Paquette, E\., and Sous, J\. \(2026\)\. Spectral lens: Activation and gradient spectra as diagnostics of LLM optimization\.*arXiv preprint*, arXiv:2605\.05683\.
Liu, E\., Sun, K\., Li, M\., Lee, I\., Tjuatja, L\., Huang, J\.\-T\., and Neubig, G\. \(2026\)\. What do language models learn and when? The implicit curriculum hypothesis\. Accepted at the*Conference on Language Modeling \(COLM 2026\)*\. arXiv:2604\.08510\.
Liu, Y\. and Gore, J\. \(2026\)\. Neural scaling universality: If exponents are fixed, time to understand coefficients\.*arXiv preprint*, arXiv:2606\.25008\.
Liu, Y\., Liu, Z\., and Gore, J\. \(2025\)\. Superposition yields robust neural scaling\. In*Advances in Neural Information Processing Systems 38 \(NeurIPS\)*\. arXiv:2505\.10465\.
Maloney, A\., Roberts, D\. A\., and Sully, J\. \(2022\)\. A solvable model of neural scaling laws\.*arXiv preprint*, arXiv:2210\.16859\.
Michaud, E\. J\., Liu, Z\., Girit, U\., and Tegmark, M\. \(2023\)\. The quantization model of neural scaling\. In*Advances in Neural Information Processing Systems 36 \(NeurIPS\)*, pages 28699–28722\.
Ngo, K\. and Ravanbakhsh, S\. \(2026\)\. Scaling laws and symmetry, evidence from neural force fields\. In*The Fourteenth International Conference on Learning Representations \(ICLR\)*\.
Nikolaou, K\., Scheunemann, J\., Krippendorf, S\., Tovey, S\., and Holm, C\. \(2026\)\. Spectral reach: Understanding neural scaling as progress into the spectral tail\.*arXiv preprint*, arXiv:2605\.31244\.
Ramani, V\. and Jain, S\. V\. \(2026\)\. On the optimizer dependence of neural scaling laws\. In*4th Workshop on High\-dimensional Learning Dynamics \(HiLD\), ICML 2026*\. arXiv:2605\.29387\.
Schaeffer, R\., Miranda, B\., and Koyejo, S\. \(2023\)\. Are emergent abilities of large language models a mirage? In*Advances in Neural Information Processing Systems 36 \(NeurIPS\)*\.
Sharma, U\. and Kaplan, J\. \(2022\)\. Scaling laws from the data manifold dimension\.*Journal of Machine Learning Research*, 23\(9\):1–34\.
Song, Z\., Ji, S\., Li, H\., Cheng, S\., and Huang, C\. \(2026\)\. Data scaling as progressive coverage of a predictive contribution spectrum\.*arXiv preprint*, arXiv:2605\.20196\.
Tay, Y\., Dehghani, M\., Abnar, S\., Chung, H\. W\., Fedus, W\., Rao, J\., Narang, S\., Tran, V\. Q\., Yogatama, D\., and Metzler, D\. \(2023\)\. Scaling laws vs model architectures: How does inductive bias influence scaling? In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 12342–12364\.
Volkova, A\., Safaryan, M\., Lampert, C\. H\., and Alistarh, D\. \(2026\)\. Towards robust scaling laws for optimizers\.*arXiv preprint*, arXiv:2602\.07712\.
Wang, S\., Zhang, G\., Luo, K\., Wu, Y\., Liu, S\., Liu, J\., Huang, W\., Yan, S\., and Li, J\. \(2026\)\. SMELT: Scaling laws for compute\-matched MoE looped Transformers\.*arXiv preprint*, arXiv:2609\.01343\.
Xiao, C\., Cai, J\., Zhao, W\., Lin, B\., Zeng, G\., Zhou, J\., Zheng, Z\., Han, X\., Liu, Z\., and Sun, M\. \(2025\)\. Densing law of LLMs\.*Nature Machine Intelligence*, 7:1823–1833\.
Yang, G\. and Hu, E\. J\. \(2021\)\. Tensor programs IV: Feature learning in infinite\-width neural networks\. In*Proceedings of the 38th International Conference on Machine Learning \(ICML\)*, PMLR 139:11727–11737\.
Zhang, J\., Liu, Z\., Yan, Z\., Zhang, Y\., Tan, G\., Liu, F\., and Cheng, D\. \(2026\)\. Mechanisms of width scaling in normalized residual networks: The effective alignment dimension\.*arXiv preprint*, arXiv:2607\.24887\.
Zou, J\., Gong, Z\., Su, Y\., Tang, H\., and Liu, Y\. \(2026\)\. Effective frontiers: A unification of neural scaling laws\.*arXiv preprint*, arXiv:2602\.02593\.Similar Articles
Unified Neural Scaling Laws
Presents a unified neural scaling law that accurately models deep neural network scaling across multiple dimensions including parameters, dataset size, training steps, and compute, validated across diverse architectures and tasks.
Unified Neural Scaling Laws
This paper presents Unified Neural Scaling Laws (UNSL), a functional form that accurately models and extrapolates deep neural network scaling behaviors as multiple dimensions such as parameters, data, and steps vary simultaneously, improving over previous scaling laws.
@lilianweng: A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason a…
Lilian Weng's blog post provides a comprehensive overview of scaling laws in deep learning, covering their derivation, compute-optimal allocation, and the debate between Kaplan et al. and Chinchilla.
Scaling Laws, Carefully (25 minute read)
A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
The paper introduces the Skaling law, a generalized neural scaling law that couples model capacity and data through an interaction exponent, reducing prediction error by 1.5-3x and enabling full-grid extrapolation using roughly 10x less compute.