相互冲突的监督影响的是模型的承诺选择,而非能力本身:在与约定无关的评分下恰好为零的 12.29σ 排列效应
摘要
本文证明,学习率调度充当了决定数据排序效应是否体现在模型参数中的平均算子:在恒定学习率下可观察到 12.29σ 的排列效应,而在余弦衰减调度下该效应恰好坍缩为零。研究还表明,冲突约定训练会改变模型承诺采用的正确答案形式,但不会改变其整体能力——这一现象是 exact-match 基准测试无法检测到的。
查看缓存全文
缓存时间: 2026/10/02 09:45
# Conflicting Supervision Moves Commitment, Not Capability A 12.29𝜎 arrangement effect that is exactly zero under a convention-agnostic score
Source: [https://arxiv.org/html/2610.00234](https://arxiv.org/html/2610.00234)
\\setmathfont
latinmodern\-math\.otf \[ Extension = \.otf, UprightFont = \*\-regular, BoldFont = \*\-bold, ItalicFont = \*\-italic, BoldItalicFont = \*\-bolditalic, \]
###### Abstract
Train a model on the same problems written under two incompatible conventions, both correct, and ask what the*ordering*of that data writes into the parameters\.
The learning\-rate schedule is not a background condition for that question\. It is the averaging operator, and it decides the answer\.We prove a bound in which the arrangement and the schedule enter the ordering effect as*separate multiplied factors*: the arrangement only as a block period, the schedule only as how much weight the endpoint can place on any one moment of the run\. A decaying schedule cannot put a large step size and an uncontracted remainder at the same moment; a constant one does exactly that at the last step\. That decay moderates ordering effects has been reported in pretraining\[[1](https://arxiv.org/html/2610.00234#bib.bib11)\]; the mechanism, the separation, and a controlled measurement of both halves are ours\. Ten orderings of one corpus, one budget, everything but the path held fixed, run twice under families differing inlr\_scheduler\_typeand nothing else: at a constant rate the interior spans0\.22210\.2221in allocation,11\.6311\.63contrast floors, monotone in how blocked the arrangement is\. Under the single cosine every published arm uses,*the same ten arms occupytwodistinguishable states where their own resolution would allow about ten*, across a54×54\\timeschange of block length; the dispersion between arms does not exceed the seed noise within them \(intraclass correlation−0\.076\-0\.076,\[−0\.409,\+0\.166\]\[\-0\.409,\+0\.166\], three seeds per arm\)\.*“Order matters” and “order does not matter” are the two ends of one knob*, which is what the divided record on ordering looks like from here\.
The theory predicted, and we falsified, the three mechanisms we had pre\-registered\. It also explains the one arm that does not move between the two schedules: at blocked training the arrangement factor is the whole run rather than a short period, so the endpoint is set by the whole of the schedule’s profile and deleting its tail does almost nothing\.
What the path writes is which convention the model commits to, and no exact\-match benchmark can see it\.Across twelve armsaccA\+accB\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}is constant to within9\.7%9\.7\\%while the allocation share runs0\.040\.04to0\.870\.87, so the12\.29σ12\.29\\sigmaarrangement switch this paper measures is*exactly zero*under a convention\-agnostic metric\. That conservation is quoted from the decayed family throughout, the constant\-rate one being a noisier place to read it, and we say where each of the two results is measured rather than merging them\. Marking the convention in the prompt collapses the switch and reaches87\.5%87\.5\\%of the union ceiling\.
## 1Newton’s Apples That Disagree
Train a model on the same problems written under two incompatible conventions, both correct, and one thing is obvious in advance: something is lost\. The interesting question is*what*\.
*The answer is that nothing need be lost from what the model can solve, and a great deal moves in which correct form it commits to\.*Twelve arrangements of one corpus holdaccA\+accB\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}constant to within9\.7%9\.7\\%while the share written under one convention runs from0\.040\.04to0\.870\.87\. The same measurement is a12\.29σ12\.29\\sigmaeffect on one coordinate and*exactly zero*on the other, and which one a benchmark reports is a property of the benchmark rather than of the model\.
*Whether that commitment survives training at all is decided by the learning\-rate schedule*, and this is the paper’s central result\. We prove that the ordering\-dependent part of an endpoint is bounded by a product of two factors that do not communicate: the arrangement enters only through the period of its alternation, and the schedule only through how much weight the endpoint can place on any one moment of the run \(Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)\)\. A schedule decaying to zero cannot put a large step size and an uncontracted remainder at the same moment; a constant one does exactly that at the last step\. The prediction is a knob, and the measurement is the knob at two settings: ten arrangements of one corpus, one budget, the same three seeds, the two families differing inlr\_scheduler\_typeand in nothing else, span2\.202\.20contrast floors under a cosine and11\.6311\.63under a constant rate\.*“Order matters” and “order does not matter” are the two ends of one knob*, which is what the divided record on data ordering looks like from here\.
> The claim in one sentence\.Conflicting supervision does not have to change what a model can solve; it changes which correct convention the model commits to, and whether that commitment survives to the endpoint is set by the schedule, not by the arrangement\.
*Where the question comes from\.*The Capability Convergence Hypothesis \(CCH\) organises*inference*around a bounded state fed by an unbounded stream, and separates a*compressive*channel that mixes the stream intoO\(1\)O\(1\)state from a*verbatim index*that pays to keep bindings addressable\[[2](https://arxiv.org/html/2610.00234#bib.bib26)\]\. Let the bindings*disagree*and the two channels stop being interchangeable: a compressive learner forced to mix has a well\-defined optimum, the mixture, while an indexed one can keep both bindings and answer either way when the query names one\. This paper asks the training\-time form of that question, with the data path as the stream and the parameters as the bounded state\.*That is where the question came from and not what the answer depends on*: every result below is derivable without any of the family’s vocabulary, the correspondence is audited in §[7](https://arxiv.org/html/2610.00234#S7), and Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)states exactly where it stops\.
#### The path space, and the limit that collapses it\.
Every arrangement of a two\-source corpus sits on one ladder: finish sourceAAentirely, thenBB\(blocked\); alternate blocks ofLLoptimiser steps \(ALBL⋯A^\{L\}B^\{L\}\\cdots\); alternate every step \(L=1L\{=\}1\); mix within each batch \(shuf\)\.*Only the order varies*, and the order is a message: withnnrows from each source it carrieslog2\(2nn\)\\log\_\{2\}\\binom\{2n\}\{n\}bits, which grows without bound\. The receiver is the parameter vector at the end of training, and the question this paper asks is how many*distinguishable endpoints*it produces\.
At the ideal end of the ladder the answer is none\. Divide infinitely,η→0\\eta\\to 0with alternation frequency→∞\\to\\infty, and the averaging theorem\[[3](https://arxiv.org/html/2610.00234#bib.bib50),[4](https://arxiv.org/html/2610.00234#bib.bib51)\]says the trajectory follows the mean\-field flow
θ˙=−∇\[12LA\+12LB\],\\dot\{\\theta\}\\;=\\;\-\\nabla\\\!\\big\[\\tfrac\{1\}\{2\}L\_\{A\}\+\\tfrac\{1\}\{2\}L\_\{B\}\\big\],\(1\)whose cross\-entropy minimiser at a contested input isq⋆=\(pA\+pB\)/2q^\{\\star\}=\(p\_\{A\}\+p\_\{B\}\)/2\.*Under conflict the infinitely divided limit installs the Bayes\-optimal mixture*, a weighted coin flip between the two conventions, and every one of thoselog2\(2nn\)\\log\_\{2\}\\binom\{2n\}\{n\}orders lands in the same place\. At the budget anyone actually trains at the receiver is not that deaf, and*the gap between the limit and the budget is what this paper measures*\.
#### How deaf it is, at the schedule the field uses\.
Across every arm that differs only in path \(L=1L\{=\}1to5454,purerand,shuf,blocked; three seeds throughout\) the allocation runs0\.4070\.407to0\.8690\.869\. Against a seed dispersion ofσ^=0\.0234\\hat\{\\sigma\}=0\.0234that range would separate about ten values at a resolution of2σ^2\\hat\{\\sigma\}\. It separatestwo: nine arms between0\.4070\.407and0\.4920\.492, andblockedalone at0\.8690\.869\. The count is not a threshold artefact, the largest gap inside the occupied group being0\.04310\.0431against0\.37700\.3770toblocked, a ratio of8\.7\\mathbf\{8\.7\}, and a second pretraining family gives two groups at20\.7\\mathbf\{20\.7\}\.
*Every arm just counted trains under a cosine, and that is the point rather than a caveat\.*The same ten paths at a constant rate occupy*four*states\. So the two\-state count is not a fact about what a path can carry; it is what survives after a decaying schedule has integrated the path away, which is Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)read as a measurement\. The wall is a resolution limit that the schedule sets, and the two mechanisms that get anything through it, a stopping phase and a query\-time key, are this paper’s other results rather than its exceptions\.
#### Contributions\.
1. 1\.The ordering effect factors, and the schedule is one of the two factors\(Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)\)\. Theπ\\pi\-dependent part of an endpoint is bounded byτL⋅supt‖w‖\\tau\_\{L\}\\cdot\\sup\_\{t\}\\\|w\\\|: the arrangement enters only as a block period, the schedule only as the weight the endpoint can place on one moment\.*We do not claim a new averaging theorem*: theτL→0\\tau\_\{L\}\\to 0corner is classical incremental\-gradient theory \(§[2](https://arxiv.org/html/2610.00234#S2)\), and we do not claim the bound predicts a magnitude\. What is new is that the two factors separate, which turns a divided empirical record into one knob and makes three of this paper’s measurements consequences of each other rather than separate findings\.
2. 2\.The knob, measured at both settings\(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\)\. Ten arrangements, one corpus, one budget, the same three seeds, families differing only inlr\_scheduler\_type:2\.202\.20contrast floors under a cosine against11\.63\\mathbf\{11\.63\}under a constant rate, monotone in how blocked the arrangement is\.blockedmoves1\.01\.0floor between the two against17\.017\.0atL=54L\{=\}54, which is the bound’s corner case \(τL=T\\tau\_\{L\}=T\) and not a coincidence\.
3. 3\.A decomposition that separates what a model can do from which form it writes\(Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)\), computable from the per\-problem scores an evaluation already produces\. Under it this paper’s largest effect and a null are one measurement on two axes, which is a warning about any benchmark whose material admits more than one correct form\.
4. 4\.The stopping phase, and a query\-time key\.blockedis not a distinct mechanism but the lowest\-frequency alternation stopped at maximum swing; a mirror test predicts a sign and a symmetry in advance and finds them on two corpora \(§[5](https://arxiv.org/html/2610.00234#S5)\)\. Marking the convention in the prompt then*collapses the arrangement effect*from5\.165\.16floors to0\.030\.03and reaches87\.5%87\.5\\%of the union ceiling\.*The point is not that a marker helps but that it removes the path’s influence*, which no schedule could \(§[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)\)\.
5. 5\.A record of what did not work\.Three pre\-registered mechanisms died \(batch purity; momentum\-window cancellation; the palindromic schedule splitting theory recommends\), one positive control inverted, one control came back void, and all are reported with the thresholds that decided them and the commits that froze them \(Table[12](https://arxiv.org/html/2610.00234#S6.T12), Appendix[D](https://arxiv.org/html/2610.00234#A4)\)\.
#### What this paper does not claim\.
Five statements a reader could reasonably extract from the above are not supported by what is measured here, and separating them from the ones that are is worth more than another result\.
- •*Not “order does not matter under decay\.”*The cosine interior spans2\.202\.20floors against a minimum detectable difference this design never reaches at any evaluation budget, so the flat interior is a statement about the instrument\. What carries the null is the pooled intraclass correlation \(−0\.076\-0\.076,\[−0\.409,\+0\.166\]\[\-0\.409,\+0\.166\]\) and the two\-state count, both weaker than “no effect\.”
- •*Not a general law of LLM training\.*The ladder, mirror, ratio and key arms are one budget on Qwen2\.5\-7B with two conflict constructions; Qwen3\-8B\-Base replicates the switch and the key and does*not*replicate the ladder’s registered criterion or the mirror \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3), §[5](https://arxiv.org/html/2610.00234#S5)\)\. Figure[12](https://arxiv.org/html/2610.00234#S5.F12)is the travel record, misses included\.
- •*Not independence of the two coordinates\.*We claim*conservation ofCC*, which is measured\.Cov\(p,s\)≠0\\mathrm\{Cov\}\(p,s\)\\neq 0is also measured and fails at5\.175\.17floors, soSSis an allocation over a population whose difficulty and convention preference are correlated \(Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)\)\.
- •*Not that conflict is harmless\.*An exact\-match benchmark reporting one convention sees a real loss, and choosing what a model*commits*to is precisely what production post\-training is for\. What is conserved is the union, and only a scorer reading both conventions recovers it\. What the measurement does contradict is the field’s instinct that*conflict destroys and arrangement decides how much*, an instinct two of our own pre\-registered mechanisms shared: nothing is destroyed, and below corpus scale nothing is even moved\.
- •*Not a quantitative theory\.*Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)is used for signs, orderings and which factor a knob enters\. The direction is derived and every magnitude is measured\.
- •*Not a statement about every coordinate\.*Everything is measured along a one\-parameter family of paths and on one behavioural coordinate, so another coordinate could carry order information we did not look at, and the checkpoints that would settle it were reclaimed after evaluation\. Two controls searched*off*that family and found nothing outside the same interval \(batch purity moves3%3\\%of the span; the mirror relocates within it\), and adding a coordinate can only*raise*the number of occupied states, so two is a floor\.
§[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)shows which coordinate separates the two states, allocation and not capability, so what the path selects is a policy and not a competence\. §[3](https://arxiv.org/html/2610.00234#S3)is where the path, the channel and the two coordinates stop being prose: it defines them, states the three propositions the measurements are read against, and shows that two of the mechanisms we pre\-registered were dead before either ran\. Figure[1](https://arxiv.org/html/2610.00234#S1.F1)is that itinerary as a map: one instrument, one wall, and the three terms measured to get past it, each with the verdict that decided it, so the three can be seen as three of a kind rather than met twenty pages apart\. Figure[2](https://arxiv.org/html/2610.00234#S1.F2)is the object itself, drawn twice, and it is where the schedule result can be seen rather than read: the same paths, the same budget and the same seeds produce two different pictures when one field of the trainer configuration changes\.
THE OBJECTOne corpus, two conventions, both correct\.An ordering of it is a monotone path; there are\(1728864\)\\binom\{1728\}\{864\}of them\. Nothing here is a defect to be repaired, so no remedy has a target to converge to, and the question becomes which coordinate the conflict moves\.THE INSTRUMENTDefinition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)splits one exact\-match score in two, at no cost, from the per\-problem scores an evaluation already produces\. capabilityCC§[4\.2](https://arxiv.org/html/2610.00234#S4.SS2) moves9\.7%9\.7\\%across twelve arms allocation share§[4\.2](https://arxiv.org/html/2610.00234#S4.SS2) moves0\.04→0\.870\.04\\to 0\.87, at12\.29σ12\.29\\sigma A benchmark reports the first and cannot express the second: this paper’s largest effect readsexactly zero\.THE WALLProposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)§[3](https://arxiv.org/html/2610.00234#S3) The endpoint is the*schedule\-weighted average*of what the path visited, so rearranging below corpus scale changes the average not at all\. Three mechanisms were pre\-registeredfor how the path might write\. The theorem predicted, and we falsified, all three\. What survives is not a mechanism list but a question: what is*not*an average?WHAT GETS THROUGH1\. the schedule§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3) constant rate:11\.63\\mathbf\{11\.63\}floors single cosine:2\.202\.20floors the two ends of one knob2\. the stopping phase§[5](https://arxiv.org/html/2610.00234#S5) 23\.89\\mathbf\{23\.89\}floors, and the only term that survives*both*schedules3\. a query\-time key§[5\.1](https://arxiv.org/html/2610.00234#S5.SS1) collapses the switch,5\.16→0\.035\.16\\to\\mathbf\{0\.03\}floors, and reaches87\.5%87\.5\\%of the union ceiling
Figure 1:The paper as a map: one instrument, one wall, and the three terms that get past it\.Read left to right\.*Every box carries a measured verdict rather than a description*, which is the property that decides whether a map of this kind is worth printing: change the experiments and this picture changes, because two of its boxes would say the opposite of what they say and the wall’s three exits would be a different three\. The instrument \(Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)\) is what makes the rest expressible, since the quantity every term here moves is invisible to the score a benchmark reports\. The three exits are not alternatives to one another: the schedule decides whether an ordering survives to the endpoint at all, the stopping phase is the one term that survives either schedule, and the key is a second channel rather than more bandwidth on this one\. Magnitudes are in contrast floors and are the body’s, not recomputed here\.rows ofAAconsumedrows ofBBconsumed00864864864864shufL=54L\{=\}54blocked\(a\) a data path*is*a monotone staircasetrainingsingle cosineconstant rateoptimiser stepη\\eta00TTconstant ratesingle cosinethe one field that differs:lr\_scheduler\_type0\.10\.10\.10\.10\.20\.20\.20\.20\.30\.30\.30\.30\.40\.40\.40\.40\.50\.50\.50\.5accA\\mathrm\{acc\}\_\{\\text\{\\scriptsize A\}\}accB\\mathrm\{acc\}\_\{\\text\{\\scriptsize B\}\}L=1…54L\{=\}1\\ldots 54,purerand,shuf: oneoccupied cellblockedBBonlyAAonlyeightcells here, none occupiedallocationruns the band9\.7%9\.7\\%\(b\) under the single cosine: two occupied states0\.10\.10\.10\.10\.20\.20\.20\.20\.30\.30\.30\.30\.40\.40\.40\.40\.50\.50\.50\.5accA\\mathrm\{acc\}\_\{\\text\{\\scriptsize A\}\}accB\\mathrm\{acc\}\_\{\\text\{\\scriptsize B\}\}interior0\.22210\.2221, 11\.6311\.63floorsL=54L\{=\}54blockedBBonlyAAonlythe*same*eight arms now walk it\(c\) at a constant rate: the band is occupied
Figure 2:The space of data paths, and its image under each of the two schedules\.\(a\)A data path*is*a monotone lattice path: at each optimiser step the trainer consumes a row fromAAor fromBB\. There are\(1728864\)\\binom\{1728\}\{864\}of them: the size of the space sampled, not a quantity anything here transmits\.shufhugs the diagonal, a ladder rung is a coarser staircase,blockedis the corner\.\(b\)and\(c\)Its image, in the read\-out’s own coordinates \(Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)\), every arm of the synthetic conflict, on the three seeds the two families share\.*The two panels differ in one field of the trainer configuration and in nothing else*: same corpus, same paths, same budget, same seeds, same step\. In both, the arms lie in a narrow capability band and spread along it, so capability is the band’s width and allocation is its length\. Under the single cosine every published arm uses, the nine ladder arms fall in*one*cell of the2σ^2\\hat\{\\sigma\}resolution of Definition[3](https://arxiv.org/html/2610.00234#Thmdefinition3)andblockedin another: two states, with eight empty cells between them where the range would support about ten, the eight interior arms spanning0\.0420\.042of share,2\.22\.2contrast floors, below anything this design can resolve\. Under a constant rate those same eight span0\.22210\.2221,11\.6311\.63floors, andL=54L\{=\}54leaves them for the corner\.*The difference between the panels is Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)*: the endpoint is the schedule\-weighted average of what the path visited, so a schedule that decays to zero has averaged the ordering away by the time anyone scores it\. Axes are square and identical across the two panels, so both bands are true45∘45^\{\\circ\}strips and the pictures can be compared by eye\. Source:constlr\_ladder\_table\.finer division⟶\\longrightarrowblocked\(L=162L\{=\}162\)all ofAA, then all ofBBshare0\.8690\.869L=54L\{=\}54one contiguous pass per epochshare0\.4920\.492L=27,18,9,6,3L\{=\}27,18,9,6,3coarse\-to\-fine alternationshare0\.4070\.407–0\.4360\.436L=1L\{=\}1/purerandpure batches, one step longshare0\.4250\.425–0\.4490\.449shufmixed inside each batchshare0\.4130\.413η→0,ω→∞\\eta\\to 0,\\ \\omega\\to\\infty*the infinite\-division limit*share0\.500\.50\(theory\)share==allocation to the minority convention,
accB/\(accA\+accB\)\\mathrm\{acc\}\_\{B\}/\(\\mathrm\{acc\}\_\{A\}\{\+\}\\mathrm\{acc\}\_\{B\}\)*the dashed group*: everything below corpus scale sits at the averaged limit,0\.4070\.407–0\.4490\.449, 2\.20 contrast floors \(offset==pretrained prior\)Figure 3:The division ladder under the single cosine, with the measured allocation at every rung\(synthetic conflict, Qwen2\.5\-7B, three seeds; §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\)\. From one optimiser step to a quarter of the epoch, arrangement does not move the allocation: the whole interior of the ladder sits at the averaging limit\.*That flatness is the schedule’s and not the path’s*: the identical ladder at a constant rate spans11\.6311\.63contrast floors in the same interior\. Under this schedule the only escapes are at the top, stopping the alternation mid\-swing \(§[5](https://arxiv.org/html/2610.00234#S5)\), and off the ladder entirely, via a query\-time key \(§[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)\)\. Source:constlr\_ladder\_table\.
#### The instrument, which outlives the result\.
Score a model on a benchmark whose answers admit more than one correct convention and exact match reports a product of two things: what the model can do, and which form it decided to write in\. Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)splits them at essentially no cost, using only the per\-problem scores an evaluation already produces, and the split is load\-bearing in a way that is easy to state and hard to unsee\.
*This paper’s largest effect is exactly zero on the coordinate that measures capability\.*The arrangement switch we measure at12\.29σ12\.29\\sigma, against its own control inert at0\.15σ0\.15\\sigma, moves the allocation share from0\.040\.04to0\.870\.87and movesaccA\+accB\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}not at all beyond9\.7%9\.7\\%arm to arm\. A twelve\-sigma result and a null are*the same measurement*read on two axes\. Any benchmark carrying contested conventions is silently reporting the first number and calling it the second, and any intervention evaluated that way \(an ordering, a schedule, a data mixture, a decoding change\) can post a large effect while changing nothing a user would call capability\. The recipe\-level member reports the same effect on a different coordinate\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]; the number here is ours, with its own same\-convention control \(Table[2](https://arxiv.org/html/2610.00234#S4.T2)\)\.
We therefore state the decomposition as a contribution in its own right rather than as apparatus for the ladder, and we state its price with it:CCis comparable only within a corpus, andSSequals the convention policy only when the covariance condition of Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)holds, which we test and which fails at5\.175\.17floors\. Both limits are smaller than the effects the coordinates separate, and neither is a reason to keep reporting the product\.
## 2Related Work
We group the literature by which side of the averaging wall its object sits on, a distinction that cuts across the usual subfield boundaries\. One assumption runs through most of it, and this paper’s result is what happens when it is dropped\.
#### The premise almost everyone shares: a conflict is an error\.
The knowledge\-conflict literature names our object exactly\.[Xu et al\. \[6\]](https://arxiv.org/html/2610.00234#bib.bib53)call it*intra\-memory*conflict, discrepancy inside the parameters traced to inconsistency in the training data, and every remedy it surveys is a repair: refine the parametric knowledge, regulate the behaviour, reweight or filter the offending rows\[[7](https://arxiv.org/html/2610.00234#bib.bib49)\]\. A repair presupposes a target, and a target presupposes that one of the conflicting forms is*wrong*\. Our natural corpus is built so that neither is \(Table[1](https://arxiv.org/html/2610.00234#S4.T1); the synthetic one stipulates a convention wrong against mathematical ground truth, and §[6](https://arxiv.org/html/2610.00234#S6)prices the difference\), and under that construction Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)makes the mixture the*loss\-minimising*policy rather than a malfunction, so there is nothing for a repair to converge to and the question becomes which coordinate the conflict moves\. The nearest work to drop the premise is[Krestnikov \[8\]](https://arxiv.org/html/2610.00234#bib.bib61), which trains small transformers on mathematics corpora carrying both correct and incorrect solutions and finds that a*coherent*alternative rule system destroys the preference for the true answer entirely, while adding a*second*competing rule restores most of it\. That is our construction reached from the truth\-tracking side, and it predicts what we measure: two coherent conventions do not degrade capability, they split allocation\. Their corpora make one form wrong and ours make neither, so their restored accuracy and our conservedCCare different quantities that happen to move together, and we know of no measurement in that line of the allocation coordinate or of a query\-time key\. This paper is a limit of that literature rather than a contribution to it: send “one of them is wrong” to zero and the remedies lose their referent while the phenomenon does not\.
The same premise, inverted, organises the evaluation side\.[Plank \[9\]](https://arxiv.org/html/2610.00234#bib.bib54)argues that human label variation is signal rather than noise and that a single gold label is inadequate where annotators legitimately differ, and the perspectivist programme that follows fixes evaluation by matching the*distribution*of human labels\. That repair also needs a target unavailable here: with two correct conventions every allocation is equally correct\. An undecomposed accuracy does not merely undercount capability in the familiar way that exact match penalises*three*against*3*\[[10](https://arxiv.org/html/2610.00234#bib.bib31),[11](https://arxiv.org/html/2610.00234#bib.bib19)\]; it cannot express the coordinate along which our arms move, which is what Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)is for\.[Schaeffer et al\. \[10\]](https://arxiv.org/html/2610.00234#bib.bib31)is the closest precedent and the closest warning: emergence turned out to be a property of the metric rather than of the model, and §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)performs the same move on a different axis when it reports its own12\.29σ12\.29\\sigmaarrangement switch as exactly zero under a convention\-agnostic score\.
Two lines reach our allocation coordinate from the metric side\.[Holtzman et al\. \[12\]](https://arxiv.org/html/2610.00234#bib.bib64)names the mechanism: several surface forms of one correct answer compete for probability mass, so a scorer reading only the highest\-probability string reports the competition rather than the knowledge\. That competition*is*ourSS, and Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)adds only that it can be divided out ofCC\.[Yeom et al\. \[13\]](https://arxiv.org/html/2610.00234#bib.bib82)measure the same split at inference time, finding1616–47%47\\%of instruct\-model hallucinations occur with substantial mass already on the correct concept, the distinguishing factor being whether that mass concentrates on one surface form or disperses across alternatives; their sharpening rises with scale and with instruction tuning, which is a training\-path property, and we read it as the query\-time image of what §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)installs\.[Janeiro et al\. \[14\]](https://arxiv.org/html/2610.00234#bib.bib65)price what it costs an evaluation: on a11–88B testbed, models trained on identical knowledge post false gaps above two points from answer phrasing alone, narrowing to under one point when several paraphrases per option are queried, and the artefact persists at7070–120120B\. Their remedy and ours point in opposite directions on purpose\. ParaEval*averages the surface\-form term away*, which is right when the phrasing is nuisance; we keep it as a coordinate, because in a conflicted corpus the phrasing is exactly what arrangement moves, and our convention\-agnostic score is ParaEval’s move applied to our own headline, duly returning zero \(§[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)\)\. A surface\-form term is nuisance when the training data agree on the convention and signal when they do not\.
#### Where Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)comes from\.
The proposition is not new as optimisation\. In the deterministic cyclic case it is the central dichotomy of the incremental\-gradient literature\[[15](https://arxiv.org/html/2610.00234#bib.bib70),[16](https://arxiv.org/html/2610.00234#bib.bib71)\]: with a step size decaying to zero the iterates converge to a minimiser of the*summed*objective and the order does not survive, while at a constant step they enter a limit cycle whose*position*depends on the order\. Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)interpolates between those regimes and Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3)is that limit cycle\. The without\-replacement line prices the orderings against each other\[[17](https://arxiv.org/html/2610.00234#bib.bib72),[18](https://arxiv.org/html/2610.00234#bib.bib73),[19](https://arxiv.org/html/2610.00234#bib.bib74),[20](https://arxiv.org/html/2610.00234#bib.bib38)\], and the constant\-step\-size bias our constant\-rate family reads is under current study in its own right\[[21](https://arxiv.org/html/2610.00234#bib.bib76)\]\. What we add is not the theorem but its transport: that literature states its results for a fixed objective, and the question here is what the same dichotomy does to a*policy*when the summed objective’s minimiser is a mixture rather than a point\. The transport supplies an allocation coordinate that moves while the loss does not, and the observation that the field’s default schedule puts nearly all of published fine\-tuning practice at one end of the dichotomy without saying so\.
#### Inside the wall: batch composition, shuffling, and per\-step gradient surgery\.
The contradiction that seeded our own v1 \(purify the batch\[[22](https://arxiv.org/html/2610.00234#bib.bib39)\], mix it\[[23](https://arxiv.org/html/2610.00234#bib.bib40)\], mix it with a theorem\[[20](https://arxiv.org/html/2610.00234#bib.bib38)\]\) is resolved sideways rather than adjudicated: at conflict, purity carries3%3\\%of the span, and below corpus scale nothing batch\-sized moves the allocation at all \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\)\. The axis those three papers dispute lies strictly inside the wall, which is why it can persist without either side being wrong about its own measurements\.[Sweeney \[24\]](https://arxiv.org/html/2610.00234#bib.bib41)shows that optimiser state makes shuffle order a first\-order*noise*source; our block sweep shows the momentum window moves the conflict*signal*not at all, and the two are consistent under Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1), buffers adding variance about a mean\-field point the time\-average sets\.[Sweeney \[25\]](https://arxiv.org/html/2610.00234#bib.bib42)proposes the sharpest positive claim we could find, that the Lie bracket of two tasks’ update operators predicts which order transfers better, and Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx3)measures a one\-shot commutator score built in its spirit and finds it inverted on our pairs, at Spearman−0\.543\-0\.543and−0\.600\-0\.600\. That is a range boundary and not a refutation, because the two experiments do not meet: their tournament scores Hessian\-vector products against a sharedθ0\\theta\_\{0\}reference and reports98\.1%/98\.9%98\.1\\%/98\.9\\%pairwise accuracy at block lengthk=1k\{=\}1falling to73\.1%/72\.2%73\.1\\%/72\.2\\%atk=20k\{=\}20, whereas our budgets are162162–324324steps per block\. Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)says why a score computed once atθ0\\theta\_\{0\}must decay withτL\\tau\_\{L\}: it is the leading term of the interior integral in Eq\. \([2](https://arxiv.org/html/2610.00234#S3.E2)\), whose neglected remainder grows with block length\. Per\-step gradient\-conflict methods\[[26](https://arxiv.org/html/2610.00234#bib.bib45),[27](https://arxiv.org/html/2610.00234#bib.bib46),[28](https://arxiv.org/html/2610.00234#bib.bib47)\]and ordered\-shuffle schemes\[[20](https://arxiv.org/html/2610.00234#bib.bib38)\]likewise operate inside the wall: they change optimisation and can change variance, but Remark[1](https://arxiv.org/html/2610.00234#Thmremark1)says they cannot change the installed allocation, because that is fixed by what a held\-out query may condition on\.
#### The wall from the other side: data mixing at corpus scale\.
Methods that reweight*proportions*\(DoReMi\[[29](https://arxiv.org/html/2610.00234#bib.bib16)\], DoGE\[[30](https://arxiv.org/html/2610.00234#bib.bib17)\], and phase\-scheduled mixtures\[[31](https://arxiv.org/html/2610.00234#bib.bib27),[32](https://arxiv.org/html/2610.00234#bib.bib18)\]\) act on exactly the quantity Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)leaves free, the source weightsnA/\(nA\+nB\)n\_\{A\}/\(n\_\{A\}\+n\_\{B\}\)in the mean field\. Our ratio experiment is the controlled version of their premise: moving1:11\{:\}1to2:12\{:\}1moves the allocation share0\.453→0\.3120\.453\\to 0\.312against mean\-field predictions0\.5000\.500and0\.3330\.333\. Corpus composition is the lever, path arrangement is not, and the boundary between them is measurable\.
#### Sequential training and forgetting\.
Catastrophic forgetting\[[33](https://arxiv.org/html/2610.00234#bib.bib3),[34](https://arxiv.org/html/2610.00234#bib.bib4)\]and its mitigations\[[35](https://arxiv.org/html/2610.00234#bib.bib5),[36](https://arxiv.org/html/2610.00234#bib.bib1),[37](https://arxiv.org/html/2610.00234#bib.bib35)\]concern capability lost when a second task overwrites a first, and the continual\-learning literature measures it as such\[[38](https://arxiv.org/html/2610.00234#bib.bib6),[39](https://arxiv.org/html/2610.00234#bib.bib2)\]\. Our decomposition separates that from what conflict does: under Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)forgetting is a movement ofCC, whereas the conflict switch is a movement ofSSatCCheld to within9\.7%9\.7\\%\.[Evron et al\. \[40\]](https://arxiv.org/html/2610.00234#bib.bib34)treat blocked linear regression as alternating projections, and the stopping phase \(Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3)\) is the nonlinear\-policy face of the same recency\. That the order survives at all has support one level down:[Krasheninnikov et al\. \[41\]](https://arxiv.org/html/2610.00234#bib.bib68)fine\-tune sequentially on six datasets and find training\-order recency*linearly encoded*in the activations, with a linear probe separating early\- from late\-learned entities at about90%90\\%\. Their read\-out is on the representation and ours on the policy, the same statement at different depths; what Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)adds is the condition under which it survives to the*endpoint*of a decayed schedule, which is where a reported score is taken\.[Conklin et al\. \[42\]](https://arxiv.org/html/2610.00234#bib.bib20)and[Gustav Olaf Yunus Laitinen\-Fredriksson Lundstrom\-Imanov \[43\]](https://arxiv.org/html/2610.00234#bib.bib22)characterise forgetting mechanistically, and neither predicts a term that switches on when two sources*disagree*while capability holds\.[Xue \[7\]](https://arxiv.org/html/2610.00234#bib.bib49)isolates internal SFT\-data inconsistency per sample, and our decomposition says what that inconsistency does: it moves the model from a deterministic to a stochastic policy at fixed capability, whose per\-sample signature is the coin\-flip fingerprint of §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)\.
#### Curriculum, ordering, and their nulls\.
The record on ordering is divided\[[44](https://arxiv.org/html/2610.00234#bib.bib14),[45](https://arxiv.org/html/2610.00234#bib.bib8),[46](https://arxiv.org/html/2610.00234#bib.bib9),[47](https://arxiv.org/html/2610.00234#bib.bib7),[48](https://arxiv.org/html/2610.00234#bib.bib13),[49](https://arxiv.org/html/2610.00234#bib.bib12),[50](https://arxiv.org/html/2610.00234#bib.bib10),[51](https://arxiv.org/html/2610.00234#bib.bib15),[1](https://arxiv.org/html/2610.00234#bib.bib11)\], and the division is what the averaging wall predicts: these are paths below corpus scale, where Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)says the endpoint is the same and only transients and stopping phases differ\.[Luo et al\. \[1\]](https://arxiv.org/html/2610.00234#bib.bib11)finds curriculum advantages over random shuffling that hold at a constant rate and diminish under standard decay, in pretraining at1\.51\.5B parameters over3030B tokens\. We reproduce that moderator under control, at77B in supervised fine\-tuning, over ten arrangements of one conflicted corpus and two families differing only inlr\_scheduler\_type: interior span2\.202\.20floors under a cosine against11\.6311\.63under a constant rate \(Table[5](https://arxiv.org/html/2610.00234#S4.T5)\)\. Their reading is that decay wastes the curriculum; ours is that decay*is*the averaging operator of Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1), so “order matters” and “order does not matter” are the two ends of one schedule knob rather than two findings to be reconciled\. The divided record should then sort by how much step size survives to the end of training, so a paper reporting an ordering effect without its schedule has not reported a moderator its own effect depends on; Eq\. \([3](https://arxiv.org/html/2610.00234#S3.E3)\) makes that a consequence rather than a caution, since the two factors multiply\. One case can be checked without new runs:[Elgaar and Amiri \[50\]](https://arxiv.org/html/2610.00234#bib.bib10)holds Pythia’s configuration fixed across orderings, decaying a cosine to a tenth of peak at every size, and reports ordering effects on stability largely gone by410410M, which is the sign Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)requires\. The recipe paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]prices the resolution at which any of these comparisons can be made at all, and[Piontkovskaia and Nikolenko \[52\]](https://arxiv.org/html/2610.00234#bib.bib48)reports pairwise order predictions degrading with block length, which Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)explains: a score computed once at initialisation is the leading term of a series whose error accumulates with the step budget\.
#### The sharpest counter\-result, and the regime boundary it draws with ours\.
[LeDoux \[53\]](https://arxiv.org/html/2610.00234#bib.bib62)reports the opposite of everything above, and reports it cleanly\. Training small networks on modular arithmetic*from scratch*, two fixed orderings reach99\.5%99\.5\\%test accuracy from a training set covering0\.3%0\.3\\%of the input space, where random ordering does not, and the learned Fourier representation’s fundamental frequency is the mathematical dual of the ordering’s own structure\. Order there is the mechanism, wide enough that the paper names it a covert information channel able to bypass content\-level auditing\.
We measure ten orderings of one corpus into two distinguishable endpoint states, on a range its own resolution would divide into about ten\. Both results are right, and what the path can write depends on how much the prior has already fixed\. From scratch the parameters carry no structure the order must compete with, so a sufficiently regular order supplies a great deal\. On a pretrained77B prior the same channel writes into parameters that already encode the answer format, the arithmetic and the convention preference, and it moves the one coordinate pretraining left underdetermined: which of two admissible conventions to commit to\. That is one bit, and §[5](https://arxiv.org/html/2610.00234#S5)shows it is the stopping phase\.*Order is a wide channel into an empty state and a narrow one into a full state*, a resource trade neither paper states alone, and it predicts that the channel narrows monotonically with pretraining scale\. The two safety readings converge from opposite regimes:[LeDoux \[53\]](https://arxiv.org/html/2610.00234#bib.bib62)names a covert channel that content auditing misses, and §[6](https://arxiv.org/html/2610.00234#S6)reports the pretrained\-model version, where an attacker controlling only data\-loader order and touching no byte of the corpus moves the allocation0\.220\.22in exact match\. Narrow is not zero, and the bit that survives is the one a user would call the model’s commitment\.
#### Order effects at LLM fine\-tuning scale\.
[Ju et al\. \[54\]](https://arxiv.org/html/2610.00234#bib.bib63)is the closest published claim to a positive result in our own setting: data order produces training imbalance in LLM SFT and degrades performance, with the proposed fix being to merge models fine\-tuned under different orderings\. Read against Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)the remedy is the informative half\. Averaging endpoints across orders explicitly constructs the quantity the infinitely divided limit installs implicitly, so a method that works by merging over orderings is evidence that the orderings differ mainly by a term the mean removes, which is what a flat interior plus a stopping phase predicts\. Their measurements are on heterogeneous instruction data where our coherence premise does not hold, so we report this as consistent rather than as replication\.
#### Theory of the training path\.
The lazy and kernel regimes\[[55](https://arxiv.org/html/2610.00234#bib.bib32),[56](https://arxiv.org/html/2610.00234#bib.bib33)\]give the setting in which the averaging argument is exact, and full\-parameter SFT is measurably not fully lazy, which is why Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)is reported as governing signs and orderings rather than magnitudes\. Averaging itself is classical\[[3](https://arxiv.org/html/2610.00234#bib.bib50),[4](https://arxiv.org/html/2610.00234#bib.bib51)\], and its modern empirical form is[Ajroldi et al\. \[57\]](https://arxiv.org/html/2610.00234#bib.bib69), who benchmark weight averaging across seven workloads, ask whether it can replace learning\-rate decay, and conclude that the two are best combined rather than exchanged\. We use the same pairing as an instrument rather than as a method: if decay and averaging were interchangeable the schedule could not be the operator that decides whether an ordering survives, and their finding that the two compose is what leaveslr\_scheduler\_typefree to be varied on its own, which is the one contrast §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)runs\. Phenomena that live in the transient rather than the endpoint, grokking\[[58](https://arxiv.org/html/2610.00234#bib.bib30)\]and double descent\[[59](https://arxiv.org/html/2610.00234#bib.bib28),[60](https://arxiv.org/html/2610.00234#bib.bib29)\], are outside our budget regime but share the moral that an endpoint measurement can be a statement about where the path stopped\. Attribution methods\[[61](https://arxiv.org/html/2610.00234#bib.bib43),[62](https://arxiv.org/html/2610.00234#bib.bib44)\]track which examples moved the parameters; the write\-channel rate of Definition[3](https://arxiv.org/html/2610.00234#Thmdefinition3)is the complementary statistic, asking how much of the path’s*order*survives at all\. Reading training as a channel with a rate has an ancestor in the information bottleneck\[[63](https://arxiv.org/html/2610.00234#bib.bib21)\], and the difference is not one of degree: the bottleneck compresses the*input*while preserving the*label*, whereas the quantity compressed here is the*order*of a fixed multiset and the receiver is the endpoint’s behaviour\.
#### Family\.
The framework, the thought experiment and the walls are[Chen et al\. \[2\]](https://arxiv.org/html/2610.00234#bib.bib26); this paper is its training\-time member, and the write\-time separability their construction assumes is here a measured quantity rather than an assumption\. The companion recipe\-level paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]measures the same corpora on a different coordinate and prices what recipe search can buy\.
## 3Method: the Path, the Channel, and the Two Coordinates
§[1](https://arxiv.org/html/2610.00234#S1)named three objects in prose: a path, a channel, and two coordinates on the endpoint\. This section defines them, so that what the rest of the paper measures are statements rather than descriptions\. It opens with the question a reader should settle before any of it, which is what*kind*of wall the paper is about, and closes with the three propositions that make two of our own pre\-registered mechanisms dead on arrival\.
###### Definition 1\(Data path and the division ladder\)\.
Fix a multiset𝒟=𝒟A⊎𝒟B\\mathcal\{D\}=\\mathcal\{D\}\_\{A\}\\uplus\\mathcal\{D\}\_\{B\}ofnA\+nBn\_\{A\}\+n\_\{B\}examples from two sources\. A*data path*is an orderingπ\\piof𝒟\\mathcal\{D\}; training isTToptimiser steps of step sizeη\\etaalongπ\\pi\. The*division ladder*is the one\-parameter familyπL\\pi\_\{L\}that alternates same\-source blocks of exactlyLLconsecutive*optimiser steps*, withL=1L=1the finest alternation andshufthe within\-batch mixture\.
*LLis counted in steps, not rows, and the two are reported in different places, so the conversion is fixed here once\.*The optimiser consumes1616rows per step, and each source contributes864864, so one source’s per\-epoch allocation is5454steps and an epoch is108108\. A rung is one arranged one\-epoch file passed over three times,324324steps, which is where its result files are read\.blockedis*not*a rung: it is two stages, three epochs onAAthen three onBBresumed from the first, so each source is one contiguous run of162162steps, the second stage’s counter starts from zero and its files are read at162162, and its effective block isL=162L=162against the ladder’s largest rung at5454\. The budget is the same324324steps either way, and that largest rung isA54B54A^\{54\}B^\{54\}three times over, not a blocked order\. Every arm differs from every other*only*inπ\\pi, at fixed𝒟\\mathcal\{D\},TT,η\\etaand seed, but forblocked’s two cosines where a rung has one, which §[6](https://arxiv.org/html/2610.00234#S6)controls\.
###### Definition 2\(Capability and allocation\)\.
Let the two evaluation sets be the same problems under mutually exclusive conventions, so a generation matches at most one gold\. Write
C=accA\+accB⏟capabilityandS=accB/C⏟allocation,\\underbrace\{C=\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}\}\_\{\\text\{capability\}\}\\qquad\\text\{and\}\\qquad\\underbrace\{S=\\mathrm\{acc\}\_\{B\}/C\}\_\{\\text\{allocation\}\},so that\(accA,accB\)=\(C\(1−S\),CS\)\(\\mathrm\{acc\}\_\{A\},\\mathrm\{acc\}\_\{B\}\)=\(C\(1\-S\),\\,CS\)is a change of coordinates, not a model\. Writep\(x\)p\(x\)for the probability the model solvesxxands\(x\)s\(x\)for the probability it then answers underBB\. Mutual exclusivity givesC=𝔼\[p\]C=\\mathbb\{E\}\[p\]for anysswhatever, which is why capability is the robust half of this pair\. The allocation isS=𝔼\[ps\]/𝔼\[p\]S=\\mathbb\{E\}\[p\\,s\]/\\mathbb\{E\}\[p\], and it equals the convention policy𝔼\[s\]\\mathbb\{E\}\[s\]*exactly whenCov\(p,s\)=0\\mathrm\{Cov\}\(p,s\)=0across problems*\. The condition is a covariance rather than an independence: it is weaker than per\-problem independence, and it is the one the estimator can test\. We measure it below, and it does not vanish\.
Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)is the paper’s instrument and also its sharpest limitation, and we state both here\. It is nearly an identity \(that is why nobody reports it\) and what it buys is that*arrangement claims become claims about which coordinate moves*, which an undecomposed accuracy cannot express\. What it costs is thatCCmeans different things in different corpora; §[6](https://arxiv.org/html/2610.00234#S6)returns to this\.
#### The independence clause is testable, and it fails\.
The sentence above is a conditional, and what matters is what happens when its antecedent is false: if hard problems fall back to the pretraining convention thenssdepends onpp, the coordinates are not orthogonal, and part of what we report as conserved capability is a property of the construction\. The per\-problem arrays settle it without new training\. Each problem is drawnk=4k\{=\}4times, so for a problem carrying any credit we observe the fraction solved,t=a\+bt=\\mathrm\{a\}\+\\mathrm\{b\}, and the convention preference,s=b/ts=\\mathrm\{b\}/t; under the independence clause𝔼\[s∣t\]\\mathbb\{E\}\[s\\mid t\]is constant intt\. It is not\. Pooling the three seeds of the interleaved arm, problems the model solves on all four draws are answered underBBwith probability0\.44780\.4478, and problems it solves on some but not all draws with probability0\.34910\.3491: a difference of0\.09880\.0988,5\.175\.17contrast floors, Mann–Whitneyz=3\.56z=3\.56, Spearmanρs=\+0\.21\\rho\_\{s\}=\+0\.21between solve rate and share\. The same sign and a comparable size appear onpurerand\(0\.10810\.1081,z=3\.67z=3\.67\) and onL=27L\{=\}27\(0\.06570\.0657,z=2\.57z=2\.57\)\.*Problems the model half\-solves drift toward the pretrained convention\.*
*Which independence this refutes, because the two are easy to conflate\.*The statistic above is computed*across*problems; the clause it refutes is therefore the across\-problem one,Cov\(p,s\)=0\\mathrm\{Cov\}\(p,s\)=0, and that is exactly the clause the decomposition needs, sinceS=𝔼\[ps\]/𝔼\[p\]S=\\mathbb\{E\}\[p\\,s\]/\\mathbb\{E\}\[p\]equals𝔼\[s\]\\mathbb\{E\}\[s\]if and only if that covariance vanishes\. What it does*not*refute is per\-problem conditional independence: a model whose convention choice is independent of solving on every single problem will still show this correlation whenever the problems differ from one another in both quantities, which they do\. So Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)’s antecedent should be read as the covariance condition and not as a statement about any individual problem, and Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)states it that way\.
Two consequences, and we separate them because they are not equally severe\. Capability conservation is unaffected:C=accA\+accBC=\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}is measured, not derived from the independence clause, so every conservation number in this paper stands as reported\. What does not stand is readingSSas a convention policy that a capability change cannot touch\.SSis an allocation*averaged over a problem population whose difficulty and convention preference are correlated*, so a manipulation that changes which problems are solved will moveSSa little even with the policy fixed\. The effect is bounded by what we measured: across the solved/half\-solved split the share moves0\.0990\.099, against the0\.46210\.4621the path family spans and the0\.42010\.4201that separatesblockedfrom the top of the interior\. It is a real coupling, it is an order of magnitude below the effects the paper reports, and it is the reason §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)’s claim is stated as conservation ofCCrather than as independence of the two coordinates\. Source:review\_statistics\.
###### Proposition 1\(The infinite\-division limit\)\.
AlongπL\\pi\_\{L\}, letθ\\thetaevolve by gradient descent on the per\-example losses\. Asη→0\\eta\\to 0withLη→0L\\eta\\to 0andTηT\\etafixed, the trajectory converges uniformly on compacts to the solution of the mean\-field flow
θ˙=−∇\[nAnA\+nBLA\+nBnA\+nBLB\],\\dot\{\\theta\}=\-\\nabla\\Bigl\[\\tfrac\{n\_\{A\}\}\{n\_\{A\}\+n\_\{B\}\}L\_\{A\}\+\\tfrac\{n\_\{B\}\}\{n\_\{A\}\+n\_\{B\}\}L\_\{B\}\\Bigr\],which depends onπ\\pionly through the source proportions\. If the losses are cross\-entropy and a contested inputxxcarriespA\(⋅∣x\)p\_\{A\}\(\\cdot\\mid x\)andpB\(⋅∣x\)p\_\{B\}\(\\cdot\\mid x\), the flow’s minimiser atxxis the mixtureq⋆=nAnA\+nBpA\+nBnA\+nBpBq^\{\\star\}=\\tfrac\{n\_\{A\}\}\{n\_\{A\}\+n\_\{B\}\}p\_\{A\}\+\\tfrac\{n\_\{B\}\}\{n\_\{A\}\+n\_\{B\}\}p\_\{B\}\.
###### Proof sketch\.
The first statement is Bogoliubov–Krylov averaging\[[3](https://arxiv.org/html/2610.00234#bib.bib50)\]applied to a piecewise\-constant vector field whose period2Lη2L\\etatends to zero: the trajectory tracks the period\-average of the field, which is the proportion\-weighted gradient\. The second is the first\-order condition forminqαCE\(pA,q\)\+\(1−α\)CE\(pB,q\)\\min\_\{q\}\\alpha\\,\\mathrm\{CE\}\(p\_\{A\},q\)\+\(1\-\\alpha\)\\,\\mathrm\{CE\}\(p\_\{B\},q\)withα=nA/\(nA\+nB\)\\alpha=n\_\{A\}/\(n\_\{A\}\+n\_\{B\}\), whose solution isαpA\+\(1−α\)pB\\alpha p\_\{A\}\+\(1\-\\alpha\)p\_\{B\}\. That second statement is about the*unconstrained*minimiser over conditional distributions; identifying it with the flow’s stationary point in parameter space additionally requires the family to be rich enough to representq⋆q^\{\\star\}at the contested inputs, which is the assumption Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)argues is comfortably met at77B and which we state here rather than leave to the prose\. ∎
###### Corollary 1\(Two mechanisms that could not have worked\)\.
Under Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1), any intervention that leaves the period\-average of the field unchanged leaves the endpoint unchanged in the limit\. Batch composition at fixed block structure and any time\-symmetric reordering of a block are two such interventions\. Both were pre\-registered as mechanisms and both are dead \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\); the averaging theorem predicted them dead before either ran\.
This is a remark rather than a proposition because its argument is one line and is not deep:π\\piorders the training stream and does not enter the conditional distribution of the convention given a held\-outII, so the Bayes\-optimal predictor is the same for everyπ\\pi, and anyπ\\pi\-dependence of the endpoint must come from failure to reach that optimum\. What it does is locate where an arrangement effect is allowed to live:*only in the escape hatch*, the distance from the optimum\. That is not a throwaway, because the escape hatch is where all three of this paper’s positive findings turn out to sit, but the content is in the escape hatch and not in the claim about the optimum\.
*What Remark[1](https://arxiv.org/html/2610.00234#Thmremark1)does and does not require\.*It requires the*optimum*to beπ\\pi\-independent, not a finite run’s*endpoint*, because a finite run sits some distance from that optimum and that distance is where every positive finding in this paper lives\. It therefore does not make the block\-length sweep’s flatness*required*rather than observed, and Table[5](https://arxiv.org/html/2610.00234#S4.T5)settles it: the same ladder at a constant rate spans11\.6311\.63floors\. What the remark licenses is the narrower statement that any ladder effect must be an escape\-hatch effect, a failure to reach the mixture\. How wide that hatch is, only Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)says: the width is how much endpoint weight any moment of the run can carry, a decaying schedule can never put a large step size and an uncontracted remainder at the same moment, and a constant one does exactly that at the last step\. A decaying schedule therefore narrows the hatch and the ladder flattens; a constant one leaves it open and the ladder resolves\. That is also what makes §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)’s palindromic failure a confirmation rather than a curiosity: a scheme designed against the discretisation cannot help, because the discretisation is not what is costing anything\.
###### Proposition 2\(The schedule sets the reach, the arrangement sets the period\)\.
Letθ\\thetafollowθ˙=−η\(t\)∇Lπ\(t\)\(θ\)\\dot\{\\theta\}=\-\\eta\(t\)\\nabla L\_\{\\pi\(t\)\}\(\\theta\)on\[0,T\]\[0,T\], and letθ¯\\bar\{\\theta\}follow the same flow with∇Lπ\(t\)\\nabla L\_\{\\pi\(t\)\}replaced by the proportion\-weighted meang¯=∇\[αLA\+\(1−α\)LB\]\\bar\{g\}=\\nabla\[\\alpha L\_\{A\}\+\(1\-\\alpha\)L\_\{B\}\],α=nA/\(nA\+nB\)\\alpha=n\_\{A\}/\(n\_\{A\}\+n\_\{B\}\)\. WriteΔ=∇LB−∇LA\\Delta=\\nabla L\_\{B\}\-\\nabla L\_\{A\}and letsπs\_\{\\pi\}be the square wave equal to1−α1\-\\alphaonAA\-blocks and−α\-\\alphaonBB\-blocks, so thatg¯−∇Lπ\(t\)=sπ\(t\)Δ\\bar\{g\}\-\\nabla L\_\{\\pi\(t\)\}=s\_\{\\pi\}\(t\)\\Deltaandsπs\_\{\\pi\}has zero mean over a periodτL\\tau\_\{L\}\. Define the*endpoint weight*of a deviation at timett,
w\(t\)=Φ\(T,t\)η\(t\)Δ\(θ¯\(t\)\),w\(t\)\\;=\\;\\Phi\(T,t\)\\,\\eta\(t\)\\,\\Delta\\bigl\(\\bar\{\\theta\}\(t\)\\bigr\),withΦ\(T,t\)\\Phi\(T,t\)the state\-transition operator of the mean flow linearised aboutθ¯\\bar\{\\theta\}\. Then to first order in the displacement, writingSπ\(t\)=∫0tsπS\_\{\\pi\}\(t\)=\\int\_\{0\}^\{t\}s\_\{\\pi\},
θ\(T\)−θ¯\(T\)=Sπ\(T\)η\(T\)Δ\(θ¯\(T\)\)⏟terminal−∫0TSπ\(t\)w˙\(t\)dt⏟interior,\\theta\(T\)\-\\bar\{\\theta\}\(T\)\\;=\\;\\underbrace\{S\_\{\\pi\}\(T\)\\,\\eta\(T\)\\,\\Delta\\bigl\(\\bar\{\\theta\}\(T\)\\bigr\)\}\_\{\\text\{terminal\}\}\\;\-\\;\\underbrace\{\\int\_\{0\}^\{T\}S\_\{\\pi\}\(t\)\\,\\dot\{w\}\(t\)\\,\\mathrm\{d\}t\}\_\{\\text\{interior\}\},\(2\)and when the path runs to a whole number of periods the terminal term vanishes and
‖θ\(T\)−θ¯\(T\)‖≤2α\(1−α\)⋅τL⏟arrangement⋅supt‖w\(t\)‖⏟schedule\.\\bigl\\\|\\theta\(T\)\-\\bar\{\\theta\}\(T\)\\bigr\\\|\\;\\leq\\;2\\,\\alpha\(1\-\\alpha\)\\;\\cdot\\;\\underbrace\{\\tau\_\{L\}\}\_\{\\text\{arrangement\}\}\\;\\cdot\\;\\underbrace\{\\textstyle\\sup\_\{t\}\\\|w\(t\)\\\|\}\_\{\\text\{schedule\}\}\.\(3\)
###### Proof sketch\.
Subtract the two flows and linearise:u=θ−θ¯u=\\theta\-\\bar\{\\theta\}obeysu˙=−η\(H¯−sπ∇Δ\)u\+ηsπΔ\+O\(∥u∥2\)\\dot\{u\}=\-\\eta\(\\bar\{H\}\-s\_\{\\pi\}\\nabla\\Delta\)u\+\\eta s\_\{\\pi\}\\Delta\+O\(\\\|u\\\|^\{2\}\)withH¯=∇2L¯\(θ¯\)\\bar\{H\}=\\nabla^\{2\}\\bar\{L\}\(\\bar\{\\theta\}\)\. The term in∇Δ\\nabla\\Deltais the instantaneous Hessian’s departure from the averaged one; it is not smaller than the term we keep, so it has to be disposed of rather than dropped\. At leading orderuuoscillates asSπηΔS\_\{\\pi\}\\,\\eta\\Delta, and⟨sπSπ⟩=τL−1∫ddt\(Sπ2/2\)=0\\langle s\_\{\\pi\}S\_\{\\pi\}\\rangle=\\tau\_\{L\}^\{\-1\}\\\!\\int\\tfrac\{\\mathrm\{d\}\}\{\\mathrm\{d\}t\}\(S\_\{\\pi\}^\{2\}/2\)=0over a whole period becauseSπS\_\{\\pi\}is periodic, so it first contributes atO\(τL2\)O\(\\tau\_\{L\}^\{2\}\)\. Variation of constants on what remains givesu\(T\)=∫0Tsπwu\(T\)=\\int\_\{0\}^\{T\}s\_\{\\pi\}w; integrating by parts withSπ\(0\)=0S\_\{\\pi\}\(0\)=0gives Eq\. \([2](https://arxiv.org/html/2610.00234#S3.E2)\)\. For the bound,SπS\_\{\\pi\}is the triangle wave ofsπs\_\{\\pi\}, sosupt\|Sπ\|=α\(1−α\)τL\\sup\_\{t\}\|S\_\{\\pi\}\|=\\alpha\(1\-\\alpha\)\\tau\_\{L\}, andwwrises and falls at most once under a monotone or single\-peaked schedule, so∫‖w˙‖≤2supt‖w‖\\int\\\|\\dot\{w\}\\\|\\leq 2\\sup\_\{t\}\\\|w\\\|\. ∎
Equation \([3](https://arxiv.org/html/2610.00234#S3.E3)\) separates the two factors, and*the separation, not either factor’s value, is the content*: the arrangement enters only through the periodτL\\tau\_\{L\}and the schedule only through how much endpoint weight any moment of the run can carry\. Three things follow, and we state the fourth thing that does*not*\.
*The interior scales with block length\.*τL∝L\\tau\_\{L\}\\propto L, so a ladder’s interior is linear inLLand vanishes asτL→0\\tau\_\{L\}\\to 0; Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)is that corner\. A ladder should therefore be*monotone inLL*, which a limit theorem alone does not predict and which the constant\-rate ladder measures \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\)\.
*The corner is not a small\-τL\\tau\_\{L\}object at all, which is why it does not move\.*Atblockedthe path is one period,τL=T\\tau\_\{L\}=T: the bound buys nothing,SπS\_\{\\pi\}is a single triangle peaking att=αTt=\\alpha Trather than a fast oscillation, and the displacement is set by the whole profile ofwwinstead of by any part of it a schedule can delete\.blockeddiffers by1\.01\.0floor between the two schedule families against17\.017\.0floors atL=54L\{=\}54\(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\)\.
*A decaying schedule has a strictly shorter reach than a constant one at the same peak rate\.*Whereη\\etais large a decaying schedule still has the rest of the run to contract through, and whereΦ≈I\\Phi\\approx Iitsη\\etahas already decayed, so it cannot have both at once; a constant schedule has both att=Tt=T, wherew\(T\)=η\(T\)Δw\(T\)=\\eta\(T\)\\Deltacarries no contraction discount at all\. Hencesupt‖w‖\\sup\_\{t\}\\\|w\\\|is attained at the endpoint and equalsη\(T\)‖Δ‖\\eta\(T\)\\\|\\Delta\\\|for a constant rate, and is strictly belowηmax‖Δ‖\\eta\_\{\\max\}\\\|\\Delta\\\|for any schedule decayed to zero, by a margin that widens with the contraction the run undergoes\.*This is the sense in which the schedule is the averaging operator*, and it is what Table[5](https://arxiv.org/html/2610.00234#S4.T5)measures at fixedτL\\tau\_\{L\}:2\.202\.20floors of interior against11\.6311\.63whenη\(T\)\\eta\(T\)is the only thing changed\.
*What the proposition does not do is predict the size of that gap\.*The reach depends onΦ\\Phias well as onη\\eta, the two families do not share aΦ\\Phi, and a bound is not an estimate: a different functional of the same profile, its total variation, orders the two schedules the other way when the run contracts little\. So the direction is derived and the magnitude is measured, and we say which is which rather than let one borrow the other’s authority\. The proposition is a first\-order statement about a linearisation and we use it for signs, orderings and which factor a knob enters, which is the standing it has in the averaging literature it comes from\[[3](https://arxiv.org/html/2610.00234#bib.bib50),[64](https://arxiv.org/html/2610.00234#bib.bib79)\]\.
*One prediction it makes that this paper has not tested\.*The terminal term in Eq\. \([2](https://arxiv.org/html/2610.00234#S3.E2)\) is the stopping phase of Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3), and it is multiplied byη\(T\)\\eta\(T\)with no contraction discount\. A decaying schedule should therefore suppress the mirror effect of §[5](https://arxiv.org/html/2610.00234#S5)in the same way it suppresses the ladder, and our mirror arms are all cosine\. That is one arm, it is listed with the others in §[6](https://arxiv.org/html/2610.00234#S6), and until it is run the terminal term is the one branch of Eq\. \([2](https://arxiv.org/html/2610.00234#S3.E2)\) we have no schedule contrast for\.
###### Proposition 3\(The stopping phase\)\.
TreatπL\\pi\_\{L\}as a periodically driven system whose state oscillates about the mean\-field point with amplitude increasing inLL\. Then \(i\) the blocked arm isπL\\pi\_\{L\}at maximalLL, stopped at maximal displacement rather than a distinct mechanism; and \(ii\) reversing which source occupies the final block relocates the endpoint to the opposite side of the cycle, moving\(accA,accB\)\(\\mathrm\{acc\}\_\{A\},\\mathrm\{acc\}\_\{B\}\)antisymmetrically while leavingCCfixed to first order\.
Part \(ii\) is a prediction with a sign and a symmetry, and §[5](https://arxiv.org/html/2610.00234#S5)reports the measurement that was frozen against it\. We label the limit\-cycle picture*post hoc*: it was formed after the block\-length sweep and before the mirror test, and only the mirror test is evidence for it\.
## 4The Instrument, and What the Path Writes
### 4\.1The Instrument
#### Corpora\.
Two conflict constructions over competition\-mathematics problems with integer answers drawn from the standard benchmarks\[[65](https://arxiv.org/html/2610.00234#bib.bib55),[66](https://arxiv.org/html/2610.00234#bib.bib23)\], each split into halves of matched size with byte\-identical prompts:*synthetic*\(cf2: one half’s boxed answers shifted\+1\+1\(contradictory by construction,864864rows per condition\) and*natural*\(nat: numeral against spelled\-out answers\) both correct,720720rows per condition\), with same\-convention controls for each\. The ratio corpora \(rr\) rebuild the natural conflict at936936rows with mixture weights1:11\{:\}1and2:12\{:\}1\. Corpora, splits and row counts are asserted by fixed\-seed build scripts before any training, and Figure[4](https://arxiv.org/html/2610.00234#S4.F4)draws thenatconstruction end to end\.*The natural construction is built, and the rate it is built at is borrowed\.*The survey\[[67](https://arxiv.org/html/2610.00234#bib.bib66)\]puts form disagreement at12\.7%12\.7\\%of shared problems on the corpus pair with the most multiplicity, rising to27\.1%27\.1\\%against externally authored answer keys; both figures are measured there and neither is re\-derived here, so a reader who wants them checked has that paper and not this one\. What they license is the choice of construction rather than any number below: they say the contested case is common enough to be worth an instrument, and every arm we report is built rather than found\.
integer\-answer rows,every skill pooledhalfAA,720720rowsnumeralshalfBB,720720rowsspelled outthe data pathcheckpoint200200held\-out problems,the*same*set for every armaccA\\mathrm\{acc\}\_\{A\}:numeralsaccB\\mathrm\{acc\}\_\{B\}:spelleddisjoint:no shared problemmutually exclusive:a sample counts onceFigure 4:The instrument: one corpus that disagrees with itself\.The two*training*halves are disjoint in problems and no single row disagrees with itself: the disagreement is a property of the corpus\. The*evaluation*is the opposite, one held\-out set scored twice against mutually exclusive conventions, so a sample can be credited to at most one\. That is what makes the sum a probability of solving and the ratio a policy\. Drawn fornat;cf2is the same construction at864864rows per half\.
#### Arms and the ladder\.
Per condition:onlyfor each half, both blocked orders, the batch\-mixed shuffle, pure batches in random order \(purerand\), and alternating same\-source blocks of exactlyLLconsecutive optimiser steps,L∈\{1,3,6,9,18,27,54\}L\\in\\\{1,3,6,9,18,27,54\\\}, trainer shuffling off so the file order is the path \(Figure[3](https://arxiv.org/html/2610.00234#S1.F3)\)\. Training is full\-parameter SFT on Qwen2\.5\-7B\[[68](https://arxiv.org/html/2610.00234#bib.bib24)\]\(AdamW,β1=0\.9\\beta\_\{1\}\{=\}0\.9,β2=0\.999\\beta\_\{2\}\{=\}0\.999\), three epochs at lr3×10−53\{\\times\}10^\{\-5\}\(the budget at which these corpora are learnable\), three seeds on every ladder rung and eight on theonly,shufandblockedarms and on the keyed corpus, run withms\-swift\[[69](https://arxiv.org/html/2610.00234#bib.bib37)\]and evaluated with vLLM\[[70](https://arxiv.org/html/2610.00234#bib.bib25)\]; the conflict switch itself is additionally measured at33B and1414B by the companion paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\], which is a pointer and not evidence a reader of this paper can check here \(§[6](https://arxiv.org/html/2610.00234#S6)\)\.
accA\\mathrm\{acc\}\_\{A\}accB\\mathrm\{acc\}\_\{B\}one armcapabilityaccA\+accB\\mathrm\{acc\}\_\{A\}\{\+\}\\mathrm\{acc\}\_\{B\}constant along each dashed lineallocationaccB/\(accA\+accB\)\\mathrm\{acc\}\_\{B\}/\(\\mathrm\{acc\}\_\{A\}\{\+\}\\mathrm\{acc\}\_\{B\}\)constant along each dotted rayFigure 5:The one thing that reads wrongly in prose\.The read\-out of Figure[4](https://arxiv.org/html/2610.00234#S4.F4)as a change of coordinates: capability is constant along the dashed anti\-diagonals, allocation along the dotted rays\. A movement*along a ray*therefore changes capability at fixed allocation and a movement*along a diagonal*does the reverse, and an undecomposed accuracy cannot tell the two apart although their meanings are opposite\.
#### WhataccA\\mathrm\{acc\}\_\{A\}andaccB\\mathrm\{acc\}\_\{B\}score\.
Two properties of the construction are easy to mis\-read and both are load\-bearing\. First, the two training halves are*disjoint in problems*: no training row disagrees with itself, and the disagreement is a property of the corpus rather than of any example\. What is contested is the convention a*held\-out*problem should be answered under, which is not a function of anything the model may condition on, which is the condition Remark[1](https://arxiv.org/html/2610.00234#Thmremark1)needs; the disjointness creates the underdetermination rather than weakening it\. Second, on the synthetic corpuscf2the two conventions are a boxed answer and that answer shifted by one, soaccA\\mathrm\{acc\}\_\{A\}andaccB\\mathrm\{acc\}\_\{B\}measure*conformance to a stipulated convention*rather than correctness against mathematical ground truth\. Every arm is scored against both, and Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx4)reports what happens when only one is installed\. The phrase “neither is wrong” is exact onnatand is a statement about the*scoring rule*oncf2, which carries the flagship arms because its two conventions are the ones the scorer separates cleanly\.
Table 1:Two corpora, and which one carries which claim\.*The philosophical claim of this paper is a claim about*nat*, and*nat*is the weaker instrument*; we print the pair rather than leave it to be assembled from four sections\. What travels between them is every*qualitative*statement the paper defends: capability conserved, allocation moved, interior flat, mirror antisymmetric at fixedCC\. What does not travel is size, and the direction is against us, so the headline figures in the abstract arecf2figures and a reader rebuilding this on a both\-correct conflict should expect two thirds the capability, and, once the step size is matched \(§[7](https://arxiv.org/html/2610.00234#S7), branchNI\-1\), about seven eighths the install\. The last block is the armsnatdoes not have; they are not claimed for it\. §[7](https://arxiv.org/html/2610.00234#S7)gives the full accounting, including the one asymmetry the construction itself creates and the budget gate that measures how much of the size gap is the step size\. Source:review\_statistics,nat\_replication\.cf2*\(flagship\)*natstatus*what the two conventions are*constructionboxed answer vs\. that\+1\+1numeral vs\. spelled out—is either wrong?yes, by constructionneitherclaim is aboutnatrows per condition864864720720—learning rate3×10−53\{\\times\}10^\{\-5\}1×10−51\{\\times\}10^\{\-5\}natran colder*how strong an instrument each is*minority install \(BB\-only\)0\.42250\.42250\.20910\.2091*unmatched rates*at lr3×10−53\{\\times\}10^\{\-5\}0\.42250\.42250\.3738\\mathbf\{0\.3738\}NI\-1:7/87/8capability, interleaved arm0\.52020\.52020\.37270\.3727two thirds, unmatchedminority form leaked untrained0\.01690\.01690\.0000\\mathbf\{0\.0000\}*different regimes**what each corpus is asked to carry*capability conservedyesyesagreesallocation movesyesyesagreesladder interior flat2\.202\.20floors \(1010arms\)1\.921\.92floors \(44rungs\)agreesmirror antisymmetric\+0\.0850/−0\.0875\+0\.0850/\{\-\}0\.0875\+0\.0183/−0\.0229\+0\.0183/\{\-\}0\.0229sign agreesmirror leavesCCfixed0\.00250\.00250\.00460\.0046agrees*what only*cf2*carries*the12\.29σ12\.29\\sigmaswitch✓not measured*cf2 only*schedule contrast,11\.6311\.63floors✓not measured*cf2 only*the key \(87\.5%87\.5\\%of ceiling\)✓not measured*cf2 only*
#### Observables\.
Because the two evaluation sets are the same problems under mutually exclusive conventions, every checkpoint yields a two\-component read\-out \(Figure[5](https://arxiv.org/html/2610.00234#S4.F5)\):
accA\+accB⏟capability:P\(solved at all\)andaccBaccA\+accB⏟allocation:P\(conventionB∣solved\)\.\\underbrace\{\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}\}\_\{\\text\{capability: \}P\(\\text\{solved at all\}\)\}\\qquad\\text\{and\}\\qquad\\underbrace\{\\frac\{\\mathrm\{acc\}\_\{B\}\}\{\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}\}\}\_\{\\text\{allocation: \}P\(\\text\{convention \}B\\mid\\text\{solved\}\)\}\.\(4\)Arrangement claims are claims about which component moves\. Evaluation isk=4k\{=\}4samples on200200held\-out problems, exact match on the boxed answer;
*Those200200are a prefix of a larger pool, and the prefix is not skill\-balanced\.*The held\-out file carries292292problems onnatand342342oncf2, sorted by skill, and the evaluation takes the first200200\. Every arm sees the identical problems, which is what the contrasts require, but the set is not the pooled corpus: algebra and geometry are complete, combinatorics is truncated \(97→4897\\to 48onnat,129→33129\\to 33oncf2\), and*calculus and physics are absent entirely*\. Every absolute capability level in this paper is therefore measured on three skills rather than five, and the difficulty coupling of Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)on the same restricted set\. Nothing here affects a between\-arm comparison and everything here affects a level\. It is reported rather than repaired because the checkpoints for most of these arms were reclaimed after evaluation, so the pool cannot be re\-scored \(§[6](https://arxiv.org/html/2610.00234#S6)\)\.
The single\-run noise floor isσ^=0\.0234\\hat\{\\sigma\}=0\.0234and the three\-seed contrast floor0\.01910\.0191, inherited with the full protocol from the companion result base; reliability of each contrast is reported as a*contrast reliability*\(an intraclass correlation over repeated measurements of the same contrast,[71](https://arxiv.org/html/2610.00234#bib.bib36)\), and the arrangement contrast’s is0\.1970\.197, which is why every claim below is stated on the decomposition of Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)rather than on a raw difference\. Every number below is written by a named analyzer to a verdict file rather than transcribed by hand, and the deciding tests’ read\-out rules were frozen in the job scripts before the runs; where a test was*redesigned*after a diagnosis \(one was: §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\), the redesign is labelled\. Registering in the script the scheduler executes is checkable rather than promised, so we make it checkable: a registration index released with the paper \(Appendix[D](https://arxiv.org/html/2610.00234#A4)\) lists every registration against the commit that froze each threshold and the timestamp of the result it decided, with the interval measured in each case: twelve rows, ten running from under three hours to nearly three days ahead of their result and two that do not, each of the two stated in the index rather than averaged away\. The palindromic schedule of §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)is one of the two, and that paragraph says so\.
#### The instrument’s first failure, and what it diagnosed\.
Every arrangement number in this paper depends on the conflict corpus actually creating a conflict the training run can feel, so the positive control \(conflict against its own same\-convention twin\) is the load\-bearing check, and on the first build it came out*backwards*\. Drawing the integer\-answer rows from algebra alone gave238238per condition at one epoch, and the read\-out is Table[2](https://arxiv.org/html/2610.00234#S4.T2)\.
Table 2:The instrument, across its three builds\.DDis the arrangement contrast on the last\-seen convention,*signed*:D=acc\(last seen\)blocked−acc\(last seen\)shufD=\\mathrm\{acc\}\(\\text\{last seen\}\)\_\{\\textsc\{blocked\}\}\-\\mathrm\{acc\}\(\\text\{last seen\}\)\_\{\\textsc\{shuf\}\}with the sources in the published order, so the switch is negative\. Elsewhere the same quantity is quoted as a*span*between two arms and is therefore positive;−0\.2346\-0\.2346here and\+0\.2275\+0\.2275in §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)are the same measurement at eight and at three seeds, not two results with different signs\. The control is the same construction with both halves under one convention, so it must stay inert; the paragraph below reads v1, where it did not\. v1 and v2 are three\-seed builds; v3 and its control carry eight\. Sources: theconflict\_switch,conflict\_switch2andconflict\_switch2b3lr3e5verdicts\.In v1 the arm that was supposed to move did not \(0\.10σ0\.10\\sigma\) and the arm that was supposed to be inert did \(2\.41σ2\.41\\sigma\), the exact inversion of the design\. The conflict arms scored0\.020\.02against the shifted golds while the same checkpoints scored0\.390\.39–0\.500\.50against unshifted ones, so the\+1\+1convention had lost to the pretrained prior andDDwas being measured at the floor of an instrument whose treatment had never been installed\. Pooling the integer\-answer rows of every skill and raising the epochs gave the conflicting convention roughly an order of magnitude more gradient steps, and v2 fired\. We report v1 rather than only the builds that worked because the12\.29σ12\.29\\sigmaof v3 is otherwise unreadable: it is contingent on the conflict being*learnable at the budget*, a property of the corpus and the budget together and not of conflict as such\.
### 4\.2Allocation Moves, Capability Does Not
Table[3](https://arxiv.org/html/2610.00234#S4.T3)is the decomposition of Eq\. \([4](https://arxiv.org/html/2610.00234#S4.E4)\) across every arm of the synthetic conflict\. The capability sum is constant to within9\.7%9\.7\\%of its mean across all twelve arms \(*including*onlyarms that never saw a row of the other convention\) while the allocation share runs from0\.040\.04to0\.870\.87\. The coherent control shows the same sum discipline at1\.011\.01with the share pinned at0\.5000\.500\. The natural conflict reproduces the pattern \(sums0\.330\.33–0\.3850\.385; shares0\.000→0\.473→0\.6910\.000\\to 0\.473\\to 0\.691\), and so do both ratio corpora \(sums0\.3350\.335–0\.3810\.381; Table[6](https://arxiv.org/html/2610.00234#S4.T6)\)\.
Table 3:The conservation decomposition\(synthetic conflict, Qwen2\.5\-7B, three seeds\)\. CapabilityaccA\+accB\\mathrm\{acc\}\_\{A\}\{\+\}\\mathrm\{acc\}\_\{B\}is flat across every arrangement; only the allocation moves\. Source: thecf2b3lr3e5/ctl2b3lr3e5result families \(33epochs, lr3×10−53\{\\times\}10^\{\-5\},n=200n\{=\}200\), named because a second conflict family at a different budget also sits in the result base, and the two must not be pooled\.Seen once, the decomposition is nearly an identity: the two evaluations partition the same problems, so the sum is the probability of solving at all and the share is the conditional convention choice\. We state it that way rather than dressing it as a discovered law: the finding is that*nobody measures this decomposition*, and that it dissolves what the undecomposed metric manufactures\. The arrangement switch that exact match reports at−0\.2346\-0\.2346,12\.29×12\.29\\timesthe noise floor and measured here \(Table[2](https://arxiv.org/html/2610.00234#S4.T2)\), is a movement of the share at fixed sum: under a metric that accepts either convention it is*exactly zero by construction*\. Exact match conflates capability with commitment, and everything recipe\-like about conflict lives in the commitment component\.
*The agnostic rule matters, and there are two\.*Crediting*either*convention makes the switch zero identically, because on a solved problem reallocatingk=4k\{=\}4samples between the two conventions conserves their sum\. Crediting each problem’s*better*convention does not: where a mixed arm splits its samples across conventions within one problem, the per\-problem maximum falls while the sum holds, and the companion paper measures that contrast on this same corpus at4\.654\.65floors\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]\. The two rules disagreeing is not a contradiction; it is Figure[6](https://arxiv.org/html/2610.00234#S4.F6)’s mixing read by a rule that penalises mixing, and it is why this paper’s zero is stated for the sum rule and for no other\.
#### The coin\-flip fingerprint\.
The mixture is visible per sample, and Figure[6](https://arxiv.org/html/2610.00234#S4.F6)is the whole of the evidence for calling it a policy rather than a deficit\. For each problem take the fraction of itsk=4k\{=\}4samples that land on whichever convention that problem favours\. Committed arms \(only,blocked\) put𝟐𝟔\\mathbf\{26\}–𝟑𝟎%\\mathbf\{30\\%\}of problems at4/44/4and4141–46%46\\%in the intermediate bins; every mixed arm \(shuf,purerand,L=27L\{=\}27\) puts𝟕\\mathbf\{7\}–𝟗%\\mathbf\{9\\%\}at4/44/4and6464–68%68\\%in the middle: the same problem answered under different conventions across four samples\. The unsolved mass is the same in both,2525–30%30\\%, which is the point: conflict has not made the model unable, it has made the*policy stochastic*, which is what the Bayes\-optimal response to contradictory supervision is\.
*One binning throughout\.*The contrast depends on how a problem is binned, and the two natural choices give different sizes: counting each problem’s best convention gives a factor of about three and a half, counting each \(problem, convention\) cell separately about three\. The figures above use the first throughout\. Mixing the two inflates the contrast to about seven, so the definition is stated rather than left to a reader to infer\.
AAonly30%30\\%blocked26%26\\%shuf9%9\\%L=27L\{=\}277%7\\%purerand8%8\\%0\.00\.00\.20\.20\.40\.40\.60\.6fraction of problemswithin each arm the five bars run0/4→4/40/4\\to 4/4: how many of thek=4k\{=\}4samples land on the convention that problem favourscommitted: deterministic, all\-or\-nothingmixed: randomised, mass in the middleFigure 6:Conflict makes the policy stochastic, not the model weak\.Per\-problem commitment atk=4k\{=\}4samples on the synthetic conflict, three seeds,200200problems each\. Bars within an arm run0/40/4to4/44/4\. The committed arms are bimodal \(the model either answers a problem the same way every time or cannot answer it\) while the mixed arms move that mass into the middle without changing the unsolved fraction \(grey, leftmost bar in each group\)\. Same corpus, same volume, same budget; only the data path differs\. Sourcecf2b3lr3e5per\-problem arrays\.
#### Qualitative structure: the same problems, re\-labelled\.
The decomposition says*how much*moves; the per\-problem arrays say*which problems*carry it, and the answer is the same ones\. Crossing fromshuftoblocked\(pooled over three seeds,600600problem instances\), the*set*of solved problems barely changes,7\.3%7\.3\\%newly solved,8\.5%8\.5\\%lost, net−1\.2%\-1\.2\\%, but among the386386instances solved in*both*arms,48\.2%48\.2\\%change which convention they favour, and the change is one\-directional:179179flipA→BA\{\\to\}Bagainst77flippingB→AB\{\\to\}A, a26×26\\timesasymmetry pointing at the source the blocked arm ends on \(Table[4](https://arxiv.org/html/2610.00234#S4.T4)\)\. Blocking does not teach different problems and does not unteach the old ones; it re\-labels the problems the model already solves\. That is what “a policy, not damage” looks like at the level of individual items, and it is the per\-problem face of the antisymmetric mirror shift of §[5](https://arxiv.org/html/2610.00234#S5)\.
Table 4:Per\-problem transitions fromshuftoblocked\(synthetic conflict, three seeds×\\times200200problems\)\. A problem’s state is the convention receiving more of itsk=4k\{=\}4samples \(AA,BB, tie\) or unsolved \(UU\)\. The solved set is nearly fixed while nearly half of it changes label,26×26\\timesmore often toward the convention trained last\. The tie row carries the33transitionsU→U\\totie and the77tie→U\\to U, which is why the prose’s newly\-solved and lost counts \(4444and5151\) exceed what the fourUUrows sum to \(4141and4444\)\. Sourcecf2b3lr3e5per\-problem arrays\.
#### How exact the conservation is\.
It is an approximation, and three measurements bound it\. The measure comes first, because the numbers are otherwise comparable only by accident: it is the*full range*of the capability sum across arms, divided by its mean\. On that measure the fifteen synthetic arms give9\.97%9\.97\\%\(a range of0\.05210\.0521about a mean of0\.52200\.5220\) and the natural arms15%15\\%\. That no single arm sits more than6\.78%6\.78\\%from the mean is also true and is the statement to ignore, since a conservation claim should be priced by its worst pair rather than its worst arm\. The range is carried by the two extreme arms, so the two dispersion measures that are not are also recorded: across the same fifteen arms the standard deviation is2\.74%2\.74\\%of the mean and the interquartile range2\.99%2\.99\\%\.*Fifteen arms and not fourteen: the census was missingB\_then\_A, the reverse of an ordering it already contained*, so it was asymmetric in exactly the variable this section is about; restoring it leaves the range where the two extreme arms had it and moves the interquartile range from2\.51%2\.51\\%to2\.99%2\.99\\%\. Third,*along*a single training trajectory the sum moves through a range of0\.14750\.1475before settling: endpoint conservation is not path conservation\. The sharpest statement we defend is §[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)’s: the key recovers74\.9%74\.9\\%of the gap to the additive ceiling, so “conserved” means*to within an eighth*, with a mechanism for the remainder\.
0\.00\.00\.20\.20\.40\.40\.60\.60\.80\.8allocationaccB/TOTAL\\;\\mathrm\{acc\}\_\{B\}/\\mathrm\{TOTAL\}sweeps0\.029→0\.7180\.029\\to 0\.718\(range0\.7080\.708\)4\.5%4\.5\\%of that range is spent by step440\.300\.300\.350\.350\.400\.400\.450\.450\.500\.5097\.4%97\.4\\%of capability’s whole excursionis spent in the first four stepscapabilityaccA\+accB\\;\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}range0\.1475=5\.5×0\.1475=5\.5\\timesits floor; band is±\\pmfloor00441212242436365454optimiser step along a singleL=54L\{=\}54run \(three seeds averaged\)Figure 7:Capability and allocation move on separate clocks, which the range alone hides\.Theblockedconstruction’s second stage: a model trained to completion onAAand then trained onBB, checkpointed every four optimiser steps over that stage’s first5454, three seeds averaged\. It is not a ladder rung, and its allocation ends at0\.7180\.718heading towardblocked’s0\.8690\.869rather than at theL=54L\{=\}54rung’s0\.4920\.492;\(top\)allocationaccB/TOTAL\\mathrm\{acc\}\_\{B\}/\\mathrm\{TOTAL\},\(bottom\)capabilityaccA\+accB\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}with a band of±\\pmits floor0\.02700\.0270around the endpoint\. The shaded column is the first four steps\. Capability spends97\.4%97\.4\\%of its total excursion there, while allocation spends4\.5%4\.5\\%of its, so the capability dip precedes the commitment shift rather than accompanying it\. The verdict file’s audit fields record that both evaluations use the same200200questions with disjoint golds\.That third bound is the one worth looking at rather than quoting, because its*shape*says something the range does not \(Figure[7](https://arxiv.org/html/2610.00234#S4.F7)\)\. Sampled every four optimiser steps through the first5454steps ofblocked’sBBstage, resuming from the completedAA\-onlycheckpoint, the two quantities move on separate clocks\. Capability falls from0\.49000\.4900to0\.34630\.3463between steps00and44:97\.4%97\.4\\%of its entire excursion, and then climbs back to0\.41910\.4191and stays\. Allocation over the same four steps goes0\.0290\.029to0\.0610\.061:4\.5%4\.5\\%of the range it eventually covers, essentially nothing\. The sweep from0\.030\.03to0\.720\.72happens later, between steps1212and3232, by which time capability has already returned to within a floor and a half of where it ends\.
Two consequences follow, the first against us\.*The dip is not the reassignment\.*It is almost entirely spent before allocation begins to move, so whatever costs capability at the start of conflicted training is a different process from the commitment shift this paper is about, and the endpoint conservation we report is the state after that process has largely undone itself\. Second, the dip is not a scoring artefact of the mixture: atk=4k\{=\}4soft scoring a model that randomised between the two conventions would leaveaccA\+accB\\mathrm\{acc\}\_\{A\}\+\\mathrm\{acc\}\_\{B\}untouched, since each problem contributes to one gold or the other\. A drop in the sum is answers matching*neither*, so conflict does more than reassign, at least transiently\. What we cannot say is which: a trajectory that recorded only accuracy cannot separate knowledge briefly lost from output the scorer briefly could not parse\. That separation needs a run that keeps generations, and it is listed as a missing cell rather than argued\.
### 4\.3The Ladder: What the Path Moves, and What This Design Can See
0\.30\.30\.50\.50\.70\.70\.90\.9the averaging wall: every arrangementhere writes the same time\-averagethe stopping phase:the one non\-average termallocationSS0\.450\.450\.500\.500\.550\.550\.600\.60capabilityCC: flat across every arm here, span0\.01830\.0183inside the0\.01910\.0191floorcapabilityCC11336699181827275454shufpurerandblockedblock lengthLL\(optimiser steps, log scale\)Figure 8:The averaging wall, measured\.Same corpus, volume, budget and seeds; only the data path differs \(Definition[1](https://arxiv.org/html/2610.00234#Thmdefinition1)\)\.\(a\)Allocation is flat fromL=1L\{=\}1toL=27L\{=\}27and acrossshufandpurerand:0\.4070\.407to0\.4490\.449over a27×27\\timeschange in block length and a change of batch composition\.L=54L\{=\}54is the first rung off that floor at0\.4920\.492andblockedsits at0\.8560\.856\.\(b\)Capability over the identical arms spans0\.01830\.0183, inside the contrast floor: the path relocates commitment without changing what can be solved\. Bars are the seed standard deviation, three seeds per point \(shufandblockedcarry eight\)\. Sourcecf2b3lr3e5\.*Two tops, two gaps, stated where the paper first needs both\.*The occupied group runs toL=54L\{=\}54at0\.49220\.4922; the*interior*of the ladder, which excludesL=54L\{=\}54because that rung is the stopping phase becoming visible, runs toL=1L\{=\}1at0\.44910\.4491\. Soblocked’s distance is0\.37700\.3770from the group and0\.42010\.4201from the interior, and both appear below with the top they are measured from named\.
Equation \([1](https://arxiv.org/html/2610.00234#S1.E1)\) makes a strong claim: below corpus scale,*how*the path is arranged is irrelevant, because every arrangement averages to the same field; the only thing the path can hand the parameters is its time\-average\. Four measurements test it, two of them post\-mortems of our own alternatives, and Figure[8](https://arxiv.org/html/2610.00234#S4.F8)is all four of them on one axis: allocation flat across a27×27\\timeschange of block length and across a change of batch composition, capability flat inside its own floor, and a single jump at the blocked end that belongs to the stopping phase rather than to the path\.
#### Batch purity is not the variable \(v1, pre\-registered, dead\)\.
The literature contradicts itself on whether batches should be pure or mixed\[[22](https://arxiv.org/html/2610.00234#bib.bib39),[23](https://arxiv.org/html/2610.00234#bib.bib40),[20](https://arxiv.org/html/2610.00234#bib.bib38)\], and our first pre\-registered mechanism sided with purity: contested directions cancel*inside a mixed batch*, so batch composition should carry the conflict effect\. The deciding cell \(pure batches in*random*order\) lands at the shuffled arm, not the blocked one: control span\+0\.0037\+0\.0037\(noise\), conflict span\+0\.2275\+0\.2275withpurerandat0\.22420\.2242againstshuf0\.21790\.2179andblocked0\.44540\.4454\(three seeds; that span is the switchDDitself, which reads−0\.2346\-0\.2346at eight seeds in §[4\.1](https://arxiv.org/html/2610.00234#S4.SS1)\)\. Purity carries3%3\\%of the span; between\-batch order carries the rest\. v1 is dead\.
#### The optimiser’s memory is not the variable either \(v2, pre\-registered, dead\)\.
One level down, AdamW’s momentum is a1010\-step low\-pass filter \(1/\(1−β1\)1/\(1\-\\beta\_\{1\}\); the second\-moment timescale exceeds the whole run\), so the contested directions could cancel*inside the optimiser state*: the stateful\-optimiser channel of[Sweeney \[24\]](https://arxiv.org/html/2610.00234#bib.bib41)\. The pre\-registered signature was a response turning on atL⋆≈10L^\{\\star\}\\approx 10; a competing basin\-escape account tied it instead to an independently measured residence timeτwrite=13\.6\\tau\_\{\\text\{write\}\}=13\.6steps\. The block\-length sweep kills both: the normalised responseM\(L\)M\(L\)reads\+0\.075\+0\.075,−0\.007\-0\.007,\+0\.037\+0\.037,\+0\.002\+0\.002,\+0\.022\+0\.022,\+0\.002\+0\.002atL=1L=1,33,66,99,1818,2727\(noiseσ\(M\)≈0\.06\\sigma\(M\)\\approx 0\.06\), flat through both predicted onsets and through a quarter of the epoch, with onlyL=54L\{=\}54rising \(0\.2050\.205\) toward the blocked anchor\. The guard passed \(control atL=9L\{=\}9:0\.52120\.5212againstshuf0\.50420\.5042\), so the flatness is a measurement, not a floored instrument\.
#### The same ladder on a second pretraining family, and a threshold set on the wrong statistic\.
The wall above is one model, and unlike the switch it had never been asked to travel\. It is also a*null*, which is the hazard: a floored instrument delivers a flat ladder for free, and §[4\.1](https://arxiv.org/html/2610.00234#S4.SS1)’s v1 is this paper’s own account of what that looks like\. So the replication was registered in stages with the gate frozen first, and the gate is the good news: on Qwen3\-8B\-Base the conflict installs*better*than on Qwen2\.5,accB\(Bonly\)=0\.4813\\mathrm\{acc\}\_\{B\}\(B\\ \\textsc\{only\}\)=0\.4813against0\.42250\.4225, with the switch at15\.015\.0floors against11\.911\.9\. Whatever the ladder says there, it cannot be dismissed as an instrument at its floor\.
*One convention, stated once, because two artefacts here disagree below the floor\.*The allocation of an arm is the mean over seeds of the per\-seed share, which is whatrate\_verdictcomputes and what every number in this paper quotes\. The registered second\-family read\-out formed the share of the seed\-mean accuracies instead, and the two differ by at most0\.00070\.0007, which is0\.040\.04contrast floors: the published flat span reads0\.04160\.0416under the registered estimator and0\.04200\.0420under the reported one,2\.182\.18against2\.202\.20floors\. Where a registered comparison is being quoted we quote the estimator that registration froze and say so; nothing in the paper turns on a difference of a twentieth of a floor, but a reader recomputing from the result base will see both and should not have to infer which\.
The registered branch isW2: the flat region spans0\.05440\.0544, which is2\.852\.85floors against a bar of two\. We report that branch, and in the same breath the thing that makes it uninterpretable as written\.*The published Qwen2\.5 ladder spans0\.04160\.0416, which is2\.182\.18floors, and does not meet the same bar either\.*That number was available before the run and we did not check the threshold against it; the fault is in the registration, not in the second family\. The two spans differ by0\.670\.67floors, less than one, and as a fraction of the excursion the blocked arm makes they are9%9\\%and12%12\\%\. Both of those comparisons are*post hoc*and neither rescues W1: what we can say is that the registered bar was mis\-specified and that the ladder is about as flat on one family as on the other, and what we cannot say is that a bar chosen after the fact was met\.
*The statistic the proposition actually predicts\.*A span in floors is a proxy for what Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)predicts, which is that the interior of the path space collapses to*one endpoint state*\. Definition[3](https://arxiv.org/html/2610.00234#Thmdefinition3)states that quantity directly, predates both second\-family runs, and needs no bar: count the clusters the ten arms occupy at the2σ^2\\hat\{\\sigma\}resolution\. On Qwen2\.5 the answer istwo\(nine arms in\[0\.407,0\.492\]\[0\.407,0\.492\],blockedat0\.8690\.869\) and on Qwen3\-8B\-Base it istwo\(nine arms in\[0\.424,0\.479\]\[0\.424,0\.479\],blockedat0\.8910\.891\)\. Both families train under a cosine, and Table[5](https://arxiv.org/html/2610.00234#S4.T5)is what the same count returns when the schedule is changed instead of the pretraining family: at the same resolution the constant\-rate ladder occupies four states\. The count replicates across models and does not survive across schedules, which is the boundary this subsection draws everywhere else\. The margin is not marginal on either: the largest gap inside the occupied cluster is0\.04310\.0431against a between\-cluster gap of0\.37700\.3770on the first family \(8\.7×8\.7\\times\) and0\.01990\.0199against0\.41210\.4121on the second \(20\.7×20\.7\\times\), unchanged whetherσ^\\hat\{\\sigma\}is the inherited0\.02340\.0234or each ladder’s own recomputed dispersion\.*The wall replicates exactly; the criterion we registered for it did not measure it\.*This reading is labelledpost hocandW2stands as the registered outcome, because a statistic chosen after seeing a miss is worth less than one chosen before it: the registered comparison was of two spans neither of which is resolvable, and the unregistered one is of two integers that agree\. Source:rate\_verdict\.
#### A second ladder, on the other corpus and at a third of the step size\.
The rungs above arecf2at lr3×10−53\\times 10^\{\-5\}\. Anatladder exists at lr1×10−51\\times 10^\{\-5\}over four rungs,L=1,9,15,45L=1,9,15,45, and its interior spans0\.03670\.0367, which is1\.921\.92floors against the published2\.202\.20\. The flatness is therefore not a property of one corpus or one step size\. It is not evidence about the schedule, because a cosine decays to zero at either learning rate\. Source:nat\_replication\.
#### The alternative this paper named, ran, and decided against\.
Every arm above trains under a single cosine, so the second half of every run has little step size left, and an interior that is flat because arrangement does nothing looks exactly like an interior that is flat because nothing much happens after the midpoint\. Those two accounts are separated by one experiment, the same ladder at a constant rate\. It has been run, and it decides for the second\.
Table 5:The same ladder under two learning\-rate schedules\.Ten arms, one corpus, one budget, the identical data path files, and*the same three seeds*: the intersection, computed rather than assumed, because the publishedshufandblockedarms carry eight and a mixed comparison would print most of the eight\-against\-three difference as a schedule effect\. The two families differ inlr\_scheduler\_typeand in nothing else, which we checked against the training arguments and the corpus checksum rather than assuming, and both run to the same step \(324324on the ladder,162162at the corner\)\. Under a cosine the interior spans0\.04200\.0420, which is2\.202\.20contrast floors and below anything this design can resolve\. Under a constant rate the same interior spans0\.2221\\mathbf\{0\.2221\}, which is11\.6311\.63floors, with an intraclass correlation of0\.8360\.836whose bootstrap interval\[0\.315,0\.926\]\[0\.315,0\.926\]lies entirely above zero\.*The corner does not move*:blockedreads0\.86920\.8692and0\.85010\.8501, one contrast floor apart, against the17\.017\.0floors the rung below it moves\. Source:constlr\_ladder\_table,constant\_lr\_verdict\.Three readings, the third of which changes what this section claims\.
*The interior is not flat; it was flat to this instrument, under this schedule\.*The registered read\-out returned branchCL\-2\. The interior spans11\.6311\.63floors against a minimum detectable difference of3\.723\.72, so it is resolved rather than merely wide, and its shape is not noise but the shape Eq\. \([3](https://arxiv.org/html/2610.00234#S3.E3)\) predicts: ordering the nine non\-corner arms by how blocked the arrangement is \(shuf, thenL=1,3,6,9,18,27L\{=\}1,3,6,9,18,27, thenL=54L\{=\}54\), the allocation is monotone inτL\\tau\_\{L\}up to22rank inversions out of3636pairs\.*The averaging wall as a claim about the path is withdrawn\.*
*What replaces it is a claim about the schedule, and it is the stronger statement\.*Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)is not refuted; it is located\. Its conclusion holds in the limit where the tail of training carries no weight, and a decaying schedule is what puts a run in that limit\. The cosine is not a neutral background against which the path was measured\.*It is the averaging operator*, integrating the arrangement away, and the constant\-rate family is the same corpus with the operator removed\. Read that way the two families are not a result and its correction but a measurement of one knob at two settings, and this paper’s own schedule control \(§[6](https://arxiv.org/html/2610.00234#S6)\), which moves the allocation0\.1160\.116by splitting one cosine into two while holding the path byte for byte, is the same knob at a third\.
*The corner is schedule\-free, which is why the paper’s central result does not move\.*blockeddiffers by1\.01\.0floor between the two families against17\.017\.0atL=54L\{=\}54, which is Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)’s corner case measured: atτL=T\\tau\_\{L\}=Tthe displacement is set by the whole profile ofwwrather than by any part of it a schedule can delete\. Every rung below has a tail the discount can act on, and every one of them moves\. The arrangement switch of §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2), the antisymmetry of §[5](https://arxiv.org/html/2610.00234#S5)and the key of §[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)are all read at or against the corner and are untouched\. What the constant\-rate family costs this paper is one sentence about the interior; what it buys is the mechanism that sentence was standing in for\.
*One thing the constant family does less well, stated because it bears on reading the table\.*Capability spans0\.06800\.0680across its ten arms against0\.02460\.0246across the cosine ten on the same three seeds, which is2\.52\.5capability floors rather than inside one\. The share is a ratio and is not mechanically driven by that spread, but the constant family is a noisier place to read a fixed\-capability claim, and §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)’s conservation result is quoted from the cosine family throughout for that reason\.
#### Flatness is a null, so we test it as one\.
A span is a description, not a test\. Three questions have to be answered in order: what can this design see, what does a proper test of the null say, and can a difference in the interior be attributed to allocation at all\. The answers point the same way and they are worth stating separately, because together they say something sharper than any one of them\.
*First, what the design can see\.*Two arms of three seeds give four degrees of freedom, so atα=0\.05\\alpha=0\.05two\-sided the smallest difference detectable with80%80\\%power is3\.723\.72contrast floors at the inheritedσ^=0\.0234\\hat\{\\sigma\}=0\.0234and5\.015\.01at each ladder’s own recomputed0\.03150\.0315; at50%50\\%power the same figures are2\.782\.78and3\.743\.74\. The ladder’s interior spans2\.20\\mathbf\{2\.20\}\.*The interior span is smaller than anything a pairwise comparison in this design could have resolved*, at either dispersion and at either power\. That is a fact about the instrument and it disposes of any reading of the interior’s shape\.
*And it is a fact about the instrument atk=4k\{=\}4, which is a weaker sentence than the one we first wrote\.*Part of the dispersion the paragraph above rests on is generation sampling rather than training stochasticity, and that part falls as1/k1/\\sqrt\{k\}for inference time and no training\.*How large a part is the whole question, and a borrowed number cannot answer it\.*The recipe\-level paper measures0\.01320\.0132of a0\.01420\.0142cross\-seed dispersion to be generation sampling,86\.4%86\.4\\%of the variance, on a different corpus, a different skill set and a different scale\. Nothing in this paper measured whether that transfers\. It does not\.
The conclusion reverses, and the row that reverses it is the one the published table did not have\.Measured on this paper’s own material by the identical closed form, the evaluation share is0\.60740\.6074rather than0\.8640\.864, and at that share*no evaluation budget resolves the interior*: the limit of infinite samples per problem still leaves anMDE80\\mathrm\{MDE\}\_\{80\}of2\.332\.33floors against a2\.202\.20\-floor span\. So a span of this size does not become resolvable atk=32k\{=\}32, and this resolution is not something a reader can buy at the inference rate at all\. The remaining gap is training noise\. It can only be bought in seeds, which is the expensive axis, and this paper does not price it\.
*A second instrument says something worse than disagreement\.*The retained constant\-rate family was re\-scored atk=16k\{=\}16over all eleven arms and three seeds: if86%86\\%of the variance were evaluation sampling, the pooled cross\-seed dispersion should have fallen from0\.02910\.0291to0\.01730\.0173\. It reads0\.03140\.0314, a ratio of1\.079\\mathbf\{1\.079\}where the model predicts0\.5940\.594, and the additive modelσ^\(k\)2=σtrain2\+σeval2⋅\(4/k\)\\hat\{\\sigma\}\(k\)^\{2\}=\\sigma\_\{\\text\{train\}\}^\{2\}\+\\sigma\_\{\\text\{eval\}\}^\{2\}\\cdot\(4/k\)can represent that only with a*negative*evaluation component,−0\.219\-0\.219\.*The two instruments agree that the borrowed number does not transfer and disagree about what replaces it*, one saying0\.610\.61and the other putting the quantity outside the model’s domain\. We take the more conservative reading: even the generous0\.60740\.6074does not resolve the interior\.
*The two instruments stopped disagreeing when the companion audited the harness, and the resolution is a mechanism*\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]\. None of this project’s evaluation drivers passes a seed to the inference engine, so the engine’s default takes effect and every run consumes the same sampling stream on the same prompts\. Evaluation sampling is therefore largely*common\-mode*across seeds: present in full in an absolute accuracy, largely cancelled in a cross\-seed dispersion\. That predicts what the second instrument measured, since raisingkkshould then leave the cross\-seed dispersion roughly unchanged, and1\.0791\.079is roughly unchanged\. It also says the first instrument’s0\.60740\.6074is not a share of the measured dispersion but an upper bound on what an evaluation drawing independently per run would face\. Both rows survive with their roles named: the measured\-share row prices a reader’s independent re\-run and thek=16k\{=\}16instrument prices this harness, on which no evaluation budget buys the interior because the termkkwould buy down is largely not in the dispersion to begin with\.
There is a second way to see that the2\.202\.20\-floor span was an instrument statement rather than a fact about arrangement, and it is untouched by any of this: the constant\-rate ladder spans11\.6311\.63floors at the samek=4k\{=\}4, comfortably above the3\.723\.72this design can resolve, so the same instrument that could not see the cosine interior sees the constant one without difficulty\. The load\-bearing claim, that the cosine interior is unresolvable at the evaluation this project ran, never depended on the exchange rate and is unaffected\. It does not touch the corner, which stands at21\.9921\.99floors\.
We cannot simply run the published arms at a largerkk\. Their checkpoints were reclaimed after evaluation \(§[4\.1](https://arxiv.org/html/2610.00234#S4.SS1)\), so those rows cannot be re\-scored at anykk; the constant\-rate family is fresh and is what both measurements above are made on\. Sources:dpd\_power\_vs\_k,evalvar\_direct,highk\_verdict\.
*Second, the null tested properly\.*Paired TOST against the interleaved arm, at an equivalence bound of2σ^=0\.04682\\hat\{\\sigma\}=0\.0468that Definition[3](https://arxiv.org/html/2610.00234#Thmdefinition3)froze rather than this paragraph chose, certifies one interior arm of seven\. Given the previous point that is what it must do: the90%90\\%intervals are two to four times wider than the bound\. Six arms are neither shown equivalent nor shown different\.
*Third, whether an interior difference would even be allocation\.*The difficulty coupling of Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)is present on all eight interior arms with the same sign, mean0\.06370\.0637, and it*varies*across them by0\.08170\.0817, which is4\.284\.28floors\. Its level cancels in a between\-arm difference; its variation does not\. So the residual coupling alone spans nearly twice the interior, and a difference of the interior’s size could not be attributed to a change of convention policy even if the design could resolve it\.
*What survives, and it is one statistic rather than any comparison\.*Treat the eight interior arms as groups and their seeds as replicates: the between\-arm intraclass correlation of the allocation share is−0\.076\-0\.076, with a bootstrap95%95\\%interval over arms of\[−0\.409,\+0\.166\]\[\-0\.409,\+0\.166\]and a permutationp=0\.61p=0\.61against the null that the rung label carries no information\. The interval covers zero, so we claim what it supports and no more:*between\-arm dispersion does not exceed within\-arm seed noise*\. It does not license the stronger sentence that arm identity explains none of the variance, and an earlier abstract of ours said exactly that\.
The corner is the one part of this figure that no power argument touches\.blockedsits21\.9921\.99floors above the top of the interior, five times the largest of the three quantities above and an order of magnitude above the interior span\.*The averaging wall as this paper can defend it is therefore one claim about a pooled statistic and one about a corner, and not a claim about the shape of the interior\.*Source:review\_statistics\.
The one substantive difference is where the amplitude starts\. On Qwen2\.5L=54L\{=\}54is the first rung off the floor, at0\.49220\.4922against a flat region of0\.4070\.407–0\.4490\.449; on Qwen3 it reads0\.45470\.4547and sits*inside*a flat region of0\.4240\.424–0\.4790\.479\. The blocked end is undiminished \(0\.89070\.8907against0\.86920\.8692\)\. The stopping phase is therefore present on both families and its onset is further along the ladder on the second, which is a statement about where the amplitude turns on rather than about whether it exists\.
#### The mixture ratio is the variable \(the deciding test, frozen\)\.
If the share is the time\-average of the path, it must track the mixture weights and nothing else\. The pre\-registered form of this test \(3:13\{:\}1\) was excluded by its own guard\. Below∼500\{\\sim\}500rows the minority convention is not learnable at all \(0\.21080\.2108at720720rows,0\.05460\.0546at468468,0\.00000\.0000at234234\), so its share would measure a learnability floor, not allocation, and was redesigned at2:12\{:\}1with the read\-out rules re\-frozen before the runs \(T0 guard; T1/T2/T3 outcomes\)\. Table[6](https://arxiv.org/html/2610.00234#S4.T6)gives the result: the shuffled arm’s share moves from0\.4530\.453at parity to0\.3120\.312at2:12\{:\}1, against mean\-field predictions0\.5000\.500and0\.3330\.333, both misses within the frozen0\.050\.05tolerance, both on the same side, the side of the pretrained preference for numerals\. Read\-out T1:*allocation tracks the mixture weights*\. The parity miss \(0\.0470\.047\) sits0\.0030\.003inside the tolerance, and is reported rather than rounding the margin up\. Table[7](https://arxiv.org/html/2610.00234#S4.T7)puts this beside the three interventions that were predicted to do nothing and did nothing, which is the shape of the claim: the mean field is weighted by source*proportions*and by nothing else the path can vary, so exactly one of four levers is allowed to move the share, and exactly one does\.
Table 6:The share tracks the mixture weights; the blocked arm does not\(natural\-conflict ratio corpora,936936rows, three seeds; read\-out rules frozen in the job script the scheduler executed: T1\)\. The3:13\{:\}1form was excluded by its own learnability guard and redesigned at2:12\{:\}1; the redesign is labelled, not laundered\.Table 7:The one lever the averaging limit leaves free\.Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)’s mean field is weighted by source*proportions*, so changing the mixture ratio must move the allocation and changing the path must not\. Both halves are measured here on the same corpus family\. Tolerance0\.050\.05was frozen with the prediction; both misses fall on the side of the pretrained prior, which the family measures independently\. \(T1\.\)The blocked arm is the counterpoint the mean\-field reading needs: its share sits at0\.731/0\.6990\.731/0\.699*regardless of the ratio*\. A mixture optimum moves with the weights; a corner state does not\. What kind of state it is, the next section measures\.
#### The designed schedule does not escape either \(pre\-registered, dead, and the sharpest of the three\)\.
The natural objection to everything above is that we rearranged the path naively\. Training onAAthenBBis, exactly and not by analogy, a*Lie–Trotter splitting*\[[72](https://arxiv.org/html/2610.00234#bib.bib77)\]ofexp\(A\+B\)\\exp\(A\{\+\}B\), which numerical analysis and NMR have studied for decades\[[73](https://arxiv.org/html/2610.00234#bib.bib80),[64](https://arxiv.org/html/2610.00234#bib.bib79)\], and that literature does not merely rank schedules but constructs better ones\. The canonical construction is palindromic\[[74](https://arxiv.org/html/2610.00234#bib.bib78)\]: runAAforT/2T/2,BBforTT,AAforT/2T/2\. Because the even\-order Magnus terms vanish for a time\-symmetric sequence\[[75](https://arxiv.org/html/2610.00234#bib.bib81)\], a Strang schedule is*second*\-order accurate where blocked is first\-order, so a practitioner forced to train in blocks would recover most of what interleaving buys at identical data and token budget\. The curriculum literature does not propose it, because it is trying to*choose*an order rather than to*design*a schedule\. The optimisation literature has since arrived at the same construction:[Nguyen et al\. \[76\]](https://arxiv.org/html/2610.00234#bib.bib75)prove that a*paired reversal*, symmetrising the epoch map exactly as a palindrome does, cancels the leading order\-dependent second\-order term and takes order sensitivity from quadratic to cubic in the step size\. That is the theorem our registration was betting on, stated more sharply than we stated it\.
We registered it in the linear model, where every arm is computable exactly and the joint arm both schedules approximate is available in closed form\. Three checks, thresholds frozen first: S1, the median\|Dstrang\|/\|Dblocked\|\|D\_\{\\text\{strang\}\}\|/\|D\_\{\\text\{blocked\}\}\|below0\.50\.5; S2, log–log slopes in the step size differing by at least0\.50\.5, so the gain is an*order*improvement rather than a constant; S3, the order effectASYM\\mathrm\{ASYM\}identically zero for a palindromic schedule, which is its own reverse, an implementation guard, not a claim\.
Table 8:The construction splitting theory recommends, and what it did\.Registered in the linear model, where every arm is computable exactly\. S3 is an implementation guard, not a claim: a palindromic schedule is its own reverse, so its order effect must vanish, and it does\. S1 and S2 are the claim and both fail: S1 by a factor of seven against its own bar, and S2*degenerately*, since refining the step by16×16\\timesmoves either penalty in the fifth decimal\. Sourcesplitting\_verdict\.checkregistered barmeasuredS3order effect of a palindromeASYM≡0\\mathrm\{ASYM\}\\equiv 000to machine precisionpassesS1median\|Dstrang\|/\|Dblocked\|\|D\_\{\\text\{strang\}\}\|/\|D\_\{\\text\{blocked\}\}\|<0\.5<0\.53\.65\\mathbf\{3\.65\}over300300draws \(IQR1\.881\.88–8\.098\.09\);3\.7%3\.7\\%of draws below the barfailsS2log–log slopes inη\\etadifferby≥0\.5\\geq 0\.5both slopes0\.00\.0to four figures acrossη=0\.4…0\.025\\eta=0\.4\\ldots 0\.025fails*the penalty across a16×16\\timesrefinement of the step, which is what “no order to improve” means*η=0\.4,0\.2,0\.1,0\.05,0\.025\\eta=0\.4,\\,0\.2,\\,0\.1,\\,0\.05,\\,0\.025blocked0\.081646,0\.081638,0\.081634,0\.081632,0\.0816310\.081646,\\ 0\.081638,\\ 0\.081634,\\ 0\.081632,\\ 0\.081631palindromic0\.32637,0\.32634,0\.32631,0\.32630,0\.326290\.32637,\\ 0\.32634,\\ 0\.32631,\\ 0\.32630,\\ 0\.32629Table[8](https://arxiv.org/html/2610.00234#S4.T8)is the result\. The palindromic schedule is roughly*four times worse*than the blocked one it was constructed to beat, and the refinement sweep says why no choice of threshold would have rescued it \(Figure[9](https://arxiv.org/html/2610.00234#S4.F9)\)\.
0\.00\.00\.10\.10\.20\.20\.30\.3palindromic \(Strang\), slope0\.00\.0blocked \(Lie–Trotter\), slope0\.00\.00\.40\.40\.20\.20\.10\.10\.050\.050\.0250\.025step sizeη\\eta\(log,16×16\\timesrange\)arrangement penalty\|D\|\|D\|\(a\)S1 and S2 both fail: the designedschedule is4×4\\timesworse and neither slope moves−0\.02\-0\.02\+0\.00\+0\.00\+0\.04\+0\.04\+0\.08\+0\.081×10−51\{\\times\}10^\{\-5\}3×10−53\{\\times\}10^\{\-5\}1×10−41\{\\times\}10^\{\-4\}learning rate\(b\)leaving the lazy regime movesDD−0\.0236→\+0\.0653\-0\.0236\\to\+0\.0653;\|D\|\|D\|grows at every stepconflict contrastDDFigure 9:Two ways of leaving the averaging wall, one designed and one accidental\.\(a\)The palindromic schedule splitting theory recommends, against the blocked one it was built to beat, in the linear model where the joint arm is available in closed form\. It is roughly four times*worse*, and refining the step by16×16\\timesmoves either penalty in the fifth decimal: there is no order to improve because the penalty is not a discretisation error\.\(b\)The one intervention that does move the contrast is leaving the lazy regime\. Raising the learning rate grows\|D\|\|D\|monotonically,0\.0236→0\.06530\.0236\\to 0\.0653, through a change of sign; the highest rate also costs capability,0\.2680\.268against0\.3400\.340, so it is the system leaving the regime the averaging argument is tight in\.The second failure is the informative one, and it is this section’s claim arriving from the other side\. Splitting theory’s guarantees are asymptotic in the step size, but refiningη\\etaby a factor of1616moves the penalty in the fifth decimal\.*The penalty is not a discretisation error at all*: it is set by the coarse structure of the arrangement and is blind to how finely that structure is resolved, which is what Eq\. \([1](https://arxiv.org/html/2610.00234#S1.E1)\) says and why no scheme designed against the discretisation can help\. The failure can therefore be located at an order\.[Nguyen et al\. \[76\]](https://arxiv.org/html/2610.00234#bib.bib75)’s guarantee is that symmetrisation deletes theO\(γ2\)O\(\\gamma^\{2\}\)order\-dependent term; our sweep says the conflict penalty is not in that term, because a quantity living atO\(γ2\)O\(\\gamma^\{2\}\)orO\(γ3\)O\(\\gamma^\{3\}\)cannot be invariant under a16×16\\timesrefinement ofγ\\gamma\. In the language of Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2), a palindrome is still*one period*: it symmetrises the block without shorteningτL\\tau\_\{L\}, so it lands where Eq\. \([3](https://arxiv.org/html/2610.00234#S3.E3)\) buys nothing, next toblockedrather than next to the interior\. The block\-length sweep found the same wall by measurement; this finds it in a model where the answer can be computed, and it kills the best\-motivated escape we could construct rather than the naive one\.
We report it as a failure because it was registered as a prediction\. It would have been the paper’s one piece of practical advice\.
One qualification about that registration\. The others in this paper are frozen in a job script committed hours to days before the run it launched, on the public history the registration index tabulates\. This check is a linear\-model computation of a few seconds whose thresholds and answer entered the repository in the same commit, so no earlier artefact carries the prediction\. S1–S3 are stated as supported\-if conditions rather than as a description of what happened, and the misses are wide enough \(3\.653\.65against0\.50\.5; both slopes zero\) that no threshold choice rescues them\. But the ordering the other tests demonstrate, this one can only assert\.
## 5The Two Non\-Average Terms: the Stopping Phase and the Key
Table 9:Every axis this project has swept, on one scale, sorted by how far it moves the allocation\.All six are measured on the allocation share of a mixed arm at one model and one budget, so the column is commensurable; the learning\-rate sweep is reported on the contrastDDrather than on the share and is deliberately absent rather than forced onto this axis\. The stopping phase is the largest term measured anywhere in the project, which is why it gets a section\. The block\-length ladder ranks fifth of six, and all three non\-path axes move the allocation further than the whole ladder does\. Source:axes\.Two of this paper’s terms are not averages of the corpus, and Table[9](https://arxiv.org/html/2610.00234#S5.T9)says which and how large\. It puts every axis the project has swept on the one scale where they are comparable, and it is sorted rather than arranged to flatter the headline\. The stopping phase is the largest term measured anywhere here,23\.8923\.89contrast floors, which is the case for spending a section on it\. The same table states the scope limit plainly: the block\-length ladder, which is the paper’s most refined path instrument, ranks fifth of six at4\.464\.46floors, and all three axes that are*not*properties of the path move the allocation further than the whole ladder does\. The path is a narrow channel and this paper’s claims are about the path, but it is neither the narrowest thing measured here nor the only input to an endpoint, and a reader who wants to predict an endpoint from the path alone should read this table before the rest of the section\. Figure[10](https://arxiv.org/html/2610.00234#S5.F10)draws the stopping phase as a position on the alternation cycle\.
\(a\)the picture, post hoc: a limit cyclemean\-field pointstop after aBBblock: share on theBBsiderebuild to end onAA:same cycle, other sideamplitude grows withLL; an endpoint measures*where the path stopped on the cycle*\(b\)the frozen prediction, measured \(M1\)0\.20\.20\.30\.30\.40\.40\.50\.5ends onBBends onAAL=54L=54ends onBBends onAAL=27L=27\+0\.0850\+0\.0850/−0\.0875\-0\.0875capabilityΔ0\.0025\\Delta 0\.0025capabilityCCaccA\\mathrm\{acc\}\_\{A\}accB\\mathrm\{acc\}\_\{B\}Figure 10:The stopping phase is a position on a cycle, and the mirror moves it to the other side\.\(a\)The reading of Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3), labelled post hoc: a driven system oscillates about the mean\-field point with amplitude growing inLL, and an endpoint measures where the path stopped\. It makes one frozen prediction\.\(b\)The measurement \(read\-out M1\): atL=54L\{=\}54the shift is antisymmetric,\+0\.0850\+0\.0850against−0\.0875\-0\.0875, with capability moving0\.00250\.0025\. AtL=27L\{=\}27both constructions sit at the floor and the mirror manufactures nothing, as it must\.TreatALBL⋯A^\{L\}B^\{L\}\\cdotsas a periodically driven system: eachAA\-block pulls the mixture weight towardAA, eachBB\-block pulls it back, and at equal weights the steady state is a*limit cycle around the mean\-field point*, with amplitude growing inLL; what an endpoint measures is*where the path stopped on the cycle*\. This picture was formed post hoc \(we label it so\) and it makes two pre\-registered predictions that then ran\.
First, amplitude: shares0\.4070\.407–0\.4490\.449forL≤27L\\leq 27\(amplitude≈0\{\\approx\}0, all within noise ofshuf’s0\.4130\.413\),0\.4920\.492atL=54L\{=\}54, and0\.8690\.869forblockedat three seeds \(0\.8560\.856at eight\), which on this reading is not a different species but*the lowest\-frequency alternation there is*, half a period that never swings back, stopped at maximum displacement\. Second, the mirror: our published sweep always ends on aBB\-block, so every measured share sits on theBB\-side of the cycle; rebuildingL=54L\{=\}54to end onAAmust relocate the same magnitude to the other side while moving capability not at all\. It does, antisymmetrically:accA\\mathrm\{acc\}\_\{A\}shifts\+0\.0850\+0\.0850whileaccB\\mathrm\{acc\}\_\{B\}shifts−0\.0875\-0\.0875, and capability moves0\.00250\.0025, an order of magnitude below either shift \(read\-out M1,dpd\_mirror\_verdict; atL=27L\{=\}27both constructions sit at the floor and the mirror manufactures nothing, as it must, the published arm anchoring ataccA=0\.3183\\mathrm\{acc\}\_\{A\}=0\.3183against the mirror’s0\.30960\.3096\)\. The residence timeτwrite=13\.6\\tau\_\{\\text\{write\}\}=13\.6steps that v2 measured as an “installation time” is, on this reading, the moment the mixture weight crosses50%50\\%: an allocation quantity, not a capacity one\.
#### The mirror on the corpus where neither convention is wrong\.
Everything above iscf2, whose second convention is a shifted answer\. The construction this paper’s claim is actually about isnat, a numeral against the same number spelled out, and a second ladder was trained on it at a third of the learning rate\. ItsL=45L\{=\}45rung carries a mirror, and the prediction is the same one: opposite signs on the two accuracies, capability unmoved\.
It holds\. Ending the alternation onAAinstead ofBBmovesaccA\\mathrm\{acc\}\_\{A\}by\+0\.0183\+0\.0183andaccB\\mathrm\{acc\}\_\{B\}by−0\.0229\-0\.0229, opposite signs, while capability moves−0\.0046\-0\.0046, which is0\.170\.17capability floors\.*Read the size before the pattern\.*Each shift is about one contrast floor, against4\.54\.5oncf2, so this is an observation consistent with the published mirror rather than an independent confirmation of it\. The smaller size is what Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx2)requires:\|D\|\|D\|grows with the step size, and this ladder runs at lr1×10−51\\times 10^\{\-5\}\. What it does establish is that the sign and the symmetry are not artefacts of a corpus in which one convention is arguably wrong\. Source:nat\_replication\.
#### The mirror on the second family manufactures nothing, and that is the prediction\.
Rebuilding theL=54L\{=\}54alternation to end onAAwas registered for Qwen3\-8B\-Base alongside the ladder, and it returnsΔaccA=\+0\.0204\\Delta\\mathrm\{acc\}\_\{A\}=\+0\.0204againstΔaccB=\+0\.0012\\Delta\\mathrm\{acc\}\_\{B\}=\+0\.0012: not antisymmetric, read\-outM2\. Taken alone that reads as the limit\-cycle picture failing to travel\. Taken with the rung it was built on, it is what the picture requires\. The mirror can only relocate an amplitude that exists, and on this familyL=54L\{=\}54sits inside the flat region rather than above it \(Table[11](https://arxiv.org/html/2610.00234#S5.T11)\), so there is nothing at that block length to move, which is exactly what we report atL=27L\{=\}27on Qwen2\.5, where both constructions sit at the floor and the mirror manufactures nothing*as it must*\. The test that would decide the picture on this family is the mirror at a rung where the amplitude has appeared, and*the ladder cannot supply one at this budget*\. We registered the bracketing rungs, the build refused them, and the registration was withdrawn having spent nothing; the reason is worth more than the experiment was, so it is given exactly rather than summarised\.
*The ladder’s rungs are quantised by the epoch, and the onset lies between two of them\.*build\_lsweepwrites a*one\-epoch*ordering, which the trainer then passes over three times, so a block must divide the5454chunks each source contributes per epoch: the divisors are1,2,3,6,9,18,27,541,2,3,6,9,18,27,54and are exactly the rungs already run\.blockedis not on that ladder at all\. It is a different construction, all ofAAfor three epochs and then all ofBB, so its effective block isL=162L\{=\}162, and the interval54<L<16254<L<162is not a finer division of an epoch but a*coarser*one that spans several\. Such a rung is constructible: a three\-epoch path file withL=81L\{=\}81is perfectly writable, so what closes this region is not an implementation limit\. What forbids it is one section further on\. Writing the path across epochs means training one pass over a tripled file, and that is precisely the construction the single\-cosine control of §[6](https://arxiv.org/html/2610.00234#S6)was built to test, which came backvoidon its own validity check because one epoch over a tripled file is not equivalent in what it learns to three epochs over the original: total accuracy0\.44420\.4442against0\.51250\.5125, a gap of2\.52\.5capability floors\.
So the two negatives this paper reports separately are one negative\.*At a fixed multiset and a fixed budget, the path space cannot be refined in the region where the stopping phase turns on*, because every route into that region changes what is learned, and a rung that is not learning\-matched is not a rung on this ladder\. That is a property of the instrument in the same sense that a diffraction limit is a property of a lens: it is read off the construction rather than discovered by failing\. We label the onset reading as we labelled the picture: post hoc, and now also untestable at this budget for a stated reason\.M2stands as the registered outcome\.
#### The mirror ofblocked, run, and what its control permits\.
The stopping\-phase reading rested on one informative mirror atL=54L\{=\}54plusblockedas an endpoint\. The most direct test is the mirror ofblockeditself, training all ofBBthen all ofAAagainst the publishedAAthenBB\. It was registered with its read\-out and run: three seeds, the same corpus, budget and evaluation\.
*The result, and then what it is worth\.*The share relocates from0\.86920\.8692to0\.03430\.0343, crossing the mean\-field point, with capability moving0\.00200\.0020, well inside the0\.02700\.0270capability floor\. The registered branch isBM\-2: the mirror goes to the other side but not by the same magnitude,\|−0\.4657\|\|\{\-\}0\.4657\|against\|\+0\.3692\|\|\{\+\}0\.3692\|, a gap of5\.055\.05contrast floors where the branch allowed two\.
*The comparability control fires, and it separates the two halves of that result\.*The published checkpoints for this corpus were reclaimed after evaluation, so the mirror had to be scored on a node whose kernel cache was cold, which compiles the sampling kernels afresh\. Before reading the mirror we re\-evaluated a*published*checkpoint that does survive, theB\_onlyarm the mirror resumes from\. It reproduces the majority accuracy,0\.42500\.4250against a published0\.42380\.4238,0\.060\.06floors, and*does not*reproduce the minority one,0\.09000\.0900against0\.05750\.0575,1\.701\.70floors\. Small accuracies moved and large ones did not\.
That is exactly the regime the mirror’s share lives in: atS=0\.0343S=0\.0343the numerator isaccB≈0\.0175\\mathrm\{acc\}\_\{B\}\\approx 0\.0175\. Propagating the control’s drift through givesSSanywhere in\[0\.000,0\.092\]\[0\.000,0\.092\]\. So the two halves of BM\-2 are not equally safe, and we separate them rather than report the branch label alone\.
*BM\-3 is excluded robustly\.*Across the control’s entire drift the share stays far below the mean\-field point, so the mirror does relocate to the other side andblockedis a position on a cycle rather than a distinct mechanism\. That is the question this cell was built to settle, it is settled, and Table[9](https://arxiv.org/html/2610.00234#S5.T9)’s largest row keeps its explanation\.
*The BM\-1 against BM\-2 distinction is not readable\.*It turns on a magnitude gap of0\.09640\.0964, and the control shows the small\-accuracy regime can move0\.03250\.0325, which is34%34\\%of it\. Whether the mirror is antisymmetric or carries a source\-order asymmetry on top is therefore open, and the cell that would close it is the same three arms scored on a node with a warm cache, or a re\-scoring of the published comparator on the same node\. We report BM\-2 as the mechanical branch label and decline to interpret it\. Source:blocked\_mirror\_verdict\.
#### The mirror at a constant rate, and a guard that could not have passed\.
Every mirror above trains under a cosine, so the branch of Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)that speaks most directly to a mirror, the terminal term carryingη\(T\)\\eta\(T\)with no contraction discount, is the one branch with no schedule contrast\. The cell was registered on a sign and a symmetry and explicitly not on a magnitude, with its read\-out frozen alongside it \(prereg/constant\_rate\_mirror\.md\)\. Six arms were trained, the comparator beside the mirror rather than reused, for a reason the registration records as an amendment: the base model the published arms trained from sits in a cache under`$HOME`, which is node\-local here and is no longer on any node the scheduler will accept a job for\.
*The registered outcome isCM\-4, void\.*Capability moves\+0\.0358\+0\.0358between the two arms,1\.331\.33capability floors, past the one\-floor bar the registration set as the condition for reading a share\. The shares are not read\.
*The guard could not have passed, which is a defect in the registration rather than a property of the result\.**Post hoc*: across the three seeds the capability difference is−0\.0238\-0\.0238,\+0\.0612\+0\.0612and\+0\.0700\+0\.0700, a sign change with a standard deviation of0\.05180\.0518andt=1\.20t=1\.20against4\.3034\.303\. The mean is smaller than its own dispersion, and that dispersion is1\.921\.92capability floors by itself\. §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)had already measured why: capability spans2\.52\.5capability floors across the constant family and under one across the cosine one, which is why §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)’s conservation result is quoted from the cosine family throughout\. The bar was set at one capability floor for a family already reported at two and a half, and a threshold no outcome of the design could have met is not stringency\.
*What the run printed, and what we decline to read\.*ΔaccA=\+0\.4150\\Delta\\mathrm\{acc\}\_\{A\}=\+0\.4150againstΔaccB=−0\.3792\\Delta\\mathrm\{acc\}\_\{B\}=\-0\.3792, which is\+21\.73\+21\.73and−19\.85\-19\.85contrast floors, the former at\+0\.4087\+0\.4087,\+0\.4187\+0\.4187and\+0\.4175\+0\.4175on the three seeds\. We print the numbers so that no reader need wonder what we saw, and we do not read them, because the arms whose difference they are are not matched in what they can do\. Closing the cell means resolving the capability difference rather than assuming it away: at the observed dispersion that is1717seeds for a95%95\\%interval inside one capability floor\. This paper does not buy them, and §[7](https://arxiv.org/html/2610.00234#S7)lists the cell at that size\. Sources:constant\_mirror\_verdict,constant\_mirror\_diagnosis\.
*One thing the co\-trained comparator does settle\.*Trained from the shared copy of the base model, it can be held against the published arm trained from the copy that has since disappeared\. The two agree to0\.01200\.0120in mean absolute accuracy, inside the0\.01910\.0191contrast floor, with the sign of the difference changing across seeds, so the substitution the amendment made is not a change this paper’s numbers can see\.
Two consequences deserve flat statement\.*Recency is not a confound of this system; it is the system’s one non\-average degree of freedom*: the stopping phase is what every recency effect in sequential fine\-tuning is made of\. And the exact\-match reader should hold this sentence against every blocked\-beats\-shuffled result, ours included:
> *blockedwins the exact\-match score precisely because it fails to optimise its own objective \(it never reaches the mixture that maximises the likelihood of its own corpus\) and exact match rewards the failure\.*
### 5\.1The Key: Training’s Index Channel
Everything above concerns a learner with only the compressive channel: no signal in the prompt says which convention applies, so the state can carry only a mixture weight\. CCH’s access\-complete move is to add the index: keep bindings addressable by key, pay the price, answer the query that names its target\. The training\-side analogue is disambiguation: rebuild the natural\-conflict corpus with a convention marker in the prompt, so the convention becomes*queryable*rather than contested, and re\-run the arms\.
*The marker is not new, and the contribution is not the marker\.*Conditioning training text on a prepended tag and then steering with it at inference is an established technique: control codes over source domains\[[77](https://arxiv.org/html/2610.00234#bib.bib56)\], conditioning on human\-preference scores\[[78](https://arxiv.org/html/2610.00234#bib.bib57)\], metadata conditioning with a cooldown so the model still runs unconditioned\[[79](https://arxiv.org/html/2610.00234#bib.bib58)\], and document identifiers injected to make knowledge attributable\[[80](https://arxiv.org/html/2610.00234#bib.bib59)\]\. What that literature has not had is a setting where the tag disambiguates*mutually contradictory*supervision, and therefore no measurement of where the technique stops working\. That is what this section supplies: the marker collapses the arrangement effect completely, and then recovers only87\.5%87\.5\\%of the union ceiling, with the shortfall traced to a specific and predictable failure, coverage of the minority convention at write time \(0\.37670\.3767against0\.00250\.0025\)\. The negative direction has independent support:[Higuchi et al\. \[81\]](https://arxiv.org/html/2610.00234#bib.bib60)find on controlled grammars that metadata conditioning*hurts*when the context does not determine the latent the tag names, which is the same boundary reached from the other side\.
The deciding read\-out was frozen with two components\.C: does the switch collapse?It does: the arrangement effect falls fromD=−0\.0985D=\-0\.0985\(5\.165\.16floors\) unmarked to−0\.0006\-0\.0006\(0\.030\.03floors, eight seeds\) marked, read\-out C1\. The switch is convention\-*selection*, and when selection is moved from the path to the query, arrangement stops mattering: the commitment reading of §[4\.2](https://arxiv.org/html/2610.00234#S4.SS2)survives its sharpest test\.L: does capability reach the union?Partially: the marked mean reaches0\.32590\.3259against the additive ceiling0\.37270\.3727,87\.5%87\.5\\%of the ceiling,74\.9%74\.9\\%of the gap recovered \(read\-out L4, against the L1 target0\.33540\.3354, short by0\.500\.50floors\)\. So conservation is an approximation \(“to within an eighth”\), not a law, and the deficit has a mechanism rather than an excuse:
\(Every number here is computed by an analyzer that also recovers the three constants the criterion was frozen against \(now0\.18630\.1863,0\.37270\.3727,0\.33540\.3354\) from the result base rather than restating them, which is what lets the section survive a change in its inputs without a hand edit: the unkeyed corpus went from three seeds to eight \(job26542654\), which moved the baseline from0\.19250\.1925to0\.18630\.1863, the ceiling from0\.38500\.3850to0\.37270\.3727and the fraction of that ceiling from84\.7%84\.7\\%to87\.5%87\.5\\%, with no hand edit and no change of verdict\.\)
> MarkedBB\-only, trained on nothing but the spelled convention, scores0\.37670\.3767on the numeral evaluation when the key asks for a numeral, against0\.12830\.1283unmarked\. MarkedAA\-only, trained on nothing but numerals and asked for a word, scores0\.00250\.0025\.*The key selects among conventions the model already holds; it cannot recover one that was never written\.*Numerals are the pretrained default; the spelled form must be acquired\.
Figure[11](https://arxiv.org/html/2610.00234#S5.F11)separates the two obstructions and Table[10](https://arxiv.org/html/2610.00234#S5.T10)gives the read\-out\.
0\.10\.10\.20\.20\.30\.3mean accuracyadditive ceiling0\.37270\.3727unkeyed0\.18630\.1863keyed0\.32590\.3259residue0\.0468\\mathbf\{0\.0468\}:12\.5%12\.5\\%of the ceiling74\.9%74\.9\\%of the gap\(a\) the key opens most of the union0\.10\.10\.20\.20\.30\.3unmarked0\.12830\.1283marked0\.3767\\mathbf\{0\.3767\}BB\-onlyasked for a*numeral*: written, so servedmarked0\.0025\\mathbf\{0\.0025\}AA\-onlyasked for a*word*: never written, so nothing to select\(b\) the residue is coverageQwen2\.50\.04680\.0468Qwen30\.1025\\mathbf\{0\.1025\}the22\-floor bar, frozen first\+2\.92\\mathbf\{\+2\.92\}floors:V1\(c\) and grows where coverage thinsFigure 11:What the key opens, and why what it leaves is a different obstruction\.\(a\)The unkeyed model sits at exactly half the additive ceiling, which is what a coin flip between two mutually exclusive golds scores; the key recovers74\.9%74\.9\\%of the gap and leaves0\.04680\.0468\.\(b\)That residue is*coverage*, and this panel is the whole of the evidence\. MarkedBB\-only, trained on nothing but the spelled convention, answers in numerals at0\.37670\.3767because numerals are the pretrained default and are therefore written; markedAA\-onlyasked for a word scores0\.00250\.0025, because there is nothing to select\.\(c\)And the residue grows on a family that writes four times less of the minority convention,\+2\.92\+2\.92floors against a bar of two frozen before the run \(V1\)\. It is a necessary\-condition test, which V2 or V3 would have refuted and which V1 cannot establish\. Sourcesunlock\_verdict,q3unlock\_verdict\.Table 10:What the key opens, and what it cannot\.The read\-out was frozen in the job script before the runs \(C on the switch, L on the level\), and the analyzer recovers the three constants it was frozen against from the result base\. The residue has a mechanism rather than an excuse: a markedBB\-onlymodel, trained on nothing but the spelled convention, answers in numerals when the key asks for numerals, but a markedAA\-onlymodel asked for a word cannot produce one it never saw\. The key*selects*among conventions already written\.read\-outquantityunkeyedkeyedverdictCarrangement switchDD−0\.0985\-0\.0985\(5\.165\.16floors\)−0\.0006\\mathbf\{\-0\.0006\}\(0\.030\.03floors\)C1: it collapsesLmean accuracy0\.18630\.18630\.3259\\mathbf\{0\.3259\}L4: partial unlockagainst ceiling0\.37270\.372750\.0%50\.0\\%87\.5%87\.5\\%74\.9%74\.9\\%of the gapagainst L1 target0\.33540\.3354—short by0\.500\.50floors*the residue, and why it is coverage rather than difficulty*markedBB\-onlyon the numeral evaluation0\.12830\.12830\.3767\\mathbf\{0\.3767\}the key selectsmarkedAA\-onlyon the spelled evaluation—0\.0025\\mathbf\{0\.0025\}it cannot createThe87\.5/12\.587\.5/12\.5split is two walls, not one\.Read through Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)the result decomposes exactly: the key removes the underdetermination term*in full*: that is the collapse ofDDfrom−0\.0985\-0\.0985to−0\.0006\-0\.0006, and it costs no capacity because it changes which policy is optimal rather than what can be stored\. What survives is a*coverage*term of a different species: the key cannot serve a convention the corpus never wrote, and the0\.37670\.3767against0\.00250\.0025asymmetry above measures precisely that residue\. Two obstacles, two mechanisms, separated by one experiment; the12\.5%12\.5\\%is not a weaker version of the87\.5%87\.5\\%but a different quantity, and a difficulty\-matched control \(§[6](https://arxiv.org/html/2610.00234#S6)\) is what would settle whether the residue is coverage or merely difficulty\.
\(a\) five registered quantities, each against the bar frozen for itQwen2\.5\-7BQwen3\-8B\-Basethe bar0\.030\.030\.10\.10\.30\.311331010measured value÷\\divthe bar it had to clearG1the conflict installsaccB\(B\-only\)≥0\.15\\mathrm\{acc\}\_\{B\}\(B\\text\{\-\}\\textsc\{only\}\)\\geq 0\.15passesG2the switch is present\|D\|\>2\|D\|\>2floorspassesK1the key collapses it\|Dkeyed\|≤1\|D\_\{\\text\{keyed\}\}\|\\leq 1floorholdsV1the residue growsresidue\>\>Qwen2\.5’s\+2\+2floorsholdsW2the ladder’s flat spanspan≤2\\leq 2floors*misses**post hoc*, chosen afterW2came back: resolvable endpoint states \(Def\.[3](https://arxiv.org/html/2610.00234#Thmdefinition3)\) are22on both families, which does not convert a miss into a hold\(b\) M2 the mirror, whose rule is a shape and not a number00224466−6\-6−4\-4−2\-20022ΔA\+ΔB=0\\Delta\_\{A\}\{\+\}\\Delta\_\{B\}\{=\}0ΔaccA\\Delta\\mathrm\{acc\}\_\{A\}\(contrast floors\)ΔaccB\\Delta\\mathrm\{acc\}\_\{B\}relocation quadrantQwen2\.5\-7BQwen3\-8B\-BaseThe mirror rebuildsL=54L\{=\}54to end onAAinstead ofBBand asks the endpoint to move the same magnitude to the other side\. On Qwen2\.5 it does:\+4\.45\+4\.45against−4\.58\-4\.58floors, on the dashed line of pure relocation, with capability moving0\.090\.09of a capability floor\. On Qwen3 both skills move*up*,\+1\.07\+1\.07and\+0\.06\+0\.06, which is a capability change of0\.800\.80capability floors and not a relocation at all\. The two families disagree about the shape rather than about a threshold, which is why this row is drawn and not scored;L=54L\{=\}54is inside the flat region on Qwen3 \(panel \(a\),W2\), so there is no amplitude there for a mirror to move\.Figure 12:Three results asked to travel, and the answers are not uniform\.\(a\)Every quantity whose threshold was frozen before the second family ran, divided by that threshold, so the bar sits at11for all of them and the arrow says which side had to be cleared\. Hollow is Qwen2\.5\-7B, filled is Qwen3\-8B\-Base, and the line between them is the travel\. The guards pass well: the conflict installs*better*on the new family, so nothing below is an instrument at its floor\. The key collapses the switch to3%3\\%of its bar on one family and66%66\\%on the other, and the residue grows as the coverage reading requires\. The ladder’s span misses*and so does the published row at1\.091\.09*, which is a fault in the registration rather than in the second family\. Rows are ordered by what they decide, not by outcome, and the post\-hoc statistic is ruled off\.\(b\)The mirror is the one registration with no scalar bar: its rule is a shape, so it is drawn as one\. Sourcesq3unlock\_verdict,q3ladder\_verdict,unlock\_verdict,rate\_verdict\.#### The same key on a second pretraining family, and what the residue is made of\.
The collapse above is one model, and a remedy can be family\-specific where a phenomenon is not: the key works by making the convention conditionable, which is a property of pretraining rather than of the conflict\. We registered the replication before running it, on Qwen3\-8B\-Base, at the unmarked family’s own budget and seeds, and it holds:DDfalls from−0\.0875\-0\.0875\(4\.584\.58floors\) to−0\.0125\-0\.0125\(0\.660\.66floors\), read\-outK1\(Table[11](https://arxiv.org/html/2610.00234#S5.T11)\)\.
The registration also carried a second branch, and it is the one worth the twenty trainings\. This family is not merely a second model: it is a second point on the axis this section’s own explanation names\. Trained on nothing but the spelled convention, Qwen3 reachesaccB=0\.0520\\mathrm\{acc\}\_\{B\}=0\.0520while still emitting0\.29830\.2983in numerals, against Qwen2\.5’s0\.20910\.2091:*four times less of conventionBBis written*, at the same corpus, budget, seeds andkk\. If the residue is coverage, a family that writes less of the minority convention must leave a larger one\. The threshold was fixed at two floors on the residue before the run, and the residue grows from0\.04680\.0468to0\.10250\.1025, a move of\+2\.92\\mathbf\{\+2\.92\}floors: read\-outV1\. The unlock falls with it, from87\.5%87\.5\\%of the ceiling to75\.5%75\.5\\%\.
*What that does and does not buy\.*It turns the87\.5/12\.587\.5/12\.5split from a number into a relationship, which is more than a single model can support and less than a proof\. The two families differ in everything a pretraining corpus can differ in, so V was registered as a*necessary\-condition*test: V2 or V3 would have refuted the coverage reading, while V1 is consistent with it and cannot establish it, because some third property of Qwen3 could move the residue the same way\. The difficulty\-matched control of §[6](https://arxiv.org/html/2610.00234#S6)therefore stays a missing cell rather than being quietly retired\.
*One instrument correction, because it fired first\.*The read\-out audits the instrument before printing any branch, and on first execution it refused: eighteen of twenty keyed arms contained problems scored correct under*both*conventions, which the audit called impossible\. The audit was wrong\. Exclusivity holds for the unmarked corpus, where one prompt carries two mutually exclusive golds; the marked evaluations are two*different*prompts, so a model that serves both scores on both, which is what this section claims the key buys\. The check was corrected against the published run:dis\_q25\_7bcarries1717–1919such problems per arm while the unmarked families carry exactly00\. The count is now recorded rather than treated as a failure, and is a coverage signal of its own:22here against1717–1919there\. Figure[12](https://arxiv.org/html/2610.00234#S5.F12)collects the three results asked to travel\.
Table 11:The key, and the ladder, asked to travel\.Both were registered before their runs with the branches and thresholds frozen, and both gates were read first\. The key replicates and its residue moves as the coverage reading requires\. The ladder’s registered branch isW2, and the row below it is why that verdict is reported together with the threshold it was read against: the published Qwen2\.5 ladder does not meet the same bar, which was checkable before the run and was not checked\. Sources:q3unlock\_verdict,q3ladder\_verdict\.quantityQwen2\.5\-7BQwen3\-8B\-Baseread\-outguardsminority convention written,accB\(Bonly\)\\mathrm\{acc\}\_\{B\}\(B\\ \\textsc\{only\}\)0\.20910\.20910\.05200\.0520—on the*synthetic*conflict0\.42250\.42250\.4813\\mathbf\{0\.4813\}G1 passesswitch present,\|D\|\|D\|on that corpus11\.911\.9fl15\.0\\mathbf\{15\.0\}flG2 passesthe keyDDunkeyed−0\.0985\-0\.0985\(5\.165\.16fl\)−0\.0875\-0\.0875\(4\.584\.58fl\)DDkeyed−0\.0006\-0\.0006\(0\.030\.03fl\)−0\.0125\-0\.0125\(0\.660\.66fl\)K1collapsesmarked mean against its ceiling87\.5%87\.5\\%75\.5%75\.5\\%L4partialresidue==ceiling−\-marked0\.04680\.0468\(2\.452\.45fl\)0\.1025\\mathbf\{0\.1025\}\(5\.375\.37fl\)V1\+2\.92\+2\.92flthe ladderflat region span,L=1L\{=\}1–2727,purerand,shuf0\.04160\.0416\(2\.182\.18fl\)0\.05440\.0544\(2\.852\.85fl\)W2at a22fl bar*which the published row also misses*differ by0\.670\.67floors*post hoc*span as a fraction of the blocked excursion9%9\\%12%12\\%*post hoc*resolvable endpoint states \(Def\.[3](https://arxiv.org/html/2610.00234#Thmdefinition3)\)𝟐\\mathbf\{2\}𝟐\\mathbf\{2\}Rreal=1R\_\{\\mathrm\{real\}\}\{=\}1bitwithin\-cluster / between\-cluster gap8\.7×8\.7\\times20\.7×20\.7\\times*post hoc*L=54L\{=\}540\.49220\.4922, off the floor0\.45470\.4547,*inside*itblocked0\.86920\.86920\.89070\.8907mirror atL=54L\{=\}54:ΔaccA\\Delta\\mathrm\{acc\}\_\{A\},ΔaccB\\Delta\\mathrm\{acc\}\_\{B\}\+0\.0850\+0\.0850,−0\.0875\-0\.0875\+0\.0204\+0\.0204,\+0\.0012\+0\.0012M2This is also the training\-time face of the inference\-time theory’s load\-bearing assumption\. CCH’s hybrid crosses its walls only under*write\-time code separability*: the anchor must be distinguishable when the binding is written, not merely when it is queried\[[2](https://arxiv.org/html/2610.00234#bib.bib26)\], which in the vocabulary of Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)is the assumption thatH\(convention∣query\)=0H\(\\text\{convention\}\\mid\\text\{query\}\)=0\. CCH assumes the stream is well posed; this paper measures what a learner does when it is not\. Our marker is exactly such an anchor, and its failure mode is exactly the assumption’s: present at write time, it opens the union up to what was written; absent \(or the content never learned\), no query\-time cleverness recovers the binding\. And the family’s super\-additivity has a measured analogue: the reachable behaviour set with the key strictly contains the set without it: calibrated mixture*and*per\-query convention control, against mixture\-or\-corner alone, at the price the walls demand \(the unlearned remainder\), not for free\.
## 6Threats to Validity
#### Every registered claim, and how it came back\.
The abstract says three mechanisms were pre\-registered and falsified\. Table[12](https://arxiv.org/html/2610.00234#S6.T12)is the whole ledger rather than those three, because a scorecard that lists only the informative failures is a selection, and because two further entries cost us something a reader should not have to reconstruct: the positive control inverted on its first build, and the control registered to discharge the learning\-rate confound came back void\. Rules are quoted from the artefact that froze them; the registration index of Appendix[D](https://arxiv.org/html/2610.00234#A4)gives the commit and the interval for each\.
Table 12:The ledger: every claim this paper registered, and what came back\.Read the verdict column as a distribution rather than a score: three mechanisms died, four predictions held, one control was void and one instrument build inverted\. The three deaths were derivable in advance from Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1), which is why they are kept rather than deleted\. Rules are quoted from the artefact that froze them; Appendix[D](https://arxiv.org/html/2610.00234#A4)gives the commit and the interval for each\.registered claimthe rule, frozen firstwhat came backverdictv1batch composition carries the switchpurerandnearblocked⇒\\Rightarrowpurity; nearshuf⇒\\Rightarroworderpurerand0\.22420\.2242againstshuf0\.21790\.2179,blocked0\.44540\.4454: purity is3%3\\%of the spandeadv2optimiser memory cancels contested directionsa response turning on atL⋆≈10L^\{\\star\}\\\!\\approx\\\!10, or atτwrite=13\.6\\tau\_\{\\text\{write\}\}\\\!=\\\!13\.6M\(L\)M\(L\)flat through both onsets,σ\(M\)≈0\.06\\sigma\(M\)\\\!\\approx\\\!0\.06; onlyL=54L\{=\}54risesdeadStranga palindromic schedule beats a blocked oneS1 median\|Dstr\|/\|Dblk\|<0\.5\|D\_\{\\text\{str\}\}\|/\|D\_\{\\text\{blk\}\}\|<0\.5; S2 slopes differ by≥0\.5\\geq 0\.5S13\.653\.65; S2 both slopes00; S3 passesdeadT1allocation tracks the mixture weightsshare within0\.050\.05of mean\-field0\.5000\.500and0\.3330\.3330\.4530\.453and0\.3120\.312: misses0\.0470\.047and0\.0210\.021, both toward the pretrained priorheldLadderevery block length lands at the interleaved allocationL=1…54L\{=\}1\\ldots 54within noise ofshuf’s own share0\.4070\.407–0\.4490\.449againstshuf’s0\.4130\.413heldM1the mirror relocates antisymmetrically at fixedCCending onAAmoves the same magnitude to the other side\+0\.0850\+0\.0850against−0\.0875\-0\.0875, capability moves0\.00250\.0025heldCthe key collapses the switch\|Dkeyed\|\|D\_\{\\text\{keyed\}\}\|within one floor of zero−0\.0985→−0\.0006\-0\.0985\\to\-0\.0006\(5\.16→0\.035\.16\\to 0\.03floors\)heldLthe key lifts capability to the unionmarked mean≥0\.90×\\geq 0\.90\\timesthe additive ceiling0\.37270\.37270\.32590\.3259:87\.5%87\.5\\%of the ceiling, short of the bar by0\.500\.50floorspartialS\-Ca single\-cosine arm discharges the schedule confoundvalidity check S\-C4 read before any allocation comparisontotal accuracy0\.44420\.4442against0\.51250\.5125, a gap of2\.52\.5capability floorsvoidControlconflict fires and its twin stays inertcontrol within noise, conflict beyond itv1*inverted*\(0\.10σ0\.10\\sigmaagainst2\.41σ2\.41\\sigma\); v3 reads0\.15σ0\.15\\sigmaagainst12\.29σ12\.29\\sigmav1 failedQ3\-keythe key collapses on a second pretraining familybranches C, L and V frozen before the first result fileDD−→−0\.0125\-0\.0875\\\!\\to\\\!\-0\.0125; unlock75\.5%75\.5\\%against87\.5%87\.5\\%; residue grows2\.922\.92floorsK1/L4/V1Q3\-wallthe averaging wall replicatesspan of the flat region≤2\\leq 2floors, gate read firstgate passes at15\.015\.0floors; span2\.852\.85floors, and the*published*row is2\.182\.18, so the bar was mis\-specifiedW2Q3\-mirrorthe stopping phase replicatesantisymmetric shift atL=54L\{=\}54, capability within a floor\+0\.0204\+0\.0204against\+0\.0012\+0\.0012;L=54L\{=\}54is inside the flat region on this familyM2high\-kkthe borrowed evaluation share transfers, soσ\\sigmafalls as0\.136\+0\.864⋅4/k\\sqrt\{0\.136\+0\.864\\cdot 4/k\}ratio\(16\)∈\[0\.475,0\.713\]\\mathrm\{ratio\}\(16\)\\in\[0\.475,0\.713\]around a predicted0\.5940\.594;*all*eleven arms required at each ofk=16k\{=\}16andk=32k\{=\}32k=16k\{=\}16complete at33/3333/33;ratio\(16\)=1\.079\\mathrm\{ratio\}\(16\)=1\.079, outside the band and above*one*, which the additive model can carry only as a negative component\. §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)’s shared engine stream predicts oneHK\-4The last row’s bookkeeping, since the verdict turns on it rather than on a measurement\.G\-HK1fails because the job ran nine arms atk=32k\{=\}32and the guard asks for eleven, so the cell is void on bookkeeping and not on measurement; the six missing evaluations were submitted and then*withdrawn unstarted*when the project’s compute closed\. They would only have moved the label toHK\-2, because §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)’s correction comes from a direct measurement that decides no branch\.
#### Three byproducts, flagged not developed\.
*Ordering poisoning:*purerandandblockedare the same multiset in different order and differ by0\.220\.22exact\-match: an attacker controlling only data\-loader order, touching no byte, decides which convention a model commits to; the existing arms are the demonstration’s skeleton\.*Theκ\\kappatracer:*a few hundred conflicting pairs planted in any large run measure that pipeline’s effective averaging in situ, without touching the main data\.*Evaluation methodology:*any benchmark whose answers admit multiple correct conventions is silently scoring commitment; the decomposition of Eq\. \([4](https://arxiv.org/html/2610.00234#S4.E4)\) separates the two at zero cost\.
#### Limitations\.
The ladder, mirror, ratio and key experiments are one model \(Qwen2\.5\-7B\) at one budget on two conflict constructions\. The switch alone is broader: the companion paper measures it at33B,77B and1414B within one pretraining family and replicates it on a second, pre\-registered before the run \(Qwen3\-8B\-Base:−0\.0875\-0\.0875at4\.584\.58floors, five of five seeds negative, control inert at0\.930\.93\)\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]\. That replication carries the caveat this paper is in the worst position to wave through: its learnability guard reads0\.0520\.052against0\.2090\.209on Qwen2\.5\-7B, so the second family fires the switch from a rung close to the floor at which §[4\.1](https://arxiv.org/html/2610.00234#S4.SS1)’s v1 build failed outright\.
Three of this paper’s own results have since been asked to travel to that family, each registered before its run, and the answers are not uniform \(Table[11](https://arxiv.org/html/2610.00234#S5.T11)\)\. Thekey replicatesand its residue moves as the coverage reading requires \(K1, V1\)\. Theladder’sregistered branch isW2, and the honest report of it is two sentences rather than one: the bar was two floors, the second family spans2\.852\.85, and the*published*first family spans2\.182\.18and misses the same bar\. We set that threshold without checking it against the row we already had\. Themirrorreturns M2 at a block length that, on this family, sits inside the flat region and therefore has no amplitude to relocate\. What travels, then, is the key and the switch; what the ladder shows is that the two families are flat to within two thirds of a floor of each other under a criterion we mis\-specified, which is weaker than the replication we registered and stronger than the failure the branch label alone suggests\. Under the criterion Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)actually implies, the number of resolvable endpoint states, both families read22with an order\-of\-magnitude margin \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\); that comparison is post hoc and does not convert W2 into W1, but it does say which of the two statistics was measuring the wall\.
*What the flatness claim is powered to say\.*§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)states it in full and it is short: the interior span is smaller than this design’s detectable difference at any power we would quote, so the averaging wall is defended by the pooled between\-arm statistic and by the corner, not by the interior’s shape\. Conservation is approximate:9\.7%9\.7\\%arm\-to\-arm on the synthetic corpus,15%15\\%on the natural one, range0\.14750\.1475along a trajectory,74\.9%74\.9\\%gap recovery under the key, and only*mutually exclusive*conflicts are tested; style conflicts, or cases where both answers can be wrong, are open\. The learning\-rate schedule was entered as a residual confound and is not one:blockedis a second cosine, the ladder a single one, and the gap it could account for is0\.85590\.8559\(the blocked arm at eight seeds\) against the ladder’s top rung at0\.49220\.4922, nineteen floors, which is not a size a reader should be asked to dismiss on a plausibility argument\. We therefore registered the single\-cosineA162B162A^\{162\}B^\{162\}control before building it, froze its four branches, ran it, andit came back void\. The validity check S\-C4 is read first by construction and it fires: total accuracy is0\.44420\.4442against the blocked arm’s0\.51250\.5125, a gap of2\.52\.5capability floors, so one epoch over a tripled file is not equivalent in learning to three epochs per stage and the two arms are not comparable\. The registered consequence is taken in full\. The allocation comparison is void, is reported as an instrument result, and is*not*repaired by adjusting the budget\.
Two things follow and neither is comfortable\. First, the confound is not discharged\. The plausibility argument that the quantity explained is a mixture weight rather than an amount learned is not available here, and the failure of the experiment we sent to settle it does not reinstate that argument\. Second, the direction of the void is not neutral\. The share it forbids us to read is0\.50890\.5089, which sits0\.870\.87floors above the ladder’s top rung, and the branch it would have landed in is S\-C3, the one that amends this paper’s abstract\. We print the number so that no reader need wonder what we saw, and we do not read it, because the arms whose difference it is are not matched in what they learned\.
The components show why the design cannot be repaired cheaply\. Under a single cosine the terminal B block trains in the decayed half of the schedule, so B is learned less and A is forgotten less: at seed4242,accA\\mathrm\{acc\}\_\{A\}rises from0\.0780\.078to0\.2150\.215whileaccB\\mathrm\{acc\}\_\{B\}falls from0\.4500\.450to0\.2240\.224\. That is the recency\-by\-learning\-rate interaction of Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx2), and it is exactly why the registration said in advance that this control measures two\-stage against one\-stage rather than the cosine alone\. That same non\-equivalence has a second consequence we did not anticipate when we ran this control, and §[5](https://arxiv.org/html/2610.00234#S5)draws it: the tripled file is also the only route to a ladder rung betweenL=54L\{=\}54andblocked, so the2\.52\.5floors below close the onset region as well as voiding this comparison\. One measurement, two negatives\. Separating them needs an arm that equalises what is learned, which is the move that registration forbids us to reach for after seeing its result\.
#### So we registered a different arm, and it found something else\.
The forbidden move was repairing S\-C; a new design frozen in advance is not that move, andprereg/schedule\_control\.mdsays so in its first section\. S\-C changed the*path*and held the schedule\. This changes the*schedule*and holds the pathbyte for byte, which is possible only because the ladder arms disable trainer shuffling: three epochs over the file is exactly the file’s order three times, so writing that order out and splitting it in half gives two stages whose concatenation is the control’s own stream\. Same rows, same order, same324324steps, same epochs’ worth of gradient; two cosines of162162steps where the control has one of324324, which isblocked’s structure\. The builder asserts the reconstruction rather than claiming it\.
The guard S\-C failed is read first and passes: capability moves0\.590\.59capability floors atL=27L\{=\}27and0\.260\.26atL=54L\{=\}54, against a bar of two\. The arms are matched in what they learned\.
*And the allocation moves a great deal\.*On the registered primary rungL=27L\{=\}27, chosen before either arm ran because its control seed dispersion is0\.01590\.0159againstL=54L\{=\}54’s0\.05950\.0595, the share falls from0\.40710\.4071to0\.29120\.2912:Δ=−0\.116\\Delta=\-0\.116,t=−10\.1t=\-10\.1on four degrees of freedom,6\.076\.07contrast floors\. Table[13](https://arxiv.org/html/2610.00234#S6.T13)carries both rungs\.
Three consequences follow, the first of which costs us a sentence\.
*The stopping phase is not the only non\-average term\.*The two arms have the identical path and therefore the identical time\-average, which is all Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)sees, and their endpoints differ by six floors\. The abstract’s claim is amended to say what is true: the only non\-average term*of the path*is the stopping phase, and the path is not the only input to the endpoint\. Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)is untouched, being a statement aboutη→0\\eta\\to 0that says nothing about where a cosine restarts\. What was too strong is the reading we hung on it: that arrangement is the interesting variable because the endpoint is otherwise fixed\.
*The confound is closed in the direction it threatened\.*A schedule term that explainedblockedwould have to push allocation*toward*it\. This pushes away, on both rungs, so the second cosine is not what putsblockedat0\.8690\.869\. That is the specific alternative §[5](https://arxiv.org/html/2610.00234#S5)’s attribution was exposed to, and it is now excluded by measurement rather than by the plausibility argument this section withdrew\.
*And the flat ladder means something different than we said\.*The whole block\-length family,L=1L\{=\}1to5454, moves the share0\.0850\.085\. One cosine restart on a fixed path moves it0\.1160\.116\. Allocation is not rigid; the path is simply not the channel that carries it, which is the sharpest form of §[1](https://arxiv.org/html/2610.00234#S1)’s two\-state reading and was not available to us before this arm ran\.
What we cannot say is*why*the schedule pushes towardAA\. The registered branch isSC\-3, whose definition is that the movement is reported as unexplained and not folded into the branch that would have closed the question, and we take that in full\. Two accounts we had ready are excluded by the data rather than by argument: the registration predicted that if a mid\-run stopping phase were the mechanism the two rungs would move in*opposite*directions, because the boundary falls on anAAblock atL=54L\{=\}54and aBBblock atL=27L\{=\}27; they move the same way\. And a bias from which source occupies the restarted cosine’s peak fails for the same reason, sinceL=54L\{=\}54’s second stage begins onBB\. This is one open cell rather than a qualification of the three above\. It sits alongside X1 and X2 of §[7](https://arxiv.org/html/2610.00234#S7)and the limit\-cycle picture’s post\-hoc origin \(its two deciding tests were frozen and passed after the picture was formed, and are labelled so\)\.
Table 13:The schedule control: identical path, one cosine against two\.The treatment consumes the control’s own stream in the same order and differs only in that the cosine restarts and the optimiser state resets at the midpoint, which isblocked’s structure\. The guard that voided the single\-cosine control is read first and passes on both rungs\.L=27L\{=\}27was fixed as primary before either arm ran, because its control dispersion is four times smaller\. Its movement is6\.076\.07contrast floors att=−10\.1t=\-10\.1\.L=54L\{=\}54moves further and does not clear its own noise: its treatment seeds span0\.400\.40, wider than the effect, and the registration recorded in advance that at this dispersion its silence is not evidence of absence\. That is why the mechanical branch labelSC\-2is reported with that sentence beside it rather than as “no effect”\. Both rungs move the same way, which the registration fixed in advance as the signature of the schedule rather than of a mid\-run stopping phase\. Sourceschedctl\_verdict\.sharerungcontrol \(one cosine\)treatment \(two\)Δ\\Deltatt\(crit2\.7762\.776\)guard\|ΔC\|\|\\Delta C\|L=27L\{=\}27*primary*0\.40710\.40710\.2912\\mathbf\{0\.2912\}−0\.116\\mathbf\{\-0\.116\}−10\.10\\mathbf\{\-10\.10\}0\.590\.59fl0\.413,0\.419,0\.3890\.413,0\.419,0\.3890\.294,0\.302,0\.2780\.294,0\.302,0\.2786\.076\.07flSC\-3passesL=54L\{=\}54*secondary*0\.49220\.49220\.20530\.2053−0\.287\-0\.287−2\.35\-2\.350\.260\.26fl0\.424,0\.526,0\.5270\.424,0\.526,0\.5270\.024,0\.424,0\.1680\.024,0\.424,0\.16815\.0215\.02flSC\-2passes*for scale: the entire block\-length ladder,L=1L\{=\}1through5454, moves the share0\.0850\.085*Two further cells are named by Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)and are the ones we would run first\.*A difficulty\-matched control for the12\.5%12\.5\\%residue*: the coverage reading of §[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)requires that the unrecovered part is a convention never written rather than one merely harder, and until that control exists the two\-wall decomposition is an interpretation with a plausible alternative\.*A predictable\-conflict corpus*: if the mixture is the optimum of an underdetermined query rather than a limit on what can be stored, then a conflict of the same size whose convention*is*a function of the input should be resolved with no key at all, at the samekk, volume, budget and architecture\. That is the sharpest test of Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)available, it costs one corpus and four arms, and a mixture there would refute the reading rather than qualify it\.
*That cell has since been run, by the family’s evaluation\-side member, and the disclosure belongs here rather than in its pages alone*\[[82](https://arxiv.org/html/2610.00234#bib.bib67)\]: a corpus keying the convention on a single\-character feature of the problem statement ran twice at this scale and decided nothing\. The pooled read\-out was voided by its own symmetry check; the paired within\-problem contrast, registered afterwards, returned a miss \(Δ=\+0\.0018\\Delta=\+0\.0018,t=0\.23t=0\.23\); and the audit of the void found the keying feature nearly collinear with per\-item scoring asymmetry, all6565problems of the discriminating stratum on one side of the feature\. The reading of Remark[2](https://arxiv.org/html/2610.00234#Thmremark2)is therefore*pending on that corpus, not supported and not refuted*, and any rebuild must first pass the orthogonality certificate that episode produced \(worst stratum’s minority share≥0\.20\\geq 0\.20, checked on the two pure arms before any mixed training\)\.
Two things about any rebuild have to be said now rather than rediscovered, because both would invalidate it; the first is this paper’s own design point and the second is that episode’s lesson, not our foresight\. First, the read\-out cannot be Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)’s\. SettingH\(conv∣I\)=0H\(\\mathrm\{conv\}\\mid I\)=0is exactly the statement that each held\-out problem has one applicable convention, so there is no allocation left to split: the measurement becomes accuracy against the applicable gold against the rate of answering under the inapplicable one, and the capability–allocation coordinates this paper is built on do not survive the construction that tests them\. Second, the keying feature must be uncorrelated with difficulty\. Assigning the convention*by skill*, which is the obvious construction and the one we would have reached for, fails that: skills differ in how often they are solved at all, so the two golds would sit on problem subsets of unequal difficulty and the comparison against the unpredictable corpus would not be matched in the one quantity \(§[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)\) it exists to separate from coverage\. The feature has to ride on the problem rather than partition the problems\.
## 7Discussion: One Stream, Two Channels
Table[14](https://arxiv.org/html/2610.00234#S7.T14)states the correspondence this paper has been using, one row per load\-bearing object\. It is a structural identification, not an analogy hunt: in both columns the object is a bounded state fed by an unbounded stream, the two channels are compression and indexed access, and disagreement \(κ\\kappa\) is the switch that makes the choice of channel visible in behaviour\.
Table 14:The CCH family: the same objects at inference time and training time\.Left column:[Chen et al\. \[2\]](https://arxiv.org/html/2610.00234#bib.bib26)\. Right column: this paper and, for the premise rows, the companion recipe\-level paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]\.The recipe\-level member of the family\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]is the premise map: within a coherent domain \(κ≈0\\kappa\\approx 0\) the compressive channel’s output does not depend on the path at all, so composition, order and arrangement sit inside single\-run noise\. The free\-recipe limit is the averaging limit’s special case at zero conflict, and its budget\-transient order effect is the low\-frequency end of the axis this paper’s ladder climbs\. The three papers measure one theory at three places: what a bounded system can*serve*at query time \(CCH\), what a recipe can*move*when premises hold or break \(FRL\), and what the path*writes*through each channel \(this paper\)\.
#### The flagship runs on the corpus where one convention is arguably wrong, and the corpus where neither is runs weaker\.
Two constructions are used throughout:*cf2*, in which one half’s boxed answers are shifted by one and are therefore contradictory by construction, and*nat*, in which one half writes a numeral and the other spells the same number out, so both are genuinely correct\. Every headline figure in this paper is measured on cf2\. On the standard install guard, the minority convention’s accuracy in the arm trained on nothing else, cf2 reads0\.42250\.4225and nat reads0\.20910\.2091,*less than half*; capability at the interleaved arm is0\.52020\.5202against0\.37270\.3727\. The philosophical claim of the paper, that a conflict need not be an error, is a claim about nat, and nat is the weaker instrument\.
This neither voids the results nor can be waved through, so here is what it costs\. It does not affect conservation or the allocation coordinate, both measured on each corpus separately and agreeing in sign and in order of magnitude\. It does mean the effect sizes quoted are the cf2 ones, and the discount a reader rebuilding this on a both\-correct conflict should apply is the one measured in the next paragraph, because the two published numbers differ in step size as well as in corpus\. And it exposes one asymmetry the construction creates: on nat theAA\-onlyarm scores*exactly*0\.00000\.0000on the spelled convention across eight seeds, so the spelled form is never produced unless it is trained, whereas cf2’s shifted form leaks at0\.01690\.0169\. A conflict between two forms the pretrained model already writes and a conflict where one form must be installed from scratch are not the same experiment, and the second is the one our philosophical framing is about\. Source:review\_statistics\.
#### One premise of that accounting, registered and tested: the size asymmetry is mostly a budget asymmetry\.
The pair above iscf2at lr3×10−53\\times 10^\{\-5\}againstnatat lr1×10−51\\times 10^\{\-5\}, so it confounds the corpus with the step size\. We registered the cheap half of the separation before running it \(prereg/nat\_budget\_gate\.mdwithreadout\_nat\_budget\.py, frozen11h 08m ahead of the verdict and untouched since, branches readNI\-4toNI\-1so that the branch closing the route is read before the branch opening it\) and reran thenatBB\-onlyarm atcf2’s learning rate with nothing else changed: same720720rows, same three epochs, same135135steps, same base model, a separate tag so no published file is written to\. Minority install rises from0\.20910\.2091\(sd0\.00890\.0089, eight seeds\) to0\.37380\.3738\(sd0\.01810\.0181, seeds4242,4343,4444\), a move of8\.628\.62contrast floors att=15\.06t=15\.06, and capability on that arm rises by5\.095\.09capability floors to0\.46880\.4688; the guard was one\-sided against a run destabilised by the larger step, and a rise is what more budget is supposed to do\. That is branchNI\-1:77\.2%77\.2\\%of the distance tocf2’s0\.42250\.4225is closed by the step size alone\. A corroborating number was already on the result base and unread: at the*published*lr1×10−51\\times 10^\{\-5\},cf2’s own install reads0\.19880\.1988againstnat’s0\.20910\.2091,0\.540\.54floors att=1\.27t=1\.27and therefore no separation, so*the*“*less than half*”*above is a statement about two learning rates before it is a statement about two corpora\.*The remainder is stated as narrowly as it is measured\. This is one arm of four and says nothing about the arrangement effect; the interleaved\-arm capability comparison,0\.52020\.5202against0\.37270\.3727, is still at unmatched budgets and stands as printed\. What is left on the arm we did run,0\.04880\.0488or2\.552\.55contrast floors, sits within a hair of the family’s two\-degree\-of\-freedom critical value \(t=4\.55t=4\.55against4\.3034\.303\) and three seeds do not resolve it\. The accounting above therefore changes in size and not in direction: on the one arm now matched, the discount is about seven eighths rather than about half\. Putting the switch itself at parity needsnatandnatcat four arms and eight seeds, roughly5656GPU\-hours, listed as future work item \(vii\) rather than bought here\. Source:nat\_budget\_verdict,nat\_budget\_context\.
#### Three claims this paper leans on and cannot check\.
Every number reported here is measured here, but three*context*claims are borrowed, and each weakens a different generalisation\.*First, scale\.*The conflict switch at33B and1414B is measured by the companion paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]and not here, so every arm in this manuscript is77B or88B and*this paper on its own establishes nothing about how the switch scales*\. That is the single largest gap in its external validity\.*Second, prevalence\.*The rate at which public corpora actually disagree about form,12\.7%12\.7\\%of shared problems rising to27\.1%27\.1\\%against externally authored keys, is the survey’s\[[67](https://arxiv.org/html/2610.00234#bib.bib66)\]and licenses the choice of construction rather than any number below; if that rate is wrong, this paper measures a real mechanism on a rare population\.*Third, the evaluation harness\.*That the sampling stream is common\-mode across seeds is diagnosed by the recipe paper\[[5](https://arxiv.org/html/2610.00234#bib.bib52)\]and is why the dispersions here are harness\-specific \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\); an external replication should expect a*higher*noise floor thanσ^=0\.0234\\hat\{\\sigma\}=0\.0234, which would widen every interval we print and shrink no effect\. None of the three is peer\-reviewed at the time of writing, and they are stated as borrowings rather than cited as support\.
#### Future work, listed at the size we think each one is\.
None is run to a readable answer\. Two have been run: the ladder at a constant learning rate \(Table[5](https://arxiv.org/html/2610.00234#S4.T5)\) and item \(v\), which returned a void and is listed at the size that would close it\. A third, item \(vii\), has had its cheap premise tested and confirmed\.\(i\) Separating the two things a constant rate leaves open\.Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)attributes the constant\-rate family’s11\.6311\.63floors to the terminal weightw\(T\)w\(T\);[Sweeney \[24\]](https://arxiv.org/html/2610.00234#bib.bib41)attributes order sensitivity instead to AdamW’s buffers advancing on step count rather than onτ=ηk\\tau=\\eta k\. Both predict a resolved ladder at a constant rate and the present design cannot tell them apart\. One arm decides it: the same ladder without bias correction, or under SGD with momentum, where the fixed clock is absent and only the terminal weight remains\.\(ii\) Eight seeds on the main rung\.Every pairwise statement in §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)is limited by three seeds rather than by the effect, which is cheap to fix\.\(iii\) A conflict outside mathematics\.Every corpus here is competition mathematics with integer answers, so the convention is surface form; whether a style or a factual conflict conserves capability the same way is untested\.\(iv\) A run that keeps generations\.The0\.14750\.1475capability excursion along a trajectory cannot presently be separated into knowledge briefly lost and scorer briefly unable to parse; retaining generations costs storage and no compute\.\(v\) The mirror at a constant rate, at enough seeds to read it\.This cell has run and returnedCM\-4, void: capability moved past the registered bar and the shares are not comparable \(§[5](https://arxiv.org/html/2610.00234#S5)\)\. It needs seeds and not arms: at the capability dispersion the constant family has,1717seeds put a95%95\\%interval on the capability difference inside one capability floor, against the three this design ran\.\(vi\) A difficulty\-unweighted allocation\.SSis app\-weighted share andCov\(p,s\)≠0\\mathrm\{Cov\}\(p,s\)\\neq 0is measured rather than assumed away \(Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)\); the unweighted𝔼^\[s\]\\hat\{\\mathbb\{E\}\}\[s\]over a fixed problem set separates policy from population, needs no new training, and should move the arms together rather than differentially, which is why the coupling is reported as a bound\.\(vii\) The switch itself, at matched budget\.The budget gate above returnedNI\-1: matching the learning rate closes77\.2%77\.2\\%of the install gap between the two corpora on theBB\-onlyarm\. That is one arm of four and says nothing about the arrangement effect\. Putting the switch at parity needsnatandnatcat four arms and eight seeds, about5656GPU\-hours; the gate is what makes the cell worth buying, and this manuscript does not buy it\.
#### The mean\-field predictions were tested as point values, and a directional test would have been the better instrument\.
The ratio test compares an installed share against0\.5000\.500and0\.3330\.333with a frozen tolerance of0\.050\.05, and both cells miss on the same side, toward the pretrained prior\. Two misses in the same direction are evidence for a model this test cannot express: mean field*plus*a prior\-bias term of unknown size\. A point test with a tolerance cannot separate that from a failure of mean field, whereas a directional test can, and the design is a third ratio: a1:11\{:\}1,2:12\{:\}1,4:14\{:\}1sweep testing monotone decrease and the slope rather than three point predictions\. We did not run it, and we flag the current test as the weaker instrument rather than reporting its two misses as though the alternative model were not on the table\.
#### Cross\-predictions, stated as missing cells\.
A unification earns its keep by predicting across members, and we state the two sharpest as pre\-registered future tests rather than claims\.X1, the mixture\-capacity wall:CCH’s capacity floor \(B/NB/Nper binding\) should have a training\-time image, withNNmutually exclusive conventions rather than two, the installed mixture’s calibration against the mixture weights should degrade withNNat fixed data per convention; a clean break would be training’s Shannon wall\.X2, replay as deferral:CCH’s index defers the horizon by1/ρ1/\\rho; mixing a fractionρ\\rhoof first\-stage data into the final stage is the recipe\-level index, so the commitment flip of §[5](https://arxiv.org/html/2610.00234#S5)should be deferred by the same coefficient, connecting this family to replay in continual learning\[[36](https://arxiv.org/html/2610.00234#bib.bib1)\]with a quantitative prediction rather than a metaphor\. Neither is run; both are cheap; either failing would cut the family at the joint we have named\.
#### Where the correspondence stops\.
Table[14](https://arxiv.org/html/2610.00234#S7.T14)identifies objects at inference time and training time and goes no further\. A third position exists and is deliberately not claimed here: a system that writes its own conclusions into a store it later retrieves from\. There the bounded state is neither the parameters nor the context window but an accumulating record; capability can be conserved by the store’s construction rather than by an averaging limit; and what fixes the allocation is a retrieval rule rather than a stopping phase\. The decomposition of Eq\. \([4](https://arxiv.org/html/2610.00234#S4.E4)\) is written for parameters trained once over a fixed corpus, and its conservation is approximate and measured\. Whether the same two coordinates describe an accumulating store is a question about a different object, and nothing in this paper is evidence either way\. Establishing that two such decompositions are the same object would itself be an experiment, not a remark\.
## 8Conclusion
An averaging theorem organises everything above, and one factorisation says when it applies\. The ordering\-dependent part of an endpoint is bounded by a product in which the arrangement enters only as a period and the schedule only as a reach \(Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)\), so the two ends of the divided record on ordering are one knob at two settings rather than two findings\. Below corpus scale, and under the decaying schedule that is the field’s default, the path is compressed to its time\-average, so batch composition, block length and optimiser buffers move nothing that this design can resolve; the time\-average installs the Bayes\-optimal mixture, so conflicting supervision trains a stochastic policy at capability conserved to within an eighth rather than damaging what the model can do; the one term of the path that is not an average is where it stops, so blocked training is a stopping phase, a commitment device whose exact\-match victory is its own objective’s defeat\. A query\-time key is a second channel rather than more bandwidth on the first, which is why it collapses the arrangement effect that no schedule could\.
The claim we would most like a reader to take is the one that costs them nothing to adopt\. Conflicting supervision moves*commitment*, and exact match reports commitment as though it were capability\. Any benchmark whose answers admit more than one correct form is affected, the correction is the two coordinates of Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2), and it needs no experiment from this paper: only the per\-problem scores an evaluation already produces\.
## Appendix ASupporting Measurements
### The stopping term in\-model: real, and not where the theory says
Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3)makes the stopping phase the onlyπ\\pi\-dependent term\. If that is right, a linear model in the lazy regime should reproduce the sign reversal we measure along the budget axis, and it does, without reproducing its*location*, which is a limitation of the theory rather than of the measurement and is reported here rather than in the main line\.
Sampling random non\-commuting task pairs inside the lazy regime and sweeping the budgetTT:ASYM\(T\)\\mathrm\{ASYM\}\(T\)changes sign in𝟕𝟐%\\mathbf\{72\\%\}of draws andD\(T\)D\(T\)in𝟕𝟎%\\mathbf\{70\\%\}, so budget\-dependent reversal is generic and not an artefact of our corpora\. The rate is stable across step sizes \(η/ηstab=0\.1\\eta/\\eta\_\{\\text\{stab\}\}=0\.1gives0\.8670\.867and0\.7250\.725; the sweep is flat to within a few points across the range\)\. But the leading Magnus/commutator proxyT⋆T^\{\\star\}, which is what a practitioner would compute once at initialisation to*predict*where the crossing sits, correlates with the measured crossing budget at Spearman0\.031\\mathbf\{0\.031\}over278278pairs with a crossing\. That is, not at all\.
Table 15:The mechanism survives every test; the one\-shot proxy for*where*it fires survives none\.Upper panel: budget\-dependent sign reversal is generic in the lazy regime, so the stopping term of Proposition[3](https://arxiv.org/html/2610.00234#Thmproposition3)is not an artefact of our corpora\. Lower panel: the leading Magnus/commutator score, the quantity a practitioner would compute once at initialisation to predict the crossing: is asked to rank three different things in three different places and fails all three, twice with the*wrong sign*\. Reported together because they are one settled negative about a single object, not three separate disappointments\.\(a\) the mechanism: reversal is generic000\.50\.5111\.51\.522000\.250\.250\.50\.50\.750\.7511η/ηstab\\eta/\\eta\_\{\\text\{stab\}\}fraction of draws reversingchanceASYM\(T\)\\mathrm\{ASYM\}\(T\):𝟕𝟐%\\mathbf\{72\\%\}D\(T\)D\(T\):𝟕𝟎%\\mathbf\{70\\%\}budgets representableAboveη/ηstab=1\.3\\eta/\\eta\_\{\\text\{stab\}\}\{=\}1\.3the sweep can no longer represent every budget, and theDDseries falls with the fraction that it can\. Drawn rather than omitted\.\(b\) the one\-shot proxy, asked where it fireswhat a usable predictor needs−1\-1−0\.5\-0\.5000\.50\.511Spearman correlation with what it ranksT⋆T^\{\\star\}vs the crossing budget \(n=278n\{=\}278\)\+0\.031\+0\.031‖\[A,B\]‖\\\|\[A,B\]\\\|vs\|ASYM\|\|\\mathrm\{ASYM\}\|,11ep \(n=6n\{=\}6\)−0\.543\-0\.543‖\[A,B\]‖\\\|\[A,B\]\\\|vs\|ASYM\|\|\\mathrm\{ASYM\}\|,33ep \(n=6n\{=\}6\)−0\.600\-0\.600Figure 13:One object, confirmed as a mechanism and failed as a predictor\.The same two halves as Table[15](https://arxiv.org/html/2610.00234#A1.T15), drawn, because both halves are claims about a distribution\.\(a\)Over random non\-commuting pairs in the lazy regime,ASYM\(T\)\\mathrm\{ASYM\}\(T\)reverses sign in72%72\\%of draws andD\(T\)D\(T\)in70%70\\%, and the rate stays above chance across a19×19\\timessweep of the step size, so the reversal is something the finite\-budget theory produces rather than a property of our corpora\. The dotted series is the fraction of budgets the sweep can still represent, which falls aboveη/ηstab=1\.3\\eta/\\eta\_\{\\text\{stab\}\}=1\.3and is why theDDseries falls with it; it is drawn rather than left out\.\(b\)Every ranking the one\-shot commutator score is asked for, on one axis\. It is uncorrelated with the in\-model crossing budget over278278pairs and*inverted*against the measured order effect at both budgets\. Sourcesasym\_budget\_verdict,commutator\_verdict\.We take the honest reading: the mechanism is confirmed in\-model, the schedule of the mechanism is not\. A single evaluation atθ0\\theta\_\{0\}is the first term of a series whose error accumulates with the budget, which is also what[Piontkovskaia and Nikolenko \[52\]](https://arxiv.org/html/2610.00234#bib.bib48)observe empirically as pairwise\-order accuracy decaying with block length\. Median rotations at the crossing are small \(11\.2∘11\.2^\{\\circ\}on the sum,11\.7∘11\.7^\{\\circ\}on the difference\), so the reversal is not a large geometric event; it is a near\-cancellation, which is exactly why its location is hard\.
### The same effect at three budgets, and at three learning rates
#### The published order effect does not survive its own budget\.
On algebra–combinatorics,ASYM\\mathrm\{ASYM\}reads\+0\.0779\+0\.0779at the published budget and−0\.0619\\mathbf\{\-0\.0619\}when the budget is raised, three seeds,s\.e\.=0\.0191\\mathrm\{s\.e\.\}=0\.0191: a reversal, not an attenuation\. The probe explains why the two budgets are different regimes at all: at one epoch and lr3×10−53\{\\times\}10^\{\-5\}, fine\-tuning*damages*three of the five skills relative to zero\-shot \(calculus0\.304→0\.2810\.304\\to 0\.281, and similarly for two others\), so the published regime is one in which training is net\-negative on part of the corpus\. An order effect measured there is a statement about that regime\. Reported as a limitation of the published number, and it is the sharpest single reason this paper reports budgets rather than recipes\.
#### Learning rate moves the conflict contrast monotonically\.
Table[16](https://arxiv.org/html/2610.00234#A1.T16)carries the sweep\.
Table 16:The switch grows with the learning rate, monotonically, and the arm that is supposed to be inert does not move with it\.Three learning rates at three epochs on the natural conflict, three seeds each;D¯\\bar\{D\}is the mean allocation contrast andacc¯only\\overline\{\\mathrm\{acc\}\}\_\{\\textsc\{only\}\}the mean single\-convention accuracy, which is what a capacity explanation would have to move and does not\. Raising the rate is the one intervention that takes the system further from the lazy regime the theory is derived in, and\|D\|\|D\|grows with it: which is the direction the suppression argument predicts\.\|D\|\|D\|grows monotonically,0\.0236→0\.06530\.0236\\to 0\.0653, as the step size grows: the direction the suppression argument requires, since raising the learning rate is the one intervention that moves the system furthest from the lazy regime in which arrangement is doubly suppressed\. Note that the highest rate also costs capability \(0\.2680\.268against0\.3400\.340\), so the growth in\|D\|\|D\|is not a free improvement: it is the system leaving the regime in which the averaging argument is tight\.
### The commutator proxy, measured and rejected
The natural predictor of which order wins is the commutator of the two tasks’ update operators, evaluated once atθ0\\theta\_\{0\}\. We measured it on Qwen2\.5\-7B activations at layer fraction0\.750\.75across the matched pairs\. It ranks\|ASYM\|\|\\mathrm\{ASYM\}\|at Spearman−0\.543\\mathbf\{\-0\.543\}at one epoch and−0\.600\\mathbf\{\-0\.600\}at three: the wrong sign, consistently, on both budgets\. Four of the six pairs flip sign between budgets, so there is no fixed ordering for a static proxy to predict \(Figure[14](https://arxiv.org/html/2610.00234#A1.F14), Table[17](https://arxiv.org/html/2610.00234#A1.T17)\)\.
801001200\.020\.020\.040\.040\.060\.060\.080\.08‖\[A,B\]‖\\\|\[A,B\]\\\|atθ0\\theta\_\{0\}, thousands\|ASYM\|\|\\mathrm\{ASYM\}\|contrast floor0\.01910\.0191measured ranking:Spearman−0\.543\-0\.543\(11ep\),−0\.600\-0\.600\(33ep\)11epoch33epochs\(a\) the predictor ranks backwards, at both budgets−0\.04\-0\.04000\.040\.040\.080\.08−0\.06\-0\.06−0\.03\-0\.0300ASYM\\mathrm\{ASYM\}at11epochASYM\\mathrm\{ASYM\}at33epochsalg–calalg–geoalg–comcal–geocal–comgeo–comthe shaded quadrants are the ones where the order effectchanged signbetween the two budgets; four of the six pairs are in them\(b\) and there is no fixed ordering to predict
Figure 14:The commutator proxy fails twice, and the two failures are different\.\(a\)The score a practitioner would compute once atθ0\\theta\_\{0\}ranks the pairs*backwards*, at both budgets \(Spearman−0\.543\-0\.543at one epoch,−0\.600\-0\.600at three\)\. A weak predictor scatters about zero; this one is consistently inverted, which is stronger than uninformative\. The dashed line is the contrast floor\.\(b\)The deeper failure: four of six pairs change the*sign*of their order effect between budgets, so there is no fixed ordering for any once\-computed quantity to predict\. Sourcecommutator\_verdict\.Table 17:The six pairs the proxy is asked to rank, and the ranking it produces\.The commutator norm is evaluated once atθ0\\theta\_\{0\}on Qwen2\.5\-7B activations at layer fraction0\.750\.75;T⋆T^\{\\star\}is the leading Magnus budget built from it;ASYM\\mathrm\{ASYM\}is the measured order effect, three seeds, contrast floor0\.01910\.0191\. Two things are fatal to a static predictor and both are visible by eye: the commutator ranks\|ASYM\|\|\\mathrm\{ASYM\}\|with the*wrong sign*at both budgets, and four of six pairs change that sign between budgets, so the quantity being ranked is not a property of the pair alone\. Sourcecommutator\_verdict\.pair‖\[A,B\]‖\\\|\[A,B\]\\\|T⋆T^\{\\star\}ASYM\\mathrm\{ASYM\}\(11ep\)ASYM\\mathrm\{ASYM\}\(33ep\)sign flips?algebra–calculus89,00289\{,\}0020\.3810\.381\+0\.0743\+0\.0743\+0\.0020\+0\.0020algebra–geometry81,52481\{,\}5240\.4590\.459\+0\.0320\+0\.0320−0\.0333\-0\.0333yesalgebra–combinatorics70,81070\{,\}8100\.2810\.281\+0\.0779\+0\.0779−0\.0619\-0\.0619yescalculus–geometry131,642131\{,\}6420\.3190\.319−0\.0396\-0\.0396\+0\.0041\+0\.0041yescalculus–combinatorics112,524112\{,\}5240\.3550\.355\+0\.0104\+0\.0104−0\.0125\-0\.0125yesgeometry–combinatorics95,17595\{,\}1750\.3610\.361\+0\.0506\+0\.0506\+0\.0091\+0\.0091Spearman,‖\[A,B\]‖\\\|\[A,B\]\\\|against\|ASYM\|\|\\mathrm\{ASYM\}\|−0\.543\\mathbf\{\-0\.543\}−0\.600\\mathbf\{\-0\.600\}4/64/6flipSpearman,T⋆T^\{\\star\}against\|ASYM\|\|\\mathrm\{ASYM\}\|—−0\.200\-0\.200Taken with Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx1)’s in\-model result \(Figure[13](https://arxiv.org/html/2610.00234#A1.F13)\) \(the same proxy at Spearman0\.0310\.031against the crossing budget, and all three rankings collected in Table[15](https://arxiv.org/html/2610.00234#A1.T15)\) the conclusion is that a one\-shot geometric score is not a usable predictor of order effects at these budgets, in the model or in the measurement\. We report this because the proxy is the obvious thing to try and because our own framework motivates it\.
*The scope of that sentence is “at these budgets”, and it is doing work\.*We measure‖\[A,B\]‖\\\|\[A,B\]\\\|once atθ0\\theta\_\{0\}over blocks of162162–324324steps;[Sweeney \[25\]](https://arxiv.org/html/2610.00234#bib.bib42)scores a different statistic and already reports its accuracy falling from98%98\\%to73%73\\%betweenk=1k\{=\}1andk=20k\{=\}20, and[Piontkovskaia and Nikolenko \[52\]](https://arxiv.org/html/2610.00234#bib.bib48)reports the same decay independently\. Nothing here contradicts either\. What the two rankings above add is that the decay does not stop at indifference: past some block length the score is*inverted*, which is the regime a practitioner choosing a fine\-tuning order is actually in\. Whether the tournament construction inverts where our scalar proxy does is a question about that construction and we have not measured it\.
### The positive control’s first build, in full
§[4\.1](https://arxiv.org/html/2610.00234#S4.SS1)summarises the inverted first build\. The full read\-out, for a reader checking whether the redesign was principled or opportunistic: v1 drew238238rows per condition from algebra alone at one epoch, and the conflict arms scored0\.020\.02against the shifted golds while the same checkpoints scored0\.390\.39–0\.500\.50against unshifted ones\. That is the diagnostic: the\+1\+1convention was never installed, soDDwas measured on an instrument whose treatment was absent, and the control’s2\.41σ2\.41\\sigmawas the only thing in the experiment with any variance to report\. The redesign changed two things, both stated before the rerun: pool the integer\-answer rows of every skill \(raising238238to864864\) and raise the epochs, together giving the conflicting convention roughly an order of magnitude more gradient steps\. Nothing about the read\-out rule changed\. Table[18](https://arxiv.org/html/2610.00234#A1.T18)shows why the first build read at the floor, and Figure[15](https://arxiv.org/html/2610.00234#A1.F15)draws the two halves of it\.
Table 18:Why v1 read at the floor: the treatment was never installed\.Every conflict arm is scored twice, against the shifted golds the\+1\+1convention defines and against the unshifted ones the pretrained model already produces\. Had the conflicting convention been learned, the first column would carry the mass\. It does not, soDDin v1 was a difference between two arms neither of which had received the treatment, and the control’s2\.41σ2\.41\\sigmawas the only variance in the experiment\. Source theconflict\_switchfamily, three seeds,238238rows per condition at one epoch\.\(a\) v1: the conflict arms, scored twice0\.10\.10\.20\.20\.30\.30\.40\.40\.50\.5exact match≈0\.02\\approx 0\.02against the*shifted*gold the treatment defines0\.390\.39–0\.500\.50against the*unshifted*gold the base already writesthe\+1\+1convention lost to the pretrained prior, soDDwas read on an instrument whose treatment was absent\(b\) the read\-out, unchanged in rule0044881212\|D\|\|D\|in seed s\.e\.noise,2s\.e\.2\\,\\mathrm\{s\.e\.\}2\.412\.410\.100\.10v1: the inert arm moved and the treated one did not0\.190\.1911\.91\\mathbf\{11\.91\}v3: the rebuild, three seeds, same rulecontrolconflictFigure 15:An instrument reporting a null because its treatment was absent\.\(a\)v1’s conflict arms scored against both golds\. The\+1\+1convention the arms were trained on is worth≈0\.02\{\\approx\}0\.02; the unshifted convention the base already writes is worth0\.390\.39–0\.500\.50\. Had the treatment installed, the first bar would carry the mass\.\(b\)What the read\-out rule then returned, in units of the seed standard error, and what the same rule returned after the redesign\. In v1 the arm registered to stay inert moved2\.412\.41and the arm registered to move returned0\.100\.10: the control is the only thing in the experiment with variance to report\. The v3 pair is the three\-seed read\-out, matched to v1’s three seeds rather than the eight\-seed pair Table[12](https://arxiv.org/html/2610.00234#S6.T12)prints\. Nothing about the rule changed between them\. Sourcesconflict\_switch\_verdict,conflict\_switch2b3lr3e5\_verdict\.
## Appendix BThe Clustering Rule, and Why It Is Not a Channel Rate
This appendix carries the endpoint\-clustering rule §[1](https://arxiv.org/html/2610.00234#S1)uses, the sweep of the one constant it depends on, and the reason the counts are reported as counts\.
###### Definition 3\(The write channel, its capacity, and its realised rate\)\.
The path carrieslog2\(nA\+nBnA\)\\log\_\{2\}\\binom\{n\_\{A\}\+n\_\{B\}\}\{n\_\{A\}\}bits of source\-order information\. LetΦ\(π\)∈ℝd\\Phi\(\\pi\)\\in\\mathbb\{R\}^\{d\}be a behavioural read\-out of the endpoint andσ^\\hat\{\\sigma\}the seed dispersion\. Two endpoints count as*distinguishable*when they are separated by2σ^2\\hat\{\\sigma\}\. The factor of two is a convention and the headline depends on it, so Table[19](https://arxiv.org/html/2610.00234#A2.T19)sweeps it and Figure[16](https://arxiv.org/html/2610.00234#A2.F16)draws the sweep\. The two rates below respond to that sweep differently: halving the resolution adds exactly one bit toRcapR\_\{\\mathrm\{cap\}\}, which is a logarithm of it, and does not toRrealR\_\{\\mathrm\{real\}\}, which counts clusters, so at1σ^1\\hat\{\\sigma\}the first family resolves a third state andRrealR\_\{\\mathrm\{real\}\}reads1\.581\.58bits rather than22\. The*write channel*is the induced mapπ↦Φ\(π\)\\pi\\mapsto\\Phi\(\\pi\), and it has two rates that must be kept apart\.
capacityRcap\\displaystyle\\text\{capacity\}\\quad R\_\{\\mathrm\{cap\}\}=log2\(rangeΦ/2σ^\),\\displaystyle=\\log\_\{2\}\\bigl\(\\operatorname\{range\}\\Phi/2\\hat\{\\sigma\}\\bigr\),realised rateRreal\\displaystyle\\text\{realised rate\}\\quad R\_\{\\mathrm\{real\}\}=log2\|Φ\(Π\)/∼2σ^\|,\\displaystyle=\\log\_\{2\}\\bigl\|\\Phi\(\\Pi\)/\{\\sim\}\_\{2\\hat\{\\sigma\}\}\\bigr\|,whereΠ\\Piis the family of paths actually run and∼2σ^\{\\sim\}\_\{2\\hat\{\\sigma\}\}is single\-linkage at the same resolution\.RcapR\_\{\\mathrm\{cap\}\}counts the distinguishable values the*range*supports; it is the rate an encoder would achieve if it could place an endpoint anywhere in that range\.RrealR\_\{\\mathrm\{real\}\}counts the distinguishable values the path family*occupies*, and it is an upper bound on the information any encoder built from this family can transmit, however many paths it uses\.
Three things this definition is not, stated before it is used\.First,log2\(nA\+nBnA\)\\log\_\{2\}\\binom\{n\_\{A\}\+n\_\{B\}\}\{n\_\{A\}\}is the entropy of the design space under a uniform measure over*all*orderings\. We ran ten\. An ensemble of ten codewords carries at mostlog210=3\.32\\log\_\{2\}10=3\.32bits, so no measurement in this paper can exhibit a compression larger than3\.32:13\.32\{:\}1, and the ratio of the design\-space entropy to a measured cluster count is not a compression measurement, and the design\-space entropy appears in this paper only as the size of the space the ladder samples\. Second, neither rate is a mutual information\.I\(π,Φ\)I\(\\pi;\\Phi\)would require an explicit prior on paths and a noise model for seed dispersion, and we estimate neither;RrealR\_\{\\mathrm\{real\}\}is a count of resolvable clusters wearing a logarithm, and the logarithm is a naming convention rather than a result\. Third,Φ\\Phihere is one scalar\. A degenerate image in one coordinate is not a degenerate image, and §[6](https://arxiv.org/html/2610.00234#S6)lists the coordinates we did not read\. The statement this paper defends needs none of this vocabulary:*ten paths spanning0\.46210\.4621in allocation, at a resolution of2σ^=0\.04682\\hat\{\\sigma\}=0\.0468that their own span would divide into about ten distinguishable values, occupy two*\. Source:review\_statistics\.
The two coincide only when the image is spread\. Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)says it is not, so the paper’s central measurement is the*gap*between them \(§[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\), and quotingRcapR\_\{\\mathrm\{cap\}\}as though it wereRrealR\_\{\\mathrm\{real\}\}was the error the gap corrects\.
Table 19:The headline pair, swept over the resolution constant it is defined at\.RrealR\_\{\\mathrm\{real\}\}counts the endpoint states the ten arms occupy under single linkage;RcapR\_\{\\mathrm\{cap\}\}is the logarithm of the range in units of the resolution and therefore gains exactly one bit per halving, whichRrealR\_\{\\mathrm\{real\}\}does not\.*Two clusters is what both families give from2σ^2\\hat\{\\sigma\}upward*, and the second family gives it from1σ^1\\hat\{\\sigma\}upward\. Below2σ^2\\hat\{\\sigma\}the first family resolves a third state, so the paper’s constant sits at the*lower edge*of the interval where the two families agree rather than in its middle, rather than presenting11bit as resolution\-free\. The interval is wide above the choice and the between\-cluster gap exceeds the largest within\-cluster gap by8\.7×8\.7\\timeson the first family and20\.7×20\.7\\timeson the second, which is why the count is stable there\. Source:rate\_sensitivity\.\(a\) the ten arm endpoints, and the window that decides whether two of them are one state0\.40\.40\.50\.50\.60\.60\.70\.70\.80\.80\.90\.9allocation share at the endpointnine arms,shuftoL=54L\{=\}54blocked2σ^=0\.04682\\hat\{\\sigma\}=0\.0468, the windowlargest gap inside the cluster,0\.04310\.0431gap between the clusters,0\.3770\.377:8\.7×8\.7\\timesthe largest one inside either\(b\) and the count does not move between2σ^2\\hat\{\\sigma\}and4σ^4\\hat\{\\sigma\}0\.50\.511223344resolution constant, in units ofσ^\\hat\{\\sigma\}11223344resolvable statesthe count is stable here22334455RcapR\_\{\\mathrm\{cap\}\}\(bits\)the paper’s2σ^2\\hat\{\\sigma\}The two rates answer the sweep differently, which is the appendix’s point\.RrealR\_\{\\mathrm\{real\}\}counts states and is flat once the window is wide enough to merge the nine arms that sit inside one: it reads11bit from2σ^2\\hat\{\\sigma\}to4σ^4\\hat\{\\sigma\}\.RcapR\_\{\\mathrm\{cap\}\}islog2\\log\_\{2\}of a volume ratio and falls with the window at every step, from5\.305\.30bits to2\.302\.30\. A quantity that moves monotonically with an arbitrary constant is not a rate of anything, which is why the counts are reported as counts\.Figure 16:The count, and the two reasons it is a count rather than a rate\.\(a\)The ten arm endpoints of the cosine family on the share axis, with the2σ^2\\hat\{\\sigma\}window that decides whether two of them are one state drawn at the size it actually is\. Nine arms fall inside a span whose largest internal gap is0\.04310\.0431; the tenth sits0\.3770\.377away,8\.78\.7times that gap\. The separation is not a knife\-edge and does not need the constant to be exactly two\.\(b\)The constant swept\. The number of resolvable states is flat from2σ^2\\hat\{\\sigma\}to4σ^4\\hat\{\\sigma\}, whileRcapR\_\{\\mathrm\{cap\}\}, which islog2\\log\_\{2\}of a volume ratio rather than a count, falls at every step of the sweep\. Sourcesrate\_verdict,rate\_sensitivity\.What the sweep does and does not rescue\. It does not make11bit resolution\-free: at1σ^1\\hat\{\\sigma\}the first family reads1\.581\.58\. It does establish that the count is22on both families over a factor of two in resolution above the chosen constant, that the choice is the*conservative*end of that interval rather than a value picked to produce a round number, and that what changes below it is one arm separating from a cluster rather than the cluster dissolving\. The claim the paper defends is therefore the ordering, that the path family occupies a small number of states against a range that would support ten, and not the specific integer\.
*The design space offers more than anything we ran carried\.*The17221722bits are the entropy of all\(1728864\)\\binom\{1728\}\{864\}orderings under a uniform measure\. Ten were executed, and ten codewords carry at mostlog210=3\.32\\log\_\{2\}10=3\.32bits, so3\.32:13\.32\{:\}1is the largest compression this experiment could exhibit and it is the one it does exhibit: ten paths, two states\. The finding is the*gap*between the ten values the range supports and the two it occupies\. The channel is not narrow because it cannot resolve; it is narrow becausethe encoder is degenerate, which is what Proposition[1](https://arxiv.org/html/2610.00234#Thmproposition1)predicts\. Sources:rate\_verdict,review\_statistics\.
*Why the body reports counts and not bits\.*Writinglog2\\log\_\{2\}of a cluster count invites a channel reading that this measurement does not support\. There is no prior over paths and no noise model, so neither rate is a mutual information; ten paths were run, so no ensemble here carries more thanlog210=3\.32\\log\_\{2\}10=3\.32bits whatever the design space offers; andΦ\\Phiis one scalar, so a degenerate image in one coordinate is not a degenerate image\. The counts and the gap ratios are the measurement\. The logarithms were a naming convention and the body no longer uses them\.
## Appendix CCorrections to Earlier Versions of This Paper
Several statements in this paper replaced earlier ones that were wrong\. The four that moved a number a reader would quote, or withdrew one, are below with their direction; the rest are listed at the end in one sentence each\.
1. 1\.The flat interior was the schedule’s, and the deciding experiment was named before it was run\.An earlier version reported the block\-length ladder’s interior flat and read that as a property of the data path, while naming the account it could not exclude: that the interior is flat because a cosine leaves no step size in the second half\. The constant\-rate ladder has now been run and decides against the reading we preferred, spanning0\.22210\.2221in the interior,11\.6311\.63contrast floors, against0\.04200\.0420and2\.202\.20under the cosine, with an intraclass correlation of0\.8360\.836whose interval excludes zero and an allocation monotone in how blocked the arrangement is \(Table[5](https://arxiv.org/html/2610.00234#S4.T5), branchCL\-2\)\.*The averaging wall as a claim about the path is withdrawn*from the abstract, §[1](https://arxiv.org/html/2610.00234#S1)and §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)\. What replaces it is not weaker: the schedule*is*the averaging operator, and the corner where every result this paper defends is read moves by one contrast floor between the two schedules against seventeen at the rung below it\. The two\-occupied\-states count goes with it: it replicates across pretraining families and does not survive a change of schedule, the constant\-rate ladder occupying four states at the same2σ^2\\hat\{\\sigma\}\.
2. 2\.The headline was a ratio of incommensurable things, and it is withdrawn\.An earlier version led with “17221722bits in,11out, a compression of∼1700×\{\\sim\}1700\\times”\. The numerator is the entropy of all\(1728864\)\\binom\{1728\}\{864\}orderings under a uniform measure; ten orderings were run, and ten codewords carry at mostlog210=3\.32\\log\_\{2\}10=3\.32bits\. The ratio divided a prior over a space against a measurement on a sample of it\. Withdrawn from the abstract, §[1](https://arxiv.org/html/2610.00234#S1), Definition[3](https://arxiv.org/html/2610.00234#Thmdefinition3)and Figure[2](https://arxiv.org/html/2610.00234#S1.F2)\.
3. 3\.Two power statements were made without the design that produced them\.An earlier Limitations paragraph quoted a detectable\-effect figure of5\.795\.79contrast floors, attributed it to the inheritedσ^=0\.0234\\hat\{\\sigma\}=0\.0234, and named no power level; it was computed at the recomputed0\.03150\.0315, on two degrees of freedom rather than four, and at50%50\\%power\. And §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)said the interior span is below anything this design could resolve, when the minimum detectable difference is computed atk=4k\{=\}4and most of the dispersion under it is generation sampling, so atk=32k\{=\}32a span of this size is resolvable\. Both are now scoped to the evaluation that was run\.
4. 4\.Two dispersion figures were reported without what they need to be read\.The between\-arm intraclass correlation was given as “arm identity explains none of the variance”; the point estimate is−0\.076\-0\.076with a bootstrap95%95\\%interval of\[−0\.409,\+0\.166\]\[\-0\.409,\+0\.166\], which supports only the weaker statement now made\. And the commitment fingerprint mixed two binnings, reporting2424–30%30\\%against2\.52\.5–6\.5%6\.5\\%by counting each problem’s best convention on one side and each \(problem, convention\) cell on the other; on one binning throughout the factor is about three and a half rather than seven\.
The remainder, one sentence each\. “Required and not merely observed” over\-read Remark[1](https://arxiv.org/html/2610.00234#Thmremark1), which constrains the optimum and not a finite run’s endpoint, corrected at both sites in §[3](https://arxiv.org/html/2610.00234#S3); that remark was a proposition and proves nothing that needs proving\. The resolution constant’s effect was stated forRcapR\_\{\\mathrm\{cap\}\}and asserted ofRrealR\_\{\\mathrm\{real\}\}, which counts clusters and reads1\.581\.58rather than22at1σ^1\\hat\{\\sigma\}on the first family\. Definition[2](https://arxiv.org/html/2610.00234#Thmdefinition2)’s antecedent said “ssindependent ofpp” where the decomposition needs onlyCov\(p,s\)=0\\mathrm\{Cov\}\(p,s\)=0across problems, and an intermediate version read the measured failure as a refutation of per\-problem conditional independence, which it is not\. The onset region was called unconstructible; it is constructible, and what closes it is the learning non\-equivalence of §[6](https://arxiv.org/html/2610.00234#S6)\. A hand\-derived marked mean of0\.32040\.3204appeared in an early draft of §[5\.1](https://arxiv.org/html/2610.00234#S5.SS1)and is reproduced by no arm or seed subset, so every number in that section is now recovered from the result base by a named analyzer\. The commutator result was described as a disagreement with a published construction; it is a range boundary, and §[2](https://arxiv.org/html/2610.00234#S2)and Appendix[A](https://arxiv.org/html/2610.00234#A1.SSx3)now say which of the two the measurement supports\. The claim that the schedule sets the width of the escape hatch was made in prose and is now Proposition[2](https://arxiv.org/html/2610.00234#Thmproposition2)\. A future\-work item still listed the constant\-rate ladder as unrun after §[4\.3](https://arxiv.org/html/2610.00234#S4.SS3)had reported it, and the abstract stated the schedule result without the provenance §[2](https://arxiv.org/html/2610.00234#S2)gives it\. The registration index corrected half a row, which Appendix[D](https://arxiv.org/html/2610.00234#A4)carries because it is evidence about the index rather than about a result\.
## Appendix DThe Registration Index: Every Threshold, and When It Was Frozen
Registering a read\-out inside the job script the scheduler executes is stronger than a markdown note in one way and weaker in another\. Stronger, because the file carrying the threshold is the file that ran\. Weaker, because a job script is committed when it is written, which is usually but not always before the run\. Table[20](https://arxiv.org/html/2610.00234#A4.T20)carries the gap, measured, rather than the assurance, and Figure[17](https://arxiv.org/html/2610.00234#A4.F17)puts every interval on one axis\.
*It also carries what each registration returned, because an index that lists only thresholds invites the reader to assume they were met\.*Of the seventeen: seven confirmed at their primary branch, two decided at a later one, two came back partial with the unreadable half named, three voided on a guard, and three returned no branch at all\. The branch numbers are per registration and not a grade:CL\-2is this paper’s largest positive result andCM\-4is a void, and nothing in the labels distinguishes them\.
*Registered*is the first commit in which the artefact carries the threshold, found withgit log \-Son a distinctive fragment of the rule, not the first commit of the file\.*Result*is the modification time of the verdict file on the result base\. Both are machine\-checkable and the commands are at the end of this appendix\.
*Two rows are evidenced differently and are marked so\.*TheCMandNIregistrations were frozen on the shared volume the scheduler reads from, and timestamped there, before the jobs that used them were submitted; they reached a commit only afterwards\. A file timestamp on a volume we can write to is weaker evidence than a commit, and the rows say which they have rather than borrowing the strength of the rows above them\. The weakness is not theoretical:CM’s registration was later given an outcome section, and that write replaced the only timestamp its interval rested on, so its\+4\+4h 52m is now an assertion this tree cannot recheck\.NI’s file was left untouched after its verdict for exactly that reason, andverify\_dpd\_registry\.pyrecomputes its interval rather than trusting it\.
Table 20:Seventeen registrations, what each returned, and the interval each one bought\.*Seven confirmed at their primary branch, two decided at another, two came back partial, three void on a guard, and three returned no branch at all\.*Fifteen were frozen from about an hour to nearly three days ahead of the result they decided\. Two are not, and both are stated rather than averaged away: the palindromic check is a linear\-model computation of a few seconds whose thresholds and whose answer entered the repository in the same commit, so it can assert an ordering it cannot demonstrate; and the amplitude\-onset registration was withdrawn before any training because the rungs it named cannot be built at this budget, which is a fact about the ladder rather than about either family\. The second\-family replications carry positive intervals and are nonetheless printed here together with their answers rather than announced in an earlier version, which is the weaker of the two guarantees a registration can offer\. The last two rows’ intervals are measured against a file timestamp on the shared volume rather than against a commit, which is weaker again and is why those cells say*no commit*instead of a hash\. Of the two onlyNIis still recomputable from this tree:CM’s registration gained an outcome section after its verdict, and writing that section overwrote the timestamp its interval was measured from\.NI’s was left byte\-frozen for that reason and its outcome is recorded here and in its verdict file instead\.claimoutcomeartefact carrying the thresholdregisteredresult \(verdict file, mtime\)gapinstrument, the positive controlconfirmedslurm\_natconflict\.sbatch65acb7a08\-01 12:52natconflict08\-03 08:23\+1\+1d 19hv1batch compositionconfirmedslurm\_batchcomp\.sbatch,decision rule2e5f5a908\-02 12:42batchcomp08\-02 15:32\+2\+2h 50mv2and the ladder*no branch fired*slurm\_pathspec\.sbatchandPLAN\.md:*predicts*L⋆≈10L^\{\\star\}\\\!\\approx\\\!10cd53c0908\-02 16:02pathspec08\-03 02:33\+10\+10h 30mM1the mirrorM1confirmedslurm\_dpd\_mirror\.sbatch:*an asymmetry falsifies the limit\-cycle picture*a85c24708\-03 15:32dpd\_mirror08\-06 07:44\+2\+2d 16hC,Lthe keyC1,L4confirmedslurm\_disambig\.sbatch, C on the switch and L on the level4830e3708\-04 06:46disambig08\-04 10:49\+4\+4h 03mT1the mixture ratioT1confirmedslurm\_ratio2\.sbatch, tolerance0\.050\.05frozen with the predictionff94a9508\-05 06:22ratio208\-05 16:20\+9\+9h 58mS\-Cthe single\-cosine controlS\-C4voidprereg/single\_cosine\_control\.md, four branches, S\-C4 first39e106d08\-08 08:41singlecos08\-08 16:53\+8\+8h 12mS1–S3the palindromic schedule*asserted, not demonstrated*verify\_splitting\.pydocstringe1eadef08\-02 11:07same commitnoneQ3\-onsetwhere the amplitude turns on*withdrawn*prereg/q3\_amplitude\_onset\.md,*withdrawn before any training*: the rungs it names lie across the epoch boundary, and the only route there is the tripled file the S\-C control measured as not learning\-matched \(§[5](https://arxiv.org/html/2610.00234#S5)\)79b05eb08\-11 01:56*none possible*—Q3\-keythe key on a second familyK1,L4,V1confirmedprereg/q3\_key\_replication\.mdwithreadout\_q3unlock\.py66eab0408\-09 14:25q3unlock08\-11 01:37\+1\+1d 11hQ3\-wallthe ladder on a second familyW2,M2partialprereg/q3\_ladder\_replication\.mdwithreadout\_q3ladder\.pya679a1b08\-09 15:41q3ladder08\-10 05:04\+13\+13h 23mSCthe schedule controlSC\-3decidedprereg/schedule\_control\.mdwithreadout\_schedctl\.py, guard G\-SC first, primary rung fixed atL=27L\{=\}27a8bcd8008\-13 00:19schedctl08\-14 02:03\+1\+1d 1hBMthe mirror ofblockedBM\-2partialprereg/blocked\_mirror\.mdwithreadout\_blocked\_mirror\.py, comparability control read before the branch3fcc1a008\-16 12:16blocked\_mirror08\-16 15:06\+2\+2h 50mCLthe ladder at a constant rateCL\-2decidedprereg/constant\_lr\_ladder\.md, four branches on the interior span against the cosine ladder’sa67e3a108\-17 06:05constlr\_ladder08\-17 14:25\+8\+8h 20mHKresolution at higherkkHK\-4voidprereg/highk\_resolution\.mdwithreadout\_highk\.py, guards G\-HK1–3 read before HK\-1–31e1daa008\-17 14:37highk08\-18 13:06\+22\+22h 29mCMthe mirror at a constant rateCM\-4voidprereg/constant\_rate\_mirror\.mdwithreadout\_constant\_mirror\.py, guardCM\-4read first, magnitude registered as exploratory and turning no branch*no commit*08\-26 05:43constant\_mirror08\-26 10:35\+4\+4h 52mNIwhethernat’s weakness is a budget artefactNI\-1confirmedprereg/nat\_budget\_gate\.mdwithreadout\_nat\_budget\.py, guardsG\-1–G\-3and a capability bar computed from the family’s own dispersion before the run*no commit*08\-26 15:28nat\_budget08\-26 16:37\+1\+1h 08mhow long before the result the threshold was frozen1 h3 h10 h1 day3 daysinterval between the freezing commit and the result file \(log\)M1the mirror\+2\+2d 16hthe instrument’s positive control\+1\+1d 19hQ3\-keythe key, second family\+1\+1d 11hSCthe schedule control\+1\+1d 1hHKresolution at higherkk\+22\+22h 29mQ3\-wallthe ladder, second family\+13\+13h 23mv2and the ladder\+10\+10h 30mT1the mixture ratio\+9\+9h 58mCLthe ladder at a constant rate\+8\+8h 20mS\-Cthe single\-cosine control\+8\+8h 12mCMthe mirror at a constant rate\+4\+4h 52mC,Lthe key\+4\+4h 03mv1batch composition\+2\+2h 50mBMthe mirror ofblocked\+2\+2h 50mNIwhethernatis budget\-limited\+1\+1h 08mS1–S3the palindromic schedulethreshold and answer in the*same commit*Q3\-onsetwhere the amplitude turns on*withdrawn before any training*: no result is possibleFigure 17:Seventeen registrations, and the interval each one bought\.Table[20](https://arxiv.org/html/2610.00234#A4.T20)in one axis: the distance between the commit that froze a threshold and the file that answered it\. Fifteen run from about an hour to nearly three days ahead of their result\. The two that are not intervals are drawn on the same axis rather than dropped from the count: the palindromic check put its thresholds and its answer in one commit, so it can assert an ordering it cannot demonstrate, and the amplitude\-onset registration was withdrawn before any training because the rungs it named cannot be built at this budget\. Sourcetab\_registry, whose gap column is computed from the commit and the verdict file’s mtime\.#### One correction this index made to our own record\.
An earlier version of it named the half\-mix job as the ladder’s registering artefact\. That is wrong: the half\-mix verdict is the volume\-matched hierarchy result, a different experiment\. The ladder’s thresholds live in the path\-spectroscopy job beside v2’s, which is correct, because the block\-length sweep*is*v2’s deciding test\.
*And the correction was applied to half the row\.*The artefact column was changed and the result column was not, so the row named the path\-spectroscopy job beside the half\-mix era’s timestamp,08\-03 15:43, which ismixture\_trajectory\_verdict’s mtime and notpathspec\_verdict’s\. The true interval is\+10\+10h 30m rather than the\+23\+23h 41m we printed\. Both are positive and neither changes a verdict, but the number was wrong and it was wrong in the direction that flattered us\. It was found by a script that rebuilds this whole table fromgitand the result base without reading it \(verify\_dpd\_registry\.py, released with the paper\), which is also why the*result*column now names its verdict file: the old column printed a bare timestamp, so a row could carry a corrected artefact and an uncorrected result and look consistent\. We report this because an index nobody audits is a claim rather than a check, and this index was not audited until it was\.
#### How to reproduce any row\.
Two commands, one per column:
git log \-S’<fragment\>’ \-\-format=’%h %ad’ \-\- <artefact\> \| tail \-1stat \-c ’%y %n’ results/<verdict\>\.json
## References
- \[1\]K\. Luo, Z\. Sun, H\. Wen, X\. Shi, J\. Cui, C\. Dang, K\. Lyu, and W\. Chen\(2025\)How learning rate decay wastes your best data in curriculum\-based LLM pretraining\.arXiv preprint\.Note:arXiv:2511\.18903Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1),[Abstract](https://arxiv.org/html/2610.00234#abstract1.2)\.
- \[2\]W\. Chen, J\. Chen, Z\. Lin, and C\. M\. Vong\(2026\)The capability convergence hypothesis: capability from access structure, not scale\.arXiv preprint\.Note:arXiv:2607\.14144External Links:[Document](https://dx.doi.org/10.5281/zenodo.21714418),[Link](https://github.com/wenhui-ml/Capability-Convergence-Hypothesis)Cited by:[§1](https://arxiv.org/html/2610.00234#S1.p5.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px10.p1.1),[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.SSS0.Px1.p5.1),[Table 14](https://arxiv.org/html/2610.00234#S7.T14)\.
- \[3\]N\. N\. Bogoliubov and Y\. A\. Mitropolsky\(1961\)Asymptotic methods in the theory of non\-linear oscillations\.Gordon and Breach\.Cited by:[§1](https://arxiv.org/html/2610.00234#S1.SS0.SSS0.Px1.p2.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1),[§3](https://arxiv.org/html/2610.00234#S3.SS0.SSS0.Px1.p12.1),[§3](https://arxiv.org/html/2610.00234#S3.SS0.SSS0.Px1.p4.1.1)\.
- \[4\]P\. L\. Kapitza\(1951\)Dynamic stability of a pendulum with an oscillating point of suspension\.Journal of Experimental and Theoretical Physics21,pp\. 588–597\.Cited by:[§1](https://arxiv.org/html/2610.00234#S1.SS0.SSS0.Px1.p2.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[5\]W\. Chen, J\. Chen, Z\. Lin, and C\. M\. Vong\(2026\)The free\-recipe limit: every measured recipe effect is a gauge of one broken premise of the ideal\.Note:Code and artefacts archived; every number cited here is checkable thereExternal Links:[Document](https://dx.doi.org/10.5281/zenodo.21905881),[Link](https://github.com/wenhui-ml/free-recipe-limit)Cited by:[§1](https://arxiv.org/html/2610.00234#S1.SS0.SSS0.Px5.p2.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px10.p1.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2610.00234#S4.SS2.p3.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px6.p7.1),[§6](https://arxiv.org/html/2610.00234#S6.SS0.SSS0.Px3.p1.1),[§7](https://arxiv.org/html/2610.00234#S7.SS0.SSS0.Px3.p1.1),[Table 14](https://arxiv.org/html/2610.00234#S7.T14),[Table 14](https://arxiv.org/html/2610.00234#S7.T14.12.9.2.1.1),[§7](https://arxiv.org/html/2610.00234#S7.p2.1)\.
- \[6\]R\. Xu, Z\. Qi, Z\. Guo, C\. Wang, H\. Wang, Y\. Zhang, and W\. Xu\(2024\)Knowledge conflicts for LLMs: a survey\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Note:arXiv:2403\.08319Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p1.1)\.
- \[7\]e\. al\. Xue\(2026\)Why supervised fine\-tuning fails to learn: a systematic study of incomplete learning in large language models\.arXiv preprint\.Note:arXiv:2604\.10079Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[8\]K\. Krestnikov\(2026\)Truth as a compression artifact in language model training\.arXiv preprint arXiv:2603\.11749\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p1.1)\.
- \[9\]B\. Plank\(2022\)The “problem” of human label variation: on ground truth in data, modeling and evaluation\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Note:arXiv:2211\.02570Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p2.1)\.
- \[10\]R\. Schaeffer, B\. Miranda, and S\. Koyejo\(2023\)Are emergent abilities of large language models a mirage?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p2.1)\.
- \[11\]F\. Kang, M\. Kuchnik, K\. Padthe, M\. Vlastelica, R\. Jia, C\. Wu, and N\. Ardalani\(2025\)Quagmires in SFT\-RL post\-training: when high SFT scores mislead and what to use instead\.arXiv preprint\.Note:arXiv:2510\.01624Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p2.1)\.
- \[12\]A\. Holtzman, P\. West, V\. Shwartz, Y\. Choi, and L\. Zettlemoyer\(2021\)Surface form competition: why the highest probability answer isn’t always right\.InProceedings of EMNLP,Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p3.1)\.
- \[13\]J\. Yeom, J\. Sok, H\. Kim, S\. Park, J\. Park, and T\. Kim\(2026\)Hallucination as commitment failure: larger LLMs misfire despite knowing the answer\.arXiv preprint\.Note:arXiv:2605\.22007Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p3.1)\.
- \[14\]J\. M\. Janeiro, M\. Videau, A\. Caciolai, B\. Piwowarski, P\. Gallinari, and L\. Barrault\(2026\)Are we evaluating knowledge or phrasing? mitigating MCQA sensitivity with ParaEval\.arXiv preprint arXiv:2606\.10657\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px1.p3.1)\.
- \[15\]D\. P\. Bertsekas\(2011\)Incremental gradient, subgradient, and proximal methods for convex optimization: a survey\.InOptimization for Machine Learning,pp\. 85–119\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[16\]A\. Nedić and D\. P\. Bertsekas\(2001\)Incremental subgradient methods for nondifferentiable optimization\.SIAM Journal on Optimization12\(1\),pp\. 109–138\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[17\]M\. Gürbüzbalaban, A\. Ozdaglar, and P\. A\. Parrilo\(2021\)Why random reshuffling beats stochastic gradient descent\.Mathematical Programming186,pp\. 49–84\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[18\]J\. Z\. HaoChen and S\. Sra\(2019\)Random shuffling beats SGD after finite epochs\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[19\]K\. Mishchenko, A\. Khaled, and P\. Richtárik\(2020\)Random reshuffling: simple analysis with vast improvements\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2006\.05988Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[20\]Y\. Lu, W\. Guo, and C\. De Sa\(2022\)GraB: finding provably better data permutations than random reshuffling\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2205\.10733Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px1.p1.1)\.
- \[21\]K\. Emmanouilidis, E\. Vlatakis\-Gkaragkounis, and R\. Vidal\(2026\)Shuffling the data, stretching the step\-size: sharper bias in constant step\-size SGD\.arXiv preprint\.Note:arXiv:2604\.10373Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px2.p1.1)\.
- \[22\]J\. Rao, X\. Liu, L\. Lian, S\. Cheng, Y\. Liao, and M\. Zhang\(2024\)CommonIT: commonality\-aware instruction tuning for large language models via data partitions\.InProceedings of EMNLP,Note:arXiv:2410\.03077Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px1.p1.1)\.
- \[23\]Y\. Dai, Y\. Huang, T\. Yang, Y\. Wang, X\. Zhang, W\. Wu, Q\. Zhao, H\. Li, Y\. Gao, K\. Yap, and S\. Li\(2026\)Demystifying data organization for enhanced LLM training\.InProceedings of ACL,Note:arXiv:2605\.30334Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px1.p1.1)\.
- \[24\]J\. Sweeney\(2026\)Optimizer memory makes shuffle order a first\-order source of fine\-tuning noise\.arXiv preprint\.Note:arXiv:2606\.29554Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px2.p1.1),[§7](https://arxiv.org/html/2610.00234#S7.SS0.SSS0.Px4.p1.1)\.
- \[25\]J\. Sweeney\(2026\)The geometry of sequential learning: Lie\-bracket prediction of transfer order\.InInternational Conference on Machine Learning \(ICML\),Note:arXiv:2606\.24993Cited by:[Appendix A](https://arxiv.org/html/2610.00234#A1.SSx3.p3.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1)\.
- \[26\]T\. Yu, S\. Kumar, A\. Gupta, S\. Levine, K\. Hausman, and C\. Finn\(2020\)Gradient surgery for multi\-task learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1)\.
- \[27\]B\. Liu, X\. Liu, X\. Jin, P\. Stone, and Q\. Liu\(2021\)Conflict\-averse gradient descent for multi\-task learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1)\.
- \[28\]A\. Navon, A\. Shamsian, I\. Achituve, H\. Maron, K\. Kawaguchi, G\. Chechik, and E\. Fetaya\(2022\)Multi\-task learning as a bargaining game\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px3.p1.1)\.
- \[29\]S\. M\. Xie, H\. Pham, X\. Dong, N\. Du, H\. Liu, Y\. Lu, P\. Liang, Q\. V\. Le, T\. Ma, and A\. W\. Yu\(2023\)DoReMi: optimizing data mixtures speeds up language model pretraining\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2305\.10429Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px4.p1.1)\.
- \[30\]S\. Fan, M\. Pagliardini, and M\. Jaggi\(2024\)DoGE: domain reweighting with generalization estimation\.arXiv preprint\.Note:arXiv:2310\.15393Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px4.p1.1)\.
- \[31\]X\. Gu, K\. Lyu, J\. Li, and J\. Zhang\(2025\)Data mixing can induce phase transitions in knowledge acquisition\.arXiv preprint arXiv:2505\.18091\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px4.p1.1)\.
- \[32\]K\. Mo, Y\. Shi, W\. Weng, Z\. Zhou, S\. Liu, H\. Zhang, and A\. Zeng\(2025\)Mid\-training of large language models: a survey\.arXiv preprint\.Note:arXiv:2510\.06826Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px4.p1.1)\.
- \[33\]M\. McCloskey and N\. J\. Cohen\(1989\)Catastrophic interference in connectionist networks: the sequential learning problem\.Psychology of Learning and Motivation24,pp\. 109–165\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[34\]R\. M\. French\(1999\)Catastrophic forgetting in connectionist networks\.Trends in Cognitive Sciences3\(4\),pp\. 128–135\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[35\]J\. Kirkpatrick, R\. Pascanu, N\. Rabinowitz,et al\.\(2017\)Overcoming catastrophic forgetting in neural networks\.Proceedings of the National Academy of Sciences114\(13\)\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[36\]D\. Lopez\-Paz and M\. Ranzato\(2017\)Gradient episodic memory for continual learning\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:1706\.08840Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1),[§7](https://arxiv.org/html/2610.00234#S7.SS0.SSS0.Px6.p1.1)\.
- \[37\]S\. Lee, S\. Goldt, and A\. Saxe\(2021\)Continual learning in the teacher\-student setup: impact of task similarity\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[38\]T\. Wu, L\. Luo, Y\. Li, S\. Pan, T\. Vu, and G\. Haffari\(2024\)Continual learning for large language models: a survey\.arXiv preprint\.Note:arXiv:2402\.01364Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[39\]N\. Díaz\-Rodríguez, V\. Lomonaco, D\. Filliat, and D\. Maltoni\(2018\)Don’t forget, there is more than forgetting: new metrics for continual learning\.arXiv preprint\.Note:arXiv:1810\.13166Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[40\]I\. Evron, E\. Moroshko, R\. Ward, N\. Srebro, and D\. Soudry\(2022\)How catastrophic can catastrophic forgetting be in linear regression?\.InConference on Learning Theory \(COLT\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[41\]D\. Krasheninnikov, R\. E\. Turner, and D\. Krueger\(2025\)Fresh in memory: training\-order recency is linearly encoded in language model activations\.arXiv preprint\.Note:arXiv:2509\.14223Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[42\]H\. C\. Conklin, T\. Hosking, T\. Yi\-Chern, J\. Gold, J\. D\. Cohen, T\. L\. Griffiths, M\. Bartolo, and S\. Goldfarb\-Tarrant\(2026\)Learning is forgetting: LLM training as lossy compression\.arXiv preprint\.Note:arXiv:2604\.07569Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[43\]Gustav Olaf Yunus Laitinen\-Fredriksson Lundstrom\-Imanov\(2026\)Mechanistic analysis of catastrophic forgetting in large language models during continual fine\-tuning\.arXiv preprint\.Note:arXiv:2601\.18699Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px5.p1.1)\.
- \[44\]Y\. Bengio, J\. Louradour, R\. Collobert, and J\. Weston\(2009\)Curriculum learning\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[45\]D\. Rohrer and K\. Taylor\(2007\)The shuffling of mathematics problems improves learning\.Instructional Science35\(6\),pp\. 481–498\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[46\]N\. Kornell and R\. A\. Bjork\(2008\)Learning concepts and categories: is spacing the “enemy of induction”?\.Psychological Science19\(6\),pp\. 585–592\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[47\]H\. Say, S\. E\. Ada, E\. Ugur, M\. Asada, and E\. Oztop\(2025\)Interleaved multitask learning with energy modulated learning progress\.arXiv preprint\.Note:arXiv:2504\.00707Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[48\]Y\. Jia, C\. Zhang, X\. Diao, X\. Yuan, Z\. Ouyang, C\. Ma, and S\. Vosoughi\(2025\)What makes a good curriculum? disentangling the effects of data ordering on LLM mathematical reasoning\.arXiv preprint\.Note:arXiv:2510\.19099Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[49\]Y\. Zhang, A\. Mohamed, H\. Abdine, G\. Shang, and M\. Vazirgiannis\(2025\)Beyond random sampling: efficient language model pretraining via curriculum learning\.arXiv preprint\.Note:arXiv:2506\.11300Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[50\]M\. Elgaar and H\. Amiri\(2026\)Curriculum learning for LLM pretraining: an analysis of learning dynamics\.arXiv preprint\.Note:arXiv:2601\.21698Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[51\]E\. Liu, K\. Sun, M\. Li, I\. Lee, L\. Tjuatja, J\. Huang, and G\. Neubig\(2026\)What do language models learn and when? the implicit curriculum hypothesis\.arXiv preprint\.Note:arXiv:2604\.08510Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[52\]I\. Piontkovskaia and S\. Nikolenko\(2026\)First\-order predictable but pairwise fragile: local task adaptation in trained transformers\.arXiv preprint\.Note:arXiv:2607\.16821Cited by:[Appendix A](https://arxiv.org/html/2610.00234#A1.SSx1.p3.1),[Appendix A](https://arxiv.org/html/2610.00234#A1.SSx3.p3.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px6.p1.1)\.
- \[53\]J\. LeDoux\(2026\)The order is the message\.arXiv preprint arXiv:2603\.25047\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px7.p1.1),[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px7.p2.1)\.
- \[54\]Y\. Ju, Z\. Ni, X\. Xing, Z\. Zeng, H\. Zhao, S\. Fan, and Z\. Zhang\(2024\)Mitigating training imbalance in LLM fine\-tuning via selective parameter merging\.arXiv preprint arXiv:2410\.03743\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px8.p1.1)\.
- \[55\]A\. Jacot, F\. Gabriel, and C\. Hongler\(2018\)Neural tangent kernel: convergence and generalization in neural networks\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[56\]L\. Chizat, E\. Oyallon, and F\. Bach\(2019\)On lazy training in differentiable programming\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[57\]N\. Ajroldi, A\. Orvieto, and J\. Geiping\(2025\)When, where and why to average weights?\.arXiv preprint\.Note:arXiv:2502\.06761Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[58\]A\. Power, Y\. Burda, H\. Edwards, I\. Babuschkin, and V\. Misra\(2022\)Grokking: generalization beyond overfitting on small algorithmic datasets\.arXiv preprint arXiv:2201\.02177\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[59\]M\. Belkin, D\. Hsu, S\. Ma, and S\. Mandal\(2019\)Reconciling modern machine\-learning practice and the classical bias–variance trade\-off\.Proceedings of the National Academy of Sciences116\(32\),pp\. 15849–15854\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[60\]P\. Nakkiran, G\. Kaplun, Y\. Bansal, T\. Yang, B\. Barak, and I\. Sutskever\(2020\)Deep double descent: where bigger models and more data hurt\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[61\]G\. Pruthi, F\. Liu, S\. Kale, and M\. Sundararajan\(2020\)Estimating training data influence by tracing gradient descent\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[62\]J\. H\. Lee, M\. Smith, M\. Adam, and J\. Hoogland\(2025\)Influence dynamics and stagewise data attribution\.arXiv preprint\.Note:arXiv:2510\.12071Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[63\]N\. Tishby and N\. Zaslavsky\(2015\)Deep learning and the information bottleneck principle\.IEEE Information Theory Workshop \(ITW\)\.Cited by:[§2](https://arxiv.org/html/2610.00234#S2.SS0.SSS0.Px9.p1.1)\.
- \[64\]E\. Hairer, C\. Lubich, and G\. Wanner\(2006\)Geometric numerical integration: structure\-preserving algorithms for ordinary differential equations\.2nd edition,Springer\.Cited by:[§3](https://arxiv.org/html/2610.00234#S3.SS0.SSS0.Px1.p12.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1)\.
- \[65\]D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt\(2021\)Measuring mathematical problem solving with the math dataset\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track \(NeurIPS\),Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px1.p1.1)\.
- \[66\]Hugging Face\(2025\)OpenR1\-Math\-220k\.Note:[https://huggingface\.co/datasets/open\-r1/OpenR1\-Math\-220k](https://huggingface.co/datasets/open-r1/OpenR1-Math-220k)Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px1.p1.1)\.
- \[67\]W\. Chen\(2026\)Which corpus supplies the gold: a free variable in exact\-match evaluation\.Note:Companion manuscript\. The measurement half of the underdetermined stream: the corpus survey, the paired read\-out, and the leaderboard inversionsCited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2610.00234#S7.SS0.SSS0.Px3.p1.1)\.
- \[68\]Qwen Team\(2024\)Qwen2\.5 technical report\.arXiv preprint\.Note:arXiv:2412\.15115Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px2.p1.1)\.
- \[69\]Y\. Zhao, J\. Huang, J\. Hu, X\. Wang,et al\.\(2025\)SWIFT: a scalable lightweight infrastructure for fine\-tuning\.InProceedings of the AAAI Conference on Artificial Intelligence,Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px2.p1.1)\.
- \[70\]W\. Kwon, Z\. Li, S\. Zhuang,et al\.\(2023\)Efficient memory management for large language model serving with PagedAttention\.InSymposium on Operating Systems Principles \(SOSP\),Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px2.p1.1)\.
- \[71\]P\. E\. Shrout and J\. L\. Fleiss\(1979\)Intraclass correlations: uses in assessing rater reliability\.Psychological Bulletin86\(2\),pp\. 420–428\.Cited by:[§4\.1](https://arxiv.org/html/2610.00234#S4.SS1.SSS0.Px4.p3.1)\.
- \[72\]H\. F\. Trotter\(1959\)On the product of semi\-groups of operators\.Proceedings of the American Mathematical Society10\(4\),pp\. 545–551\.Cited by:[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1)\.
- \[73\]R\. I\. McLachlan and G\. R\. W\. Quispel\(2002\)Splitting methods\.Acta Numerica11,pp\. 341–434\.Cited by:[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1)\.
- \[74\]G\. Strang\(1968\)On the construction and comparison of difference schemes\.SIAM Journal on Numerical Analysis5\(3\),pp\. 506–517\.Cited by:[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1)\.
- \[75\]S\. Blanes, F\. Casas, J\. A\. Oteo, and J\. Ros\(2009\)The Magnus expansion and some of its applications\.Physics Reports470\(5–6\),pp\. 151–238\.Cited by:[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1)\.
- \[76\]L\. M\. Nguyen, D\. T\. Phan, and J\. Kalagnanam\(2026\)Learning to shuffle: block reshuffling and reversal schemes for stochastic optimization\.arXiv preprint\.Note:arXiv:2604\.00260Cited by:[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p1.1),[§4\.3](https://arxiv.org/html/2610.00234#S4.SS3.SSS0.Px8.p4.1)\.
- \[77\]N\. S\. Keskar, B\. McCann, L\. R\. Varshney, C\. Xiong, and R\. Socher\(2019\)CTRL: a conditional transformer language model for controllable generation\.arXiv preprint arXiv:1909\.05858\.Cited by:[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.p2.1)\.
- \[78\]T\. Korbak, K\. Shi, A\. Chen, R\. Bhalerao, C\. L\. Buckley, J\. Phang, S\. R\. Bowman, and E\. Perez\(2023\)Pretraining language models with human preferences\.InProceedings of the 40th International Conference on Machine Learning \(ICML\),Cited by:[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.p2.1)\.
- \[79\]T\. Gao, A\. Wettig, L\. He, Y\. Dong, S\. Malladi, and D\. Chen\(2025\)Metadata conditioning accelerates language model pre\-training\.arXiv preprint arXiv:2501\.01956\.Cited by:[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.p2.1)\.
- \[80\]M\. Khalifa, D\. Wadden, E\. Strubell, H\. Lee, L\. Wang, I\. Beltagy, and H\. Peng\(2024\)Source\-aware training enables knowledge attribution in language models\.InConference on Language Modeling \(COLM\),Cited by:[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.p2.1)\.
- \[81\]R\. Higuchi, R\. Kawata, N\. Nishikawa, K\. Oko, S\. Yamaguchi, S\. Kobayashi, S\. Tokui, K\. Hayashi, D\. Okanohara, and T\. Suzuki\(2025\)When does metadata conditioning \(NOT\) work for language model pre\-training? a study with context\-free grammars\.arXiv preprint arXiv:2504\.17562\.Cited by:[§5\.1](https://arxiv.org/html/2610.00234#S5.SS1.p2.1)\.
- \[82\]W\. Chen\(2026\)Writing a convention into weights: a dose response in marker reliability, and one undecided capacity row\.Note:Companion manuscript\. The training\-side half: the key’s information against its presence, the dose law, and the capacity row its corpus could not decideCited by:[§6](https://arxiv.org/html/2610.00234#S6.SS0.SSS0.Px4.p10.1)\.相似文章
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
This paper proves that using error-penalized scoring rules with abstention as a discrete action can kill both the reward gradient and the KL anchor, causing models to collapse toward refusing everything. It proposes a structural repair — training a mandatory confidence report — and validates the mechanism with simulations and language model experiments.
低方差奖励下群组相对优化中的优势尺度校准不平衡:诊断与有界恢复
本文针对低方差奖励下群组相对优化中的优势尺度校准问题,提出了如Reward-Resolution Protocol和MaxNorm-AC等方法,以过滤亚分辨率抖动并为可信差距提供有界恢复。
相同损失,相同噪声,相反调度:噪声结构与优化器归一化共同决定学习率冷却是否有益
本文从理论上证明,在WSD调度中学习率冷却是否有益取决于梯度噪声的结构以及优化器是否对其更新进行归一化,从而解释了为什么冷却对SGD无效但对归一化方法却是必要的。
梯度冲突能预测理解—生成权衡吗?统一多模态模型中冲突指标有效性的受控审计
本文审计了梯度冲突指标(余弦相似度、冲突率)是否真正能够预测统一多模态模型中的理解—生成权衡。作者构建了一个受控测试平台(GridUMM),在其中真实权衡是可计算的。在涵盖 63 种配置、372 个检查点的实验中,没有任何方向性冲突指标能可靠地与权衡相关;通过剂量反应干预抑制冲突后,权衡表现依然不变,而 eff_rank 和训练损失在诊断能力上优于冲突几何指标。
当预处理指数转为负数:学习率耦合与跨环境泛化
本文探讨了自适应优化器中预处理指数与学习率的耦合,表明最大化跨环境准确率的指数随对数学习率降低,并发现源域验证与环境变化鲁棒性之间存在冲突。