A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
Summary
The paper demonstrates that using a shared learning rate as a control in selective on-policy distillation experiments is not neutral, leading to varying performance and conclusions, and advocates for reporting full learning-rate matrix comparisons for fair evaluation.
View Cached Full Text
Cached at: 09/22/26, 09:11 AM
# A Shared Learning Rate Is Not a Neutral Controlin Selective On-Policy Distillation
Source: [https://arxiv.org/html/2609.22109](https://arxiv.org/html/2609.22109)
###### Abstract
Selective on\-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate—a control chosen to be neutral\. We show it is not\. Under LoRA adaptation, across a8×8\\timeslearning\-rate grid on GSM8K with a Qwen2\.5\-1\.5B student and 7B teacher, dense supervision is statistically flat \(swing1\.8pp1\.8\\,\\text\{pp\},p=0\.26p\{=\}0\.26\) while every selective arm we test moves with the rate:5\.4pp5\.4\\,\\text\{pp\}for a*random*5%5\\%subset,6\.7pp6\.7\\,\\text\{pp\}for a total\-variation selector,11\.711\.7–17\.7pp17\.7\\,\\text\{pp\}for a teachability selector\. The asymmetry has a direct consequence for how these methods are compared: the dense\-versus\-selective verdict reads10\.1pp10\.1\\,\\text\{pp\}atη=10−4\\eta\{=\}10^\{\-4\}and5\.1pp5\.1\\,\\text\{pp\}at5⋅10−55\{\\cdot\}10^\{\-5\}—the same comparison, differing by2\.0×2\.0\\times, decided by a parameter the protocol treats as scenery\. Among the selectors themselves, rankings stay stable in our setting but two of six pairwise significance calls flip between adjacent rates—the protocol changes what a paper concludes without any rank inversion\. We call this*selector–rate entanglement*and trace it to selection itself rather than to step size: AdamW update magnitudes track the rate to within2\.2%2\.2\\%despite15\.5×15\.5\\timesgradient\-norm differences across arms\. A frozen\-scoring ablation—selection scored by the initial student, with criterion, budget, and on\-policy rollouts unchanged—isolates how much of the entanglement comes from selection reading the model it is training: in a preregistered test at 12 seeds per cell, live scoring adds3\.79±1\.69pp3\.79\\pm 1\.69\\,\\text\{pp\}of rate sensitivity \(p=0\.035p\{=\}0\.035\) and produces a strictly separated selection\-drift trajectory, while the frozen arm remains significantly entangled itself \(p=0\.015p\{=\}0\.015\)—the loop aggravates the phenomenon rather than causing it\. The added sensitivity comes from the cool end of the grid, where live scoring is2\.98pp2\.98\\,\\text\{pp\}*better*\(p=0\.001p\{=\}0\.001\) rather than worse, so the very same ablation reads as harmful or beneficial depending on which single rate an experimenter fixes\. Under full fine\-tuning at the rates this literature actually uses \(10−610^\{\-6\}–10−510^\{\-5\}\), the pattern survives in graded form and grows: dense itself swings19\.8pp19\.8\\,\\text\{pp\}, the selective arm49\.5pp49\.5\\,\\text\{pp\}\(2\.5×2\.5\\timesmore\), and the dense\-versus\-selective verdict ranges from a non\-significant\+3\.6pp\+3\.6\\,\\text\{pp\}at2⋅10−62\{\\cdot\}10^\{\-6\}—the published operating point—to\+34pp\+34\\,\\text\{pp\}\(p=0\.005p\{=\}0\.005\) one notch hotter\. On MATH\-500 the rate dependence does not reproduce under LoRA, scoping that result, while the cost of selective training there does \(∼10pp\{\\sim\}10\\,\\text\{pp\}\)\. We prescribe reporting the arm×\\timesrate matrix, not a shared\-rate column, as a precondition for selector comparisons\.
## 1Introduction
Figure 1:Dense supervision is flat across the learning\-rate grid; every selective arm is not\.Accuracy change over the untrained student \(GSM8K, mean±\\pmseed SD; dots are seeds\) versus learning rate\. Swings: dense1\.8pp1\.8\\,\\text\{pp\}\(t=1\.40t\{=\}1\.40, n\.s\.\), random5%5\\%subset5\.45\.4, frozen\-scored TV2\.92\.9\(p=0\.015p\{=\}0\.015\), live\-scored TV6\.76\.7\(p<0\.001p\{<\}0\.001\)\. The consequence for comparisons: the dense\-versus\-selective verdict is10\.1pp10\.1\\,\\text\{pp\}atη=10−4\\eta\{=\}10^\{\-4\}and5\.1pp5\.1\\,\\text\{pp\}at5⋅10−55\{\\cdot\}10^\{\-5\}—a factor of2\.02\.0decided by a parameter the protocol treats as neutral\.To compare token selectors, the selective\-distillation literature holds everything else fixed: same data, same teacher, same token budget, and one shared learning rate\([Wang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib17);[Xu et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib18);[Jiang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib7);[Huang et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib6);[Koo et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib8)\)\. The protocol embodies an assumption so natural it is never stated: that a selector is a static filter—an exogenous lens held up to the data—so that whatever separates two arms trained under identical hyperparameters is the selectors’ contribution\.
The assumption has a testable consequence: if the learning rate is neutral, then arms should respond to it alike, and the ranking between them should be stable across the grid\. Figure[1](https://arxiv.org/html/2609.22109#S1.F1)shows that they do not and it is not\. Under LoRA adaptation, across an8×8\\timesgrid, dense supervision is statistically flat \(swing1\.8pp1\.8\\,\\text\{pp\},t=1\.40t\{=\}1\.40\) while every selective arm moves: a*random*5%5\\%subset swings5\.4pp5\.4\\,\\text\{pp\}, a total\-variation selector6\.7pp6\.7\\,\\text\{pp\}, a teachability selector11\.711\.7–17\.7pp17\.7\\,\\text\{pp\}\. Restricting the loss support—even to a subset chosen without looking at anything—makes the outcome a function of the learning rate in a way dense training is not\. We call this asymmetry*selector–rate entanglement*\. It is not a LoRA artifact: under full fine\-tuning at the rates this literature actually publishes with, everything becomes rate\-sensitive and the asymmetry survives in graded form—the selective arm swings2\.5×2\.5\\timesmore than dense, and the gap between them moves from statistically invisible at the published rates to34pp34\\,\\text\{pp\}one notch hotter \(Section[5\.5](https://arxiv.org/html/2609.22109#S5.SS5)\)\.
The consequence for the comparison protocol is direct\. The dense\-versus\-selective verdict reads10\.1pp10\.1\\,\\text\{pp\}atη=10−4\\eta\{=\}10^\{\-4\}and5\.1pp5\.1\\,\\text\{pp\}at5⋅10−55\{\\cdot\}10^\{\-5\}: the same two arms, the same data, the same budget, a factor of2\.02\.0between the two answers, and nothing to choose between the rates except convention\. A comparison run at one shared rate does not report a property of the selector; it reports a property of the selector*at that rate*, and the literature does not currently vary the rate \(Table[6](https://arxiv.org/html/2609.22109#S6.T6)\)\.
An obvious mundane explanation must die first\. Selective losses have very different gradient scales, so perhaps the arms simply take different effective steps at the same nominal rate\. They do not: gradient norms differ by15\.5×15\.5\\timesacross arms, yet the AdamW update magnitude per unit rate is constant to within2\.2%2\.2\\%\(Section[4](https://arxiv.org/html/2609.22109#S4)\)\. The arms take equal\-sized steps in different directions; no rescaling of the rate can repair the comparison\.
Where does the asymmetry come from? Selection in these methods is recomputed from the*current*student at every rollout refresh—a selection frozen at initialization would be a different method—so the selector reads the variable it is helping to move\. Section[5](https://arxiv.org/html/2609.22109#S5)tests whether that dependence matters by freezing the scoring model atθ0\\theta\_\{0\}while holding criterion, budget, and on\-policy rollouts fixed\. It changes the*dynamics of selection*unambiguously: frozen\-scored selections drift monotonically away from the initial selection while live\-scored ones turn around and re\-converge, with no overlap across seeds\. On accuracy, a preregistered test at 12 seeds per cell puts the interaction at3\.79±1\.69pp3\.79\\pm 1\.69\\,\\text\{pp\}\(p=0\.035p\{=\}0\.035\): live scoring adds about four points of rate sensitivity on top of what restricting the loss support already costs\. The frozen arm stays entangled itself \(2\.89pp2\.89\\,\\text\{pp\},p=0\.015p\{=\}0\.015\), so the scoring loop aggravates the phenomenon without being its origin—and the added sensitivity sits at the cool end of the grid, where live scoring is2\.98pp2\.98\\,\\text\{pp\}better rather than worse\. We arrive at this number the long way: an earlier version of this work argued the same point from “significant in one arm, not the other,” which is not a valid inference\([Gelman and Stern 2006](https://arxiv.org/html/2609.22109#bib.bib4)\), and the correction required quadrupling the sample \(Appendix[A](https://arxiv.org/html/2609.22109#A1), P7\)\.
This paper contributes a measurement, an exclusion, and a protocol change\. We measure selector–rate entanglement on a duration\-homogeneous arm×\\timesrate matrix spanning four selection rules, two datasets, two adaptation regimes, and 133 training runs; we exclude the step\-size explanation by logging actual parameter displacement; and we show what the entanglement does to the comparisons this literature reports, prescribing the arm×\\timesrate matrix in place of a shared\-rate column \(Section[6](https://arxiv.org/html/2609.22109#S6)\)\. Every decision criterion was fixed before its data, including the three that went against us \(Appendix[A](https://arxiv.org/html/2609.22109#A1)\)\. The summary fits in one sentence: a selector comparison that fixes the learning rate does not control for it\.
## 2Setup: Selection as a Model\-Dependent Component
### 2\.1Setup
We study on\-policy distillation of Qwen2\.5\-1\.5B\-Instruct from Qwen2\.5\-7B\-Instruct on GSM8K\([Cobbe et al\. 2021](https://arxiv.org/html/2609.22109#bib.bib2);[Qwen Team 2024](https://arxiv.org/html/2609.22109#bib.bib14)\)\. Training proceeds in*refresh rounds*: in each round the student generates 512 fresh rollouts from training prompts \(768 max new tokens, temperature0\.70\.7, top\-pp0\.90\.9\), the teacher scores every generated position, a selector keeps the top5%5\\%of positions, and the student takes 512 optimizer steps on the distillation loss restricted to the selected positions \(LoRA rank 32, AdamW, gradient clipping at1\.01\.0; three rounds, 1,536 steps total\)\. Selection is recomputed at the start of every round from live student and teacher scores—the granularity every selective\-distillation method in this literature shares, since a selection frozen at initialization would be a different method from the one proposed\. We evaluate greedily on all 1,319 GSM8K test problems and report accuracy change over the untrained student \(64\.97%64\.97\\%\); greedy evaluation removes sampling noise from an already noise\-limited contrast\. Full training details are in Appendix[B](https://arxiv.org/html/2609.22109#A2)\.
### 2\.2The loop
The selector we make closed\-loop scores each positioniiby the total variation distance between the student’s and teacher’s next\-token distributions,
st\(i\)=TV\(pθt\(⋅∣xi\),q\(⋅∣xi\)\),Mt=top−5%\(st\),s\_\{t\}\(i\)\\;=\\;\\operatorname\{TV\}\\\!\\big\(p\_\{\\theta\_\{t\}\}\(\\cdot\\mid x\_\{i\}\),\\;q\(\\cdot\\mid x\_\{i\}\)\\big\),\\qquad M\_\{t\}\\;=\\;\\operatorname\{top\-5\\%\}\(s\_\{t\}\),\(1\)and the update rule trains only onMtM\_\{t\}:θt\+1=𝒜\(θt,∇ℒ\|Mt\)\\theta\_\{t\+1\}=\\mathcal\{A\}\\big\(\\theta\_\{t\},\\nabla\\mathcal\{L\}\|\_\{M\_\{t\}\}\\big\)\. The composition is the point of this paper:θ\\thetadeterminesss,ssdeterminesMM, andMMdetermines the nextθ\\theta\. The selector is a thermostat that also heats the room—it measures the very variable it is helping to move\. In control terms the training system is a closed loop, and the learning rateη\\etais its gain: it sets how far the controlled variable moves between two re\-measurements\.
Two arms in our design never readθ\\thetaand therefore run open\-loop:Fulltrains on every position \(Mt=M\_\{t\}=all\), andRandomdrawsMtM\_\{t\}uniformly at the same5%5\\%budget\. Both change other factors relative to the TV arm as well \(budget and criterion respectively\); Section[5](https://arxiv.org/html/2609.22109#S5)constructs the comparison that changes nothing else\.
### 2\.3Measuring the loop: drift observables
If the loop matters, its action should be visible on the selection itself, not only on end\-task accuracy\. At every round we score the round’s rollout positions twice—once with the current modelθt\\theta\_\{t\}and once with the frozen initial modelθ0\\theta\_\{0\}—and compare the two selections they induce: the Spearman correlationρt\\rho\_\{t\}of the two score vectors, and the Jaccard overlapJtJ\_\{t\}of the two top\-5%5\\%sets\. Because rollouts are regenerated each round, positions have no identity across rounds;JtJ\_\{t\}asks,*on today’s data, doθt\\theta\_\{t\}andθ0\\theta\_\{0\}still agree about what to select?*
JtJ\_\{t\}is the primary observable\. The top\-5%5\\%boundary is thin: after a single round of training,ρ\\rhofalls only2\.6%2\.6\\%\(to0\.9740\.974\) whileJJfalls4343–48%48\\%\(to0\.520\.52–0\.570\.57\), because small score perturbations flip many positions near the threshold\. Rank correlation is too blunt an instrument for a5%5\\%tail; set churn is not\. These within\-run observables average over∼512\{\\sim\}512chains per round, giving them an effective sample size that accuracy contrasts atn≤5n\{\\leq\}5seeds cannot approach\.
## 3Selector–Rate Entanglement
Table[1](https://arxiv.org/html/2609.22109#S3.T1)is the arm×\\timesrate matrix underlying Figure[1](https://arxiv.org/html/2609.22109#S1.F1)\. Every cell is duration\-homogeneous \(1,536 optimizer steps, selection recomputed once per round\); an earlier version of this matrix mixed training durations across columns and was discarded by audit \(Appendix[A](https://arxiv.org/html/2609.22109#A1)\)\.
Table 1:Arm×\\timeslearning\-rate matrix \(duration\-homogeneous\)\.Accuracy change \(pp\) over the untrained baseline \(64\.97%64\.97\\%, greedy, 1,319 problems\), mean±\\pmseed SD; parenthesized counts are seeds\. Swing is the difference between the arm’s best and worst cell means\.TV\-frozenis the frozen\-scoring arm of Section[5](https://arxiv.org/html/2609.22109#S5)\.#### Selective training is rate\-entangled even without feedback—our own strong hypothesis died here\.
We preregisteredRandomas a falsification test with frozen thresholds: if the fully state\-independent random selection swings more than4pp4\\,\\text\{pp\}across the grid, then rate sensitivity is not exclusive to the feedback loop, and our initial hypothesis—that the loop is the sole source—is falsified\. It fired\.Randomswings5\.4pp5\.4\\,\\text\{pp\}\(every celln=3n\{=\}3\), while dense supervision is statistically flat \(1\.8pp1\.8\\,\\text\{pp\},t=1\.40t\{=\}1\.40,p=0\.26p\{=\}0\.26\)\. The criterion was written on swing magnitude; we report the corresponding endpoint test as well \(\+5\.36±1\.95\+5\.36\\pm 1\.95,t=2\.74t\{=\}2\.74,p=0\.097p\{=\}0\.097\)—above the preregistered threshold, below conventional significance at three seeds\. Merely restricting the loss to a5%5\\%subset—*any*subset—couples the outcome to the learning rate\. We report this as a first\-class finding: the entanglement has a loop\-independent layer, and any account \(including our original one\) that attributes it entirely to feedback is wrong\.
#### Coupling depth orders the damage at the shared rate\.
The four arms differ in how much the selection reads the training state:Fullnot at all;Randomrestricts support but ignores state;TV\-frozenapplies a model\-based criterion frozen atθ0\\theta\_\{0\};TV\-liveapplies the same criterion live\. Atη=10−4\\eta\{=\}10^\{\-4\}the cell means fall in exactly this order—\+1\.87→−3\.64→−7\.37→−8\.18pp\+1\.87\\to\-3\.64\\to\-7\.37\\to\-8\.18\\,\\text\{pp\}—but only the endpoint contrast is established \(Fullvs\.TV\-live:10\.05pp10\.05\\,\\text\{pp\},t=6\.09t\{=\}6\.09,p<0\.001p\{<\}0\.001\)\. The individual rungs are not:\+5\.51\+5\.51\(p=0\.08p\{=\}0\.08\),\+3\.73\+3\.73\(p=0\.17p\{=\}0\.17\), and, once the TV arms reachn=12n\{=\}12,\+0\.81\+0\.81\(p=0\.60p\{=\}0\.60\)\. The last rung in particular is indistinguishable from zero, so at the hot rate the frozen and live scoring arms are equally damaged\. We report the ordering as a description of the means and rest the claim on the endpoints\. Section[5](https://arxiv.org/html/2609.22109#S5)shows where the two TV arms*do*separate: at the cooler rate, not this one\.
#### The verdict a shared\-rate comparison returns depends on which rate is shared\.
Atη=10−4\\eta\{=\}10^\{\-4\}dense supervision beats the live\-scored TV selector by10\.05pp10\.05\\,\\text\{pp\}\(t=6\.09t\{=\}6\.09\); at5⋅10−55\{\\cdot\}10^\{\-5\}the same contrast is5\.14pp5\.14\\,\\text\{pp\}\. Both are answers to “how much does this selective recipe cost?” computed on identical data with identical budgets; they differ by1\.96×1\.96\\times\(seed\-level bootstrap95%95\\%CI\[1\.27,2\.76\]\[1\.27,2\.76\], 20k resamples\)\. Neither rate is privileged—both lie inside the grid, and the published LoRA range brackets them \(Table[6](https://arxiv.org/html/2609.22109#S6.T6)\)\.
We state precisely what this is and is not\. It is*not*an estimate of how much per\-arm tuning would shrink the gap: our per\-cell means come from the same test set used to report accuracy, with no held\-out split for selecting a rate, and in fact both arms happen to peak at the same cell \(5⋅10−55\{\\cdot\}10^\{\-5\}\), so the two numbers above are two shared\-rate columns rather than a tuned\-versus\-untuned contrast\. Quantifying a tuning protocol properly requires a validation split and a declared tuning budget, which we did not run\. What the pair of numbers does show is narrower and still consequential: the answer a single\-shared\-rate comparison returns is a function of the rate chosen, by a factor of2\.02\.0here\. In a literature where selector\-versus\-selector margins are routinely∼1\{\\sim\}1–3pp3\\,\\text\{pp\}, that dependence is large enough to reorder methods, which is why we prescribe reporting the matrix rather than a column\.
#### What we retracted along the way\.
Two earlier versions of this section’s claims did not survive our own audits, and we record them because both failure modes are generic\. Atn=1n\{=\}1we observed an apparent*ranking inversion*between arms at their best cells; it vanished atn=3n\{=\}3\. And our first matrix produced a headline ratio of5\.2×5\.2\\timesthat was an artifact of mixed training durations across columns—the shared\-rate column came from8×8\\timesshorter runs than the rest\. Duration\-homogeneous, the ratio is2\.0×2\.0\\times\. A third retraction, of a causal claim about the scoring path, is recorded in Appendix[A](https://arxiv.org/html/2609.22109#A1)\(P7\)\. Comparisons in this area are fragile enough that the paper diagnosing them made three errors of the kind it diagnoses\.
## 4What It Is Not: Step Size
Before attributing the entanglement to a feedback loop, we must dispose of a far more mundane candidate\. A loss restricted to5%5\\%of positions has a different gradient scale than a dense loss; if arms take different effective step sizes at the same nominalη\\eta, then the “entanglement” would be nothing but mis\-scaled steps, and the fix would be a per\-arm rescaling ofη\\etaby gradient norm\.
The premise is real and the conclusion is false, and both halves are measurements\. Pre\-clipping gradient norms do differ dramatically across arms—by a factor of15\.515\.5\(0\.350\.35to5\.425\.42\)—so the concern is not a straw man\. But the quantity that reaches the weights is not the gradient; it is the AdamW update, and we log the actual per\-step parameter displacement for every run\. The ratio of update norm toη\\etais constant across arms to within2\.2%2\.2\\%: AdamW’s per\-coordinate normalization erases the gradient\-scale differences before they touch a single weight\. All arms take the same size steps\. They differ in*where*the steps point—which is determined by what the selector reads\.
This measurement matters beyond hygiene: it eliminates the entire family of ‘‘rescale the learning rate by the gradient norm’’ remedies, which would otherwise be the obvious response to our results\. No rescaling of step size can fix a confound that is not in the step size\.111This was the first of five mechanism hypotheses we tested; four died\. \(i\) “12×12\\timeslarger gradients take12×12\\timeslarger steps”—killed by the2\.2%2\.2\\%dispersion above\. \(ii\) “dense supervision wins because the global objective is unbiased”—Randomranks mid\-field, explaining only part of the gap\. \(iii\) “discarding95%95\\%of rollouts is the main cost”—killed by a direct budget\-matched comparison\. \(iv\) “high\-gradient\-norm arms are hurt more at largeη\\eta”—killed by the same update/η\\etaconstancy\. \(v\) The drift\-magnitude version of the loop hypothesis is killed in Section[5](https://arxiv.org/html/2609.22109#S5); what survives is the drift\-dynamics version\.
## 5Does the Model\-Dependence of Selection Matter?
### 5\.1The intervention: freeze the scoring model
One cannot freeze the selection*set*: rollouts are regenerated every round, so positions have no identity across rounds and there is no set to carry forward\. What can be frozen is the*scoring model*\. Our frozen\-scored arm,TV\-frozen, computes the same score as Eq\.[1](https://arxiv.org/html/2609.22109#S2.E1)but withθ0\\theta\_\{0\}in place ofθt\\theta\_\{t\};TV\-liveis the method the literature describes\. Rollouts remain on\-policy in both arms, the teacher is fixed in both, and criterion and budget are identical\. Implementation is free with LoRA: disabling the adapter recoversθ0\\theta\_\{0\}exactly\.
#### What this ablation removes, and what it leaves\.
It cuts the*direct*path by which training rewrites its own selection scores\. It does not remove the model\-dependence of selection entirely: rollouts are still generated byθt\\theta\_\{t\}, so which positions are available to be selected keeps tracking the current model, in both arms\. We name the armTV\-frozenrather thanTV\-frozenfor exactly this reason, and scope every claim below to the direct scoring path\. Among the comparisons available to us it remains the tightest:FullversusTVchanges budget*and*scoring dependence;RandomversusTVchanges criterion*and*scoring dependence;TV\-frozenversusTV\-livechanges the scoring dependence alone\. We run both arms atη∈\{10−4,5⋅10−5\}\\eta\\in\\\{10^\{\-4\},5\{\\cdot\}10^\{\-5\}\\\}, the two rates spanning the live arm’s swing\.
### 5\.2Preregistration
All decision criteria were frozen as executable assertions before the data arrived \(Appendix[A](https://arxiv.org/html/2609.22109#A1)gives them verbatim\)\. Drift was declared the primary observable; the accuracy contrast was declared underpowered in advance \(seed SDs of22–4pp4\\,\\text\{pp\}atn≤5n\{\\leq\}5\) and demoted to directional support\. One deviation is disclosed: after then=3n\{=\}3closed\-vs\-open contrast atη=10−4\\eta\{=\}10^\{\-4\}landed att=−2\.77t\{=\}\{\-\}2\.77against a threshold of2\.782\.78, we extended the two closed cells—symmetrically, under rules frozen before the new data: same Welch test, both sample sizes reported, no sampling beyondn=5n\{=\}5regardless of outcome\. Section[5\.2](https://arxiv.org/html/2609.22109#S5.SS2.SSS0.Px4)reports what the extension did to the estimate, which is itself instructive\.
#### Live scoring adds rate entanglement, by3\.79±1\.69pp3\.79\\pm 1\.69\\,\\text\{pp\}\.
The quantity at issue is an interaction—does the arm’s rate effect depend on which model scores the selection—and we report it directly, because the tempting shortcut of comparing one arm’s significance to the other’s is invalid\([Gelman and Stern 2006](https://arxiv.org/html/2609.22109#bib.bib4)\)\. Atn=12n\{=\}12per cell \(Table[2](https://arxiv.org/html/2609.22109#S5.T2)\) the live\-scored arm’s rate effect isΔlive=6\.68±1\.31pp\\Delta\_\{\\text\{live\}\}=6\.68\\pm 1\.31\\,\\text\{pp\}\(p<0\.001p\{<\}0\.001\), the frozen\-scored arm’s isΔfrozen=2\.89±1\.07pp\\Delta\_\{\\text\{frozen\}\}=2\.89\\pm 1\.07\\,\\text\{pp\}\(p=0\.015p\{=\}0\.015\), and their difference—the interaction—is3\.79±1\.69pp\\mathbf\{3\.79\\pm 1\.69\\,\\text\{pp\}\}\(t=2\.24t\{=\}2\.24,p=0\.035p\{=\}0\.035,95%95\\%CI\[\+0\.29,\+7\.30\]\[\+0\.29,\+7\.30\]\)\. Reading the loop into selection scores therefore costs roughly3\.8pp3\.8\\,\\text\{pp\}of additional rate sensitivity beyond what restricting the loss support already costs\.
#### The interaction is carried by the cool rate, not the hot one\.
Decomposing it by rate is instructive and, we think, the most interesting thing in this section\. Atη=10−4\\eta\{=\}10^\{\-4\}the two arms are indistinguishable \(−0\.81±1\.52pp\-0\.81\\pm 1\.52\\,\\text\{pp\},p=0\.60p\{=\}0\.60\): when the rate is hot enough to hurt, it hurts both equally\. At5⋅10−55\{\\cdot\}10^\{\-5\}live scoring is*better*by2\.98±0\.75pp2\.98\\pm 0\.75\\,\\text\{pp\}\(p=0\.001p\{=\}0\.001\)\. Letting selection track the model is therefore not a liability that a well\-chosen rate mitigates; it is a feature that only pays off at a well\-chosen rate, and the same intervention would be written up as harmful, neutral, or beneficial depending on which single rate an experimenter happened to fix—this paper’s own thesis, applied to its own ablation\. \(These per\-rate contrasts are descriptive: the confirmatory test we preregistered is the interaction, and an earlier per\-rate analysis was capped atn=5n\{=\}5by its own stopping rule, Appendix[A](https://arxiv.org/html/2609.22109#A1)P3\.\)
Three qualifications belong with the interaction estimate\. First, the confidence interval only just excludes zero; the effect is established, not pinned down\. Second, the frozen arm is itself significantly rate\-entangled \(p=0\.015p\{=\}0\.015\), so the scoring path is an aggravating factor and not the origin of the phenomenon—consistent withRandomin Section[3](https://arxiv.org/html/2609.22109#S3)\. Third, this test exists because an earlier version of this paper made the invalid inference: at our originaln=5/3n\{=\}5/3the interaction was3\.60±2\.843\.60\\pm 2\.84\(p=0\.25p\{=\}0\.25\), which establishes nothing, and we argued the claim from significance\-in\-one\-arm instead\. The preregistered replication \(Appendix[A](https://arxiv.org/html/2609.22109#A1), P7\) fixed the sample size atn=12n\{=\}12in advance with no top\-ups\. What makes the confirmation credible is not thepp\-value but the stability of the point estimate asnngrew:3\.60→3\.793\.60\\to 3\.79, with the interval shrinking around it\.
#### The loop changes the shape of selection drift, not its magnitude\.
Figure[2](https://arxiv.org/html/2609.22109#S5.F2)shows theJtJ\_\{t\}trajectories, and they falsify our own first mechanism hypothesis\. The naive causal chain—largerη\\etamovesθ\\thetafarther, drift grows, instability follows—predicts that drift*magnitude*tracksη\\eta\. It does not: after one round,J1≈0\.35J\_\{1\}\\approx 0\.35in every arm at every rate\. What separates the arms is the second round\. Open\-loop selections keep drifting away fromθ0\\theta\_\{0\}\(JJ:1\.00→0\.35→0\.271\.00\\to 0\.35\\to 0\.27–0\.290\.29, both rates, indistinguishable\)\. Closed\-loop selections turn around:JJrecovers, and the recovery strength is monotone in the rate across all four grid points \(J2=0\.39J\_\{2\}=0\.39,0\.440\.44,0\.440\.44,0\.480\.48fromη=10−4\\eta\{=\}10^\{\-4\}down to1\.25⋅10−51\.25\{\\cdot\}10^\{\-5\}; the two lowest\-rate points are single runs\)\. Atn=12n\{=\}12the two scoring sources do not overlap on a single seed \(t=17\.6t\{=\}17\.6,p<10−13p\{<\}10^\{\-13\}\), which makes this the most sharply separated measurement in the paper—and a reminder that a within\-run observable can be decisive where an end\-task contrast at the same seed count is not\. The loop does not amplify drift; it couples drift*dynamics*—recovery versus continued divergence—to the learning rate\. This is the observable on which the confound acts, and it is invisible to any protocol that logs only end\-task accuracy\.
#### A directional observation: the sign of the loop’s net effect flips with the rate\.
Atη=10−4\\eta\{=\}10^\{\-4\}the loop hurts \(closed−\-open=−2\.40±2\.58pp=\-2\.40\\pm 2\.58\\,\\text\{pp\},n=5n\{=\}5;−5\.43±1\.96\-5\.43\\pm 1\.96atn=3n\{=\}3\); at5⋅10−55\{\\cdot\}10^\{\-5\}it helps \(\+1\.21±1\.19pp\+1\.21\\pm 1\.19\\,\\text\{pp\}\)\. The signs are opposite at both sample sizes, but neither end is individually significant, and by our preregistered rule this is recorded permanently as directional support, not a finding\. Then=3→n=5n\{=\}3\\to n\{=\}5shrinkage at10−410^\{\-4\}\(−5\.43→−2\.40\-5\.43\\to\-2\.40\) is worth pausing on: it is regression to the mean, caught because the extension rule was frozen before the data and forced symmetric sampling\. Had we extended only the near\-threshold cell, or stopped on significance, this paper would be reporting an inflated effect—the same failure mode, at the meta level, that it diagnoses in selector comparisons\.
Table 2:Scoring selection with the live model adds3\.79pp3\.79\\,\\text\{pp\}of rate sensitivity\.Accuracy change \(pp\) over the untrained baseline, mean±\\pmseed SD\. Both arms: TV criterion,5%5\\%budget, on\-policy rollouts, 512 steps/round, selection recomputed every round; the arms differ only in which model scores the selection\.Δ\(η\)\\Delta\(\\eta\)is the within\-arm difference across rates \(Welch\)\. Then=12n\{=\}12block is the preregistered powered test \(P7\); the smaller\-nnblock is the original, underpowered version, reported as preregistered\.ArmScoringη=10−4\\eta\{=\}10^\{\-4\}η=5⋅10−5\\eta\{=\}5\{\\cdot\}10^\{\-5\}Δ\(η\)\\Delta\(\\eta\)*Preregistered powered test,n=12n\{=\}12per cell*TV\-frozenθ0\\theta\_\{0\}−7\.37\-7\.37−4\.48\-4\.482\.89±1\.072\.89\\pm 1\.07\(p=0\.015p\{=\}0\.015\)TV\-liveθt\\theta\_\{t\}−8\.18\-8\.18−1\.50\-1\.506\.68±1\.316\.68\\pm 1\.31\(p<0\.001p\{<\}0\.001\)interaction3\.79±1\.69\\mathbf\{3\.79\\pm 1\.69\}\(p=0\.035p\{=\}0\.035\)*Original underpowered version*TV\-frozen\(n=3n\{=\}3\)θ0\\theta\_\{0\}−5\.99±2\.84\-5\.99\\pm 2\.84−3\.36±1\.08\-3\.36\\pm 1\.082\.632\.63\(p=0\.25p\{=\}0\.25\)TV\-live\(n=5n\{=\}5\)θt\\theta\_\{t\}−8\.39±4\.46\-8\.39\\pm 4\.46−2\.15±2\.25\-2\.15\\pm 2\.256\.236\.23\(p=0\.032p\{=\}0\.032\)interaction3\.60±2\.843\.60\\pm 2\.84\(p=0\.25p\{=\}0\.25\)Figure 2:The loop changes the shape of selection drift, not its magnitude—with a dose–response in the rate\.Top\-5% JaccardJtJ\_\{t\}betweenθt\\theta\_\{t\}\- andθ0\\theta\_\{0\}\-scored selections on the same positions \(mean over seeds; bands are seed ranges;n=12n\{=\}12per TV cell at the final round\)\. After one round every arm sits atJ1≈0\.35J\_\{1\}\\approx 0\.35regardless of rate or scoring source\. Then they separate: frozen\-scored selections keep drifting away \(→0\.28\\to 0\.28–0\.290\.29, both rates\), live\-scored ones recover toward the initial selection, with the recovery monotone in the rate \(J2=0\.39J\_\{2\}=0\.39,0\.440\.44,0\.440\.44,0\.480\.48fromη=10−4\\eta\{=\}10^\{\-4\}down to1\.25⋅10−51\.25\{\\cdot\}10^\{\-5\}\)\. The frozen and live ranges do not overlap at any seed \(max0\.314\\max 0\.314vs\.min0\.354\\min 0\.354;t=17\.6t\{=\}17\.6\)\.
### 5\.3Replication with other criteria
Does the loop effect generalize beyond the TV criterion? We preregistered the sharpest version of the question \(Appendix[A](https://arxiv.org/html/2609.22109#A1)\): a student\-only*entropy*selector also readsθt\\theta\_\{t\}, so if the loop hypothesis holds as stated, it must entangle more when closed than open\. We ran the full open/closed design for Entropy and a closed\-only arm for the teachability criterion\([Wang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib17)\)\(22rates×\\times33seeds each; Table[3](https://arxiv.org/html/2609.22109#S5.T3)\)\.
Table 3:S2 replication\.Accuracy change \(pp\), mean±\\pmseed SD,n=3n\{=\}3per cell\.J1J\_\{1\}is the first\-round selection drift \(smaller = the criterion’s scores move more as the model moves\)\.#### The preregistered verdict fired against us\.
Entropy entangles \(Δ\(η\)=3\.16\\Delta\(\\eta\)=3\.16,p=0\.043p\{=\}0\.043\)—but identically with the loop open \(4\.304\.30\): the interaction is absent \(−1\.14±1\.34\-1\.14\\pm 1\.34\), and the drift\-recovery signature does not reproduce \(closedJ2\>J1J\_\{2\}\>J\_\{1\}in only1/31/3seeds at either rate, versus5/55/5and3/33/3cells for TV\)\. By our frozen criteria this is recorded asrevise\-loop:*whatever the frozen\-scoring ablation of Section[5](https://arxiv.org/html/2609.22109#S5)is detecting, it is specific to the TV criterion and does not generalise to model\-dependent selection in general*\. Meanwhile the teachability selector swings11\.68pp11\.68\\,\\text\{pp\}\(t=6\.04t\{=\}6\.04\)—the largest of any arm we measured, with a catastrophic−15\.5pp\-15\.5\\,\\text\{pp\}at the shared rate\.
#### A coupling account, stated post hoc and then tested once\.
The account was formed after the Entropy verdict: the criteria differ in how much selection*churn*training induces, measured by first\-round drift \(J1=0\.35J\_\{1\}=0\.35,0\.620\.62,0\.860\.86for TV, Entropy, Teach\), and churn is the carrier of the loop effect—so the open/closed interaction should shrink along that ordering\. This makes a prediction for the one cell we had not yet run: Teach, with the most stable selection, should show an interaction indistinguishable from zero and below TV’s\+3\.60\+3\.60\. We froze that criterion in a hash\-stamped script and then ranTeach\-frozen\. Measured interaction:−6\.07±5\.06\-6\.07\\pm 5\.06\(t=−1\.20t\{=\}\{\-\}1\.20; consistent with zero, below3\.603\.60\)—the account survives its first out\-of\-sample test, with the caveat that the interval is wide\. Strikingly,Teach\-frozenitself swings17\.74pp17\.74\\,\\text\{pp\}, reaching−22\.1pp\-22\.1\\,\\text\{pp\}at the shared rate: the criterion’s rate fragility needs no loop at all\. Two further observations stand regardless of the account: drift magnitude does not predict damage \(Teach has the most stable selection and the largest swings\), and every selective criterion we tested is rate\-entangled while dense supervision is not\.
### 5\.4Replication on a second dataset: the damage transfers, the entanglement does not
Both results so far come from GSM8K\. We repeated the decisive contrast on MATH\([Hendrycks et al\. 2021](https://arxiv.org/html/2609.22109#bib.bib5)\)—the numeric\-answer subset, so the answer checker and the accuracy definition are reused byte\-identically and the measuring instrument is not a variable—training on MATH prompts and evaluating on the 325 numeric\-answer problems of MATH\-500\([Lightman et al\. 2024](https://arxiv.org/html/2609.22109#bib.bib10)\)\. Grid:\{Full,TV\-live\}×\{10−4,5⋅10−5\}×3\\\{\\textsc\{Full\},\\textsc\{TV\-live\}\\\}\\times\\\{10^\{\-4\},5\{\\cdot\}10^\{\-5\}\\\}\\times 3seeds, everything else unchanged; the untrained student scores49\.23%49\.23\\%on this set\. The verdict rule was fixed before the runs: entanglement counts as reproduced only ifTV\-liveshows a significant rate effect that also exceedsFull’s\.
Table 4:MATH\-500 replication\.Greedy accuracy \(%\) on the 325 numeric\-answer problems, mean±\\pmseed SD,n=3n\{=\}3; the untrained student scores49\.23%49\.23\\%on the same set, and the bracketed values are changes over it\.#### Verdict: not reproduced—the rate entanglement is scoped to GSM8K\.
On MATH,TV\-live’s rate effect is1\.74pp1\.74\\,\\text\{pp\}\(p=0\.48p\{=\}0\.48\) and does not exceedFull’s2\.56pp2\.56\\,\\text\{pp\}\(p=0\.33p\{=\}0\.33\); neither arm is rate\-sensitive over this range\. By the frozen rule this is recorded asnot reproduced, and every entanglement claim in this paper is scoped to the GSM8K setting accordingly\. We flag one caveat that cuts against over\-reading the negative: with325325evaluation problems and seed SDs of2\.72\.7–3\.5pp3\.5\\,\\text\{pp\}, this grid could not have detected a22–3pp3\\,\\text\{pp\}rate effect, so it bounds the effect rather than excluding it\.
#### What did transfer, and it is the more useful half\.
The dense\-vs\-selective damage replicates almost exactly:10\.0510\.05and10\.87pp10\.87\\,\\text\{pp\}at the two rates on MATH, against10\.05pp10\.05\\,\\text\{pp\}at the shared rate on GSM8K, all significant\. Measured against each dataset’s own untrained student the two settings are near\-isomorphic: dense supervision is neutral\-to\-slightly\- positive \(−1\.2\-1\.2to\+1\.3pp\+1\.3\\,\\text\{pp\}on MATH,\+1\.9\+1\.9to\+3\.6\+3\.6on GSM8K\) while the5%5\\%TV\-selected budget costs about ten points \(−11\.3\-11\.3to−9\.5\-9\.5on MATH,−8\.2\-8\.2to−1\.5\-1\.5on GSM8K\)\. Combined with the flatness ofFullon both datasets, this says the protocol prescription of Section[6](https://arxiv.org/html/2609.22109#S6)is not a GSM8K artifact even though the entanglement magnitude is: whichever dataset one works on, the selective arm is the fragile one and the dense arm is the stable reference, so the two must not be compared at a single shared rate without checking\.
### 5\.5Replication under full fine\-tuning, at the literature’s own rates
Everything so far is LoRA at rates1010–100×100\\timeshotter than the full\-fine\-tuning rates the literature publishes with \(Table[6](https://arxiv.org/html/2609.22109#S6.T6)\)—the gap our own scope paragraph flagged as the most consequential open question\. We close it here: full fine\-tuning of the same student \(fp32 weights, standard fp32 AdamW, batch and step budget unchanged\), on a grid\{10−5,2⋅10−6,10−6\}\\\{10^\{\-5\},2\{\\cdot\}10^\{\-6\},10^\{\-6\}\\\}that brackets the exact published rates \(2⋅10−62\{\\cdot\}10^\{\-6\}in[Jiang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib7);10−610^\{\-6\}in[Xu et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib18)\),\{Full,TV\}×3\\\{\\textsc\{Full\},\\textsc\{TV\}\\\}\\times 3seeds, verdict rule frozen in advance \(Appendix[A](https://arxiv.org/html/2609.22109#A1), P8\)\. Two engineering notes matter for validity: at these rates a single update moves weights by less than bf16’s resolution, so fp32 master weights are mandatory and probed parameter displacements are asserted nonzero at run time \(Appendix[B](https://arxiv.org/html/2609.22109#A2)\); and these runs live on different hardware with its own re\-measured baseline and a passed reproduction gate \(Appendix[A](https://arxiv.org/html/2609.22109#A1), P8\-anchor\)\.
Figure 3:Under full fine\-tuning the asymmetry survives in graded form and grows\.Accuracy change over the untrained student \(A800 baseline, mean±\\pmseed SD,n=3n\{=\}3\) at rates bracketing those the literature publishes with\. Dense swings19\.8pp19\.8\\,\\text\{pp\}; the TV\-selective arm swings49\.5pp49\.5\\,\\text\{pp\}\(2\.5×2\.5\\timesmore\)\. The dense\-versus\-selective verdict is a non\-significant\+3\.6pp\+3\.6\\,\\text\{pp\}at the published operating point \(2⋅10−62\{\\cdot\}10^\{\-6\}\) and\+34\.0pp\+34\.0\\,\\text\{pp\}\(p=0\.005p\{=\}0\.005\) one notch hotter\.Table 5:Full fine\-tuning grid\.Accuracy change \(pp\) over the untrained student on the same hardware \(63\.46%63\.46\\%; see Appendix[A](https://arxiv.org/html/2609.22109#A1)for why this baseline differs from the LoRA tables’\), mean±\\pmseed SD,n=3n\{=\}3\.Armη=10−5\\eta\{=\}10^\{\-5\}2⋅10−62\{\\cdot\}10^\{\-6\}10−610^\{\-6\}SwingFull−14\.30±1\.9\-14\.30\{\\pm\}1\.9\+4\.22±0\.9\+4\.22\{\\pm\}0\.9\+5\.53±1\.2\+5\.53\{\\pm\}1\.219\.8419\.84\(p<0\.001p\{<\}0\.001\)TV−48\.32±5\.6\-48\.32\{\\pm\}5\.6\+0\.61±1\.9\+0\.61\{\\pm\}1\.9\+1\.21±3\.5\+1\.21\{\\pm\}3\.549\.5349\.53\(p=0\.001p\{=\}0\.001\)gap\+34\.02±3\.4\+34\.02\{\\pm\}3\.4\(p=0\.005p\{=\}0\.005\)\+3\.61±1\.2\+3\.61\{\\pm\}1\.2\(p=0\.060p\{=\}0\.060\)\+4\.32±2\.1\+4\.32\{\\pm\}2\.1\(p=0\.153p\{=\}0\.153\)#### The entanglement reproduces and is7×7\\timeslarger\.
The TV arm’s swing is49\.53pp49\.53\\,\\text\{pp\}\(p=0\.001p\{=\}0\.001\) against its LoRA counterpart’s6\.76\.7; atη=10−5\\eta\{=\}10^\{\-5\}full fine\-tuning with a5%5\\%TV selection destroys the model \(−48pp\-48\\,\\text\{pp\}, final accuracy∼15%\{\\sim\}15\\%\) while dense training merely suffers \(−14pp\-14\\,\\text\{pp\}\)\. By the preregistered rule—TV swing significant and larger thanFull’s—the verdict isreproduced\.
#### An honest revision: dense is not flat here\.
Under LoRA, dense supervision was statistically flat and we leaned on that flatness\. Under full fine\-tuning it is not \(19\.84pp19\.84\\,\\text\{pp\},p<0\.001p\{<\}0\.001\): at10−510^\{\-5\}*everything*degrades\. What survives, and what we now state as the regime\-independent form of the claim, is*graded*asymmetry: in every regime we measured, selective arms respond to the learning rate more than dense ones—2\.5×2\.5\\timesmore here, unboundedly more under LoRA where the dense response was indistinguishable from zero\.
#### Published comparisons sit one notch from a cliff\.
At the rates this literature actually uses, our data shows selective TV training as roughly neutral \(\+0\.6\+0\.6to\+1\.2pp\+1\.2\\,\\text\{pp\}\) and statistically indistinguishable from dense atn=3n\{=\}3\(\+3\.6pp\+3\.6\\,\\text\{pp\},p=0\.06p\{=\}0\.06;\+4\.3pp\+4\.3\\,\\text\{pp\},p=0\.15p\{=\}0\.15\)—though at these cells’ variance,n=3n\{=\}3has80%80\\%power only for effects of≈4\.4\{\\approx\}4\.4and7\.9pp7\.9\\,\\text\{pp\}respectively, so these are bounds, not equivalence claims\. At5×5\\timesthat rate, the same comparison reads\+34\.0pp\+34\.0\\,\\text\{pp\}\(p=0\.005p\{=\}0\.005\)\. Both facts matter: the published operating points are not obviously misleading about selective methods’ quality—but the verdict a shared\-rate comparison returns spans an order of magnitude within a10×10\\timesrate window, and nothing in the published protocols would detect this, because none of them vary the rate\. This is the strongest form of the paper’s thesis, obtained in the literature’s own regime\.
## 6Implications for Selector Comparisons
The prescription follows directly from the diagnosis, and it is two lines long\. First, tune the learning rate per arm and report the arm×\\timesrate matrix, not a shared\-rate column: a selector comparison that fixes the learning rate does not control for it, because every selective arm we measured responds to the rate while dense training does not\. Second, report the selection driftJtJ\_\{t\}alongside accuracy\. It costs one extra forward pass of the initial model per round, it is the observable on which the confound acts, and aJtJ\_\{t\}trajectory that differs across arms flags that selection is responding to training\. We offer it as a descriptive diagnostic: its link to rate sensitivity is correlational in our data \(J1J\_\{1\}orders with swing across criteria\), not a validated predictor\.
#### Would a shared\-rate selector comparison have been misled?
The thesis concerns selector\-versus\-selector comparisons, so we run one, both ways, on our own arms\. Across the four selective criteria on the LoRA grid the*ranking*is stable—Entropy\>\>Random\>\>TV\>\>Teach at both10−410^\{\-4\}and5⋅10−55\{\\cdot\}10^\{\-5\}—so in this setting the entanglement stretches margins rather than reordering methods, and we report that as a bounded negative\. The stretching, however, is systematic and conclusion\-relevant: all six pairwise margins are1\.51\.5–3\.1×3\.1\\timeslarger at the hotter rate, and two of the six flip significance\. Entropy beats Random by\+2\.10±0\.42pp\+2\.10\\pm 0\.42\\,\\text\{pp\}\(p=0\.035p\{=\}0\.035\) at5⋅10−55\{\\cdot\}10^\{\-5\}but by a non\-significant\+4\.02±2\.01\+4\.02\\pm 2\.01\(p=0\.16p\{=\}0\.16\) at10−410^\{\-4\}; Random–TV behaves the same way\. A paper run at one shared rate concludes “the selector does not significantly matter”; the same paper at the other rate concludes it does\. No rank inversion is needed for the protocol to change what a paper reports—the significance calls move first\. \(Scope: four criteria, one dataset,n=3n\{=\}3per cell except TV’sn=12n\{=\}12\.\)
#### How the literature actually sets the rate\.
We checked the five selective\-distillation papers closest to ours \(Table[6](https://arxiv.org/html/2609.22109#S6.T6)\)\. The pattern is uniform: where a learning rate is reported at all, one value is shared by every selector arm, varied only by model pair or training stage—never by method\.[Xu et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib18)index their hyperparameter table by model pair;[Jiang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib7)give one rate per training stage and state that no other hyperparameters were tuned;[Koo et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib8)search a grid but by model size, and do not report which value was selected\. Two papers report no usable rate for their LLM experiments at all\.*Nobody tunes the rate per selector arm*, which is precisely the practice our results argue is unsafe\.
Table 6:Learning\-rate practice in selective on\-policy distillation\.Rates as reported in each paper; “shared” means one value across all selector arms compared in that paper\.
#### Is our grid in the right place?
Two of these rates are one to two orders of magnitude below ours, and we state plainly why the numbers are not directly comparable: those runs are full fine\-tunes, ours is LoRA, and LoRA is conventionally trained at1010–100×100\\timesthe full\-fine\-tuning rate\. The like\-for\-like anchor is[Koo et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib8), the one LoRA\-based comparison that reports a range: its grid tops out at5⋅10−55\{\\cdot\}10^\{\-5\}, interior to ours\. Our lower two grid points thus sit inside published LoRA practice, and our upper point is one notch hotter\. And the full\-fine\-tuning regime is no longer an open question: Section[5\.5](https://arxiv.org/html/2609.22109#S5.SS5)runs the decisive contrast at the published rates themselves, where the entanglement is7×7\\timeslarger than under LoRA and the dense\-versus\-selective verdict moves from statistically invisible to34pp34\\,\\text\{pp\}within a10×10\\timesrate window\.
Read through this lens, the existing evidence base needs re\-weighting rather than dismissal\. Our results do not say these selectors are without value; they say a margin reported at one rate is one sample from a range that, in our setting, spans a factor of2\.02\.0\. The claims that survive this re\-weighting are the ones established across multiple rates; we found none in this literature that report more than one\.
#### Loop frequency is a second dial\.
An unplanned observation supports the same conclusion from another direction\. Two generations of our own pipeline differ in how often selection is recomputed—every optimizer batch versus once per 512\-step round—with criterion, budget, rate, and steps held equal\. Atη=5⋅10−5\\eta\{=\}5\{\\cdot\}10^\{\-5\}the per\-batch variant lands at\+2\.12pp\+2\.12\\,\\text\{pp\}and the per\-round variant at−2\.15pp\-2\.15\\,\\text\{pp\}: a∼4\.3pp\{\\sim\}4\.3\\,\\text\{pp\}swing from the re\-measurement frequency alone\. Seeds are unpaired across pipeline versions, so we report this as an observation, not a claim; but it is exactly what a control\-loop reading predicts \(re\-measurement frequency changes loop behavior\) and hard to explain under the static\-filter reading, where rescoring unchanged data should be nearly idempotent\.
#### Scope\.
Our evidence is two datasets \(GSM8K throughout; MATH\-500 for the decisive contrast, where the entanglement did*not*reproduce\), one model pair \(Qwen2\.5 1\.5B student, 7B teacher—a same\-family configuration that[Li et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib9)identify as failure\-prone, which is arguably the regime where selector choice matters most\), and two adaptation regimes \(LoRA throughout; full fine\-tuning for the decisive dense\-versus\-selective contrast at the literature’s published rates, Section[5\.5](https://arxiv.org/html/2609.22109#S5.SS5)\)\. Three selector criteria are covered \(TV, entropy, teachability; Sections[5](https://arxiv.org/html/2609.22109#S5)and[5\.3](https://arxiv.org/html/2609.22109#S5.SS3)\), with the live/frozen decomposition available for all three\. The protocol\-level entanglement claim holds for every selective arm we tested; the additional contribution of the live scoring path is established for TV atn=12n\{=\}12\(Appendix[A](https://arxiv.org/html/2609.22109#A1), P7\) and, by the S2 replication, does not extend to an entropy criterion\.
## 7Related Work
#### Hyperparameter tuning as a fairness control\.
That method comparisons can be decided by tuning effort rather than method quality is an established concern outside distillation:[Sivaprasad et al\. 2020](https://arxiv.org/html/2609.22109#bib.bib16)show that optimizer rankings depend on the hyperparameter\-tuning protocol, and argue benchmarks must specify it\. Our contribution is to demonstrate the same failure inside selective distillation, to locate it in an asymmetry between dense and selective training rather than in tuning effort per se, and to give the diagnostic \(the arm×\\timesrate matrix\) that makes it visible\.
#### Selective and token\-level on\-policy distillation\.
A growing line selects which positions of an on\-policy rollout deserve teacher supervision: by teachability\([Wang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib17)\), token importance\([Xu et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib18)\), difficulty profile\([Jiang et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib7)\), weighted token subsets\([Huang et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib6)\), or scheduled teacher involvement\([Koo et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib8)\); multi\-teacher and dual formulations extend the recipe\([Ma et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib13);[Yu et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib19)\)\. All of these recompute selection from the live student—they are closed\-loop in our sense—and all evaluate selectors at a shared learning rate \(Table[6](https://arxiv.org/html/2609.22109#S6.T6)\)\. None vary the rate per arm or measure selection drift\. Our contribution is orthogonal to each specific criterion: it concerns the comparison protocol they share\.
#### Phenomenology of on\-policy distillation\.
[Li et al\. 2026](https://arxiv.org/html/2609.22109#bib.bib9)map when on\-policy distillation fails, flagging same\-family small\-gap pairs as a failure configuration, and[Ma 2026a](https://arxiv.org/html/2609.22109#bib.bib11)analyze confounded supervision signals within rollouts, and[Ma 2026b](https://arxiv.org/html/2609.22109#bib.bib12)revisit what selective knowledge distillation actually buys\. Our finding that a*random*5%5\\%subset is already rate\-entangled, and that all our selective arms lose to dense supervision, is consistent with that sceptical line; what we add is that the size of the loss depends on the rate, so the comparison protocol has to be part of the discussion\. We add a confound at the level of the experimental protocol rather than the training signal, and our arm×\\timesrate matrix offers a re\-reading of failure reports at a single rate: some “failures” may be loop instability at the chosen rate rather than properties of the method\.
#### Training systems that measure what they move\.
Feedback pathologies of this shape are documented elsewhere: recursive training on model\-generated data collapses the data distribution\([Shumailov et al\. 2024](https://arxiv.org/html/2609.22109#bib.bib15);[Alemohammad et al\. 2024](https://arxiv.org/html/2609.22109#bib.bib1)\), and optimizing against a learned reward degrades as the policy consumes the signal that evaluates it\([Gao et al\. 2023](https://arxiv.org/html/2609.22109#bib.bib3)\)\. Selective distillation has not previously been placed in this family, and we stop short of claiming it belongs there: our frozen\-scoring ablation isolates the direct scoring path only, and its effect on end\-task accuracy is not established at our sample sizes\.
## 8Limitations and Conclusion
#### Limitations\.
The strongest limitation is breadth: one model pair, and an entanglement result that held on GSM8K but not on MATH—so the confound is demonstrated to exist and to be large where it occurs, not to be universal\. Whether the difference is task\-driven or power\-driven \(the MATH grid, at 325 evaluation problems, could not have detected a22–3pp3\\,\\text\{pp\}rate effect\) is unresolved, and a larger MATH evaluation set is the direct test\. The replication bounded one of our own claims—whatever the frozen\-scoring ablation detects is TV\-specific—and the coupling account of Section[5\.3](https://arxiv.org/html/2609.22109#S5.SS3), though it survived one frozen out\-of\-sample test, has been tested on exactly one criterion with a wide interval\. The interaction between scoring source and rate is established but not precisely located:3\.79±1\.69pp3\.79\\pm 1\.69\\,\\text\{pp\}with a confidence interval whose lower end is0\.290\.29, so “live scoring roughly doubles the rate effect” is the strongest honest reading and a tighter estimate would need more seeds than the 48 we ran\. The load\-bearing evidence in this paper is the arm×\\timesrate matrix and the within\-run drift dynamics, and we have kept the two categories separate throughout\. The sign\-flip result is a directional observation, permanently capped atn=5n\{=\}5by our own stopping rule\. The full fine\-tuning replication \(Section[5\.5](https://arxiv.org/html/2609.22109#S5.SS5)\) removes the LoRA\-only caveat for the central contrast, though its frozen\-scoring and drift instrumentation remain LoRA\-only\. Finally, one sampling decision was made after seeing a near\-threshold statistic; its rules were frozen before the new data, and both sample sizes appear in every affected table\.
#### Conclusion\.
Selective distillation belongs to a family that machine learning keeps rediscovering the hard way: training procedures that measure the thing they move\. The selector is not a lens held up to the data; it is a component inside the loop whose stability the learning rate governs\. The remedy is not a better selector—it is a protocol that treats the loop as part of the method: rates tuned per arm, drift reported alongside accuracy, and margins read as what they are, upper bounds\. A selector comparison that fixes the learning rate does not control for it\.
## References
- Alemohammad et al\. \[2024\]Sina Alemohammad, Josue Casco\-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G\. Baraniuk\.Self\-consuming generative models go MAD\.In*International Conference on Learning Representations*, 2024\.
- Cobbe et al\. \[2021\]Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\.Training verifiers to solve math word problems\.In*arXiv preprint arXiv:2110\.14168*, 2021\.
- Gao et al\. \[2023\]Leo Gao, John Schulman, and Jacob Hilton\.Scaling laws for reward model overoptimization\.In*International Conference on Machine Learning*, 2023\.
- Gelman and Stern \[2006\]Andrew Gelman and Hal Stern\.The difference between “significant” and “not significant” is not itself statistically significant\.*The American Statistician*, 60\(4\):328–331, 2006\.
- Hendrycks et al\. \[2021\]Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the MATH dataset\.In*NeurIPS Datasets and Benchmarks Track*, 2021\.
- Huang et al\. \[2025\]Haiduo Huang, Jiangcheng Song, Yadong Zhang, and Pengju Ren\.SelecTKD: Selective token\-weighted knowledge distillation for LLMs\.*arXiv preprint arXiv:2510\.24021*, 2025\.
- Jiang et al\. \[2026\]Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta, Dawei Li, and Zhao Yang\.Cornerstones or stumbling blocks? deciphering the rock tokens in on\-policy distillation\.*arXiv preprint arXiv:2605\.09253*, 2026\.
- Koo et al\. \[2025\]Jahyun Koo, Yerin Hwang, Yongil Kim, Taegwan Kang, Hyunkyung Bae, and Kyomin Jung\.SWITCH: Studying with teacher for knowledge distillation of large language models\.In*Findings of the Association for Computational Linguistics: NAACL 2025*, 2025\.arXiv:2410\.19503\.
- Li et al\. \[2026\]Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan\-ang Gao, Wenkai Yang, Zhiyuan Liu, and Ning Ding\.Rethinking on\-policy distillation of large language models: Phenomenology, mechanism, and recipe\.*arXiv preprint arXiv:2604\.13016*, 2026\.
- Lightman et al\. \[2024\]Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*International Conference on Learning Representations*, 2024\.
- Ma \[2026a\]Guoqing Ma\.Outcome\-confounded local supervision in on\-policy distillation\.*arXiv preprint arXiv:2607\.23731*, 2026a\.
- Ma \[2026b\]Guoqing Ma\.Rethinking selective knowledge distillation\.*arXiv preprint arXiv:2602\.01395*, 2026b\.
- Ma et al\. \[2026\]Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, and Fuli Luo\.MOPD: Multi\-teacher on\-policy distillation for capability integration in LLM post\-training\.*arXiv preprint arXiv:2606\.30406*, 2026\.
- Qwen Team \[2024\]Qwen Team\.Qwen2\.5 technical report\.*arXiv preprint arXiv:2412\.15115*, 2024\.
- Shumailov et al\. \[2024\]Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal\.AI models collapse when trained on recursively generated data\.*Nature*, 631:755–759, 2024\.
- Sivaprasad et al\. \[2020\]Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret\.Optimizer benchmarking needs to account for hyperparameter tuning\.In*International Conference on Machine Learning*, 2020\.
- Wang et al\. \[2026\]Yuanyi Wang, Su Lu, Yanggan Gu, and Pengkai Wang\.Not all disagreement is learnable: Token teachability in on\-policy distillation\.*arXiv preprint arXiv:2605\.26844*, 2026\.
- Xu et al\. \[2026\]Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, and Alborz Geramifard\.TIP: Token importance in on\-policy distillation\.*arXiv preprint arXiv:2604\.14084*, 2026\.
- Yu et al\. \[2026\]Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, and Shuicheng Yan\.DOPD: Dual on\-policy distillation\.*arXiv preprint arXiv:2606\.30626*, 2026\.
- Yuan et al\. \[2025\]Jiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie, Yu\-Jhe Li, Wentian Zhao, Kun Wan, Jing Shi, Xia Hu, and Zirui Liu\.Understanding and mitigating numerical sources of nondeterminism in LLM inference\.*arXiv preprint arXiv:2506\.09501*, 2025\.
## Appendix APreregistration and Deviations
This appendix is the audit trail: every decision criterion, frozen before its data, and every deviation, with its cause\.
#### P1 \(Random falsification test\)—fired\.
Frozen before theRandomrate\-grid runs: if the fully state\-independentRandomarm swings more than4pp4\\,\\text\{pp\}across the rate grid, the hypothesis that the feedback loop is the*sole*source of rate sensitivity is falsified; a swing below2\.5pp2\.5\\,\\text\{pp\}is consistent\. On the original grid the measured swing was1\.47pp1\.47\\,\\text\{pp\}and the hypothesis survived—but a subsequent config audit found that grid mixed training durations across columns \(64 vs\. 512 steps per round\), voiding the measurement\. On the duration\-homogeneous grid, with both endpoint cells extended ton=3n\{=\}3under rules frozen before the new data \(thresholds unchanged; no further sampling\), the swing is5\.08pp5\.08\\,\\text\{pp\}:*the criterion fired, and the strong hypothesis is recorded as falsified*\. The surviving two\-layer account \(support restriction entangles; the loop deepens it\) is what Sections 3 and 5 report\.
#### P2 \(frozen\-scoring control, primary/secondary split\)\.
Frozen before theTV\-frozenruns: drift trajectories are the primary observable \(within\-run,∼512\{\\sim\}512chains per round\); the accuracy contrast was declared underpowered in advance \(predicted SE≈2pp\{\\approx\}2\\,\\text\{pp\}\) and demoted to directional support regardless of outcome\. The prediction: live\-scoredJtJ\_\{t\}dynamics depend onη\\eta, frozen\-scored dynamics do not\.
#### P3 \(sample extension\)\.
After the live\-vs\-frozen contrast atη=10−4\\eta\{=\}10^\{\-4\}landed att=−2\.77t\{=\}\{\-\}2\.77against a critical value of2\.782\.78atn=3n\{=\}3, we decided to extend—an optional\-stopping situation, disclosed as such\. Rules frozen before the new data: \(i\) symmetric extension of both closed cells ton=5n\{=\}5, never only the near\-threshold cell; \(ii\) the identical Welch test, with critical values at the new degrees of freedom; \(iii\) bothn=3n\{=\}3andn=5n\{=\}5reported everywhere; \(iv\) no sampling beyondn=5n\{=\}5regardless of outcome\. Outcome: the near\-threshold effect shrank \(−5\.43→−2\.40\-5\.43\\to\-2\.40\), and the verdict “directional support” is permanent\.
#### P4 \(S2 prediction\)—fired\.
Frozen before the S2 runs, as executable assertions in a hash\-stamped verdict script: a student\-only entropy selector readsθt\\theta\_\{t\}too; if the scoring\-dependence hypothesis holds as stated, its live arm must entangle more than its frozen arm \(E1\), and the drift\-recovery signature must reproduce \(E2\)\. Outcome: E1revise\-loop\(live3\.163\.16,p=0\.043p\{=\}0\.043; frozen4\.304\.30,p=0\.059p\{=\}0\.059; interaction−1\.14±1\.34\-1\.14\\pm 1\.34, absent\), E2not reproduced\(1/31/3seeds\)\. The loop hypothesis as stated was wrong; the criterion\-scoped version and a post\-hoc coupling account appear in Section[5\.3](https://arxiv.org/html/2609.22109#S5.SS3)\. T1 \(teachability entangles\):11\.68±1\.9311\.68\\pm 1\.93,p=0\.016p\{=\}0\.016, confirmed\. All S2 cells are capped atn=3n\{=\}3; no top\-ups were made\.
#### P5 \(coupling account, out\-of\-sample test\)—passed\.
After the P4 verdict we stated the churn\-carrier account and froze its prediction for the one unrun cell, in a hash\-stamped script, before the data: the Teach open/closed interaction must be indistinguishable from zero \(\|I\|≤2SE\|I\|\\leq 2\\,\\mathrm\{SE\}\) and below TV’s\+3\.60\+3\.60; an interaction≥3\.60\\geq 3\.60witht\>2t\{\>\}2refutes the account\. Measured:−6\.07±5\.06\-6\.07\\pm 5\.06\(t=−1\.20t\{=\}\{\-\}1\.20\)—criterion met, account survives; the interval is wide and we claim survival, not confirmation\. One process note: our first informal statement of this prediction had the sign of the ordering wrong \(“Teach’s interaction should be largest”\); writing the frozen criterion caught the error before any data were collected\.
#### P6 \(dataset replication\)—not reproduced\.
Rule fixed before the MATH runs: entanglement counts as reproduced only ifTV\-liveshows a significant rate effect that also exceedsFull’s; ifTV\-live’s effect is non\-significant*and*no larger thanFull’s, the claim is scoped to GSM8K\. Outcome:1\.74pp1\.74\\,\\text\{pp\}\(p=0\.48p\{=\}0\.48\) versus2\.56pp2\.56\\,\\text\{pp\}\(p=0\.33p\{=\}0\.33\)⇒\\Rightarrowscoped \(Section[5\.4](https://arxiv.org/html/2609.22109#S5.SS4)\)\. Unlike P4 and P5, this rule was recorded in the project log rather than hash\-stamped in a script before execution; we note the weaker provenance rather than claim otherwise\.
#### P7 \(powered interaction test\)—running\.
An earlier version of this paper argued that freezing the scoring model removes rate entanglement, on the grounds that the rate effect was significant in the live arm \(p=0\.032p\{=\}0\.032\) and not in the frozen arm \(p=0\.25p\{=\}0\.25\)\. That is the difference\-of\-significance fallacy\[[Gelman and Stern 2006](https://arxiv.org/html/2609.22109#bib.bib4)\]; the interaction it stands in for is3\.60±2\.84pp3\.60\\pm 2\.84\\,\\text\{pp\}\(t=1\.27t\{=\}1\.27,p=0\.31p\{=\}0\.31\) and establishes nothing\. The claim was withdrawn and the test it needed was preregistered in a hash\-stamped script, frozen before any of its data existed: all four cells of\{live,frozen\}×\{10−4,5⋅10−5\}\\\{\\text\{live\},\\text\{frozen\}\\\}\\times\\\{10^\{\-4\},5\{\\cdot\}10^\{\-5\}\\\}extended ton=12n\{=\}12\(the size at which an effect of the observed magnitude would reacht≈2t\{\\approx\}2\), Welch interaction test,p<0\.05p\{<\}0\.05with positive sign confirms, anything else demotes the effect permanently to the drift level, no top\-ups regardless of outcome\.Outcome: confirmed—3\.79±1\.69pp3\.79\\pm 1\.69\\,\\text\{pp\},t=2\.24t\{=\}2\.24,p=0\.035p\{=\}0\.035,95%95\\%CI\[\+0\.29,\+7\.30\]\[\+0\.29,\+7\.30\], on 32 additional training runs\. The point estimate moved from3\.603\.60to3\.793\.79asnnwent from5/35/3to1212; it is the stability of that estimate, more than thepp\-value, that we regard as the evidence\. The claim is restored to the paper in the scoped form Section[5](https://arxiv.org/html/2609.22109#S5)states, and the invalid inference that preceded it is left on the record here\.
#### Statistical accounting\.
The preregistered decision criteria above \(P1–P6\) are the confirmatory tests; every othertt\-statistic in the paper—the ladder contrasts of Section[3](https://arxiv.org/html/2609.22109#S3), the post\-hoc interaction term, the per\-cell comparisons—is descriptive and reported without multiplicity correction\. We state this rather than apply a correction after the fact, since the confirmatory set was fixed in advance and is small\.
#### P8 \(full fine\-tuning replication\)—reproduced\.
Rules frozen in the project log before any full\-fine\-tuning data existed, verdict script written before the final 2 of 18 runs completed: grid\{Full,TV\}×\{10−5,2⋅10−6,10−6\}×3\\\{\\textsc\{Full\},\\textsc\{TV\}\\\}\\times\\\{10^\{\-5\},2\{\\cdot\}10^\{\-6\},10^\{\-6\}\\\}\\times 3seeds;reproducediff TV’s swing is significant \(Welchp<0\.05p\{<\}0\.05\)*and*exceedsFull’s;n=3n\{=\}3final, no top\-ups\. Outcome: TV49\.53pp49\.53\\,\\text\{pp\}\(p=0\.001p\{=\}0\.001\) vs\.Full19\.84pp19\.84\\,\\text\{pp\}\(p<0\.001p\{<\}0\.001\)—reproduced\(Section[5\.5](https://arxiv.org/html/2609.22109#S5.SS5)\)\.
#### P8\-anchor \(hardware reproduction gate\)—passed\.
The full fine\-tuning runs live on different hardware \(A800 80 GB; all earlier runs: RTX 4090\)\. Gate, frozen in advance: the LoRA TV arm re\-run on the new machine must swing≥4pp\\geq 4\\,\\text\{pp\}across\{10−4,5⋅10−5\}\\\{10^\{\-4\},5\{\\cdot\}10^\{\-5\}\\\}, or rate sensitivity is machine\-dependent and the full\-FT results are uninterpretable\. Measured: swing6\.046\.04\(−8\.19±2\.0\-8\.19\\pm 2\.0vs\.−2\.15±1\.5\-2\.15\\pm 1\.5,n=3n\{=\}3\), against6\.686\.68\(−8\.18\-8\.18,−1\.50\-1\.50\) on the original machine—the training effect reproduces across GPUs almost exactly\.
#### Hardware sensitivity of the evaluation itself\.
The new machine required its own baseline: the identical checkpoint, prompt \(SHA\-verified\), generation config \(SHA\-verified\), dataset \(hash\-verified\), torch and transformers versions scores63\.46%63\.46\\%on the A800 against64\.97%64\.97\\%on the RTX 4090—a1\.51pp1\.51\\,\\text\{pp\}difference from the GPU alone\. We traced it to single\-ULP bf16 logit differences at near\-tied top\-2 positions, which greedy decoding then amplifies autoregressively;[Yuan et al\. 2025](https://arxiv.org/html/2609.22109#bib.bib20)document the same mechanism at up to9%9\\%across GPU types\. Two consequences for this paper: all full\-FT numbers are reported against the A800 baseline and never numerically differenced against 4090 results; and the observation itself reinforces the thesis, since a hardware swap moves the measured score by more than a typical published selector margin\.
#### Deviations and recorded errors\.
Three times during this project a criterion written in prose was translated into analysis code incorrectly \(a drift sanity check whose “passing” value was actually the failure signature; a three\-branch verdict encoded as two branches; a threshold mis\-transcribed\)\. All three were caught by cross\-checking code against the frozen prose\. The resulting practice, adopted midway: criteria are frozen*as executable assertions*, not as prose to be translated later\. We report this because protocols fail at the translation step more often than at the design step\.
#### Data hygiene\.
Two result files from early runs were overwritten by later runs before an output\-tagging flag existed; both were detected by config\-fingerprint audits of every artifact \(the values survive in logs\) and neither enters any table in this paper\. Runs that mixed training durations \(64 vs\. 512 steps per round\) across cells of the rate matrix were likewise detected by audit and excluded\. Every number in this paper comes from the 133 duration\-homogeneous runs \(1,536 steps each\) whose configs are fingerprinted in the released artifacts\. One further bug is disclosed because it shaped an intermediate analysis: the training script computes “change over untrained” against a single cached baseline file, which is dataset\-specific; the MATH runs therefore initially reported deltas against the GSM8K baseline\. All MATH numbers in this paper are raw accuracies, which the bug does not touch, and all MATH contrasts are within\-dataset differences, in which the constant cancels\.
## Appendix BTraining and Measurement Details
#### Training\.
LoRA rank 32 on all attention and MLP projections; AdamW; gradient clipping at1\.01\.0; batch size 2; 512 rollouts per round from GSM8K training prompts; 3 refresh rounds×\\times512 optimizer steps; distillation loss on teacher top\-512 support \(computed aslogit−logsumexp\\mathrm\{logit\}\-\\mathrm\{logsumexp\}without materializing full\-vocabulary log\-softmax\)\. Generation: 768 max new tokens, temperature0\.70\.7, top\-pp0\.90\.9, matching the configuration under which all diagnostic surfaces were measured\. The teacher is released from memory after scoring each round\.
#### Evaluation\.
Greedy decoding on all 1,319 GSM8K test problems, final\-number extraction\. Sampled evaluation differs from greedy by up to25pp25\\,\\text\{pp\}at this model scale and is unusable for contrasts of a few pp\.
#### Update\-size measurement\.
For every optimizer step we log the pre\-clip gradient norm and the actual parameter displacement∥θk\+1−θk∥\\lVert\\theta\_\{k\+1\}\-\\theta\_\{k\}\\rVert\. The2\.2%2\.2\\%figure in Section[4](https://arxiv.org/html/2609.22109#S4)is the across\-arm dispersion of displacement/η/\\eta\. Master weights are fp32; bf16 master weights silently swallow displacements three orders of magnitude below their resolution, a failure mode we hit and instrumented against \(an assertion verifies that applied steps change the weights\)\.
#### Scoring\-source assertions\.
The open/closed switch is validated structurally, not by the drift statistic: the training mask must equal the mask induced by the declared scoring source, and once the model has moved, must differ from the mask induced by the other source\. Both checks are hard failures at run time\. \(The tempting alternative—checking that frozen\-scored drift stays at1\.01\.0—is backwards: drift comparesθt\\theta\_\{t\}toθ0\\theta\_\{0\}scores regardless of which one selects, and a drift pinned at1\.01\.0is precisely the signature of a switch that silently does nothing\.\)
#### Full fine\-tuning specifics\.
fp32 master weights are mandatory, not preferential: atη=10−6\\eta=10^\{\-6\}a single AdamW step moves a weight by∼10−6\{\\sim\}10^\{\-6\}against magnitudes of∼10−2\{\\sim\}10^\{\-2\}—a relative change of10−410^\{\-4\}, two orders of magnitude below bf16’s∼4⋅10−3\{\\sim\}4\{\\cdot\}10^\{\-3\}resolution, so under bf16 every update would be silently rounded away and the run would produce a clean\-looking flat result\. Guards: parameter displacement is probed on a fixed 8\-tensor subset every 32 steps \(cloning all1\.51\.5B parameters per step, as the LoRA path does with its3737M, would cost6\.26\.2GB per step\), any probed step with zero movement is a hard failure, and a run ending with no observed movement aborts rather than reporting\. Optimizer states are fp32 AdamW \(no quantization\); peak memory32\.232\.2GB on an 80 GB A800\. The frozen\-scoring arm is unavailable under full fine\-tuning \(θ0\\theta\_\{0\}is not recoverable without a resident copy\) and the drift statistic is recorded as undefined rather than approximated\.
#### Compute\.
All runs on a single RTX 4090 \(24 GB\); a full three\-round run takes∼75\{\\sim\}75minutes\. The complete evidence base of this paper is∼205\{\\sim\}205GPU\-hours across two machines, including audits and discarded grids\.Similar Articles
On-Policy Distillation (5 minute read)
This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
This paper presents a comprehensive empirical study on on-policy distillation for large language models, identifying failure mechanisms like distribution mismatch and optimization instability, and proposing fixes such as stop-gradient objectives and RLVR-adapted teachers.
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
This paper introduces SA-OPD, a spurious-signal-aware on-policy distillation framework that filters misleading token-level teacher supervision based on input-grounding and optimization impact, improving LLM and VLM distillation performance.
Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation
This paper analyzes off-policy distillation for LLM pre-training, characterizing how training objectives shape token-level supervision and downstream capabilities, and proposes adaptive objective routing that applies different objectives to different data domains, reframing pre-training as a data-conditional supervision design problem.
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
This paper introduces a training-free diagnostic framework to analyze per-token distillation signals for reasoning models, revealing that guidance is more beneficial on incorrect rollouts and depends on student capacity and task context.