Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
Summary
Identifies Supervision Fidelity Decay (SFD) in on-policy distillation, where teacher supervision degrades as student sequences lengthen, and proposes Lookahead Group Reward (LGR) to mitigate SFD, improving performance on math and code benchmarks.
View Cached Full Text
Cached at: 06/01/26, 09:29 AM
# Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
Source: [https://arxiv.org/html/2605.30833](https://arxiv.org/html/2605.30833)
Yanjiang Liu1,2Jie Lou3Xinyan Guan1,2Yuqiu Ji3Hongyu Lin2Ben He1,2 Xianpei Han2Le Sun2Xing Yu3Yaojie Lu2 1University of Chinese Academy of Sciences 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China 3Xiaohongshu
###### Abstract
On\-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token\-level feedback from a teacher\. However, we identify a critical bottleneck,Supervision Fidelity Decay \(SFD\): as student\-generated prefixes lengthen, the teacher’s next\-token distribution becomes less confident and less discriminative\. Consequently, the teacher\-dependent corrective signal in reverse\-KL distillation weakens, causing student drift to compound across long reasoning chains\. To mitigate SFD, we introduceLookahead Group Reward \(LGR\)\. Building on the insight that next\-step teacher confidence reflects the discriminative strength of future reverse\-KL supervision, LGR evaluates the student’s top\-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group\-normalized reward\. To maintain computational efficiency, we further design an entropy\-triggered tree\-attention mechanism\. Across six math and code benchmarks, LGR improves mean@8 by2\.57points over OPD for a 7B student, with gains increasing in longer\-generation and reaching \+4\.92points on AIME\-26 at 39k tokens\.
††footnotetext:Email:liuyanjiang22@mails\.ucas\.ac\.cn,\{luyaojie,hongyu,xianpei,sunle\}@iscas\.ac\.cn††footnotetext:Codes are available at[https://github\.com/zui\-jiang/LGR](https://github.com/zui-jiang/LGR)\.## 1Introduction
Large language models \(LLMs\) with reasoning capabilities have achieved remarkable performance on complex mathematical and coding tasks\. Recent advance models\[[27](https://arxiv.org/html/2605.30833#bib.bib2),[11](https://arxiv.org/html/2605.30833#bib.bib1),[7](https://arxiv.org/html/2605.30833#bib.bib3)\]demonstrate that extended reasoning chains can unlock capabilities previously thought to require much larger models\. On\-policy distillation \(OPD\), where a student model generates its own reasoning trajectories and learns from the teacher’s token\-level feedback, has emerged as a primary paradigm for transferring these capabilities to efficient, deployable models\[[1](https://arxiv.org/html/2605.30833#bib.bib4),[26](https://arxiv.org/html/2605.30833#bib.bib10),[10](https://arxiv.org/html/2605.30833#bib.bib5),patiño2025\_unlocking\_on\_policy\_distillation\_for\_any\_model\_family,[32](https://arxiv.org/html/2605.30833#bib.bib6),[30](https://arxiv.org/html/2605.30833#bib.bib7),[34](https://arxiv.org/html/2605.30833#bib.bib9)\]\.
However, current OPD methods\[[1](https://arxiv.org/html/2605.30833#bib.bib4),[26](https://arxiv.org/html/2605.30833#bib.bib10),[10](https://arxiv.org/html/2605.30833#bib.bib5),patiño2025\_unlocking\_on\_policy\_distillation\_for\_any\_model\_family,[20](https://arxiv.org/html/2605.30833#bib.bib14)\]implicitly treat the teacher as a*static oracle*whose supervision quality is unaffected by what the student generates\. We challenge this assumption and reveal a critical failure mode: as the student generates increasingly long sequences, its outputs progressively deviate from the teacher’s training distribution, causing the teacher’s supervision quality to*degrade monotonically*, which we refer to asSupervision Fidelity Decay \(SFD\)\.
Does teacher supervision quality remain reliable throughout long reasoning chains — and if not, how can we actively maintain it?
As an initial observation, training OPD with varying maximum generation lengths \(Figure[2](https://arxiv.org/html/2605.30833#S1.F2)\) reveals that performance improves from 3k to 9k tokens and plateaus around 16k, followed by a significant decline at 39k\. This trend indicates a potential failure mode during extremely long generations\. To isolate the cause, we design a controlled prefix\-completion experiment \(Figure[2](https://arxiv.org/html/2605.30833#S1.F2)\): as the student prefix grows longer, both the teacher’s downstream task accuracy and peak next\-token probability decay monotonically, while the teacher on its own prefixes maintains significantly higher confidence, which indicating that SFD stems from student drift\. The subplots further show that teacher confidence jumps immediately when the teacher takes over, confirming that different token choices lead to different teacher confidence one step ahead\.
Figure 1:Performance of different generation length in OPD\.AIME24 over training tokens for two model pairs\. Performance improves from 3k to 9k, plateaus around 16k, and degrades at 39k\.
Figure 2:Supervision Fidelity Decay\.Main:Teacher completion accuracy decays with student prefix length\.Insets:Teacher confidence \(max/sampled prob\) at the student\-to\-teacher handoff for∼\\sim2k \(left\) and∼\\sim14k \(right\) prefixes\. Confidence jumps at the handoff \(dashed line\), confirming token choices affect next\-step teacher confidence\.
Through theoretical analysis, we find that the declining max\-prob directly collapses the reverse\-KL gradient\. As teacher confidence falls, its log\-probability varies less across token choices, effectively reducing the learning signal to a student\-only signal that reinforces existing modes without correction\. Yet observations from the subplot suggest a solution; even at the same out\-of\-distribution context, teacher confidence at the subsequent position differs across token choices\. By the same gradient analysis, higher next position confidence implies a more discriminative future signal: choosing the token that maximizes the next position’s confidence directly preserves future supervision quality\. We operationalize this asLookahead Group Reward \(LGR\), a group\-normalized reward over the student’s top\-KKcandidates, where this group normalization removes the high\-variance absolute confidence level and retains only the relative ranking across candidates\. To maintain computational efficiency, we further design a tree\-attention mechanism triggered by entropy\.
Our key contributions are:
- •We identify SFD as a fundamental failure mode of OPD\.We show that teacher accuracy and peak confidence decay as student prefix length increases, a process that we prove collapses the reverse KL gradient into a signal that reinforces itself based solely on student outputs\.
- •We propose a principled remedy that looks one step ahead\.Since higher teacher confidence at the next position provides more discriminative future supervision, LGR selects tokens using group normalized rewards and an efficient tree attention mechanism triggered by entropy\.
- •We show LGR’s gains grow with reasoning length\.LGR significantly outperforms OPD and alternative distillation methods, with pronounced gains on long reasoning tasks\.
## 2Supervision Fidelity Decays Along Student Trajectories
### 2\.1On\-Policy Reverse\-KL as Policy Gradient
LetπT\\pi\_\{T\}denote a teacher model andπθ\\pi\_\{\\theta\}a student model parameterized byθ\\theta\. Given a prompt𝐜\\mathbf\{c\}, the student autoregressively generates a sequence𝐱=\(x1,x2,…,xL\)\\mathbf\{x\}=\(x\_\{1\},x\_\{2\},\\ldots,x\_\{L\}\)\. On\-policy reverse\-KL distillation minimizes\[[1](https://arxiv.org/html/2605.30833#bib.bib4),[31](https://arxiv.org/html/2605.30833#bib.bib13)\]:
ℒR\-KL\(θ\)=𝔼𝐱∼πθ\(⋅\|𝐜\)\[∑t=1Llogπθ\(xt\|𝐱<t,𝐜\)πT\(xt\|𝐱<t,𝐜\)\]\.\\mathcal\{L\}\_\{\\text\{R\-KL\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{c\}\)\}\\left\[\\sum\_\{t=1\}^\{L\}\\log\\frac\{\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\},\\mathbf\{c\}\)\}\{\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\},\\mathbf\{c\}\)\}\\right\]\.\(1\)Following standard practice in on\-policy distillation\[[26](https://arxiv.org/html/2605.30833#bib.bib10)\], we detach the generated sequence𝐱\\mathbf\{x\}from the computation graph during loss computation\. Under this stop\-gradient assumption, the per\-position KL gradient decomposes independently, yielding the policy gradient form:
∇θℒR\-KL=𝔼𝐱∼πθ\[∑t=1L∇θlogπθ\(xt\|𝐱<t\)⋅At\],\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\text\{R\-KL\}\}=\\mathbb\{E\}\_\{\\mathbf\{x\}\\sim\\pi\_\{\\theta\}\}\\left\[\\sum\_\{t=1\}^\{L\}\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\\cdot A\_\{t\}\\right\],\(2\)where the per\-token advantage is:
At=1\+logπθ\(xt\|𝐱<t\)−logπT\(xt\|𝐱<t\)\.A\_\{t\}=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\-\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\.\(3\)This reveals that on\-policy reverse\-KL is equivalent to a policy gradient in the style of REINFORCE\. Equivalently maximizing the per\-token rewardrt=logπT\(xt\|𝐱<t\)−logπθ\(xt\|𝐱<t\)r\_\{t\}=\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\-\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\[[29](https://arxiv.org/html/2605.30833#bib.bib11),[28](https://arxiv.org/html/2605.30833#bib.bib12)\]\. We refer to this aslocal supervision, in which the teacher provides a distributional target at each position\.
### 2\.2The Supervision Capability Functional and Supervision Fidelity Decay
Supervision as a functional\.The student policyπθ\\pi\_\{\\theta\}generating a sequence of lengthLLinduces a*state visitation distribution*ρπθL\(c\)\\rho\_\{\\pi\_\{\\theta\}\}^\{L\}\(c\)over prefix contextscc\. The teacher’s supervision capability is then afunctionalof the student policy:
𝒞\[πθ\]=𝔼c∼ρπθL\[fT\(c\)\],wherefT\(c\)=maxv∈𝒱πT\(v\|c\)\\mathcal\{C\}\[\\pi\_\{\\theta\}\]=\\mathbb\{E\}\_\{c\\sim\\rho\_\{\\pi\_\{\\theta\}\}^\{L\}\}\\\!\\left\[f\_\{T\}\(c\)\\right\],\\quad\\text\{where\}\\quad f\_\{T\}\(c\)=\\max\_\{v\\in\\mathcal\{V\}\}\\pi\_\{T\}\(v\|c\)\(4\)is the teacher’s local supervision quality at statecc\. This functional captures a key asymmetry\. While the teacher remains unchanged\[[1](https://arxiv.org/html/2605.30833#bib.bib4),[26](https://arxiv.org/html/2605.30833#bib.bib10),[10](https://arxiv.org/html/2605.30833#bib.bib5),patiño2025\_unlocking\_on\_policy\_distillation\_for\_any\_model\_family,[32](https://arxiv.org/html/2605.30833#bib.bib6),[30](https://arxiv.org/html/2605.30833#bib.bib7)\], the student’s state visitation distributionρπθL\\rho\_\{\\pi\_\{\\theta\}\}^\{L\}that shifts, dragging the integration domain into regions wherefTf\_\{T\}is low\.
Position\-dependent SFD curve\.We define the position\-dependent version:
𝒞\(t\)\[πθ\]=𝔼c∼ρπθt\[maxvπT\(v\|c\)\],\\mathcal\{C\}^\{\(t\)\}\[\\pi\_\{\\theta\}\]=\\mathbb\{E\}\_\{c\\sim\\rho\_\{\\pi\_\{\\theta\}\}^\{t\}\}\\\!\\left\[\\max\_\{v\}\\pi\_\{T\}\(v\|c\)\\right\],\(5\)so that𝒞\(t\)\[πθ\]\\mathcal\{C\}^\{\(t\)\}\[\\pi\_\{\\theta\}\]as a function oftttraces the SFD curve\. The aggregate𝒞\[πθ\]=𝔼t\[𝒞\(t\)\[πθ\]\]\\mathcal\{C\}\[\\pi\_\{\\theta\}\]=\\mathbb\{E\}\_\{t\}\[\\mathcal\{C\}^\{\(t\)\}\[\\pi\_\{\\theta\}\]\]summarizes total supervision quality\. WhenLLis short,ρπθL\\rho\_\{\\pi\_\{\\theta\}\}^\{L\}remains within the teacher’s training manifold andfT\(c\)f\_\{T\}\(c\)stays high; whenLLis long, autoregressive drift causesρπθL\\rho\_\{\\pi\_\{\\theta\}\}^\{L\}to shift into the teacher’s OOD region, wherefT\(c\)f\_\{T\}\(c\)collapses\.
Supervision fidelity as a downstream metric\.For a specific student\-generated prefix𝐱<t\\mathbf\{x\}\_\{<t\}, the teacher’s*supervision fidelity*at positionttis:
ℱ\(t\)≔maxv∈𝒱πT\(v\|𝐱≤t,𝐜\),\\mathcal\{F\}\(t\)\\coloneqq\\max\_\{v\\in\\mathcal\{V\}\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{\\leq t\},\\mathbf\{c\}\),\(6\)i\.e\., the teacher’s peak next\-token probability after observing the student’s prefix up tott\. Note that𝒞\(t\)\[πθ\]=𝔼\[ℱ\(t\)\]\\mathcal\{C\}^\{\(t\)\}\[\\pi\_\{\\theta\}\]=\\mathbb\{E\}\[\\mathcal\{F\}\(t\)\]is the expectation ofℱ\(t\)\\mathcal\{F\}\(t\)over student trajectories\. At the task level, we can also measure supervision fidelity asP\(teacher completes correctly from positiont∣𝐱<tθ\)P\(\\text\{teacher completes correctly from position \}t\\mid\\mathbf\{x\}\_\{<t\}^\{\\theta\}\), which correlates strongly withℱ\(t\)\\mathcal\{F\}\(t\), thereby validating that max\-ppis an effective proxy for downstream supervision quality\.
Empirical observation\.Using DeepSeek\-R1\-Distill\-Qwen\-1\.5B\[[11](https://arxiv.org/html/2605.30833#bib.bib1)\]as the student and 32B\[[11](https://arxiv.org/html/2605.30833#bib.bib1)\]as the teacher on AIME\[[36](https://arxiv.org/html/2605.30833#bib.bib20)\], we generate student prefixes of varying lengths and hand off to the teacher\. Figure[2](https://arxiv.org/html/2605.30833#S1.F2)shows: \(1\) teacher completion accuracy drops monotonically with prefix length; \(2\)ℱ¯\\bar\{\\mathcal\{F\}\}decreases correspondingly, validating max\-ppas a proxy; \(3\) the teacher on its own prefixes maintains significantly higher fidelity, confirming SFD stems from student drift\. The inset subplots further show that at the handoff point, teacher confidence*jumps immediately*, which demonstrates that token choices at positionttdo affect teacher confidence att\+1t\{\+\}1, even under severe SFD\. The jump shrinks at longer prefixes \(∼14\{\\sim\}14k vs\.∼2\{\\sim\}2k\) and after OPD optimization, consistent with accumulated drift\. This directly motivates Section[3\.2](https://arxiv.org/html/2605.30833#S3.SS2)\.
### 2\.3Theoretical Analysis: Signal Collapse and Compounding Drift
We formalize the impact of SFD on gradient quality\. As the teacher’s distribution becomes diffuse, its discriminative contribution vanishes; consequently, this single\-position failure compounds across positions into a self\-reinforcing drift\.
###### Proposition 1\(Teacher signal vanishing under SFD\)\.
DefineΔT\(t\)≔Varxt∼πθ\[logπT\(xt\|𝐱<t\)\]\\Delta\_\{T\}\(t\)\\coloneqq\\mathrm\{Var\}\_\{x\_\{t\}\\sim\\pi\_\{\\theta\}\}\[\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\]\. ThenΔT\(t\)≤log2\|𝒱\|−Ent\(πT\(⋅\|𝐱<t\)\)2\\Delta\_\{T\}\(t\)\\leq\\log^\{2\}\|\\mathcal\{V\}\|\-\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)^\{2\}, so as the teacher’s distribution becomes diffuse under SFD,ΔT\(t\)→0\\Delta\_\{T\}\(t\)\\to 0andSNRT\(t\)=O\(ΔT\(t\)\)→0\\mathrm\{SNR\}\_\{T\}\(t\)=O\(\\Delta\_\{T\}\(t\)\)\\to 0: the teacher’s discriminative contribution to the gradient vanishes entirely, leaving a student\-only signal that reinforces existing modes without correction\.*\(Proof: Appendix[F\.1](https://arxiv.org/html/2605.30833#A6.SS1)\.\)*
Once the teacher signal vanishes at positiontt, the student selects tokens from its own sharpened distribution, pushing the next context further out\-of\-distribution, thereby further degrading the signal att\+1t\{\+\}1\.
###### Proposition 2\(Self\-reinforcing drift under reverse\-KL\)\.
Letdt≔D\(πT\(⋅\|𝐱<tθ\),πT\(⋅\|𝐱<t∗\)\)d\_\{t\}\\coloneqq D\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\\theta\}\),\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\*\}\)\)be the distributional drift at positiontt, and assumeΔT\(t\)\\Delta\_\{T\}\(t\)is non\-increasing indtd\_\{t\}\(greater drift degrades teacher discriminability, as implied by SFD\)\. WhenΔT\(t\)<δcrit\\Delta\_\{T\}\(t\)<\\delta\_\{\\mathrm\{crit\}\}, the teacher\-free advantage reinforces the student’s existing modes, causing𝔼\[dt\+1\|dt\]≥dt\\mathbb\{E\}\[d\_\{t\+1\}\|d\_\{t\}\]\\geq d\_\{t\}: drift compounds across positions, creating a positive feedback loop that degrades supervision irreversibly\. Forward\-KL avoids this by construction \(dt=0d\_\{t\}=0always\) but introduces exposure bias\.*\(Proof sketch: Appendix[F\.2](https://arxiv.org/html/2605.30833#A6.SS2)\.\)*
### 2\.4Training Dynamics: The Vicious Cycle and Supervision Boundary Contraction
The functional perspective reveals a deeper consequence\. Define the*effective learning horizon*teff\(k\)=sup\{t:𝒞\(t\)\[πθk\]\>Cmin\}t\_\{\\mathrm\{eff\}\}\(k\)=\\sup\\\{t:\\mathcal\{C\}^\{\(t\)\}\[\\pi\_\{\\theta\_\{k\}\}\]\>C\_\{\\min\}\\\}as the furthest position where teacher supervision exceeds a minimum useful threshold at training stepkk\.
Under the self\-reinforcing drift established in Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2), training induces avicious cycle: positions beyondtefft\_\{\\mathrm\{eff\}\}receive no useful supervision \(ℓ≈0\\ell\\approx 0\) and therefore do not improve; the student’s uncorrected behavior at these positions continues to push the teacher further out\-of\-distribution, which in turn may causetefft\_\{\\mathrm\{eff\}\}to*shrink*at the next training step\. This creates a “learning desert” that expands over training, establishing areasoning length ceilingt∗t^\{\*\}beyond which the student can never improve under standard on\-policy distillation\.
The𝒞\(t\)\\mathcal\{C\}^\{\(t\)\}curve follows a sigmoidal decay: the early portion maintains high supervision quality, the late portion collapses, and the transition sharpens progressively over training\. We verify this empirically: fitting a logistic sigmoid to both curves in Figure[2](https://arxiv.org/html/2605.30833#S1.F2)yieldsR2\>0\.997R^\{2\}\>0\.997, and the fitted parameters show that OPD training contracts the supervision boundary \(t∗t^\{\*\}:7\.43→7\.047\.43\\to 7\.04k\) and steepens the transition slope \(ε\\varepsilon:0\.21→0\.260\.21\\to 0\.26\)\. Notably, all four parameter shifts align with the direction predicted by the vicious cycle \(see Appendix[C](https://arxiv.org/html/2605.30833#A3)for full fit results\)\.
This compounding effect is fundamental to on\-policy reverse\-KL and cannot be resolved by simply adjusting learning rates or adding regularization, as these measures only slow the drift rateε\\varepsilonwithout changing the structural problem\. Instead, it motivates directly optimizing the supervision functional𝒞\[πθ\]\\mathcal\{C\}\[\\pi\_\{\\theta\}\]itself, which provides a structurally different signal that steers the student toward states where the teacher can provide high\-quality supervision\.
## 3Lookahead Confidence as a Supervision Signal
### 3\.1From Gradient Failure to One\-Step Lookahead
The preceding analysis shows that SFD collapses the gradient: asπT\(⋅\|𝐱<t\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)becomes diffuse,At=1\+logπθ−logπTA\_\{t\}=1\+\\log\\pi\_\{\\theta\}\-\\log\\pi\_\{T\}loses discriminability and reduces to a student\-only signal \(Proposition[1](https://arxiv.org/html/2605.30833#Thmproposition1)\)\. The standard objective \(Eq\.[1](https://arxiv.org/html/2605.30833#S2.E1)\) is thusagnostic to state quality: it still forces the student to match the teacher’s now\-uninformative distribution\.
One\-step\-ahead retains discriminability\.Even whenπT\(⋅\|𝐱<t\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)is near\-uniform, different token choicesxt\(k\)x\_\{t\}^\{\(k\)\}steer the context into different next\-states, and the teacher’s distributions att\+1t\{\+\}1can differ substantially across candidates\. Instead of asking what the teacher wants at positiontt\(unanswerable under SFD\), we ask which candidate causes the*least further drift*, a question that remains informative even when local supervision has collapsed\.
Breaking the vicious cycle\.This signal directly targets Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2): by favoring tokens that maintain higher teacher confidence att\+1t\{\+\}1, the student slows the per\-step drift rate, preventingtefft\_\{\\mathrm\{eff\}\}from contracting\. The concrete realization is a per\-token one\-step\-ahead confidence reward \(Section[3\.2](https://arxiv.org/html/2605.30833#S3.SS2)\), which provides a greedy one\-step approximation to∇θ𝒞\[πθ\]\\nabla\_\{\\theta\}\\mathcal\{C\}\[\\pi\_\{\\theta\}\], sufficient to slow drift without requiring full multi\-step lookahead\. We now formalize when this approximation retains useful discriminability\.
###### Proposition 3\(One\-step\-ahead discriminability survives local supervision failure\)\.
Define the one\-step\-ahead discriminability at positionttas:
Dahead\(t\)≔maxk,k′\|maxvπT\(v\|𝐱<t,xt\(k\)\)−maxvπT\(v\|𝐱<t,xt\(k′\)\)\|,D\_\{\\text\{ahead\}\}\(t\)\\coloneqq\\max\_\{k,k^\{\\prime\}\}\\left\|\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k\)\}\)\-\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k^\{\\prime\}\)\}\)\\right\|,\(7\)which measures the maximum difference in teacher confidence att\+1t\{\+\}1across candidate token choices attt\. ThenDahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)isindependentof the teacher’s local entropyℋ\(πT\(⋅\|𝐱<t\)\)\\operatorname\{\\mathcal\{H\}\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\): even whenπT\(⋅\|𝐱<t\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)is uniform \(maximum SFD,ΔT\(t\)=0\\Delta\_\{T\}\(t\)=0\),Dahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)can be arbitrarily large\.*\(Proof: Appendix[F\.3](https://arxiv.org/html/2605.30833#A6.SS3)\.\)*
This proposition is the theoretical foundation for our confidence reward: even when local supervision \(Proposition[1](https://arxiv.org/html/2605.30833#Thmproposition1)\) has completely failed, the one\-step\-ahead comparison can still extract useful guidance from the teacher\.
### 3\.2Confidence Reward Design
We operationalize the one\-step\-ahead comparison as a reward signal\. At each positiontt, after the student samplesxt∼πθ\(⋅\|𝐱<t\)x\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\), we evaluate the teacher’s supervision fidelity at positiont\+1t\{\+\}1:
rraw\(xt\)=maxv∈𝒱πT\(v\|𝐱<t,xt,𝐜\)\.r\_\{\\text\{raw\}\}\(x\_\{t\}\)=\\max\_\{v\\in\\mathcal\{V\}\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\},\\mathbf\{c\}\)\.\(8\)This measures the teacher’s peak next\-token probability*after*observing the student’s actionxtx\_\{t\}\. The critical distinction is temporal: local supervision \(R\-KL\) evaluates the teacher’s state atttgiven context𝐱<t\\mathbf\{x\}\_\{<t\}, while the confidence reward evaluates the teacher’s state att\+1t\{\+\}1given context\(𝐱<t,xt\)\(\\mathbf\{x\}\_\{<t\},x\_\{t\}\)\. By comparing this quantity acrossKKcandidates, we identify which token choice atttcauses the least further degradationin teacher supervision quality att\+1t\{\+\}1\. This selection seeks not to recover an in\-distribution state, but rather to minimize the next step of SFD progression\.
###### Proposition 4\(Max\-ppas relative drift indicator\)\.
LetPT∗P^\{\*\}\_\{T\}denote the teacher’s next\-token distribution conditioned on an in\-distribution prefix, and letPT\(𝐱\)P\_\{T\}^\{\(\\mathbf\{x\}\)\}denote the distribution conditioned on a student\-generated prefix𝐱\\mathbf\{x\}\. If the teacher isβ\\beta\-smooth \(small perturbations in context produce bounded distributional shifts\), then:
maxvPT∗\(v\)−maxvPT\(𝐱\)\(v\)≤‖PT∗−PT\(𝐱\)‖∞≤β⋅d\(𝐱,𝒳T\),\\max\_\{v\}P^\{\*\}\_\{T\}\(v\)\-\\max\_\{v\}P\_\{T\}^\{\(\\mathbf\{x\}\)\}\(v\)\\leq\\\|P^\{\*\}\_\{T\}\-P\_\{T\}^\{\(\\mathbf\{x\}\)\}\\\|\_\{\\infty\}\\leq\\beta\\cdot d\(\\mathbf\{x\},\\mathcal\{X\}\_\{T\}\),\(9\)whered\(𝐱,𝒳T\)d\(\\mathbf\{x\},\\mathcal\{X\}\_\{T\}\)measures the distance from the teacher’s in\-distribution manifold𝒳T\\mathcal\{X\}\_\{T\}\. That is,higher max\-ppimplies closer proximity to the teacher’s competent region\.*\(Proof: Appendix[F\.4](https://arxiv.org/html/2605.30833#A6.SS4)\.\)*
Group normalization\.Sincerrawr\_\{\\text\{raw\}\}varies in absolute magnitude across positions and tasks, we normalize within the top\-KKstudent candidates \(ranked byπθ\\pi\_\{\\theta\}\), inspired by GRPO\[[11](https://arxiv.org/html/2605.30833#bib.bib1)\]:
rconf\(xt\)=rraw\(xt\)−μKσK\+ϵ,r\_\{\\text\{conf\}\}\(x\_\{t\}\)=\\frac\{r\_\{\\text\{raw\}\}\(x\_\{t\}\)\-\\mu\_\{K\}\}\{\\sigma\_\{K\}\+\\epsilon\},\(10\)whereμK\\mu\_\{K\},σK\\sigma\_\{K\}are the group mean and std over\{rraw\(k\)\}k=1K\\\{r\_\{\\text\{raw\}\}^\{\(k\)\}\\\}\_\{k=1\}^\{K\}\. This removes absolute\-scale variance while preserving relative ranking across candidates \(Appendix[F\.5](https://arxiv.org/html/2605.30833#A6.SS5)\)\.
Combined loss\.Adding the confidence reward to the per\-token advantage \(Eq\. \([3](https://arxiv.org/html/2605.30833#S2.E3)\)\) gives the LGR loss:
ℒtLGR=At\+γ⋅rconf\(xt\)\.\\mathcal\{L\}\_\{t\}^\{LGR\}=A\_\{t\}\+\\gamma\\cdot r\_\{\\text\{conf\}\}\(x\_\{t\}\)\.\(11\)Since we already compute teacher confidence for the top\-KKcandidates, we extend to a multi\-sample estimator\. LettingAt\(k\)=1\+logπθ\(xt\(k\)\|𝐱<t\)−logπT\(xt\(k\)\|𝐱<t\)A\_\{t\}^\{\(k\)\}=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\|\\mathbf\{x\}\_\{<t\}\)\-\\log\\pi\_\{T\}\(x\_\{t\}^\{\(k\)\}\|\\mathbf\{x\}\_\{<t\}\), the LGR \(topk\) loss weights each candidate by its student probability:
ℒtLGR\(topk\)=∑k=1Kπθ\(xt\(k\)\|𝐱<t\)⋅\[At\(k\)\+γ⋅rconf\(xt\(k\)\)\]\.\\mathcal\{L\}\_\{t\}^\{LGR\(topk\)\}=\\sum\_\{k=1\}^\{K\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\|\\mathbf\{x\}\_\{<t\}\)\\cdot\\left\[A\_\{t\}^\{\(k\)\}\+\\gamma\\cdot r\_\{\\text\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)\\right\]\.\(12\)
### 3\.3Efficient Computation
Entropy trigger\.Computing the confidence reward for allLLpositions withKKcandidates each would be prohibitively expensive\. We observe that at positions where the*student*has low generation entropy, the top\-1 token dominates the probability mass, and the remaining candidates carry negligible weight in the policy gradient\. We trigger the confidence reward only at positions𝒮=\{t:ℋ\(πθ\(⋅\|𝐱<t\)\)\>τ\}\\mathcal\{S\}=\\\{t:\\operatorname\{\\mathcal\{H\}\}\(\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)\>\\tau\\\}; elsewhere, only the standard reverse\-KL loss is applied\. This typically selects\|𝒮\|≈0\.15L–0\.25L\|\\mathcal\{S\}\|\\approx 0\.15L\\text\{\-\-\}0\.25Lpositions early in training, stabilizing around0\.10L0\.10Las the student distribution sharpens, reducing computational overhead by4–74\\text\{\-\-\}7times\.
The triggering uses*student*entropy rather than teacher entropy, avoiding circularity: teacher entropy atttis exactly what degrades under SFD, so triggering on it would disable the reward precisely where it is most needed\. This selection is near\-lossless \(Appendix[D](https://arxiv.org/html/2605.30833#A4)\)\.
Tree attention\.Naively evaluatingKKcandidates at eacht∈𝒮t\\in\\mathcal\{S\}requiresK⋅\|𝒮\|K\\cdot\|\\mathcal\{S\}\|teacher forward prefill for student’s prefix\. We instead construct an extended sequence \(main sequence𝐱\\mathbf\{x\}followed by all candidates as branches\) with a tree\-structured mask𝐌\\mathbf\{M\}\[[5](https://arxiv.org/html/2605.30833#bib.bib16),[24](https://arxiv.org/html/2605.30833#bib.bib17),[23](https://arxiv.org/html/2605.30833#bib.bib18)\]: main branch tokens attend causally; each candidatext\(k\)x\_\{t\}^\{\(k\)\}attends to𝐱≤t−1\\mathbf\{x\}\_\{\\leq t\-1\}but not to other candidates \(see Figure[12](https://arxiv.org/html/2605.30833#A5.F12)for an example\)\. In practice, GPU memory limits the total sequence length, so candidates are processed inNNsegments rather than all at once\. Appendix[E](https://arxiv.org/html/2605.30833#A5)gives the full cost analysis\.
Algorithm 1LGR: Lookahead Group Reward0:Student
πθ\\pi\_\{\\theta\}, Teacher
πT\\pi\_\{T\}, entropy threshold
τ\\tau, top\-
KK, reward weight
γ\\gamma
1:foreach training stepdo
2:Sample prompt
𝐜\\mathbf\{c\}from dataset
3:Generate
𝐱=\(x1,…,xL\)∼πθ\(⋅\|𝐜\)\\mathbf\{x\}=\(x\_\{1\},\\ldots,x\_\{L\}\)\\sim\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{c\}\)// Student rollout
4:Compute student logits and entropy
ℋt\\operatorname\{\\mathcal\{H\}\}\_\{t\}for all positions
5:Identify high\-entropy set
𝒮=\{t:ℋt\>τ\}\\mathcal\{S\}=\\\{t:\\operatorname\{\\mathcal\{H\}\}\_\{t\}\>\\tau\\\}
6:Extract top\-
KKcandidate tokens at each
t∈𝒮t\\in\\mathcal\{S\}
7:Construct tree\-attention input: main sequence \+ all candidates
8:Construct multi\-branch tree mask
𝐌\\mathbf\{M\}
9:Run teacher forward pass with tree mask
𝐌\\mathbf\{M\}// Single pass for all positions
10:Extract
rraw\(k\)r\_\{\\text\{raw\}\}^\{\(k\)\}for all candidates; compute
rconfr\_\{\\text\{conf\}\}via Eq\. \([10](https://arxiv.org/html/2605.30833#S3.E10)\)
11:foreach position
t=1,…,Lt=1,\\ldots,Ldo
12:if
t∈𝒮t\\in\\mathcal\{S\}then
13:
ℒt←ℒtLGR\\mathcal\{L\}\_\{t\}\\leftarrow\\mathcal\{L\}\_\{t\}^\{LGR\}via Eq\. \([12](https://arxiv.org/html/2605.30833#S3.E12)\)// Local \+ lookahead
14:else
15:
ℒt←logπθ\(xt\|𝐱<t\)πT\(xt\|𝐱<t\)\\mathcal\{L\}\_\{t\}\\leftarrow\\log\\frac\{\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\}\{\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\}// Standard R\-KL
16:endif
17:endfor
18:Update
θ\\thetawith
∇θ∑tℒt\\nabla\_\{\\theta\}\\sum\_\{t\}\\mathcal\{L\}\_\{t\}
19:endfor
## 4Experiments
### 4\.1Experimental Setup
Full training and evaluation details are provided in Appendix[A](https://arxiv.org/html/2605.30833#A1)\.
Models\.We evaluate two teacher–student configurations sharing the same vocabulary, each with a single teacher: \(1\)1\.5B student:DeepSeek\-R1\-Distill\-Qwen\-1\.5B\[[11](https://arxiv.org/html/2605.30833#bib.bib1)\]as the student, trained separately for math \(teacher: SkyWork\-OR1\-Math\-7B\[[12](https://arxiv.org/html/2605.30833#bib.bib15)\]\) and code \(teacher: DeepCoder\-14B\[[2](https://arxiv.org/html/2605.30833#bib.bib19)\]\)\. \(2\)7B student:DeepSeek\-R1\-Distill\-Qwen\-7B as the student, with DeepSeek\-R1\-Distill\-Qwen\-32B as the teacher\.
Benchmarks\.We evaluate our method on several rigorous datasets\. For mathematical reasoning, we utilize AIME\-24\[[36](https://arxiv.org/html/2605.30833#bib.bib20)\], AIME–25\[[37](https://arxiv.org/html/2605.30833#bib.bib21)\], and AIME\-26\[[38](https://arxiv.org/html/2605.30833#bib.bib22)\], as well as the February sessions of HMMT\-25 and HMMT\-26\[[8](https://arxiv.org/html/2605.30833#bib.bib23)\]\. For evaluation of code generation, we employ the LiveCodeBench v6 suite\[[15](https://arxiv.org/html/2605.30833#bib.bib24)\]\.
Baselines\.We compare LGR against: \(1\)GRPO\[[11](https://arxiv.org/html/2605.30833#bib.bib1)\]: Group Relative Policy Optimization with outcome\-level reward; \(2\)OPD: on\-policy distillation with standard reverse\-KL, plus a top\-KKmulti\-sample variant \(OPD topk\); \(3\)JSD\[[1](https://arxiv.org/html/2605.30833#bib.bib4)\]: Jensen–Shannon divergence distillation; and \(4\)REOPOLD\[[20](https://arxiv.org/html/2605.30833#bib.bib14)\]: Relaxed On\-Policy Distillation, which stabilizes RKL training via mixture\-based reward clipping \(to handle heavy\-tailed negative rewards\) and entropy\-guided token\-level dynamic sampling \(to filter near\-zero reward tokens\)\.
### 4\.2Main Results
Table 1:Comparison of distillation methods on mathematical reasoning and code benchmarks \(mean@8 / pass@8\)\.Δ\\Deltadenotes the absolute difference between our best method and OPD \(9k\)\.Table[1](https://arxiv.org/html/2605.30833#S4.T1)presents the main comparison across both teacher–student configurations\.
LGR delivers superior aggregate performance across benchmarks\.In the 1\.5B student setting, LGR achieves an average of 33\.18% mean@8, which is an increase of 1\.61% absolute over OPD 9k\. These gains are consistent across all six evaluation sets, with the largest improvements occurring on AIME\-26 and LCB\_v6\. For the larger 7B student model, LGR widens the lead to an average improvement of 2\.57%\. While it performs competitively but slightly behind the baseline on AIME\-244, LGR yields significant gains on the remaining five benchmarks, particularly on AIME\-25 and LCB\_v6 where improvements exceed 4\.1%\.
The advantage is more pronounced at larger scale\.In the 7B student setting, LGR achieves 48\.35% average mean@8, a \+2\.57% improvement over OPD 9k\. The gains are particularly large on AIME\-25 \(\+4\.16%\), HMMT\-25 \(\+4\.17%\), and LCB\_v6 \(\+4\.35%\), suggesting that LGR is especially effective when the student has sufficient capacity to leverage the improved supervision signal\.
LGR vs\. LGR \(topk\)\.The top\-KKextension provides complementary but not strictly additive gains\. In the 1\.5B setting, LGR \(topk\) achieves notably higher pass@8 on several benchmarks \(e\.g\., 56\.67% vs\. 50\.00% on AIME\-25\), suggesting that the multi\-sample estimator helps with diversity\. In the 7B setting, LGR without top\-KKgenerally outperforms, indicating that the single\-sample estimator with confidence reward is sufficient at larger scale\.
Stabilization alone does not address SFD\.JSD and REOPOLD achieve comparable or lower performance than OPD, confirming that the core issue is not the choice of divergence or optimization instability, but the degradation of teacher supervision quality on student\-generated contexts, a problem that demands a structurally different solution\.
### 4\.3LGR Addresses Supervision Fidelity Decay
Figure 3:Training dynamics across distillation methods\.Comparison of metrics over training steps\. LGR maintains higher teacher log\-probability and more stable entropy\.Training dynamics\.Figure[3](https://arxiv.org/html/2605.30833#S4.F3)tracks distill KL, teacher log\-probability, student log\-probability and entropy during training\. Compared to OPD variants, LGR maintains higher teacher log\-probability throughout training, indicating that the student’s generated contexts remain closer to the teacher’s in\-distribution manifold\. The entropy curves show that LGR maintains more stable generation diversity rather than collapsing into the mode\-sharpening behavior predicted by Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2)\.
Table 2:OPD vs\. LGR across max generation lengths\.Mean@8/Pass@8 on AIME benchmarks \(1\.5B student\)\.LGR’s advantage grows with maximum generation length\.Table[2](https://arxiv.org/html/2605.30833#S4.T2)compares OPD and LGR at different maximum generation lengths \(3k, 9k, 16k, 39k\)\. At short lengths \(3k\), where SFD is minimal, LGR provides no advantage because the teacher’s local supervision is already reliable\. As the maximum length increases, LGR increasingly outperforms OPD\. At 39k, LGR achieves the strongest gains across all three AIME benchmarks \(e\.g\., 36\.75 vs\. 32\.08 on AIME\-25, 34\.92 vs\. 30\.00 on AIME\-26\)\. This pattern is precisely what our SFD analysis predicts: the confidence reward becomes increasingly valuable as the teacher’s local supervision degrades at longer positions\.
Figure 4:LGR confidence reward dynamics\.Top:The average confidence reward increases over training\.Bottom:The applied ratio stabilizes around 10%\.Confidence reward dynamics\.Figure[4](https://arxiv.org/html/2605.30833#S4.F4)tracks the LGR confidence reward over training\. The average confidence reward increases over training \(top panel\), indicating that the student progressively learns to generate tokens that lead to higher teacher confidence at the next position\. The applied ratio \(bottom panel\) stabilizes around 10%, showing that the entropy\-triggered activation identifies a consistent fraction of high\-entropy decision points\. This stable activation ratio confirms that the entropy trigger effectively identifies positions where lookahead guidance is most needed, without requiring manual tuning over the course of training\.
### 4\.4Ablation Studies
Confidence metric comparison\.We compare four candidate confidence metrics for the group\-normalized reward \(Table[3](https://arxiv.org/html/2605.30833#S4.T3)a\)\. Max\-ppconsistently outperforms the alternatives, achieving 46\.67 on AIME\-24 compared to 43\.75 for Sampled\-pp, 42\.08 for entropy, and 37\.92 for PPL\. The strong advantage of Max\-ppaligns with Proposition[4](https://arxiv.org/html/2605.30833#Thmproposition4): it directly measures the proximity to the teacher’s in\-distribution regime, while PPL and entropy are noisier proxies that can be inflated by irrelevant low\-probability tokens\.
Table 3:Ablation studieson AIME benchmarks \(Mean@8/Pass@8, 1\.5B student\)\.Left:Confidence metric comparison\.Right:Sensitivity to confidence weightγ\\gamma\.\(a\) Confidence metric
\(b\) Confidence weightγ\\gamma
Sensitivity to confidence weightγ\\gamma\.Table[3](https://arxiv.org/html/2605.30833#S4.T3)b illustrates how the choice ofγ\\gammaaffects the distillation outcomes\. Whenγ\\gammais as low as 0\.1, the influence of the lookahead signal is negligible, and the performance remains close to the standard OPD baseline\. The optimal results are achieved in the range of 1\.0 to 1\.5, where the confidence reward effectively guides the student toward states that maintain high supervision quality\. In contrast, settingγ\\gammato 10\.0 causes a significant decline in accuracy\. In this high weight regime, the optimization becomes susceptible to reward hacking, as the student model tends to generate repetitive token sequences that artificially inflate teacher confidence but lack the logical substance required to solve the task correctly\.
## 5Related Work
On\-Policy Distillation\.Knowledge distillation\[[14](https://arxiv.org/html/2605.30833#bib.bib25),[21](https://arxiv.org/html/2605.30833#bib.bib26),[19](https://arxiv.org/html/2605.30833#bib.bib28),[35](https://arxiv.org/html/2605.30833#bib.bib29),hübotter2026reinforcementlearningselfdistillation,[17](https://arxiv.org/html/2605.30833#bib.bib27)\]transfers capabilities from teacher to student\. For LLMs, training on teacher\-generated \(off\-policy\) data creates a distribution mismatch at inference time\[[3](https://arxiv.org/html/2605.30833#bib.bib32),[13](https://arxiv.org/html/2605.30833#bib.bib33),[18](https://arxiv.org/html/2605.30833#bib.bib30)\]\. MiniLLM\[[10](https://arxiv.org/html/2605.30833#bib.bib5)\]addresses this by minimizing reverse\-KL on student\-generated rollouts to avoid the mode\-averaging problem of forward\-KL; GKD\[[1](https://arxiv.org/html/2605.30833#bib.bib4)\]further demonstrates that on\-policy generated data consistently outperforms off\-policy training\. These works establish OPD as the standard paradigm for reasoning model compression\. Subsequent work has explored several directions\. On*training stability*: REOPOLD\[[20](https://arxiv.org/html/2605.30833#bib.bib14)\]stabilizes reverse\-KL training via mixture\-based reward clipping and entropy\-guided token\-level dynamic sampling\. On*objective design*: G\-OPD\[[33](https://arxiv.org/html/2605.30833#bib.bib34)\]formalizes OPD as KL\-constrained RL and proposes reward extrapolation to learn beyond the teacher ceiling; TSD\-KD\[[16](https://arxiv.org/html/2605.30833#bib.bib35)\]applies KL loss selectively on high\-entropy tokens where the student is genuinely uncertain, reducing noise from low\-uncertainty positions\.
Two concurrent works are most closely related\. Revisiting OPD\[[9](https://arxiv.org/html/2605.30833#bib.bib36)\]identify that OPD becomes unreliable on student generations and propose teacher top\-KKsupport matching\. Rethinking OPD\[[22](https://arxiv.org/html/2605.30833#bib.bib37)\]study OPD phenomenology and find that reward quality degrades with trajectory depth—consistent with our SFD analysis\. Our work differs in two respects: \(1\) we provide a formal analysis of*why*and*where*supervision degrades—characterizing SFD as a position\-dependent functional that worsens monotonically along student trajectories and establishes a reasoning length ceiling; and \(2\) we propose a one\-step lookahead reward remedy that directly optimizes teacher confidence at the future position, orthogonal to divergence modification and applicable on top of OPD objective\.
## 6Conclusion
We proposed Lookahead Group Reward, which augments standard reverse\-KL distillation with a group\-normalized confidence reward that directly optimizes the teacher’s supervision capability functional\. Unlike masking or truncation strategies that passively avoid weak\-supervision regions, LGR actively steers the student toward trajectories where the teacher maintains high supervision fidelity\. The entropy\-triggered tree\-attention mechanism makes this approach computationally practical\. Achieving a 1000×\\timesspeedup compared to a naïve implementation\.
Limitations and future work\.First, LGR requires white\-box access to the teacher model’s logits, which may not always be available\. Second, the confidence reward relies on an implicit assumption that the teacher’s*relative ranking*of candidate tokens remains informative even when its absolute predictions are unreliable\. While our group normalization design and empirical results support this assumption, it may weaken for extremely out\-of\-distribution student trajectories or poorly calibrated teachers\. Third, the current design uses a fixed number of candidatesKKand a static entropy thresholdτ\\tau; dynamically adjusting these based on training progress could improve both efficiency and effectiveness\. Finally, extending the confidence reward framework to multi\-modal reasoning settings represents a promising direction\.
## References
- \[1\]\(2024\)On\-policy distillation of language models: learning from self\-generated mistakes\.External Links:2306\.13649,[Link](https://arxiv.org/abs/2306.13649)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§1](https://arxiv.org/html/2605.30833#S1.p2.1),[§2\.1](https://arxiv.org/html/2605.30833#S2.SS1.p1.5),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p1.7),[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p4.1),[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[2\]M\. Balog, A\. L\. Gaunt, M\. Brockschmidt, S\. Nowozin, and D\. Tarlow\(2017\)DeepCoder: learning to write programs\.External Links:1611\.01989,[Link](https://arxiv.org/abs/1611.01989)Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p2.1)\.
- \[3\]S\. Bengio, O\. Vinyals, N\. Jaitly, and N\. Shazeer\(2015\)Scheduled sampling for sequence prediction with recurrent neural networks\.External Links:1506\.03099,[Link](https://arxiv.org/abs/1506.03099)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[4\]R\. Bhatia and C\. Davis\(2000\)A better bound on the variance\.The american mathematical monthly107\(4\),pp\. 353–357\.Cited by:[§F\.1](https://arxiv.org/html/2605.30833#A6.SS1.3.p3.1)\.
- \[5\]T\. Cai, Y\. Li, Z\. Geng, H\. Peng, J\. D\. Lee, D\. Chen, and T\. Dao\(2024\)Medusa: simple llm inference acceleration framework with multiple decoding heads\.External Links:2401\.10774,[Link](https://arxiv.org/abs/2401.10774)Cited by:[§3\.3](https://arxiv.org/html/2605.30833#S3.SS3.p3.8)\.
- \[6\]T\.M\. Cover and J\.A\. Thomas\(2012\)Elements of information theory\.Wiley\.External Links:ISBN 9781118585771,LCCN 2005047799,[Link](https://books.google.com.sg/books?id=VWq5GG6ycxMC)Cited by:[§F\.1](https://arxiv.org/html/2605.30833#A6.SS1.3.p3.13)\.
- \[7\]DeepSeek\-AI\(2026\)DeepSeek\-v4: towards highly efficient million\-token context intelligence\.Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1)\.
- \[8\]J\. Dekoninck, N\. Jovanović, T\. Gehrunger, K\. Rögnvalddson, I\. Petrov, C\. Sun, and M\. Vechev\(2026\)Beyond benchmarks: matharena as an evaluation platform for mathematics with llms\.External Links:2605\.00674,[Link](https://arxiv.org/abs/2605.00674)Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p3.1)\.
- \[9\]Y\. Fu, H\. Huang, K\. Jiang, Y\. Zhu, and D\. Zhao\(2026\)Revisiting on\-policy distillation: empirical failure modes and simple fixes\.arXiv preprint arXiv:2603\.25562\.Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p2.1)\.
- \[10\]Y\. Gu, L\. Dong, F\. Wei, and M\. Huang\(2026\)MiniLLM: on\-policy distillation of large language models\.External Links:2306\.08543,[Link](https://arxiv.org/abs/2306.08543)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§1](https://arxiv.org/html/2605.30833#S1.p2.1),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p1.7),[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[11\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Ding, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Chen, J\. Yuan, J\. Tu, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. You, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. Zhang\(2025\-sept\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.External Links:ISSN 1476\-4687,[Link](http://dx.doi.org/10.1038/s41586-025-09422-z),[Document](https://dx.doi.org/10.1038/s41586-025-09422-z)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p4.6),[§3\.2](https://arxiv.org/html/2605.30833#S3.SS2.p2.3),[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p4.1)\.
- \[12\]J\. He, J\. Liu, C\. Y\. Liu, R\. Yan, C\. Wang, P\. Cheng, X\. Zhang, F\. Zhang, J\. Xu, W\. Shen, S\. Li, L\. Zeng, T\. Wei, C\. Cheng, B\. An, Y\. Liu, and Y\. Zhou\(2025\)Skywork open reasoner 1 technical report\.External Links:2505\.22312,[Link](https://arxiv.org/abs/2505.22312)Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p2.1)\.
- \[13\]T\. He, J\. Zhang, Z\. Zhou, and J\. Glass\(2021\)Exposure bias versus self\-recovery: are distortions really incremental for autoregressive text generation?\.External Links:1905\.10617,[Link](https://arxiv.org/abs/1905.10617)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[14\]G\. Hinton, O\. Vinyals, and J\. Dean\(2015\)Distilling the knowledge in a neural network\.External Links:1503\.02531,[Link](https://arxiv.org/abs/1503.02531)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[15\]N\. Jain, K\. Han, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. Stoica\(2024\)LiveCodeBench: holistic and contamination free evaluation of large language models for code\.External Links:2403\.07974,[Link](https://arxiv.org/abs/2403.07974)Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p3.1)\.
- \[16\]M\. Kim and S\. J\. Baek\(2026\)Explain in your own words: improving reasoning via token\-selective dual knowledge distillation\.External Links:2603\.13260,[Link](https://arxiv.org/abs/2603.13260)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[17\]T\. Kim, J\. Oh, N\. Kim, S\. Cho, and S\. Yun\(2021\)Comparing kullback\-leibler divergence and mean squared error loss in knowledge distillation\.External Links:2105\.08919,[Link](https://arxiv.org/abs/2105.08919)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[18\]Y\. Kim, D\. Shin, M\. Kang, B\. Na, and I\. Moon\(2026\)Distillation of large language models via concrete score matching\.External Links:2509\.25837,[Link](https://arxiv.org/abs/2509.25837)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[19\]Y\. Kim and A\. M\. Rush\(2016\)Sequence\-level knowledge distillation\.CoRRabs/1606\.07947\.External Links:[Link](http://arxiv.org/abs/1606.07947),1606\.07947Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[20\]J\. Ko, S\. Abdali, Y\. J\. Kim, T\. Chen, and P\. Cameron\(2026\)Scaling reasoning efficiently via relaxed on\-policy distillation\.External Links:2603\.11137,[Link](https://arxiv.org/abs/2603.11137)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p2.1),[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p4.1),[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[21\]J\. Ko, T\. Chen, S\. Kim, T\. Ding, L\. Liang, I\. Zharkov, and S\. Yun\(2025\)DistiLLM\-2: a contrastive approach boosts the distillation of llms\.External Links:2503\.07067,[Link](https://arxiv.org/abs/2503.07067)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[22\]Y\. Li, Y\. Zuo, B\. He, J\. Zhang, C\. Xiao, C\. Qian, T\. Yu, H\. Gao, W\. Yang, Z\. Liu, and N\. Ding\(2026\)Rethinking on\-policy distillation of large language models: phenomenology, mechanism, and recipe\.External Links:2604\.13016,[Link](https://arxiv.org/abs/2604.13016)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p2.1)\.
- \[23\]Y\. Li, F\. Wei, C\. Zhang, and H\. Zhang\(2025\)EAGLE\-3: scaling up inference acceleration of large language models via training\-time test\.External Links:2503\.01840,[Link](https://arxiv.org/abs/2503.01840)Cited by:[§3\.3](https://arxiv.org/html/2605.30833#S3.SS3.p3.8)\.
- \[24\]Y\. Li, F\. Wei, C\. Zhang, and H\. Zhang\(2025\)EAGLE: speculative sampling requires rethinking feature uncertainty\.External Links:2401\.15077,[Link](https://arxiv.org/abs/2401.15077)Cited by:[§3\.3](https://arxiv.org/html/2605.30833#S3.SS3.p3.8)\.
- \[25\]A\. Lin, J\. Wohlwend, H\. Chen, and T\. Lei\(2020\-11\)Autoregressive knowledge distillation through imitation learning\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 6121–6133\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.494/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.494)Cited by:[§B\.1](https://arxiv.org/html/2605.30833#A2.SS1.p1.1)\.
- \[26\]K\. Lu and T\. M\. Lab\(2025\)On\-policy distillation\.Thinking Machines Lab: Connectionism\.Note:https://thinkingmachines\.ai/blog/on\-policy\-distillationExternal Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§1](https://arxiv.org/html/2605.30833#S1.p2.1),[§2\.1](https://arxiv.org/html/2605.30833#S2.SS1.p1.6),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p1.7)\.
- \[27\]OpenAI, :, A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney, A\. Iftimie, A\. Karpenko, A\. T\. Passos, A\. Neitz, A\. Prokofiev, A\. Wei, A\. Tam, A\. Bennett, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Duberstein, A\. Kondrich, A\. Mishchenko, A\. Applebaum, A\. Jiang, A\. Nair, B\. Zoph, B\. Ghorbani, B\. Zhang, B\. Rossen, B\. Sokolowsky, B\. Barak, B\. McGrew, B\. Minaiev, B\. Hao, B\. Baker, B\. Houghton, B\. McKinzie, B\. Eastman, C\. Lugaresi, C\. Bassin, C\. Hudson, C\. M\. Li, C\. de Bourcy, C\. Voss, C\. Shen, C\. Zhang, C\. Koch, C\. Orsinger, C\. Hesse, C\. Fischer, C\. Chan, D\. Roberts, D\. Kappler, D\. Levy, D\. Selsam, D\. Dohan, D\. Farhi, D\. Mely, D\. Robinson, D\. Tsipras, D\. Li, D\. Oprica, E\. Freeman, E\. Zhang, E\. Wong, E\. Proehl, E\. Cheung, E\. Mitchell, E\. Wallace, E\. Ritter, E\. Mays, F\. Wang, F\. P\. Such, F\. Raso, F\. Leoni, F\. Tsimpourlas, F\. Song, F\. von Lohmann, F\. Sulit, G\. Salmon, G\. Parascandolo, G\. Chabot, G\. Zhao, G\. Brockman, G\. Leclerc, H\. Salman, H\. Bao, H\. Sheng, H\. Andrin, H\. Bagherinezhad, H\. Ren, H\. Lightman, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. Osband, I\. C\. Gilaberte, I\. Akkaya, I\. Kostrikov, I\. Sutskever, I\. Kofman, J\. Pachocki, J\. Lennon, J\. Wei, J\. Harb, J\. Twore, J\. Feng, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Q\. Candela, J\. Palermo, J\. Parish, J\. Heidecke, J\. Hallman, J\. Rizzo, J\. Gordon, J\. Uesato, J\. Ward, J\. Huizinga, J\. Wang, K\. Chen, K\. Xiao, K\. Singhal, K\. Nguyen, K\. Cobbe, K\. Shi, K\. Wood, K\. Rimbach, K\. Gu\-Lemberg, K\. Liu, K\. Lu, K\. Stone, K\. Yu, L\. Ahmad, L\. Yang, L\. Liu, L\. Maksin, L\. Ho, L\. Fedus, L\. Weng, L\. Li, L\. McCallum, L\. Held, L\. Kuhn, L\. Kondraciuk, L\. Kaiser, L\. Metz, M\. Boyd, M\. Trebacz, M\. Joglekar, M\. Chen, M\. Tintor, M\. Meyer, M\. Jones, M\. Kaufer, M\. Schwarzer, M\. Shah, M\. Yatbaz, M\. Y\. Guan, M\. Xu, M\. Yan, M\. Glaese, M\. Chen, M\. Lampe, M\. Malek, M\. Wang, M\. Fradin, M\. McClay, M\. Pavlov, M\. Wang, M\. Wang, M\. Murati, M\. Bavarian, M\. Rohaninejad, N\. McAleese, N\. Chowdhury, N\. Chowdhury, N\. Ryder, N\. Tezak, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, P\. Chao, P\. Ashbourne, P\. Izmailov, P\. Zhokhov, R\. Dias, R\. Arora, R\. Lin, R\. G\. Lopes, R\. Gaon, R\. Miyara, R\. Leike, R\. Hwang, R\. Garg, R\. Brown, R\. James, R\. Shu, R\. Cheu, R\. Greene, S\. Jain, S\. Altman, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Hernandez, S\. Baker, S\. McKinney, S\. Yan, S\. Zhao, S\. Hu, S\. Santurkar, S\. R\. Chaudhuri, S\. Zhang, S\. Fu, S\. Papay, S\. Lin, S\. Balaji, S\. Sanjeev, S\. Sidor, T\. Broda, A\. Clark, T\. Wang, T\. Gordon, T\. Sanders, T\. Patwardhan, T\. Sottiaux, T\. Degry, T\. Dimson, T\. Zheng, T\. Garipov, T\. Stasi, T\. Bansal, T\. Creech, T\. Peterson, T\. Eloundou, V\. Qi, V\. Kosaraju, V\. Monaco, V\. Pong, V\. Fomenko, W\. Zheng, W\. Zhou, W\. Zhan, W\. McCabe, W\. Zaremba, Y\. Dubois, Y\. Lu, Y\. Chen, Y\. Cha, Y\. Bai, Y\. He, Y\. Zhang, Y\. Wang, Z\. Shao, and Z\. Li\(2026\)OpenAI o1 system card\.External Links:2412\.16720,[Link](https://arxiv.org/abs/2412.16720)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1)\.
- \[28\]R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn\(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 53728–53741\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/a85b405ed65c6477a4fe8302b5e06ce7-Paper-Conference.pdf)Cited by:[§2\.1](https://arxiv.org/html/2605.30833#S2.SS1.p1.7)\.
- \[29\]R\. S\. Sutton and A\. G\. Barto\(2018\)Reinforcement learning: an introduction\.Second edition,The MIT Press\.External Links:[Link](http://incompleteideas.net/book/the-book-2nd.html)Cited by:[§2\.1](https://arxiv.org/html/2605.30833#S2.SS1.p1.7)\.
- \[30\]C\. Team, B\. Xiao, B\. Xia, B\. Yang, B\. Gao, B\. Shen, C\. Zhang, C\. He, C\. Lou, F\. Luo, G\. Wang, G\. Xie, H\. Zhang, H\. Lv, H\. Li, H\. Chen, H\. Xu, H\. Zhang, H\. Liu, J\. Duo, J\. Wei, J\. Xiao, J\. Dong, J\. Shi, J\. Hu, K\. Bao, K\. Zhou, L\. Li, L\. Zhao, L\. Zhang, P\. Li, Q\. Chen, S\. Liu, S\. Yu, S\. Cao, S\. Chen, S\. Yu, S\. Liu, T\. Zhou, W\. Su, W\. Wang, W\. Ma, X\. Deng, B\. Mao, B\. Ye, C\. Cai, C\. Wang, C\. Zhu, C\. Ma, C\. Chen, C\. Li, D\. Zhu, D\. Xiao, D\. Zhang, D\. Zhang, F\. Liu, F\. Yang, F\. Shi, G\. Wang, H\. Tian, H\. Wu, H\. Qu, H\. Yi, H\. An, H\. Guan, X\. Zhang, Y\. Song, Y\. Yan, Y\. Zhao, Y\. Lai, Y\. Gao, Y\. Cheng, Y\. Tian, Y\. Wang, Z\. Tang, Z\. Tang, Z\. Wen, Z\. Song, Z\. Zheng, Z\. Jiang, J\. Wen, J\. Sun, J\. Li, J\. Xue, J\. Xia, K\. Fang, M\. Zhu, N\. Chen, Q\. Tu, Q\. Zhang, Q\. Wang, R\. Li, R\. Ma, S\. Zhang, S\. Wang, S\. Li, S\. Gu, S\. Ren, S\. Deng, T\. Guo, T\. Lu, W\. Zhuang, W\. Zhang, W\. Xiong, W\. Huang, W\. Yang, X\. Zhang, X\. Yong, X\. Wang, X\. Xie, Y\. Jiang, Y\. Yang, Y\. He, Y\. Tu, Y\. Dong, Y\. Liu, Y\. Ma, Y\. Yu, Y\. Xiang, Z\. Huang, Z\. Lin, Z\. Xu, Z\. Chen, Z\. Deng, Z\. Zhang, and Z\. Yue\(2026\)MiMo\-v2\-flash technical report\.External Links:2601\.02780,[Link](https://arxiv.org/abs/2601.02780)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p1.7)\.
- \[31\]T\. Xiao, Y\. Yuan, M\. Li, Z\. Chen, and V\. G\. Honavar\(2025\)On a connection between imitation learning and rlhf\.External Links:2503\.05079,[Link](https://arxiv.org/abs/2503.05079)Cited by:[§2\.1](https://arxiv.org/html/2605.30833#S2.SS1.p1.5)\.
- \[32\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu\(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1),[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p1.7)\.
- \[33\]W\. Yang, W\. Liu, R\. Xie, K\. Yang, S\. Yang, and Y\. Lin\(2026\)Learning beyond teacher: generalized on\-policy distillation with reward extrapolation\.External Links:2602\.12125,[Link](https://arxiv.org/abs/2602.12125)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[34\]Z\. Yang, Z\. Liu, Y\. Chen, W\. Dai, B\. Wang, S\. Lin, C\. Lee, Y\. Chen, D\. Jiang, J\. He, R\. Pi, G\. Lam, N\. Lee, A\. Bukharin, M\. Shoeybi, B\. Catanzaro, and W\. Ping\(2026\)Nemotron\-cascade 2: post\-training llms with cascade rl and multi\-domain on\-policy distillation\.External Links:2603\.19220,[Link](https://arxiv.org/abs/2603.19220)Cited by:[§1](https://arxiv.org/html/2605.30833#S1.p1.1)\.
- \[35\]T\. Ye, L\. Dong, Z\. Chi, X\. Wu, S\. Huang, and F\. Wei\(2026\)Black\-box on\-policy distillation of large language models\.External Links:2511\.10643,[Link](https://arxiv.org/abs/2511.10643)Cited by:[§5](https://arxiv.org/html/2605.30833#S5.p1.1)\.
- \[36\]Y\. Zhang and T\. Math\-AI\(2024\)American invitational mathematics examination \(aime\) 2024\.Cited by:[§2\.2](https://arxiv.org/html/2605.30833#S2.SS2.p4.6),[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p3.1)\.
- \[37\]Y\. Zhang and T\. Math\-AI\(2025\)American invitational mathematics examination \(aime\) 2025\.Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p3.1)\.
- \[38\]Y\. Zhang and T\. Math\-AI\(2026\)American invitational mathematics examination \(aime\) 2026\.Cited by:[§4\.1](https://arxiv.org/html/2605.30833#S4.SS1.p3.1)\.
- \[39\]L\. Zheng, L\. Yin, Z\. Xie, C\. Sun, J\. Huang, C\. H\. Yu, S\. Cao, C\. Kozyrakis, I\. Stoica, J\. E\. Gonzalez, C\. Barrett, and Y\. Sheng\(2024\)SGLang: efficient execution of structured language model programs\.External Links:2312\.07104,[Link](https://arxiv.org/abs/2312.07104)Cited by:[Appendix A](https://arxiv.org/html/2605.30833#A1.p1.1)\.
- \[40\]Z\. Zhu, C\. Xie, X\. Lv, and slime Contributors\(2025\)Slime: an llm post\-training framework for rl scaling\.Note:[https://github\.com/THUDM/slime](https://github.com/THUDM/slime)GitHub repository\. Corresponding author: Xin LvCited by:[Appendix A](https://arxiv.org/html/2605.30833#A1.p1.1)\.
## Technical Appendices and Supplementary Material
Table of Contents
1. A\.[Training and Evaluation Details](https://arxiv.org/html/2605.30833#A1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A](https://arxiv.org/html/2605.30833#A1)
2. B\.[Additional Experimental Results](https://arxiv.org/html/2605.30833#A2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B](https://arxiv.org/html/2605.30833#A2) 1. B\.1\.[Effect of Renormalization in Top\-KKTraining](https://arxiv.org/html/2605.30833#A2.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.2](https://arxiv.org/html/2605.30833#A2.SS2) 2. B\.2\.[Effect of Rollout Temperature](https://arxiv.org/html/2605.30833#A2.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.3](https://arxiv.org/html/2605.30833#A2.SS3) 3. B\.3\.[Per\-Token Top\-KKLogit Visualization](https://arxiv.org/html/2605.30833#A2.SS4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.4](https://arxiv.org/html/2605.30833#A2.SS4) 4. B\.4\.[Future\-KL with GAE Weighting](https://arxiv.org/html/2605.30833#A2.SS5)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.5](https://arxiv.org/html/2605.30833#A2.SS5) 5. B\.5\.[Comparison of KL Divergence Objectives](https://arxiv.org/html/2605.30833#A2.SS6)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.6](https://arxiv.org/html/2605.30833#A2.SS6)
3. C\.[Sigmoid Fit of the SFD Curve](https://arxiv.org/html/2605.30833#A3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C](https://arxiv.org/html/2605.30833#A3)
4. D\.[Entropy\-Triggered Activation is Near\-Lossless](https://arxiv.org/html/2605.30833#A4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D](https://arxiv.org/html/2605.30833#A4)
5. E\.[Tree Attention: Cost Analysis and Practical Segmentation](https://arxiv.org/html/2605.30833#A5)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[E](https://arxiv.org/html/2605.30833#A5)
6. F\.[Proofs of Theoretical Results](https://arxiv.org/html/2605.30833#A6)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F](https://arxiv.org/html/2605.30833#A6) 1. F\.1\.[Proof of Proposition1\(Teacher Signal Vanishing under SFD\)](https://arxiv.org/html/2605.30833#A6.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F\.1](https://arxiv.org/html/2605.30833#A6.SS1) 2. F\.2\.[Proof Sketch of Proposition2\(Self\-Reinforcing Drift under Reverse\-KL\)](https://arxiv.org/html/2605.30833#A6.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F\.2](https://arxiv.org/html/2605.30833#A6.SS2) 3. F\.3\.[Proof of Proposition3\(One\-Step\-Ahead Discriminability\)](https://arxiv.org/html/2605.30833#A6.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F\.3](https://arxiv.org/html/2605.30833#A6.SS3) 4. F\.4\.[Proof of Proposition4\(Max\-ppas Relative Drift Indicator\)](https://arxiv.org/html/2605.30833#A6.SS4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F\.4](https://arxiv.org/html/2605.30833#A6.SS4) 5. F\.5\.[Group Normalization: Design Rationale and Formal Properties](https://arxiv.org/html/2605.30833#A6.SS5)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F\.5](https://arxiv.org/html/2605.30833#A6.SS5)
## Appendix ATraining and Evaluation Details
Training framework\.All models are trained on 4 nodes of 8×\\timesH20 GPUs \(32 GPUs total\) with the SLIME framework\[[40](https://arxiv.org/html/2605.30833#bib.bib38)\], using SGLang\[[39](https://arxiv.org/html/2605.30833#bib.bib39)\]as the inference backend\. We made two adaptations to SGLang: \(1\) we extended it to support tree\-structured attention masks, enabling the single\-pass multi\-candidate teacher evaluation described in Section[3\.3](https://arxiv.org/html/2605.30833#S3.SS3); \(2\) we patched SGLang to return temperature\-scaled logits for the prefill portion, which was required for experiments exploring non\-unit rollout temperatures \(see below\)\.
Training data\.For math training we use the Polaris dataset; for code training we use the DeepScaler code dataset\. This applies uniformly across all student model sizes\.
Rollout temperature\.All experiments use rollout temperature=1\.0=1\.0for both student and teacher\. We also evaluated two non\-unit configurations: \(1\)tstudent=1,tteacher=0\.2t\_\{\\mathrm\{student\}\}=1,t\_\{\\mathrm\{teacher\}\}=0\.2\(sharpened teacher distribution\); \(2\)tstudent=0\.8,tteacher=0\.8t\_\{\\mathrm\{student\}\}=0\.8,t\_\{\\mathrm\{teacher\}\}=0\.8\(symmetric low temperature\)\. Both configurations degraded performance\. We therefore fix temperature=1\.0=1\.0across all methods and settings\.
Baseline\-specific configurations\.
- •GRPO:Rather than training GRPO from scratch, we report results from publicly available RL\-trained checkpoints to ensure a fair and reproducible comparison\. For the1\.5B studentsetting, we use DeepScaler\-1\.5B, which was obtained by applying GRPO to DeepSeek\-R1\-Distill\-Qwen\-1\.5B on a large\-scale math dataset\. For the7B studentsetting, we use SkyWork\-OR1\-Math\-7B, which was obtained by applying GRPO to DeepSeek\-R1\-Distill\-Qwen\-7B\. Both checkpoints share the same base model as the corresponding distillation experiments, making the comparison directly controlled for initialization\.
- •OPD \(topk\) and LGR \(topk\):The per\-token top\-KKcandidate set is the union of the student’s top\-10 and teacher’s top\-10 tokens, renormalized to a valid probability distribution\. Renormalization is essential: without it the joint candidate set has inconsistent probability mass and training diverges\. The union construction also guarantees that teacher\-preferred tokens are always included even when the student assigns them low probability\.
- •REOPOLD:We follow the original training configuration but set staleness=1=1\(fully on\-policy\)\. The original paper uses staleness=4=4; we found this setting to be unstable in our experiments, likely because the larger policy lag interacts poorly with the entropy\-triggered reward clipping mechanism\.
- •JSD:We use mixture coefficientβ=0\.2\\beta=0\.2\(i\.e\.,πmix=0\.2πT\+0\.8πθ\\pi\_\{\\mathrm\{mix\}\}=0\.2\\,\\pi\_\{T\}\+0\.8\\,\\pi\_\{\\theta\}\)\.
Table 4:Training hyperparameters for the two student configurations\.Evaluation\.All models are evaluated with a maximum generation length of 39k tokens for math benchmarks and 36k tokens for code benchmarks \(code prompts are longer, leaving less budget for generation\), at temperature=1=1\. We report mean@8 and pass@8 across 8 responses per problem\.
On the choice of student models\.Our main experiments use DeepSeek\-R1\-Distill models as students \(1\.5B and 7B\), which are initialized from a base model via supervised fine\-tuning on reasoning traces\. We also experimented with applying OPD to Qwen3 reasoning models \(specifically, distilling from Qwen3\-32B to Qwen3\-1\.7B\), but found that training is highly unstable: performance initially improves but then degrades progressively rather than converging—a pattern visible in Figures[7](https://arxiv.org/html/2605.30833#A2.F7)and[8](https://arxiv.org/html/2605.30833#A2.F8), where even the best\-configured runs show initial gains followed by gradual degradation\. We attribute this to the nature of the Qwen3 reasoning models, which are the result of extensive integrated multi\-domain RL fusion applied to a broad mixture of tasks—yielding a well\-balanced student that has already converged on the teacher’s distribution across diverse contexts\. When such a model is used as a student in further OPD, the training signal is dominated by the small residual distribution mismatch rather than systematic supervision failures, making the optimization landscape highly sensitive and difficult to stabilize\. Note that our results differ significantly from those reported by Thinking Machine Lab on similar model families: their experiments use a base model or an SFT\-from\-base model as the student, rather than a fully trained reasoning model, which explains the more stable training dynamics they observe\.
In contrast, DeepSeek\-R1\-Distill models are derived from a base model through reasoning\-focused supervised learning, without the breadth of integrated multi\-domain RL fusion\. This leaves meaningful room for OPD to provide corrective supervision and makes the SFD phenomenon clearly observable—as the student has not already adapted to the teacher’s distribution on student\-generated prefixes\. Using this model family therefore provides a cleaner testbed for diagnosing and addressing SFD, and the instability observed on Qwen3 further underscores the importance of student model selection in OPD experiments\.
## Appendix BAdditional Experimental Results
### B\.1Homogeneous vs\. Heterogeneous Teacher–Student Pairs
In on\-policy distillation, the teacher and student can either share the same base model \(*homogeneous*\) or originate from different model families \(*heterogeneous*\)\[[25](https://arxiv.org/html/2605.30833#bib.bib42),patiño2025\_unlocking\_on\_policy\_distillation\_for\_any\_model\_family\]\. In the homogeneous setting, the teacher is obtained by further training the student base model via RLVR, so the two models share identical tokenization and pretraining priors—the student’s on\-policy distribution is close to the teacher’s from the outset\. In the heterogeneous setting \(used throughout this paper\), the student and teacher come from different base models, introducing a structural distribution gap that is present even before any distillation training begins\.
Figures[5](https://arxiv.org/html/2605.30833#A2.F5)and[6](https://arxiv.org/html/2605.30833#A2.F6)compare the joint distribution of per\-token student log\-probability and teacher log\-probability at training steps 0, 40, and 80 for the two settings\. In the homogeneous case, the scatter plots show a tight correlation between student and teacher log\-probabilities throughout training: the student’s on\-policy tokens remain well within the teacher’s in\-distribution region, and this alignment persists even as training progresses\. In the heterogeneous case, the scatter is substantially wider at all steps, with the student frequently visiting token positions where the teacher assigns low probability—precisely the regime in which supervision fidelity degrades\. This structural distribution gap is why the SFD phenomenon is clearly observable in the heterogeneous setting, and why our experiments adopt this configuration as the primary testbed\.
\(a\)Step 0
\(b\)Step 40
\(c\)Step 80
Figure 5:Homogeneous teacher–student pair: joint distribution of per\-token student log\-probability vs\. teacher log\-probability at training steps 0, 40, and 80\. The teacher is obtained by further RLVR training from the same student base model, so the two models share pretraining priors and the student tokens remain well within the teacher’s in\-distribution region throughout training\.\(a\)Step 0
\(b\)Step 40
\(c\)Step 80
Figure 6:Heterogeneous teacher–student pair: joint distribution of per\-token student log\-probability vs\. teacher log\-probability at training steps 0, 40, and 80\. The teacher and student originate from different base model families\. The scatter is substantially wider than the homogeneous case, indicating that the student frequently generates tokens in low\-probability regions of the teacher’s distribution—the primary driver of supervision fidelity decay\.
### B\.2Effect of Renormalization in Top\-KKTraining
In OPD \(topk\) and LGR \(topk\), each token’s distribution is restricted to the union of the student’s top\-10 and teacher’s top\-10 candidates, and the resulting truncated distribution is renormalized to sum to one before computing the KL loss\. Here we ablate the necessity of this renormalization step by training the same models without it—i\.e\., computing the KL directly over the raw unnormalized top\-KKunion probabilities\.
Figure[7](https://arxiv.org/html/2605.30833#A2.F7)compares OPD \(topk\) with and without renormalization on the Qwen3\-1\.7B student configuration across AIME benchmarks\. We make three observations\.
Unnormalized training shows higher teacher and student log\-probabilities\.Counterintuitively, models trained without renormalization exhibit*higher*teacher and student log\-probabilities throughout training\. We attribute this to an artifact of the truncation: without renormalization, the top\-KKprobabilities do not sum to one, which effectively inflates the absolute probability of each candidate token\. During rollout sampling this inflation makes it*easier*to draw low\-probability tokens, distorting the on\-policy distribution and masking the true quality of the supervision signal\.
Performance improves with larger top\-KKbut never matches renormalized training\.Among the unnormalized variants, increasing the total top\-KKsize \(top\-5→\\totop\-10\) improves performance, suggesting that a wider candidate set provides more signal\. However, even at the same total top\-KKcount, unnormalized training consistently underperforms its renormalized counterpart\. This is because without renormalization, all probability mass assigned to tokens*outside*the top\-KKunion is completely invisible to the KL loss—the model can shift mass to out\-of\-vocabulary positions without incurring any penalty\.
Renormalization closes an optimization loophole\.Without renormalization, the model can learn a degenerate strategy: push probability mass outside the top\-KKsupport \(escaping KL penalty entirely\), while simultaneously lowering the absolute probabilities of tokens within the top\-KKset, making the measured KL appear small\. With renormalization, any mass that leaks outside the top\-KKunion is implicitly redistributed back by the normalization step—the normalized probabilities within the top\-KKset rise whenever out\-of\-support mass increases, removing any incentive for this exploit and forcing the student to genuinely match the teacher on the retained candidates\.
Figure 7:Effect of renormalization in top\-KKdistillation training \(Qwen3\-1\.7B student\)\. Training curves and AIME\-24 performance for OPD \(topk\) with renormalization vs\. three unnormalized variants \(top\-5, top\-10, student\-only top\-5\)\. Without renormalization, training is unstable and final performance degrades substantially\.
### B\.3Effect of Rollout Temperature
All main experiments fix the rollout temperature to1\.01\.0for both the student and the teacher\. We additionally evaluate two non\-unit configurations: \(1\)tstudent=1,tteacher=0\.2t\_\{\\mathrm\{student\}\}=1,t\_\{\\mathrm\{teacher\}\}=0\.2\(sharpened teacher distribution\); \(2\)tstudent=0\.8,tteacher=0\.8t\_\{\\mathrm\{student\}\}=0\.8,t\_\{\\mathrm\{teacher\}\}=0\.8\(symmetric low temperature\)\.
tstudent=1,tteacher=0\.2t\_\{\\mathrm\{student\}\}=1,t\_\{\\mathrm\{teacher\}\}=0\.2\.A lower teacher temperature concentrates teacher probability mass sharply on its top tokens\. Intuitively this sharpens the supervision signal, but it also increases the mismatch between the temperature\-conditioned teacher logits and the student’s on\-policy distribution, causing training instability\.
tstudent=0\.8,tteacher=0\.8t\_\{\\mathrm\{student\}\}=0\.8,t\_\{\\mathrm\{teacher\}\}=0\.8\.Lowering both temperatures symmetrically reduces generation diversity and dilutes the discriminative signal between high\- and low\-confidence token choices, weakening the confidence reward\.
Figure[8](https://arxiv.org/html/2605.30833#A2.F8)shows training curves and final evaluation scores under both settings compared to the default temperature\-1\.01\.0baseline\. Both non\-unit configurations degrade performance, confirming that the symmetric unit temperature is the best operating point and that the SGLang prefill temperature\-scaling patch \(Appendix[A](https://arxiv.org/html/2605.30833#A1)\) is not needed in the final system\.
Figure 8:Effect of rollout temperature on training dynamics and final AIME\-24 performance \(Qwen3\-1\.7B student\)\. Conditions:tstudent=1,tteacher=1t\_\{\\mathrm\{student\}\}=1,t\_\{\\mathrm\{teacher\}\}=1\(default\),tstudent=1,tteacher=0\.2t\_\{\\mathrm\{student\}\}=1,t\_\{\\mathrm\{teacher\}\}=0\.2\(sharpened teacher\), andtstudent=0\.8,tteacher=0\.8t\_\{\\mathrm\{student\}\}=0\.8,t\_\{\\mathrm\{teacher\}\}=0\.8\(symmetric low temperature\)\. Both non\-unit configurations degrade performance, confirming that symmetric unit temperature is the optimal operating point\.
### B\.4Per\-Token Top\-KKLogit Visualization
To understand*why*per\-token RKL remains stubbornly high at certain positions even after an OPD gradient step, we visualize the full top\-KKlogit distribution of the student at individual tokens, comparing the distributions before and after each optimization step\.
High\-RKL tokens have dispersed student top\-KKlogits\.Figure[9](https://arxiv.org/html/2605.30833#A2.F9)shows representative tokens from a student rollout late in training\. For tokens where the RKL is large and remains large after the update step, the student’s top\-KKprobability mass is spread relatively uniformly across many candidates—the student is genuinely uncertain, and no single token dominates\. The teacher, by contrast, concentrates mass sharply on one or two tokens\. The gradient step reduces the KL slightly but cannot collapse the student distribution in a single step given the flat landscape\.
Low\-RKL tokens have logits concentrated at top\-1\.Tokens that achieve low RKL exhibit the opposite pattern: the student distribution is already sharply peaked at the same top\-1 token as the teacher\. These tokens contribute near\-zero loss and near\-zero gradient\.
Persistent high\-RKL late in training\.As training progresses, the proportion of low\-RKL \(peaked\) tokens grows, but a residual population of high\-RKL \(dispersed\) tokens persists and proves resistant to further optimization\. Crucially, in the standard per\-token average loss, the large and growing majority of low\-RKL tokens effectively act as a low\-magnitude denominator that dilutes the gradient contribution of the remaining high\-RKL tokens\. Each high\-RKL token’s gradient is upweighted in absolute terms but its relative influence in the batch average decreases as low\-RKL tokens accumulate—creating a natural but undesirable gradient imbalance\.




Figure 9:Per\-token top\-KKstudent logit distributions before and after an OPD update step for four representative tokens from a student rollout late in training \(DeepSeek\-R1\-Distill\-Qwen\-1\.5B student\)\. Each panel shows the token context, student top\-KKlogits and probabilities, teacher top\-KKlogits, and the per\-token KL divergence before and after the gradient step\. High\-RKL tokens exhibit dispersed student probability mass across many candidates, while the teacher concentrates sharply on one or two tokens; the gradient step reduces the KL only marginally, consistent with the flat optimization landscape described in Section[B\.4](https://arxiv.org/html/2605.30833#A2.SS4)\.
### B\.5Future\-KL with GAE Weighting
Motivated by the credit assignment literature in RL, we explore an alternative to the local per\-token RKL loss: rather than penalizing the student only for the divergence at positiontt, we compute a*future\-KL*signal by aggregating the RKL over all future tokenst′\>tt^\{\\prime\}\>tand weighting them with a generalized advantage estimate \(GAE,λ\\lambda\-return\) decay:
ℒfuture\-KL\(t\)=∑t′=tT\(ρλ\)t′−tRKL\(πθ\(⋅∣x<t′\)∥πT\(⋅∣x<t′\)\),\\mathcal\{L\}\_\{\\mathrm\{future\\text\{\-\}KL\}\}^\{\(t\)\}=\\sum\_\{t^\{\\prime\}=t\}^\{T\}\(\\rho\\lambda\)^\{t^\{\\prime\}\-t\}\\,\\mathrm\{RKL\}\\\!\\left\(\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{<t^\{\\prime\}\}\)\\,\\\|\\,\\pi\_\{T\}\(\\cdot\\mid x\_\{<t^\{\\prime\}\}\)\\right\),\(13\)whereρ\\rhois a discount factor andλ\\lambdais the GAE trace\-decay parameter\.Note that thisλ\\lambdais the GAE eligibility\-trace coefficient and is distinct from the confidence reward weightγ\\gammain LGR; the two hyperparameters play entirely different roles\.Intuitively, this encourages the student to make choices at positionttthat lead to lower KL*throughout*the trajectory, not just locally—which could in principle counteract the self\-reinforcing drift described in Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2)\.
Figure[10](https://arxiv.org/html/2605.30833#A2.F10)compares future\-KL distillation against standard OPD under a sweep of\(ρ,λ\)\(\\rho,\\lambda\)pairs\. Despite the appealing intuition, future\-KL consistently fails to improve over standard per\-token RKL\. We hypothesize that the difficulty lies in the credit assignment itself: in long reasoning chains, the future\-KL signal at early positions is dominated by the noise accumulated over hundreds of subsequent tokens, making the gradient at positiontteffectively uninformative about the local decision quality\. This suggests that the right way to address long\-horizon supervision degradation is through the one\-step\-ahead lookahead used in LGR, rather than through discounted future aggregation\.
Figure 10:Future\-KL distillation with GAE weighting across discount factorsγ∈\{0\.9999,0\.999,0\.995,0\.99\}\\gamma\\in\\\{0\.9999,0\.999,0\.995,0\.99\\\}, compared to standard per\-token OPD \(DeepSeek\-R1\-Distill\-Qwen\-1\.5B student\)\. Despite varying the discount factor across a wide range, future\-KL consistently fails to improve over the per\-token RKL baseline on AIME\-24, with more aggressive discounting \(γ=0\.999,0\.995\\gamma=0\.999,0\.995\) causing entropy collapse and severe performance degradation\.
### B\.6Comparison of KL Divergence Objectives
We compare three on\-policy distillation objectives that differ in the direction of the KL divergence: reverse\-KL \(OPD/RKLD\), forward\-KL \(FKLD\), and Jensen–Shannon divergence \(JSD\)\. Figure[11](https://arxiv.org/html/2605.30833#A2.F11)tracks training dynamics and AIME\-24 performance across all three\.
Forward\-KL collapses\.FKLD training is highly unstable: teacher log\-probability degrades sharply after∼\\sim50 steps, entropy explodes, and AIME\-24 performance drops toward zero and does not recover\. This confirms the exposure bias problem of forward\-KL in the on\-policy setting—the student is forced to cover the full teacher distribution using its own generated prefixes, leading to mode\-covering behavior that pushes the student out of distribution\.
JSD provides moderate stability but underperforms OPD\.JSD achieves intermediate stability: entropy grows moderately and teacher log\-probability declines more gradually than FKLD\. However, final AIME\-24 performance remains below OPD\(RKLD\), consistent with the main table results\. The mixture objective partially inherits the forward\-KL instability without the full benefits of reverse\-KL’s mode\-seeking behavior\.
OPD \(RKLD\) is the strongest baseline\.Reverse\-KL maintains the lowest distill KL, highest teacher log\-probability, and stable entropy throughout training, confirming it as the appropriate base objective—and motivating LGR as an enhancement of OPD rather than a replacement of the divergence choice\.
Figure 11:Training dynamics and AIME\-24 performance for three KL divergence objectives \(DeepSeek\-R1\-Distill\-Qwen\-1\.5B student\): OPD \(reverse\-KL\), forward\-KL \(FKLD\), and Jensen–Shannon divergence \(JSD\)\. FKLD collapses due to exposure bias; JSD is moderately stable but underperforms OPD; reverse\-KL achieves the best training stability and final performance\.
## Appendix CSigmoid Fit of the SFD Curve
Section[2\.4](https://arxiv.org/html/2605.30833#S2.SS4)predicts that the SFD curve𝒞\(t\)\\mathcal\{C\}^\{\(t\)\}follows a sigmoidal decay, and that OPD training causes the transition point to shift leftward \(supervision boundary contraction\)\. We verify this empirically by fitting the teacher completion accuracy curves from Figure[2](https://arxiv.org/html/2605.30833#S1.F2)to a parametric sigmoid:
𝒞\(t\)=A1\+eε\(t−t∗\)\+b,\\mathcal\{C\}^\{\(t\)\}=\\frac\{A\}\{1\+e^\{\\,\\varepsilon\\,\(t\-t^\{\*\}\)\}\}\+b,\(14\)whereAAis the decay amplitude,ε\\varepsilonis the transition steepness,t∗t^\{\*\}is the midpoint of the transition \(supervision boundary\), andbbis the residual supervision floor\.
Fitted parameters\.Table[5](https://arxiv.org/html/2605.30833#A3.T5)reports the fitted parameters for the base student \(before OPD\) and the OPD\-trained student \(after OPD\), along with goodness\-of\-fitR2R^\{2\}\.
Table 5:Sigmoid fit parameters for the SFD curve before and after OPD training\.AAε\\varepsilont∗t^\{\*\}\(k tokens\)bbUpper asymptoteR2R^\{2\}Before OPD \(\+\+\)0\.34120\.21457\.430\.41330\.750\.9980After OPD \(×\\times\)0\.54520\.26087\.040\.22940\.770\.9977Δ\\Delta\+0\.20\+0\.20\+0\.046\+0\.046−0\.39\-0\.39−0\.18\-0\.18——Interpretation\.All four parameter shifts are consistent with the vicious cycle prediction from Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2):
- •t∗t^\{\*\}shifts left\(7\.43→7\.047\.43\\to\\mathbf\{7\.04\}k,Δ=−0\.39\\Delta=\-0\.39k\): the supervision boundary contracts after OPD training, confirming optimized trajectories push the teacher into OOD regions earlier\.
- •ε\\varepsilonincreases\(0\.21→0\.260\.21\\to\\mathbf\{0\.26\},Δ=\+0\.046\\Delta=\+0\.046\): the transition steepens, meaning supervision quality collapses more abruptly once the boundary is crossed\.
- •Floorbbdrops\(0\.41→0\.230\.41\\to\\mathbf\{0\.23\},Δ=−0\.18\\Delta=\-0\.18\): residual supervision quality at long prefixes degrades substantially after OPD training\.
- •AmplitudeAAincreases\(0\.34→0\.550\.34\\to\\mathbf\{0\.55\},Δ=\+0\.20\\Delta=\+0\.20\): total supervision loss is larger post\-training, indicating deeper drift into the teacher’s incompetent region\.
These two snapshots \(before/after training\) do not constitute a full training trajectory, so we do not claim precise estimates of the logistic growth parameters\. However, all four parameter shifts are in the direction predicted by the vicious cycle analysis, providing empirical support for the sigmoidal SFD structure described in Section[2\.4](https://arxiv.org/html/2605.30833#S2.SS4)\.
## Appendix DEntropy\-Triggered Activation is Near\-Lossless
###### Observation 1\(Entropy\-triggered activation is near\-lossless\)\.
At low\-entropy positions \(ℋ\(πθ\(⋅\|𝐱<t\)\)≤τ\\operatorname\{\\mathcal\{H\}\}\(\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)\\leq\\tau\), the student’s top\-1 candidate dominates:πθ\(xt\(1\)\)≫πθ\(xt\(k\)\)\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\gg\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\)fork≥2k\\geq 2\. The combined loss in Eq\. \([12](https://arxiv.org/html/2605.30833#S3.E12)\) is:
ℒtLGR=∑k=1Kπθ\(xt\(k\)\|𝐱<t\)⋅\[At\(k\)\+γ⋅rconf\(xt\(k\)\)\]\.\\mathcal\{L\}\_\{t\}^\{LGR\}=\\sum\_\{k=1\}^\{K\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\|\\mathbf\{x\}\_\{<t\}\)\\cdot\\left\[A\_\{t\}^\{\(k\)\}\+\\gamma\\cdot r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)\\right\]\.Whenπθ\(xt\(1\)\)≈1\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\approx 1andπθ\(xt\(k\)\)≈0\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\)\\approx 0fork≥2k\\geq 2, this reduces toℒtLGR≈At\(1\)\+γ⋅rconf\(xt\(1\)\)\\mathcal\{L\}\_\{t\}^\{LGR\}\\approx A\_\{t\}^\{\(1\)\}\+\\gamma\\cdot r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(1\)\}\)\. Sincerconf\(xt\(1\)\)r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(1\)\}\)is a single\-sample reward with near\-zero gradient contribution whenπθ\(xt\(1\)\)≈1\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\approx 1\(the student has already committed to this token with probability 1\), the confidence reward term adds negligible signal\. Formally, the gradient of the confidence reward term with respect toθ\\thetais:
∇θ\[πθ\(xt\(1\)\)⋅rconf\(xt\(1\)\)\]=rconf\(xt\(1\)\)⋅∇θπθ\(xt\(1\)\)≈rconf\(xt\(1\)\)⋅0=0,\\nabla\_\{\\theta\}\\left\[\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\cdot r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(1\)\}\)\\right\]=r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(1\)\}\)\\cdot\\nabla\_\{\\theta\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\approx r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(1\)\}\)\\cdot 0=0,sinceπθ\(xt\(1\)\)≈1\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\approx 1implies∇θπθ\(xt\(1\)\)≈0\\nabla\_\{\\theta\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(1\)\}\)\\approx 0\(the probability is saturated\)\. Disabling the confidence reward at low\-entropy positions therefore incurs near\-zero information loss while savingKKteacher forward\-pass evaluations per position\.
## Appendix ETree Attention: Cost Analysis and Practical Segmentation
Figure 12:Tree attention for efficient confidence\-reward computation\.*Left*: the naïve approach requiresK×\|𝒮\|K\\times\|\\mathcal\{S\}\|separate teacher forward passes—one per candidate per high\-entropy position\.*Middle*: tree\-attention construction\. At each high\-entropy positiont∈𝒮t\\in\\mathcal\{S\}, the top\-KKcandidate tokens are appended as branches off the main sequence\. Main\-branch tokens attend causally; each candidate attends only to its own prefix𝐱<t\\mathbf\{x\}\_\{<t\}\(shaded region\)\.*Right*: the resulting sparse attention mask𝐌\\mathbf\{M\}, reducingK×\|𝒮\|K\\times\|\\mathcal\{S\}\|passes to a single teacher forward pass\.Baseline cost \(naïve prefill\)\.Without tree attention, evaluating candidatext\(k\)x\_\{t\}^\{\(k\)\}requires prefilling the contextx1,…,xt−1,xt\(k\)x\_\{1\},\\ldots,x\_\{t\-1\},x\_\{t\}^\{\(k\)\}from scratch—the main sequence KV cache cannot be reused because positionst∈𝒮t\\in\\mathcal\{S\}have different prefix lengths\. The prefill cost for a sequence of lengthttisO\(t2d\)O\(t^\{2\}d\), giving total cost:
Cnaïve=K∑t∈𝒮t2d≈K⋅\|𝒮\|⋅L23⋅d=0\.2K3L3d,C\_\{\\text\{naïve\}\}=K\\sum\_\{t\\in\\mathcal\{S\}\}t^\{2\}d\\;\\approx\\;K\\cdot\|\\mathcal\{S\}\|\\cdot\\frac\{L^\{2\}\}\{3\}\\cdot d\\;=\\;\\frac\{0\.2K\}\{3\}\\,L^\{3\}d,where we assume positions in𝒮\\mathcal\{S\}are roughly uniform over\[0,L\]\[0,L\], giving∑t∈𝒮t2≈\|𝒮\|⋅L2/3\\sum\_\{t\\in\\mathcal\{S\}\}t^\{2\}\\approx\|\\mathcal\{S\}\|\\cdot L^\{2\}/3\. This isO\(L3\)O\(L^\{3\}\), cubically expensive\.
Tree attention construction\.Given𝐱=\(x1,…,xL\)\\mathbf\{x\}=\(x\_\{1\},\\ldots,x\_\{L\}\)and𝒮\\mathcal\{S\}, we append allK\|𝒮\|K\|\\mathcal\{S\}\|candidate tokens after the main sequence and apply a tree\-structured mask𝐌\\mathbf\{M\}:
- •Main tokens\(1,…,L1,\\ldots,L\): standard causal mask\.
- •Candidatext\(k\)x\_\{t\}^\{\(k\)\}: attends to𝐱≤t−1\\mathbf\{x\}\_\{\\leq t\-1\}and itself only\.
The main sequence is prefilled once \(costO\(L2d\)O\(L^\{2\}d\)\); each candidate is then a single decode step attending tot−1t\-1cached keys \(costO\(td\)O\(t\\,d\)per candidate\)\. Total:
Ctree=L2d\+K∑t∈𝒮td≈L2d\+K⋅\|𝒮\|⋅L2⋅d=L2d\(1\+0\.1K\)\.C\_\{\\text\{tree\}\}=L^\{2\}d\+K\\sum\_\{t\\in\\mathcal\{S\}\}t\\,d\\;\\approx\\;L^\{2\}d\+K\\cdot\|\\mathcal\{S\}\|\\cdot\\frac\{L\}\{2\}\\cdot d\\;=\\;L^\{2\}d\\,\(1\+0\.1K\)\.This isO\(L2\)O\(L^\{2\}\)—one order of magnitude cheaper than the naïve baseline\.
Speedup\.
Speedup=CnaïveCtree=0\.2KL/31\+0\.1K\.\\text\{Speedup\}=\\frac\{C\_\{\\text\{naïve\}\}\}\{C\_\{\\text\{tree\}\}\}=\\frac\{0\.2KL/3\}\{1\+0\.1K\}\.The speedup grows*linearly*withLLbecause naïve scales asO\(L3\)O\(L^\{3\}\)while tree attention scales asO\(L2\)O\(L^\{2\}\)\. ForK=8K=8,L=16kL=16\\text\{k\}: Speedup≈𝟒𝟕𝟒𝟏×\\approx\\mathbf\{4741\\times\}\.
Practical segmentation\.In practice, the total sequence lengthL\+K\|𝒮\|/NL\+K\|\\mathcal\{S\}\|/Nmust fit within GPU memoryLmaxGPUL\_\{\\max\}^\{\\text\{GPU\}\}, requiring at leastN∗=⌈K\|𝒮\|/\(LmaxGPU−L\)⌉N^\{\*\}=\\lceil K\|\\mathcal\{S\}\|/\(L\_\{\\max\}^\{\\text\{GPU\}\}\-L\)\\rceilsegments\. WithNNsegments, the main sequence is re\-prefilledNNtimes:
Ctree,N=N⋅L2d\+K∑t∈𝒮td≈L2d\(N\+0\.1K\),SpeedupN=0\.2KL/3N\+0\.1K\.C\_\{\\text\{tree\},N\}=N\\cdot L^\{2\}d\+K\\sum\_\{t\\in\\mathcal\{S\}\}t\\,d\\approx L^\{2\}d\\,\(N\+0\.1K\),\\qquad\\text\{Speedup\}\_\{N\}=\\frac\{0\.2KL/3\}\{N\+0\.1K\}\.In our experiments \(L=16kL=16\\text\{k\},K=8K=8,\|𝒮\|≈3\.2k\|\\mathcal\{S\}\|\\approx 3\.2\\text\{k\}\), we useN=8N=8segments to stay within GPU memory, giving:
SpeedupN=8=0\.2⋅8⋅16000/38\+0\.1⋅8≈85338\.8≈970×\.\\text\{Speedup\}\_\{N=8\}=\\frac\{0\.2\\cdot 8\\cdot 16000/3\}\{8\+0\.1\\cdot 8\}\\approx\\frac\{8533\}\{8\.8\}\\approx 970\\times\.Even with segmentation, the speedup remains large because theO\(L3\)O\(L^\{3\}\)vs\.O\(L2\)O\(L^\{2\}\)gap dominates\.
## Appendix FProofs of Theoretical Results
### F\.1Proof of Proposition[1](https://arxiv.org/html/2605.30833#Thmproposition1)
###### Proposition\(Teacher signal vanishing under SFD, restated\)\.
Decompose the per\-token advantage asAt=\(1\+logπθ\(xt\|𝐱<t\)\)−logπT\(xt\|𝐱<t\)A\_\{t\}=\(1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\)\-\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\. Define the teacher’s discriminative signal asΔT\(t\)≔Varxt∼πθ\[logπT\(xt\|𝐱<t\)\]\\Delta\_\{T\}\(t\)\\coloneqq\\mathrm\{Var\}\_\{x\_\{t\}\\sim\\pi\_\{\\theta\}\}\[\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\]\. Then:
1. 1\.WhenπT\(⋅\|𝐱<t\)=Uniform\(𝒱\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)=\\mathrm\{Uniform\}\(\\mathcal\{V\}\), we haveΔT\(t\)=0\\Delta\_\{T\}\(t\)=0andAtA\_\{t\}depends only on the student\.
2. 2\.ΔT\(t\)≤log2\|𝒱\|−Ent\(πT\(⋅\|𝐱<t\)\)2\\Delta\_\{T\}\(t\)\\leq\\log^\{2\}\|\\mathcal\{V\}\|\-\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)^\{2\}\.
3. 3\.SNRT\(t\)=O\(ΔT\(t\)\)\\mathrm\{SNR\}\_\{T\}\(t\)=O\(\\Delta\_\{T\}\(t\)\)and decreases monotonically under SFD\.
###### Proof\.
Part \(1\)\.WhenπT\(⋅\|𝐱<t\)=Uniform\(𝒱\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)=\\mathrm\{Uniform\}\(\\mathcal\{V\}\), we haveπT\(xt\|𝐱<t\)=1/\|𝒱\|\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)=1/\|\\mathcal\{V\}\|for everyxt∈𝒱x\_\{t\}\\in\\mathcal\{V\}, sologπT\(xt\|𝐱<t\)=−log\|𝒱\|\\log\\pi\_\{T\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)=\-\\log\|\\mathcal\{V\}\|is a constant independent ofxtx\_\{t\}\. ThereforeΔT\(t\)=Varxt\[−log\|𝒱\|\]=0\\Delta\_\{T\}\(t\)=\\mathrm\{Var\}\_\{x\_\{t\}\}\[\-\\log\|\\mathcal\{V\}\|\]=0, and:
At=1\+logπθ\(xt\|𝐱<t\)−\(−log\|𝒱\|\)=1\+logπθ\(xt\|𝐱<t\)\+log\|𝒱\|\.A\_\{t\}=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\-\(\-\\log\|\\mathcal\{V\}\|\)=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\+\\log\|\\mathcal\{V\}\|\.This depends only onπθ\\pi\_\{\\theta\}and the constantlog\|𝒱\|\\log\|\\mathcal\{V\}\|, providing no teacher correction signal\.
Part \(2\)\.LetP=πT\(⋅\|𝐱<t\)P=\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)be a distribution over𝒱\\mathcal\{V\}\. SinceP\(v\)∈\(0,1\]P\(v\)\\in\(0,1\]for allvv, we havelogP\(v\)∈\[−log\|𝒱\|,0\]\\log P\(v\)\\in\[\-\\log\|\\mathcal\{V\}\|,0\]\. Define the random variableZ=logP\(xt\)Z=\\log P\(x\_\{t\}\)wherext∼πθ\(⋅\|𝐱<t\)x\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\. We boundVar\[Z\]\\mathrm\{Var\}\[Z\]by bounding the second moment\.
By the Bhatia–Davis inequality\[[4](https://arxiv.org/html/2605.30833#bib.bib41)\], for a bounded random variableZ∈\[a,b\]Z\\in\[a,b\]:
Var\[Z\]≤\(b−𝔼\[Z\]\)\(𝔼\[Z\]−a\)\.\\mathrm\{Var\}\[Z\]\\leq\(b\-\\mathbb\{E\}\[Z\]\)\(\\mathbb\{E\}\[Z\]\-a\)\.Herea=−log\|𝒱\|a=\-\\log\|\\mathcal\{V\}\|,b=0b=0\. The expected value satisfies:
𝔼xt∼πθ\[logP\(xt\)\]=∑vπθ\(v\)logP\(v\)\.\\mathbb\{E\}\_\{x\_\{t\}\\sim\\pi\_\{\\theta\}\}\[\\log P\(x\_\{t\}\)\]=\\sum\_\{v\}\\pi\_\{\\theta\}\(v\)\\log P\(v\)\.In the worst case \(maximizing variance\), this equals the cross\-entropyH\(πθ,P\)H\(\\pi\_\{\\theta\},P\)\. Applying the AM\-GM inequality to the Bhatia–Davis bound:
Var\[Z\]≤\(b−a\)24=log2\|𝒱\|4\.\\mathrm\{Var\}\[Z\]\\leq\\frac\{\(b\-a\)^\{2\}\}\{4\}=\\frac\{\\log^\{2\}\|\\mathcal\{V\}\|\}\{4\}\.A tighter bound exploiting the entropy ofPPproceeds as follows\. Note that whenPPis uniform,logP\(v\)=−log\|𝒱\|\\log P\(v\)=\-\\log\|\\mathcal\{V\}\|for allvv, soVar\[Z\]=0\\mathrm\{Var\}\[Z\]=0regardless ofπθ\\pi\_\{\\theta\}\. AsPPbecomes more peaked \(lower entropy\), the spread oflogP\(v\)\\log P\(v\)over the vocabulary increases, allowingΔT\(t\)\\Delta\_\{T\}\(t\)to grow\. Using the variance\-entropy relationship for log\-probabilities \(a standard result in information theory\[[6](https://arxiv.org/html/2605.30833#bib.bib40)\]\):
Varxt∼πθ\[logP\(xt\)\]≤log2\|𝒱\|−Ent\(P\)2,\\mathrm\{Var\}\_\{x\_\{t\}\\sim\\pi\_\{\\theta\}\}\[\\log P\(x\_\{t\}\)\]\\leq\\log^\{2\}\|\\mathcal\{V\}\|\-\\mathrm\{Ent\}\(P\)^\{2\},where equality holds whenPPis supported on a single token \(maximum peakedness\)\. AsEnt\(P\)→log\|𝒱\|\\mathrm\{Ent\}\(P\)\\to\\log\|\\mathcal\{V\}\|\(P becomes uniform\),ΔT\(t\)≤log2\|𝒱\|−log2\|𝒱\|=0\\Delta\_\{T\}\(t\)\\leq\\log^\{2\}\|\\mathcal\{V\}\|\-\\log^\{2\}\|\\mathcal\{V\}\|=0, confirmingΔT\(t\)→0\\Delta\_\{T\}\(t\)\\to 0\.
Part \(3\)\.The gradient direction attributable to the teacher is determined by the variation oflogπT\(xt\)\\log\\pi\_\{T\}\(x\_\{t\}\)across token choices\. Specifically, the teacher’s contribution to the policy gradient is:
gT=𝔼xt∼πθ\[∇θlogπθ\(xt\)⋅logπT\(xt\)\]\.g\_\{T\}=\\mathbb\{E\}\_\{x\_\{t\}\\sim\\pi\_\{\\theta\}\}\\\!\\left\[\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(x\_\{t\}\)\\cdot\\log\\pi\_\{T\}\(x\_\{t\}\)\\right\]\.The signal\-to\-noise ratio of this term scales as the standard deviation oflogπT\(xt\)\\log\\pi\_\{T\}\(x\_\{t\}\)divided by the standard deviation of∇θlogπθ\(xt\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(x\_\{t\}\), i\.e\.,SNRT\(t\)∝ΔT\(t\)\\mathrm\{SNR\}\_\{T\}\(t\)\\propto\\sqrt\{\\Delta\_\{T\}\(t\)\}\. Under SFD, as the student’s prefix drifts further from the teacher’s training distribution,Ent\(πT\(⋅\|𝐱<t\)\)\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)increases monotonically \(the teacher’s distribution becomes more diffuse\), soΔT\(t\)\\Delta\_\{T\}\(t\)decreases monotonically by Part \(2\), andSNRT\(t\)=O\(ΔT\(t\)\)→0\\mathrm\{SNR\}\_\{T\}\(t\)=O\(\\Delta\_\{T\}\(t\)\)\\to 0\. ∎
### F\.2Proof Sketch of Proposition[2](https://arxiv.org/html/2605.30833#Thmproposition2)
###### Proposition\(Self\-reinforcing drift under reverse\-KL, restated\)\.
Define distributional driftdt≔D\(πT\(⋅\|𝐱<tθ\),πT\(⋅\|𝐱<t∗\)\)d\_\{t\}\\coloneqq D\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\\theta\}\),\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\*\}\)\), and assumeΔT\(t\)\\Delta\_\{T\}\(t\)is non\-increasing indtd\_\{t\}\(greater drift degrades teacher discriminability\)\. Under reverse\-KL with stop\-gradient:
1. 1\.WhenΔT\(t\)=0\\Delta\_\{T\}\(t\)=0, the advantageAt=1\+logπθ\(xt\)\+log\|𝒱\|A\_\{t\}=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\)\+\\log\|\\mathcal\{V\}\|reinforces the student’s existing mode without teacher correction\.
2. 2\.WhenΔT\(t\)<δcrit\\Delta\_\{T\}\(t\)<\\delta\_\{\\mathrm\{crit\}\},𝔼\[dt\+1\|dt\]≥dt\\mathbb\{E\}\[d\_\{t\+1\}\|d\_\{t\}\]\\geq d\_\{t\}, creating a positive feedback loop\.
3. 3\.Forward\-KL avoids SFD by construction but introduces exposure bias\.
###### Proof sketch\.
Part \(1\)\.From Part \(1\) of Proposition[1](https://arxiv.org/html/2605.30833#Thmproposition1), whenΔT\(t\)=0\\Delta\_\{T\}\(t\)=0:
At=1\+logπθ\(xt\|𝐱<t\)\+log\|𝒱\|\.A\_\{t\}=1\+\\log\\pi\_\{\\theta\}\(x\_\{t\}\|\\mathbf\{x\}\_\{<t\}\)\+\\log\|\\mathcal\{V\}\|\.Under gradient descent onℒrkl\\mathcal\{L\}\_\{\\mathrm\{rkl\}\}, the update toπθ\(xt\)\\pi\_\{\\theta\}\(x\_\{t\}\)is proportional to−At\-A\_\{t\}\. SinceAtA\_\{t\}is an increasing function oflogπθ\(xt\)\\log\\pi\_\{\\theta\}\(x\_\{t\}\), the gradient is negative \(decreasingπθ\(xt\)\\pi\_\{\\theta\}\(x\_\{t\}\)\) whenπθ\(xt\)\>e−\(1\+log\|𝒱\|\)=\(e⋅\|𝒱\|\)−1\\pi\_\{\\theta\}\(x\_\{t\}\)\>e^\{\-\(1\+\\log\|\\mathcal\{V\}\|\)\}=\(e\\cdot\|\\mathcal\{V\}\|\)^\{\-1\}\. For any tokenxtx\_\{t\}assigned probability above this threshold \(which holds for the student’s top tokens whenever the distribution is non\-uniform\), the gradient reinforces the student’s high\-probability tokens\. Specifically, high\-probability tokens haveAt≫0A\_\{t\}\\gg 0, receiving strong gradient, while low\-probability tokens haveAt<0A\_\{t\}<0and are further suppressed—sharpening the distribution toward the student’s existing mode without any teacher guidance\.
Part \(2\)\.By assumption,ΔT\(t\)≤f\(dt\)\\Delta\_\{T\}\(t\)\\leq f\(d\_\{t\}\)withffdecreasing\. WhenΔT\(t\)<δcrit\\Delta\_\{T\}\(t\)<\\delta\_\{\\mathrm\{crit\}\}, the teacher contributes negligible correction signal \(Proposition[1](https://arxiv.org/html/2605.30833#Thmproposition1), Part 3\)\. Therefore, the tokenxtx\_\{t\}selected at stepttis drawn from the student’s sharpened distributionπθ\(⋅\|𝐱<t\)\\pi\_\{\\theta\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)rather than being guided toward the teacher’s preferred continuation\. Letxt∗∈argmaxvπT\(v\|𝐱<t∗\)x\_\{t\}^\{\*\}\\in\\arg\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\}^\{\*\}\)be the teacher’s most likely token at the corresponding teacher\-generated position\. Since the student selects from its own mode, we havext≠xt∗x\_\{t\}\\neq x\_\{t\}^\{\*\}with probability bounded away from zero\. Appending the divergent tokenxtx\_\{t\}to the context shifts the student’s prefix further from the teacher’s distribution:
dt\+1=D\(πT\(⋅\|𝐱≤tθ\),πT\(⋅\|𝐱≤t∗\)\)≥dt\+δd\_\{t\+1\}=D\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{\\leq t\}^\{\\theta\}\),\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{\\leq t\}^\{\*\}\)\)\\geq d\_\{t\}\+\\deltafor someδ\>0\\delta\>0depending on the token divergence\. This establishes the positive feedback loop:dtd\_\{t\}increases wheneverΔT\(t\)\\Delta\_\{T\}\(t\)is below the critical threshold, and increasingdtd\_\{t\}further reducesΔT\(t\)\\Delta\_\{T\}\(t\), sustaining the loop\.
Part \(3\)\.Under forward\-KL \(off\-policy\), the training sequences are teacher\-generated:𝐱∼πT\(⋅\|𝐜\)\\mathbf\{x\}\\sim\\pi\_\{T\}\(\\cdot\|\\mathbf\{c\}\)\. Therefore𝐱<tθ=𝐱<t∗\\mathbf\{x\}\_\{<t\}^\{\\theta\}=\\mathbf\{x\}\_\{<t\}^\{\*\}for alltt, anddt=D\(πT\(⋅\|𝐱<t∗\),πT\(⋅\|𝐱<t∗\)\)=0d\_\{t\}=D\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\*\}\),\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}^\{\*\}\)\)=0throughout\. SFD does not arise by construction\. The cost is that at inference time the student generates𝐱∼πθ\\mathbf\{x\}\\sim\\pi\_\{\\theta\}, but during training it always conditioned on teacher\-generated prefixes𝐱∼πT\\mathbf\{x\}\\sim\\pi\_\{T\}—a train\-test mismatch known as exposure bias\. ∎
### F\.3Proof of Proposition[3](https://arxiv.org/html/2605.30833#Thmproposition3)
###### Proposition\(One\-step\-ahead discriminability, restated\)\.
DefineDahead\(t\)≔maxk,k′\|maxvπT\(v\|𝐱<t,xt\(k\)\)−maxvπT\(v\|𝐱<t,xt\(k′\)\)\|D\_\{\\mathrm\{ahead\}\}\(t\)\\coloneqq\\max\_\{k,k^\{\\prime\}\}\\left\|\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k\)\}\)\-\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k^\{\\prime\}\)\}\)\\right\|\. ThenDahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)is independent ofEnt\(πT\(⋅\|𝐱<t\)\)\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\): even whenπT\(⋅\|𝐱<t\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)is uniform \(ΔT\(t\)=0\\Delta\_\{T\}\(t\)=0\),Dahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)can be arbitrarily large\.
###### Proof\.
We prove by explicit construction\. Let𝒱=\{a,b,c\}\\mathcal\{V\}=\\\{a,b,c\\\}and supposeπT\(⋅\|𝐱<t\)=\(1/3,1/3,1/3\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)=\(1/3,1/3,1/3\), i\.e\., perfectly uniform at positionttwithEnt\(πT\(⋅\|𝐱<t\)\)=log3\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)=\\log 3\(maximum entropy,ΔT\(t\)=0\\Delta\_\{T\}\(t\)=0\)\. Assign the following teacher distributions at positiont\+1t\{\+\}1:
πT\(⋅\|𝐱<t,a\)\\displaystyle\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\},a\)=\(1−2ε,ε,ε\),maxv=1−2ε,\\displaystyle=\(1\-2\\varepsilon,\\;\\varepsilon,\\;\\varepsilon\),\\quad\\max\_\{v\}=1\-2\\varepsilon,πT\(⋅\|𝐱<t,b\)\\displaystyle\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\},b\)=\(1/3,1/3,1/3\),maxv=1/3,\\displaystyle=\(1/3,\\;1/3,\\;1/3\),\\quad\\max\_\{v\}=1/3,πT\(⋅\|𝐱<t,c\)\\displaystyle\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\},c\)=\(1/3,1/3,1/3\),maxv=1/3,\\displaystyle=\(1/3,\\;1/3,\\;1/3\),\\quad\\max\_\{v\}=1/3,for anyε∈\(0,1/3\)\\varepsilon\\in\(0,1/3\)\. Then:
Dahead\(t\)=\(1−2ε\)−13=23−2ε\.D\_\{\\mathrm\{ahead\}\}\(t\)=\(1\-2\\varepsilon\)\-\\frac\{1\}\{3\}=\\frac\{2\}\{3\}\-2\\varepsilon\.Asε→0\\varepsilon\\to 0,Dahead\(t\)→2/3D\_\{\\mathrm\{ahead\}\}\(t\)\\to 2/3, which is arbitrarily large relative to the uniform bound\. This construction is valid for any vocabulary size\|𝒱\|≥2\|\\mathcal\{V\}\|\\geq 2and any level of local entropyEnt\(πT\(⋅\|𝐱<t\)\)=log\|𝒱\|\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\)=\\log\|\\mathcal\{V\}\|\(maximum entropy at positiontt\)\. ThusDahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)is not bounded byEnt\(πT\(⋅\|𝐱<t\)\)\\mathrm\{Ent\}\(\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)\), confirming independence\. The intuition is thatπT\(⋅\|𝐱<t\)\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\_\{<t\}\)is a marginal distribution obtained by integrating over the next token; its entropy characterizes uncertainty*about the current position*, whileDahead\(t\)D\_\{\\mathrm\{ahead\}\}\(t\)captures*differences across branching futures*—orthogonal quantities\. ∎
### F\.4Proof of Proposition[4](https://arxiv.org/html/2605.30833#Thmproposition4)
###### Proposition\(Max\-ppas relative drift indicator, restated\)\.
LetPT∗P\_\{T\}^\{\*\}be the teacher’s distribution on an in\-distribution prefix, andPT\(𝐱\)P\_\{T\}^\{\(\\mathbf\{x\}\)\}on a student\-generated prefix\. If the teacher isβ\\beta\-smooth, then:
maxvPT∗\(v\)−maxvPT\(𝐱\)\(v\)≤‖PT∗−PT\(𝐱\)‖∞≤β⋅d\(𝐱,𝒳T\)\.\\max\_\{v\}P\_\{T\}^\{\*\}\(v\)\-\\max\_\{v\}P\_\{T\}^\{\(\\mathbf\{x\}\)\}\(v\)\\leq\\\|P\_\{T\}^\{\*\}\-P\_\{T\}^\{\(\\mathbf\{x\}\)\}\\\|\_\{\\infty\}\\leq\\beta\\cdot d\(\\mathbf\{x\},\\mathcal\{X\}\_\{T\}\)\.
###### Proof\.
First inequality\.For any two distributionsf,gf,gover a finite set𝒱\\mathcal\{V\}:
maxvf\(v\)−maxvg\(v\)\\displaystyle\\max\_\{v\}f\(v\)\-\\max\_\{v\}g\(v\)≤maxv\[f\(v\)−g\(v\)\]≤maxv\|f\(v\)−g\(v\)\|=‖f−g‖∞\.\\displaystyle\\leq\\max\_\{v\}\[f\(v\)\-g\(v\)\]\\leq\\max\_\{v\}\|f\(v\)\-g\(v\)\|=\\\|f\-g\\\|\_\{\\infty\}\.The first step usesmaxvf\(v\)=f\(v∗\)≤f\(v∗\)−g\(v∗\)\+maxvg\(v\)\\max\_\{v\}f\(v\)=f\(v^\{\*\}\)\\leq f\(v^\{\*\}\)\-g\(v^\{\*\}\)\+\\max\_\{v\}g\(v\)for the maximizerv∗=argmaxvf\(v\)v^\{\*\}=\\arg\\max\_\{v\}f\(v\), givingmaxvf\(v\)−maxvg\(v\)≤f\(v∗\)−g\(v∗\)\\max\_\{v\}f\(v\)\-\\max\_\{v\}g\(v\)\\leq f\(v^\{\*\}\)\-g\(v^\{\*\}\)\. The second step replaces the signed difference with the absolute value, and the last step is the definition ofℓ∞\\ell^\{\\infty\}norm\.
Second inequality\.Byβ\\beta\-smoothness of the teacher model, small perturbations in input context produce bounded output distribution shifts: there existsβ\>0\\beta\>0such that for any two prefixes𝐱\\mathbf\{x\}and𝐱∗\\mathbf\{x\}^\{\*\}:
∥πT\(⋅\|𝐱\)−πT\(⋅\|𝐱∗\)∥∞≤β⋅d\(𝐱,𝐱∗\)\.\\\|\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}\)\-\\pi\_\{T\}\(\\cdot\|\\mathbf\{x\}^\{\*\}\)\\\|\_\{\\infty\}\\leq\\beta\\cdot d\(\\mathbf\{x\},\\mathbf\{x\}^\{\*\}\)\.Setting𝐱∗∈𝒳T\\mathbf\{x\}^\{\*\}\\in\\mathcal\{X\}\_\{T\}\(in\-distribution prefix minimizing distance from𝐱\\mathbf\{x\}\) gives the result withd\(𝐱,𝒳T\)=min𝐱∗∈𝒳Td\(𝐱,𝐱∗\)d\(\\mathbf\{x\},\\mathcal\{X\}\_\{T\}\)=\\min\_\{\\mathbf\{x\}^\{\*\}\\in\\mathcal\{X\}\_\{T\}\}d\(\\mathbf\{x\},\\mathbf\{x\}^\{\*\}\)\.
Implication for relative comparison\.Critically, this proposition is used in a*relative*sense: we compare max\-ppacrossKKcandidates\{xt\(k\)\}\\\{x\_\{t\}^\{\(k\)\}\\\}at the same position\. For candidateskkandk′k^\{\\prime\}:
maxvπT\(v\|𝐱<t,xt\(k\)\)−maxvπT\(v\|𝐱<t,xt\(k′\)\)≈β⋅\[d\(𝐱<t⋅xt\(k\),𝒳T\)−d\(𝐱<t⋅xt\(k′\),𝒳T\)\]\.\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k\)\}\)\-\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}^\{\(k^\{\\prime\}\)\}\)\\approx\\beta\\cdot\[d\(\\mathbf\{x\}\_\{<t\}\\cdot x\_\{t\}^\{\(k\)\},\\mathcal\{X\}\_\{T\}\)\-d\(\\mathbf\{x\}\_\{<t\}\\cdot x\_\{t\}^\{\(k^\{\\prime\}\)\},\\mathcal\{X\}\_\{T\}\)\]\.Thus a higher max\-ppatt\+1t\{\+\}1for candidatekkimplies candidatekkhas caused less drift from the teacher’s in\-distribution manifold—not a return to in\-distribution, but less*additional*drift\. This is the “relative drift indicator” interpretation used in Section[3\.2](https://arxiv.org/html/2605.30833#S3.SS2)\. ∎
### F\.5Group Normalization: Design Rationale and Formal Properties
Why group normalization is necessary\.The raw confidencerraw\(xt\)=maxvπT\(v\|𝐱<t,xt\)r\_\{\\mathrm\{raw\}\}\(x\_\{t\}\)=\\max\_\{v\}\\pi\_\{T\}\(v\|\\mathbf\{x\}\_\{<t\},x\_\{t\}\)varies substantially across positions and tasks for reasons unrelated to the relative quality of token choices: \(i\) token frequency effects \(common tokens systematically attract higher teacher probability\), \(ii\) context difficulty \(some prefixes are inherently harder to continue regardless of token choice\), and \(iii\) vocabulary size variation across model configurations\. Using raw max\-ppdirectly as a reward would introduce a high\-variance baseline that dominates the gradient signal\. Group normalization subtracts the group meanμK\\mu\_\{K\}computed over the student’s top\-KKcandidates at the*same position and context*, canceling all position\- and task\-level confounds\. Only the*relative ranking*across candidates at the same position survives, which is exactly the signal we need: which token choice causes the least additional drift\.
Connection to GRPO\.This normalization is inspired by Group Relative Policy Optimization \(GRPO\), which normalizes rewards within a group of sampled responses\. Our group is defined over the top\-KKcandidates at each position rather than over full response samples, enabling a per\-token signal rather than a sequence\-level reward\.
###### Proposition\(Properties of group normalization\)\.
The group\-normalized confidence rewardrconf\(xt\(k\)\)=\(rraw\(k\)−μK\)/\(σK\+ϵ\)r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)=\(r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}\)/\(\\sigma\_\{K\}\+\\epsilon\)satisfies:
1. 1\.Zero\-mean:∑k=1Kπθ\(xt\(k\)\)⋅rconf\(xt\(k\)\)≈0\\sum\_\{k=1\}^\{K\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\)\\cdot r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)\\approx 0when top\-KKprobabilities are approximately equal\.
2. 2\.Graceful degradation:WhenσK→0\\sigma\_\{K\}\\to 0,rconf\(xt\(k\)\)→0r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)\\to 0for allkk\.
3. 3\.Scale invariance:The ranking byrconfr\_\{\\mathrm\{conf\}\}is invariant to affine transformations ofrrawr\_\{\\mathrm\{raw\}\}\.
###### Proof\.
Part \(1\)\.By the definition ofμK\\mu\_\{K\}:
∑k=1K\(rraw\(k\)−μK\)=∑k=1Krraw\(k\)−KμK=0\.\\sum\_\{k=1\}^\{K\}\(r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}\)=\\sum\_\{k=1\}^\{K\}r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-K\\mu\_\{K\}=0\.Whenπθ\(xt\(k\)\)≈1/K\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\)\\approx 1/Kfor allkk\(approximately uniform top\-KK\):
∑k=1Kπθ\(xt\(k\)\)⋅rconf\(xt\(k\)\)≈1K\(σK\+ϵ\)∑k=1K\(rraw\(k\)−μK\)=0\.\\sum\_\{k=1\}^\{K\}\\pi\_\{\\theta\}\(x\_\{t\}^\{\(k\)\}\)\\cdot r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)\\approx\\frac\{1\}\{K\(\\sigma\_\{K\}\+\\epsilon\)\}\\sum\_\{k=1\}^\{K\}\(r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}\)=0\.For non\-uniformπθ\\pi\_\{\\theta\}, the weighted sum is not exactly zero but remains small: it equalsCovk∼πθ\(1,rconf\(k\)\)=0\\mathrm\{Cov\}\_\{k\\sim\\pi\_\{\\theta\}\}\(1,r\_\{\\mathrm\{conf\}\}^\{\(k\)\}\)=0by the zero\-mean property ofrraw\(k\)−μKr\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}under uniform weighting\.
Part \(2\)\.When all candidates lead to the same teacher confidence,rraw\(k\)=cr\_\{\\mathrm\{raw\}\}^\{\(k\)\}=cfor allkk, soμK=c\\mu\_\{K\}=candσK=0\\sigma\_\{K\}=0\. The numeratorrraw\(k\)−μK=0r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}=0for allkk, givingrconf\(xt\(k\)\)=0/ϵ=0r\_\{\\mathrm\{conf\}\}\(x\_\{t\}^\{\(k\)\}\)=0/\\epsilon=0\. More generally, asσK→0\\sigma\_\{K\}\\to 0, the numerators approach zero while the denominator is bounded below byϵ\>0\\epsilon\>0, sorconf→0r\_\{\\mathrm\{conf\}\}\\to 0\. This is the desired graceful degradation: at positions where the teacher cannot discriminate between candidates \(all lead to equally uncertain teacher states\), the confidence reward automatically suppresses itself without requiring external gating\.
Part \(3\)\.Letr~raw\(k\)=αrraw\(k\)\+β\\tilde\{r\}\_\{\\mathrm\{raw\}\}^\{\(k\)\}=\\alpha r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\+\\betafor constantsα\>0\\alpha\>0,β∈ℝ\\beta\\in\\mathbb\{R\}\. Then:
μ~K=αμK\+β,σ~K=\|α\|σK=ασK\(sinceα\>0\)\.\\tilde\{\\mu\}\_\{K\}=\\alpha\\mu\_\{K\}\+\\beta,\\quad\\tilde\{\\sigma\}\_\{K\}=\|\\alpha\|\\sigma\_\{K\}=\\alpha\\sigma\_\{K\}\\quad\(\\text\{since \}\\alpha\>0\)\.Therefore:
r~conf\(k\)=\(αrraw\(k\)\+β\)−\(αμK\+β\)ασK\+ϵ=α\(rraw\(k\)−μK\)ασK\+ϵ\.\\tilde\{r\}\_\{\\mathrm\{conf\}\}^\{\(k\)\}=\\frac\{\(\\alpha r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\+\\beta\)\-\(\\alpha\\mu\_\{K\}\+\\beta\)\}\{\\alpha\\sigma\_\{K\}\+\\epsilon\}=\\frac\{\\alpha\(r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}\)\}\{\\alpha\\sigma\_\{K\}\+\\epsilon\}\.For largeσK\\sigma\_\{K\}\(whereϵ\\epsilonis negligible\),r~conf\(k\)≈rconf\(k\)\\tilde\{r\}\_\{\\mathrm\{conf\}\}^\{\(k\)\}\\approx r\_\{\\mathrm\{conf\}\}^\{\(k\)\}, and the ranking is preserved\. More precisely, for anyk,k′k,k^\{\\prime\}:r~conf\(k\)\>r~conf\(k′\)\\tilde\{r\}\_\{\\mathrm\{conf\}\}^\{\(k\)\}\>\\tilde\{r\}\_\{\\mathrm\{conf\}\}^\{\(k^\{\\prime\}\)\}iffrraw\(k\)−μK\>rraw\(k′\)−μKr\_\{\\mathrm\{raw\}\}^\{\(k\)\}\-\\mu\_\{K\}\>r\_\{\\mathrm\{raw\}\}^\{\(k^\{\\prime\}\)\}\-\\mu\_\{K\}\(sinceα\>0\\alpha\>0\), iffrraw\(k\)\>rraw\(k′\)r\_\{\\mathrm\{raw\}\}^\{\(k\)\}\>r\_\{\\mathrm\{raw\}\}^\{\(k^\{\\prime\}\)\}, iffrconf\(k\)\>rconf\(k′\)r\_\{\\mathrm\{conf\}\}^\{\(k\)\}\>r\_\{\\mathrm\{conf\}\}^\{\(k^\{\\prime\}\)\}\. This makes the reward robust to systematic shifts in absolute confidence level across positions and tasks\. ∎
## NeurIPS Paper Checklist
1. 1\.Claims
2. Question: Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope?
3. Answer:\[Yes\]
4. Justification: The claims in the abstract and introduction section strictly follow the paper’s contributions and scope\.
5. Guidelines: - •The answer NA means that the abstract and introduction do not include the claims made in the paper\. - •The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations\. A No or NA answer to this question will not be perceived well by the reviewers\. - •The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings\. - •It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper\.
6. 2\.Limitations
7. Question: Does the paper discuss the limitations of the work performed by the authors?
8. Answer:\[Yes\]
9. Justification: We discuss the limitations of the work in limitations\.
10. Guidelines: - •The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper\. - •The authors are encouraged to create a separate "Limitations" section in their paper\. - •The paper should point out any strong assumptions and how robust the results are to violations of these assumptions \(e\.g\., independence assumptions, noiseless settings, model well\-specification, asymptotic approximations only holding locally\)\. The authors should reflect on how these assumptions might be violated in practice and what the implications would be\. - •The authors should reflect on the scope of the claims made, e\.g\., if the approach was only tested on a few datasets or with a few runs\. In general, empirical results often depend on implicit assumptions, which should be articulated\. - •The authors should reflect on the factors that influence the performance of the approach\. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting\. Or a speech\-to\-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon\. - •The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size\. - •If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness\. - •While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that aren’t acknowledged in the paper\. The authors should use their best judgment and recognize that individual actions in favor of transparency play an important role in developing norms that preserve the integrity of the community\. Reviewers will be specifically instructed to not penalize honesty concerning limitations\.
11. 3\.Theory assumptions and proofs
12. Question: For each theoretical result, does the paper provide the full set of assumptions and a complete \(and correct\) proof?
13. Answer:\[Yes\]
14. Justification: We provide proofs in the Appendix\.
15. Guidelines: - •The answer NA means that the paper does not include theoretical results\. - •All the theorems, formulas, and proofs in the paper should be numbered and cross\-referenced\. - •All assumptions should be clearly stated or referenced in the statement of any theorems\. - •The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition\. - •Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material\. - •Theorems and Lemmas that the proof relies upon should be properly referenced\.
16. 4\.Experimental result reproducibility
17. Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper \(regardless of whether the code and data are provided or not\)?
18. Answer:\[Yes\]
19. Justification: We summarize all the information for experimental reproduction in Appendix\.
20. Guidelines: - •The answer NA means that the paper does not include experiments\. - •If the paper includes experiments, a No answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not\. - •If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable\. - •Depending on the contribution, reproducibility can be accomplished in various ways\. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model\. In general\. releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model \(e\.g\., in the case of a large language model\), releasing of a model checkpoint, or other means that are appropriate to the research performed\. - •While NeurIPS does not require releasing code, the conference does require all submissions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution\. For example 1. \(a\)If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm\. 2. \(b\)If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully\. 3. \(c\)If the contribution is a new model \(e\.g\., a large language model\), then there should either be a way to access this model for reproducing the results or a way to reproduce the model \(e\.g\., with an open\-source dataset or instructions for how to construct the dataset\)\. 4. \(d\)We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility\. In the case of closed\-source models, it may be that access to the model is limited in some way \(e\.g\., to registered users\), but it should be possible for other researchers to have some path to reproducing or verifying the results\.
21. 5\.Open access to data and code
22. Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material?
23. Answer:\[Yes\]
24. Justification: The source code is provided in the anonymized link\.
25. Guidelines: - •The answer NA means that paper does not include experiments requiring code\. - • - •While we encourage the release of code and data, we understand that this might not be possible, so “No” is an acceptable answer\. Papers cannot be rejected simply for not including code, unless this is central to the contribution \(e\.g\., for a new open\-source benchmark\)\. - •The instructions should contain the exact command and environment needed to run to reproduce the results\. See the NeurIPS code and data submission guidelines \([https://nips\.cc/public/guides/CodeSubmissionPolicy](https://nips.cc/public/guides/CodeSubmissionPolicy)\) for more details\. - •The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc\. - •The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines\. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why\. - •At submission time, to preserve anonymity, the authors should release anonymized versions \(if applicable\)\. - •Providing as much information as possible in supplemental material \(appended to the paper\) is recommended, but including URLs to data and code is permitted\.
26. 6\.Experimental setting/details
27. Question: Does the paper specify all the training and test details \(e\.g\., data splits, hyperparameters, how they were chosen, type of optimizer, etc\.\) necessary to understand the results?
28. Answer:\[Yes\]
29. Justification: We provide details in the Appendix\.
30. Guidelines: - •The answer NA means that the paper does not include experiments\. - •The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them\. - •The full details can be provided either with the code, in appendix, or as supplemental material\.
31. 7\.Experiment statistical significance
32. Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments?
33. Answer:\[No\]
34. Justification: We reported Avg@K and Pass@k for evaluations\.
35. Guidelines: - •The answer NA means that the paper does not include experiments\. - •The authors should answer "Yes" if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper\. - •The factors of variability that the error bars are capturing should be clearly stated \(for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions\)\. - •The method for calculating the error bars should be explained \(closed form formula, call to a library function, bootstrap, etc\.\) - •The assumptions made should be given \(e\.g\., Normally distributed errors\)\. - •It should be clear whether the error bar is the standard deviation or the standard error of the mean\. - •It is OK to report 1\-sigma error bars, but one should state it\. The authors should preferably report a 2\-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified\. - •For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range \(e\.g\. negative error rates\)\. - •If error bars are reported in tables or plots, The authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text\.
36. 8\.Experiments compute resources
37. Question: For each experiment, does the paper provide sufficient information on the computer resources \(type of compute workers, memory, time of execution\) needed to reproduce the experiments?
38. Answer:\[Yes\]
39. Justification: We provide them in the Appendix\.
40. Guidelines: - •The answer NA means that the paper does not include experiments\. - •The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage\. - •The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute\. - •The paper should disclose whether the full research project required more compute than the experiments reported in the paper \(e\.g\., preliminary or failed experiments that didn’t make it into the paper\)\.
41. 9\.Code of ethics
43. Answer:\[Yes\]
44. Justification: We follow every aspect of the NeurIPS Code of Ethics in this research\.
45. Guidelines: - •The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics\. - •If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics\. - •The authors should make sure to preserve anonymity \(e\.g\., if there is a special consideration due to laws or regulations in their jurisdiction\)\.
46. 10\.Broader impacts
47. Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed?
48. Answer:\[Yes\]
49. Justification: We discuss the broader impact in Limitations\.
50. Guidelines: - •The answer NA means that there is no societal impact of the work performed\. - •If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact\. - •Examples of negative societal impacts include potential malicious or unintended uses \(e\.g\., disinformation, generating fake profiles, surveillance\), fairness considerations \(e\.g\., deployment of technologies that could make decisions that unfairly impact specific groups\), privacy considerations, and security considerations\. - •The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments\. However, if there is a direct path to any negative applications, the authors should point it out\. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate deepfakes for disinformation\. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster\. - •The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from \(intentional or unintentional\) misuse of the technology\. - •If there are negative societal impacts, the authors could also discuss possible mitigation strategies \(e\.g\., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML\)\.
51. 11\.Safeguards
52. Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse \(e\.g\., pretrained language models, image generators, or scraped datasets\)?
53. Answer:\[N/A\]
54. Justification: The paper poses no such risks\.
55. Guidelines: - •The answer NA means that the paper poses no such risks\. - •Released models that have a high risk for misuse or dual\-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters\. - •Datasets that have been scraped from the Internet could pose safety risks\. The authors should describe how they avoided releasing unsafe images\. - •We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort\.
56. 12\.Licenses for existing assets
57. Question: Are the creators or original owners of assets \(e\.g\., code, data, models\), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected?
58. Answer:\[Yes\]
59. Justification: We cite every paper of the existing assets we used\.
60. Guidelines: - •The answer NA means that the paper does not use existing assets\. - •The authors should cite the original paper that produced the code package or dataset\. - •The authors should state which version of the asset is used and, if possible, include a URL\. - •The name of the license \(e\.g\., CC\-BY 4\.0\) should be included for each asset\. - •For scraped data from a particular source \(e\.g\., website\), the copyright and terms of service of that source should be provided\. - •If assets are released, the license, copyright information, and terms of use in the package should be provided\. For popular datasets,[paperswithcode\.com/datasets](https://arxiv.org/html/2605.30833v1/paperswithcode.com/datasets)has curated licenses for some datasets\. Their licensing guide can help determine the license of a dataset\. - •For existing datasets that are re\-packaged, both the original license and the license of the derived asset \(if it has changed\) should be provided\. - •If this information is not available online, the authors are encouraged to reach out to the asset’s creators\.
61. 13\.New assets
62. Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets?
63. Answer:\[N/A\]
64. Justification: The paper does not release new assets\.
65. Guidelines: - •The answer NA means that the paper does not release new assets\. - •Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates\. This includes details about training, license, limitations, etc\. - •The paper should discuss whether and how consent was obtained from people whose asset is used\. - •At submission time, remember to anonymize your assets \(if applicable\)\. You can either create an anonymized URL or include an anonymized zip file\.
66. 14\.Crowdsourcing and research with human subjects
67. Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation \(if any\)?
68. Answer:\[N/A\]
69. Justification: The paper does not involve crowdsourcing nor research with human subjects\.
70. Guidelines: - •The answer NA means that the paper does not involve crowdsourcing nor research with human subjects\. - •Including this information in the supplemental material is fine, but if the main contribution of the paper involves human subjects, then as much detail as possible should be included in the main paper\. - •According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector\.
71. 15\.Institutional review board \(IRB\) approvals or equivalent for research with human subjects
72. Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board \(IRB\) approvals \(or an equivalent approval/review based on the requirements of your country or institution\) were obtained?
73. Answer:\[N/A\]
74. Justification: The paper does not involve crowdsourcing nor research with human subjects\.
75. Guidelines: - •The answer NA means that the paper does not involve crowdsourcing nor research with human subjects\. - •Depending on the country in which research is conducted, IRB approval \(or equivalent\) may be required for any human subjects research\. If you obtained IRB approval, you should clearly state this in the paper\. - •We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution\. - •For initial submissions, do not include any information that would break anonymity \(if applicable\), such as the institution conducting the review\.
76. 16\.Declaration of LLM usage
77. Question: Does the paper describe the usage of LLMs if it is an important, original, or non\-standard component of the core methods in this research? Note that if the LLM is used only for writing, editing, or formatting purposes and does not impact the core methodology, scientific rigorousness, or originality of the research, declaration is not required\.
78. Answer:\[Yes\]
79. Justification: We use LLM for writing, editing and formatting\.
80. Guidelines: - •The answer NA means that the core method development in this research does not involve LLMs as any important, original, or non\-standard components\. - •Similar Articles
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
This paper introduces SA-OPD, a spurious-signal-aware on-policy distillation framework that filters misleading token-level teacher supervision based on input-grounding and optimization impact, improving LLM and VLM distillation performance.
Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
The paper shows that Direct-On-Policy Distillation's token-level log-ratio reward can remain unchanged even when the teacher checkpoints' divergences vanish, motivating S²D-OPD, a method that masks supervision at low-divergence states using teacher-reference JSD. S²D-OPD improves held-out accuracy on AIME and HMMT benchmarks in 7 of 8 teacher-student settings across models from 1.7B to 8B parameters without extra forward passes.
On the Off-Policy Teacher in On-Policy Distillation
The paper identifies an off-policy asymmetry in on-policy distillation, where the teacher must supervise student-generated prefixes it was not trained on, and proposes SCOUT, a co-training framework that adapts the teacher via RL with verifiable rewards, consistently improving distillation across configurations, scales, and reasoning domains.
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
SAKI introduces a supervision allocation method for on-policy distillation that uses KL-constrained teacher-guided rollouts and maximal coupling to route token-level supervision, improving performance on mathematical reasoning benchmarks for small models.