When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation

arXiv cs.CL Papers

Summary

This paper analyzes why pointwise clipping of the forward-KL objective in on-policy self-distillation (OPSD) can backfire: it proves and empirically shows that clipped objectives can push token probabilities away from the teacher, causing persistent token repetitions in reasoning outputs.

arXiv:2609.38995v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged information. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math-related tokens, and that pointwise clipping of the forward KL objective stabilizes training. Pointwise clipping caps each vocabulary-wise forward KL term at a fixed threshold before summing over the vocabulary. Follow-up studies have adopted this clipping, but its effect on training has not been directly examined. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to the end of the response than their unclipped counterparts. We trace this failure to the clipped objective. We prove that the clipped objective can fail to correct the student toward the teacher and can instead push clipped and unclipped token probabilities away from its teacher. Our training runs agree with this analysis: inside repetitions, the clipped student places less probability than its teacher on leaving the repetition, and more on continuing it, whereas the unclipped runs stay close to their teachers.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:46 AM

# When Clipping Reverses Correction: Failure Dynamics of Pointwise Forward-KL On-Policy Self-Distillation
Source: [https://arxiv.org/html/2609.38995](https://arxiv.org/html/2609.38995)
Hao LiAffiliation:Department of Computer ScienceYixin ChenAffiliation:Department of Computer ScienceFuhai LiAffiliation:Department of Computer ScienceAffiliation:Department of PediatricsWashington University in St\. Louis \*Correspondence:fuhai\.li@wustl\.edu

###### Abstract

On\-policy self\-distillation \(OPSD\) trains a student on its own generated responses using feedback from the same model conditioned on privileged information\. On mathematical reasoning, the original OPSD study finds that stylistic tokens can dominate the training signal over math\-related tokens, and that pointwise clipping of the forward KL objective stabilizes training\. Pointwise clipping caps each vocabulary\-wise forward KL term at a fixed threshold before summing over the vocabulary\. Follow\-up studies have adopted this clipping, but its effect on training has not been directly examined\. In matched training runs differing only in whether clipping is applied, we observe that clipped runs produce substantially more repetitions that persist to the end of the response than their unclipped counterparts\. We trace this failure to the clipped objective\. We prove that the clipped objective can fail to correct the student toward the teacher and can instead push clipped and unclipped token probabilities away from its teacher\. Our training runs agree with this analysis: inside repetitions, the clipped student places less probability than its teacher on leaving the repetition, and more on continuing it, whereas the unclipped runs stay close to their teachers\.

## 1Introduction

Outcome rewards give a reasoning model one scalar for thousands of tokens\. On\-policy distillation instead provides a per\-token target from a teacher evaluated on the student’s own responses\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38995#bib.bib1);[Lu & Lab, 2025](https://arxiv.org/html/2609.38995#bib.bib14)\)\. Self\-distillation uses the same network as both teacher and student, replacing a stronger teacher model with access to additional privilege information that is withheld from the student\([Zhao et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib27);[Hübotter et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib8);[Shenfeld et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib16)\)\. Self\-distillation provides dense supervision, but how the student learns from it depends on the training objective\.

On\-Policy Self\-Distillation \(OPSD\) trains the student with a pointwise\-clipped forward KL objective\([Zhao et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib27)\)\. At each response position, this objective clips each vocabulary\-wise forward KL term at a fixed threshold before summing over the vocabulary\. The authors introduced this pointwise clipping to keep stylistic tokens from dominating the math\-related tokens in the training signal, and many subsequent methods retain it\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3);[Li et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib11);[Liu et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib13);[Shrestha & Tessier, 2026](https://arxiv.org/html/2609.38995#bib.bib18)\)\.

Studies using this recipe report response length inflation\([Yang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib23)\), responses that fail to terminate within the generation budget\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3);[Ichihara et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib9)\), redundant reasoning chain\([Gu et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib5)\), and accuracy that declines after early gains\([Zhao et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib27);[Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3);[Ichihara et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib9);[Pan et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib15);[Zhang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib24)\)\. Proposed explanations include a frozen teacher that cannot adapt to the student\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3)\), shifts in the teacher’s style or reasoning pattern induced by the privileged context\([Ichihara et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib9);[Pan et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib15);[Yang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib23);[Gu et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib5)\)\. No study reports that pointwise clipped objective contribute to the text degeneration, leaving this possibility open for further investigation\.

Prior work offers a partial mathematical description of this clipped objective: a clipped term becomes constant and loses its direct gradient\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3)\), and the clipped sum is no longer a divergence\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3);[Feng et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib4)\)\. These observations do not establish which student distributions the clipped objective favors, in particular whether minimizing it still brings a clipped token’s probability closer to the teacher’s\.

We trained the pointwise\-clipped forward KL OPSD recipe with a frozen reference\-conditioned teacher and encountered a failure mode of text degeneration: generations ended in periodic tails that never terminated\. When we removed only the clipping, nearly all of these loops disappeared\. The teacher and its privileged context are identical in the two runs, so neither accounts for the difference\. We therefore ask whether, and through what mechanism, this clipped objective contributed to this failure\.

Our contributions therefore follow three questions: how the clipped objective changes the logit gradient, where its minimum lies, and what it does to training\.The logit gradient\(Section[4\.1](https://arxiv.org/html/2609.38995#S4.SS1)\)\. We show that, relative to exact forward KL, the clipped objective reverses the direction in which gradient descent moves the logits, on every clipped token and on unclipped token whose student probability lies within a bounded range above the teacher’s\.The minimizer\(Section[4\.2](https://arxiv.org/html/2609.38995#S4.SS2)\)\. We show that, at a response position with the unclipped and clipped token sets held fixed, the clipped objective is minimized only when the clipped tokens’ probability is reduced to zero rather than restored toward the teacher’s, and each unclipped token receives more probability than the teacher assigns it\. A lower total probability on the clipped tokens always permits a lower objective value\.Training runs\(Sections[4\.3](https://arxiv.org/html/2609.38995#S4.SS3)\)\. We train four matched pairs of runs that differ only in whether clipping is applied and vary whether the student and the teacher think\. The clipped runs produce far more repetitions that continue to the end of the response than their unclipped runs\. Inside repetitions, the clipped student places less probability than its teacher on leaving, and more on continuing, whereas the unclipped runs stay close to their teachers\. These observations are consistent with the logit\-gradient reversals we derived\.

## 2Related Work

#### Pointwise\-clipped forward KL\.

OPSD introduces the pointwise clipping with thresholdτ\\tau\([Zhao et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib27)\)on each vocabulary\-wise forward KL term\. Later studies default thresholdτ=0\.05\\tau=0\.05with training methods deviations\([Tan & Hong, 2026b](https://arxiv.org/html/2609.38995#bib.bib20);[Hou et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib7);[Zhang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib24);[Liu et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib13);[Wang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib21);[Chen et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib2);[Ichihara et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib9)\), at other values and/or with training methods deviations\([Chen et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib3);[Ichihara et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib9);[Shrestha & Tessier, 2026](https://arxiv.org/html/2609.38995#bib.bib18);[Tan & Hong, 2026a](https://arxiv.org/html/2609.38995#bib.bib19);[Liang et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib12)\), or without stating one\([Li et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib11);[Zhao et al\., 2026b](https://arxiv.org/html/2609.38995#bib.bib28)\)\.[Chen et al\. \(2026b\)](https://arxiv.org/html/2609.38995#bib.bib3)observe that a clipped term no longer supplies a gradient, so the cap stops the student from following large, mostly stylistic terms, and that a sum of capped signed terms is not a divergence\.[Feng et al\. \(2026\)](https://arxiv.org/html/2609.38995#bib.bib4)notes that the clipped objective equals exact forward KL wherever no term exceeds the threshold, so the two share their gradient and Hessian around the student that matches the teacher, but that elsewhere it can be negative\.

#### Failures reported under the clipped objective\.

OPSD reports its best checkpoint, yet its comparison of objectives scores lower at 100 updates than at 50\([Zhao et al\., 2026a](https://arxiv.org/html/2609.38995#bib.bib27)\)\.[Chen et al\. \(2026b\)](https://arxiv.org/html/2609.38995#bib.bib3)finds that its OPSD baseline loses accuracy from 100 to 200 updates while truncation doubles, and also stating frozen teacher that cannot respond to what the student rejects as one of the limitations\.[Zhang et al\. \(2026\)](https://arxiv.org/html/2609.38995#bib.bib24)find that their three SmolLM3 students fall below the base model by 100 updates under multiple different configurations\.[Ichihara et al\. \(2026\)](https://arxiv.org/html/2609.38995#bib.bib9)finds that, a teacher with a math question and a solution to a physics question leaves many generations repetitive and non\-terminating, as compared to one with another math question’s solution instead, attribute this to the context without isolating which property is responsible, and leave deterioration under longer training unexplained\.[Gu et al\. \(2026\)](https://arxiv.org/html/2609.38995#bib.bib5)observe that “OPSD trajectories often recompute the same intermediate quantities or revise earlier steps without new information, leading to long and redundant reasoning chains,” and attribute this to a teacher conditioned on a single reference solution\.

## 3Experimental Setup

#### Training design\.

Our goal is to reproduce the training design of OPSD, which we reimplement in the verl framework\([Sheng et al\., 2024](https://arxiv.org/html/2609.38995#bib.bib17)\)\. Each run starts from Qwen3\-4B\([Yang et al\., 2025](https://arxiv.org/html/2609.38995#bib.bib22)\), with a trainable full\-parameter student and a frozen teacher initialized from the same checkpoint\. The teacher additionally receives a reference solution and provides its next\-token distribution at every position of the student’s responses\.

We train for 200 updates on the same 30k OpenThoughts\([Guha et al\., 2026](https://arxiv.org/html/2609.38995#bib.bib6)\)dataset used to train OPSD\. We construct four matched clipped/unclipped pairs\. Within each pair, the data, sampling setup, and model configuration are identical; only whether pointwise clipping is applied differs and we call the unclipped run the*twin*of the clipped run\. Table[1](https://arxiv.org/html/2609.38995#S3.T1)summarizes the run\-specific configurations\. The complete training hyperparameters and chat template are given in Appendix[C](https://arxiv.org/html/2609.38995#A3.SS0.SSS0.Px1)\. We refer to each pair by its letter in Table[1](https://arxiv.org/html/2609.38995#S3.T1)\. F, M, K and L differ only in which side thinks\. The letters index the order in which the configurations entered our run series, each added to isolate one variable from an earlier run\.

#### Distillation objective\.

OPSD computes its loss over the full vocabulary\. Our teacher runs as a separate inference service under SGLang\([Zheng et al\., 2024](https://arxiv.org/html/2609.38995#bib.bib29)\), and sending a full\-vocabulary distribution for every response position is impractical, so our interface returns the teacher’s top\-128 token log probabilities at each response position\.

At response positiontt, let support𝒮t\\mathcal\{S\}\_\{t\}denote set of the teacher’s top\-128 tokens, and letat,ia\_\{t,i\}andbt,ib\_\{t,i\}be the teacher’s and the student’s logits for tokenii\. Both distributions are formed by a softmax over𝒮t\\mathcal\{S\}\_\{t\}at temperatureT=1\.1T=1\.1,

pt,i=exp⁡\(at,i/T\)∑j∈𝒮texp⁡\(at,j/T\),qt,i=exp⁡\(bt,i/T\)∑j∈𝒮texp⁡\(bt,j/T\),i∈𝒮t,p\_\{t,i\}=\\frac\{\\exp\(a\_\{t,i\}/T\)\}\{\\sum\_\{j\\in\\mathcal\{S\}\_\{t\}\}\\exp\(a\_\{t,j\}/T\)\},\\qquad q\_\{t,i\}=\\frac\{\\exp\(b\_\{t,i\}/T\)\}\{\\sum\_\{j\\in\\mathcal\{S\}\_\{t\}\}\\exp\(b\_\{t,j\}/T\)\},\\qquad i\\in\\mathcal\{S\}\_\{t\},\(1\)so that the teacher distributionptp\_\{t\}and the student distributionqtq\_\{t\}each sum to one on𝒮t\\mathcal\{S\}\_\{t\}\. We then define the vocabulary\-wise forward KL term of tokenii:

dt,i=pt,i​log⁡pt,iqt,i\.d\_\{t,i\}=p\_\{t,i\}\\log\\frac\{p\_\{t,i\}\}\{q\_\{t,i\}\}\.\(2\)Heredt,id\_\{t,i\}is positive where the student assigns tokeniiless probability than the teacher and negative where it assigns more\. Exact forward KL sumsdt,id\_\{t,i\}directly, whereas the pointwise\-clipped forward KL of OPSD first replaces anydt,id\_\{t,i\}aboveτ\\taubyτ\\tau, leaving smaller and negativedt,id\_\{t,i\}unchanged:

DFKL,t=DKL\(pt∥qt\)=∑j∈𝒮tdt,j,Dclip,t=∑j∈𝒮tmin\(dt,j,τ\),τ=0\.05\.D\_\{\\mathrm\{FKL\},t\}=D\_\{\\mathrm\{KL\}\}\(p\_\{t\}\\\|q\_\{t\}\)=\\sum\_\{j\\in\\mathcal\{S\}\_\{t\}\}d\_\{t,j\},\\qquad D\_\{\\mathrm\{clip\},t\}=\\sum\_\{j\\in\\mathcal\{S\}\_\{t\}\}\\min\(d\_\{t,j\},\\tau\),\\qquad\\tau=0\.05\.\(3\)Despite the symbol, which follows OPSD’s notation,Dclip,tD\_\{\\mathrm\{clip\},t\}is not a divergence, since it can be negative\. The training loss is the mean ofDclip,tD\_\{\\mathrm\{clip\},t\}in the clipped runs, or ofDFKL,tD\_\{\\mathrm\{FKL\},t\}in the unclipped runs, over all valid response positions in the update\. We call the tokens in the teacher support𝒮t\\mathcal\{S\}\_\{t\}the*candidate tokens*at positiontt, to distinguish them from the token the student emits there, and a candidate token withdt,i\>τd\_\{t,i\}\>\\tauan*over\-threshold*token\. In a clipped run, the over\-threshold tokens are the clipped tokens\.

Table 1:Paired training families\. All use Qwen3\-4B, teacher top\-128 support,T=1\.1T=1\.1, 30 problems per update with one rollout per problem, learning rate of1×10−61\\times 10^\{\-6\}, forward\-KL distillation, and one trajectory per run\. Clipped runs useτ=0\.05\\tau=0\.05and their unclipped runs do not\.
#### Trajectory analysis and evaluation\.

At every response position of every update, we record the teacher’s and the student’s log probabilities on the candidate tokens\. We can therefore compute both the clipped objective and exact forward KL at the same response positions in both runs of each pair\. Full capture details are provided in Appendix[C](https://arxiv.org/html/2609.38995#A3.SS0.SSS0.Px3)\.

We evaluate checkpoints saved every 25 training updates on AIME 2025\([Zhang & Math\-AI, 2025](https://arxiv.org/html/2609.38995#bib.bib26)\)and AIME 2024\([Zhang & Math\-AI, 2024](https://arxiv.org/html/2609.38995#bib.bib25)\)\. The main text reports AIME 2025 only, with two metrics: per\-sample accuracy \(avg@12\) and the terminal loop rate \(generation ended in periodic tails that never terminated\)\. The complete AIME 2024 and AIME 2025 results, and evaluation configuration are given in Appendix[D](https://arxiv.org/html/2609.38995#A4)\.

## 4Failure Dynamics of Pointwise Clipping

### 4\.1Pointwise clipping reverses the logit gradient

For an over\-threshold candidate token, the exact forward KL and the pointwise\-clipped forward KL produce gradients on that token’s logit with opposite signs\. Exact KL gives a negative gradient, whereas the clipped objective gives a positive gradient\. For a candidate token that is not over\-threshold, the two gradients again have opposite signs exactly when the student’s probability exceeds the teacher’s but stays below the teacher’s probability divided by the teacher’s total probability on the tokens that are not over\-threshold: exact KL then gives a positive gradient, whereas the clipped objective gives a negative one\.

Fix one response position and suppress the indextt\. The teacher is frozen, soppand its support𝒮\\mathcal\{S\}do not depend on the student’s logits\. We usedi=pi​log⁡\(pi/qi\)d\_\{i\}=p\_\{i\}\\log\(p\_\{i\}/q\_\{i\}\)from Equation[2](https://arxiv.org/html/2609.38995#S3.E2)\. Letzi=bi/Tz\_\{i\}=b\_\{i\}/Tbe the temperature\-scaled student logit, so thatqi=exp⁡\(zi\)/∑j∈𝒮exp⁡\(zj\)q\_\{i\}=\\exp\(z\_\{i\}\)/\\sum\_\{j\\in\\mathcal\{S\}\}\\exp\(z\_\{j\}\)\. We call the derivative with respect toziz\_\{i\}as its*logit gradient*for candidate tokenii\. Define

A=\{j∈𝒮:dj≤τ\},C=𝒮∖A,PA=∑j∈Apj\.A=\\\{j\\in\\mathcal\{S\}:d\_\{j\}\\leq\\tau\\\},\\qquad C=\\mathcal\{S\}\\setminus A,\\qquad P\_\{A\}=\\sum\_\{j\\in A\}p\_\{j\}\.Here,AAandCCare the active and clipped candidate token sets, andPAP\_\{A\}is the teacher’s total probability mass on the active set\. Equation[3](https://arxiv.org/html/2609.38995#S3.E3)becomes

Dclip=∑j∈Apj​log⁡pjqj\+\|C\|​τ\.D\_\{\\mathrm\{clip\}\}=\\sum\_\{j\\in A\}p\_\{j\}\\log\\frac\{p\_\{j\}\}\{q\_\{j\}\}\+\|C\|\\tau\.\(4\)
###### Proposition 1\(Logit gradient reversal on clipped candidate tokens\)\.

For any thresholdτ\>0\\tau\>0and any student distribution withdj≠τd\_\{j\}\\neq\\taufor everyj∈𝒮j\\in\\mathcal\{S\}, the gradient ofDclipD\_\{\\mathrm\{clip\}\}with respect toziz\_\{i\}is

∂Dclip∂zi=PAqi−pi𝟏\{i∈A\}\.\\frac\{\\partial D\_\{\\mathrm\{clip\}\}\}\{\\partial z\_\{i\}\}=P\_\{A\}q\_\{i\}\-p\_\{i\}\\mathbf\{1\}\\\{i\\in A\\\}\.\(5\)Every clipped candidate tokeni∈Ci\\in Csatisfies

∂DFKL∂zi=qi−pi<0,∂Dclip∂zi=PA​qi\>0\.\\frac\{\\partial D\_\{\\mathrm\{FKL\}\}\}\{\\partial z\_\{i\}\}=q\_\{i\}\-p\_\{i\}<0,\\qquad\\frac\{\\partial D\_\{\\mathrm\{clip\}\}\}\{\\partial z\_\{i\}\}=P\_\{A\}q\_\{i\}\>0\.\(6\)

Appendix[A](https://arxiv.org/html/2609.38995#A1)gives the proof\. Both inequalities in Equation[6](https://arxiv.org/html/2609.38995#S4.E6)are strict, and clipping is what makes their gradient signs differ\. For a clipped candidate tokeni∈Ci\\in C,di\>τ\>0d\_\{i\}\>\\tau\>0impliespi\>qip\_\{i\}\>q\_\{i\}, so the exact KL logit gradientqi−piq\_\{i\}\-p\_\{i\}is negative\. Pointwise clipping instead caps the termdjd\_\{j\}of everyj∈Cj\\in Cat the constantτ\\tau, which removes the gradient term−pi\-p\_\{i\}in Equation[5](https://arxiv.org/html/2609.38995#S4.E5)\. The remaining gradientPA​qiP\_\{A\}q\_\{i\}is positive becauseqi\>0q\_\{i\}\>0under softmax andPA\>0P\_\{A\}\>0\. Sinceppandqqare both normalized over the same support, there exists somejjwithqj≥pjq\_\{j\}\\geq p\_\{j\}\. Hencedj≤0<τd\_\{j\}\\leq 0<\\tau, soj∈Aj\\in A, and thereforePA≥pj\>0P\_\{A\}\\geq p\_\{j\}\>0, since Equation[1](https://arxiv.org/html/2609.38995#S3.E1)givespj\>0p\_\{j\}\>0\.

Clipping also changes the logit gradient on the active candidate tokens when the student probability lies within a bounded range above the teacher’s\. When the clipped set is nonempty,PA<1P\_\{A\}<1, sopi/PA\>pip\_\{i\}/P\_\{A\}\>p\_\{i\}, fori∈Ai\\in A\. By equation[5](https://arxiv.org/html/2609.38995#S4.E5), if the student’s probability in the active set satisfies

pi<qi<pi/PA,p\_\{i\}<q\_\{i\}<p\_\{i\}/P\_\{A\},the exact KL gradientqi−piq\_\{i\}\-p\_\{i\}is positive while the clipped gradientPA​qi−pi<0P\_\{A\}q\_\{i\}\-p\_\{i\}<0is negative\.

### 4\.2Clipped objective moves probability from clipped to active tokens

Proposition[1](https://arxiv.org/html/2609.38995#Thmproposition1)shows that, for a clipped candidate token, which already has less student than the teacher probability, the logit gradient of the clipped objective points toward lowering its logit, whereas that of exact KL points toward raising it\. This reversal does not by itself determine whether the token’s probability is still restored toward the teacher’s, as under exact KL, because softmax probabilities depend on all logits\. We therefore consider the clipped candidate tokens jointly and ask what happens at the minimizer of the clipped objective at a fixed response position: does an optimized student favor restoring the probability toward the teacher’s, or reduced even further?

At a fixed token position, for fixed nonempty active and clipped setsAAandCC, we consider the constrained problem:

minimize𝑞\\displaystyle\\underset\{q\}\{\\operatorname\{minimize\}\}Dclip​\(q\)=∑j∈𝒮min⁡\(dj,τ\)\\displaystyle D\_\{\\mathrm\{clip\}\}\(q\)=\\sum\_\{j\\in\\mathcal\{S\}\}\\min\(d\_\{j\},\\tau\)subject to\\displaystyle\\text\{subject to\}qi≥0for everyi∈𝒮,∑j∈𝒮qj=1,\\displaystyle q\_\{i\}\\geq 0\\ \\text\{ for every \}i\\in\\mathcal\{S\},\\qquad\\sum\_\{j\\in\\mathcal\{S\}\}q\_\{j\}=1,di≤τfor everyi∈A,di\>τfor everyi∈C\.\\displaystyle d\_\{i\}\\leq\\tau\\ \\text\{ for every \}i\\in A,\\qquad d\_\{i\}\>\\tau\\ \\text\{ for every \}i\\in C\.The feasible set of student distributions, denoted the regionRR, is the*fixed clipping region*of\(A,C\)\(A,C\)\. Unlike in Section[4\.1](https://arxiv.org/html/2609.38995#S4.SS1), whereqqis a softmax of finite logits and everyqi\>0q\_\{i\}\>0, we allowqi=0q\_\{i\}=0so that a minimizer exists\. For such a candidate tokenii,di=\+∞d\_\{i\}=\+\\infty, the capped term is stillτ\\tau, so it stays clipped\. Everywhere inRR,DclipD\_\{\\mathrm\{clip\}\}is given by Equation[4](https://arxiv.org/html/2609.38995#S4.E4)with these fixed sets\.

###### Proposition 2\(Fixed\-region minimizer of the clipped objective\)\.

The clipped objective has a unique minimizer over the fixed clipping regionRR,

qi⋆=\{pi/PA,i∈A,0,i∈C,Dclip​\(q⋆\)=PA​log⁡PA\+\|C\|​τ\.q\_\{i\}^\{\\star\}=\\begin\{cases\}p\_\{i\}/P\_\{A\},&i\\in A,\\\\ 0,&i\\in C,\\end\{cases\}\\qquad D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)=P\_\{A\}\\log P\_\{A\}\+\|C\|\\tau\.\(7\)LetQC=∑j∈CqjQ\_\{C\}=\\sum\_\{j\\in C\}q\_\{j\}be the student’s total probability mass on the clipped set\. For every value ofQCQ\_\{C\}attained inRR, if its value is further added to the constraints without inconsistencies with the other constraints, the minimum ofDclipD\_\{\\mathrm\{clip\}\}over the distributions ofqq’s inR⁡\(QC\)R\(Q\_\{C\}\)with that value is strictly increasing inQCQ\_\{C\}:

Dclip⋆​\(QC\)=PA​log⁡PA−PA​log⁡\(1−QC\)\+\|C\|​τ,d​Dclip⋆d​QC=PA1−QC\>0\.D\_\{\\mathrm\{clip\}\}^\{\\star\}\(Q\_\{C\}\)=P\_\{A\}\\log P\_\{A\}\-P\_\{A\}\\log\(1\-Q\_\{C\}\)\+\|C\|\\tau,\\qquad\\frac\{dD\_\{\\mathrm\{clip\}\}^\{\\star\}\}\{dQ\_\{C\}\}=\\frac\{P\_\{A\}\}\{1\-Q\_\{C\}\}\>0\.\(8\)

Appendix[B](https://arxiv.org/html/2609.38995#A2)gives the proof\. At the minimizerq⋆q^\{\\star\}of the clipped objective, the clipped candidate tokens’ total probability is not restored toward the teacher’s total1−PA1\-P\_\{A\}on these tokens but reduced further, to zero\. All probability then lies onAA, and the student’s probability on each active candidate token equals the teacher’s probabilitypip\_\{i\}multiplied by the same factor1/PA\>11/P\_\{A\}\>1\. Clipping therefore does more than limit the influence of dominant terms: inside a fixed clipping region, its minimizer assigns zero student probability to clipped candidate tokens, even though they already have less student probability than teacher probability, and gives each active candidate token more student probability than the teacher probability\.

A softmax of finite logits, however, never attainsq⋆q^\{\\star\}, because it gives every clipped token positive probability\. HenceQC\>0Q\_\{C\}\>0, andDclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)is an unattained infimum\. For such a student, Equation[8](https://arxiv.org/html/2609.38995#S4.E8)characterizes the minimum attainable objective value for eachQCQ\_\{C\}\. The lower the student’s total probability on the clipped candidate tokens, the closer the objective can come to this infimum, and restoringQCQ\_\{C\}toward the teacher’s1−PA1\-P\_\{A\}only moves this minimum attainable value farther from it\.

### 4\.3Clipped updates turn started repetition into persistent loops

In matched runs that differ only in whether the objective is clipped, clipped training produces substantially more terminal loops than exact KL training \(Figure[1](https://arxiv.org/html/2609.38995#S4.F1); Table[3](https://arxiv.org/html/2609.38995#S4.T3)\)\. We defined the repetition\-related terminologies in Table[2](https://arxiv.org/html/2609.38995#S4.T2)and use them throughout the section\. The preceding two sections give a candidate cause: the clipped objective reverses the logit gradient of exact KL on every clipped token and on active tokens within a bounded range above the teacher’s probability\. Within a fixed clipping region, its minimizer sets the clipped tokens’ probability to zero and gives each active token more probability than the teacher, instead of correcting the student toward the teacher\. If, in training, the teacher’s exit tokens are clipped and the copy token lies in this range, the clipped objective would push the student to leave a repetition less often and to continue it more often than the teacher, so repetitions and terminal loops would be more than in the exact KL twin, although the size of this effect may differ across families\. Both results describe one response position with a given clipped set, whereas in training the clipped set changes across sampled responses and positions, so whether the student’s probabilities move this way must be measured\. We compare student and teacher exit mass at the first copy, which every repetition passes through, and the copy token’s logit gradient under both objectives before maturity\. These comparisons can support the logit gradient reversal as the path from clipping to loops but cannot establish that path\.

Table 2:Definitions used to characterize token repetition and analyze its continuation\. The upper block defines periodic segments, repetitions, mature repetitions, and terminal loops\. The lower block defines the first copy, copy token, and exit mass\.Figure 1:Loop outcome over training, one trajectory per run\. Each point is the number of the 30 rollouts sampled at that update that end in a terminal loop\. Clipped runs in red, unclipped twins in blue\. Shaded windows are three phases: early \(updates 1–25, gray\), ignition \(F 51–75, M 76–100, K 151–175, L 101–125, orange\), and late \(updates 176–200, pink\)\.#### Outcome\.

We define three 25\-update phases: an early phase at the start of training, an ignition window at each clipped run’s loop onset, and a late phase at the end of training\. In the early phase, terminal loops are nearly absent from every run\. The exact KL runs remain there for the whole run, never exceeding 2 of 30 rollouts at any update\. The clipped runs do not\. Each run has a family\-specific onset after which terminal loops recur, and over the full run it produces 25 to 450 times as many of them as its twin\. The result is shown in Figure[1](https://arxiv.org/html/2609.38995#S4.F1)\. The thinking\-mode configuration sets both the scale and the ignition time\.

#### Trend\.

Table[3](https://arxiv.org/html/2609.38995#S4.T3)follows repetitions \(N≥2N\\geq 2\) through three phases\. In the early phase the two runs of every pair are indistinguishable except run L\. They generate repetition and mature repetition at the same rate\. From ignition on, two amplifications hold\. First, a repetition grows into a mature repetition \(N≥3N\\geq 3andL≥30L\\geq 30\) more often under clipping than in the twin\. The countN⁡\(mat\.∣N≥2\)N\(\\text\{mat\.\}\\mid N\\geq 2\)is 8\.6 to over 400 times the twin’s, and the rateP⁡\(mat\.∣N≥2\)P\(\\text\{mat\.\}\\mid N\\geq 2\)4\.5 to over 150 times\. Second, far more repetitions end as terminal loops: 26 to 381 per phase in the clipped runs, against at most two in any twin\.

#### Family\-specific loop patterns\.

With student thinking off and teacher thinking on \(F\), the student writes answer\-mode derivations that the teacher scores where its own thinking block would begin\. The student chains the repeated “⇒\\Rightarrow” where the teacher prefers\\text\. This two\-token cycle supplies most mature repetitions at ignition, and by the late phase all of them are terminal loops\. With both modes off \(K\), student and teacher share the answer register, and most of its terminal loops are answer\-closing formatting such as nested “\\boxed\{”\. The teacher prefers a content command such as\\textor\\frac\. With both modes on \(L\), most mature repetitions lie inside the thinking block and repeat an equation or a reasoning frame such as “Let me compute:”, where the teacher’s preferred alternative varies, most often a space, an opening parenthesis, or a different word such as “take”\. Most of its mature repetitions exit rather than become terminal\. With student thinking on and teacher thinking off \(M\), the teacher’s prompt contains a closed, empty thinking block, so the teacher scores the student’s thinking as answer text, and both runs stop emitting the closing tag\. At ignition, M’s terminal loops repeat an answer\-closing mark such as a check mark inside the unclosed thinking block, where the teacher prefers the end\-of\-response token<\|im\_end\|\>\.

#### Exit mass at the first copy\.

We dive deep into the training signal in each runs and how clipped objective applied to it\. We examine whether clipping suppresses exit mass at the first copy of a repetition, terms defined in Table[2](https://arxiv.org/html/2609.38995#S4.T2)\. Letℱu\\mathcal\{F\}\_\{u\}denotes the window containing first\-copy positions of all detected repetitions with periods2≤ℓ≤502\\leq\\ell\\leq 50across responses generated at updateuu\. For each nonoverlapping five\-update blockBB, we compute the student/teacher exit\-mass ratio by summing each distribution’s exit mass over these positions:

RB=∑u∈B∑t∈ℱu\(1−qu,t,copy\)∑u∈B∑t∈ℱu\(1−pu,t,copy\)\.R\_\{B\}=\\frac\{\\sum\_\{u\\in B\}\\sum\_\{t\\in\\mathcal\{F\}\_\{u\}\}\(1\-q\_\{u,t,\\mathrm\{copy\}\}\)\}\{\\sum\_\{u\\in B\}\\sum\_\{t\\in\\mathcal\{F\}\_\{u\}\}\(1\-p\_\{u,t,\\mathrm\{copy\}\}\)\}\.We then measure the clipped ratio of the teacher’s exit mass that lies on clipped tokens\. For the unclipped twins, we report the corresponding fraction on over\-threshold tokens \(Figure[2](https://arxiv.org/html/2609.38995#S4.F2)a\)\. Pooled over the early phase, the student/teacher exit\-mass ratio is \.79 to 1\.04\. Under exact KL, the ratio is close to 1 across the whole run, even though 33 to 45 percent of the teacher’s exit mass still lies on over\-threshold tokens\. Under clipping, however, the exit mass ratio falls to \.41 to \.70 in the late phase, while 71 to 84 percent of the teacher’s exit mass is clipped\. Thus, at the first copy, the clipped student places less probability than its teacher on leaving the repetition, and most of the teacher’s exit mass lies on the tokens whose logit gradient Proposition[1](https://arxiv.org/html/2609.38995#Thmproposition1)reverses\.

Table 3:Progression of repetition by training phase \(period 2–50\)\. Rows:n⁡\(N≥2\)n\(N\\geq 2\), the number of repetitions;N⁡\(mat\.∣N≥2\)N\(\\text\{mat\.\}\\mid N\\geq 2\), the mature repetitions among repetitions;P⁡\(mat\.∣N≥2\)P\(\\text\{mat\.\}\\mid N\\geq 2\)the proportion of repetitions that mature;N⁡\(term\.∣mat\.\)N\(\\text\{term\.\}\\mid\\text\{mat\.\}\), the terminal loops among mature repetition\. Each cell reads early→\\toignition→\\tolate, the three phases of Figure[1](https://arxiv.org/html/2609.38995#S4.F1), 750 rollouts each\.Figure 2:The copy token inside repetitions \(period 2–50\), with the student’s probabilities taken before the sampler’s truncation\. \(a\) Student/teacher exit\-mass ratio, one point per block of five updates \(1–5, …, 196–200\)\. Lines without markers give the ratio of the teacher’s exit mass on clipped tokens, or on over\-threshold tokens in the twins \(right axis\)\. \(b\) Copied positions between the first copy and maturity: fraction withqcopy\>pcopyq\_\{\\mathrm\{copy\}\}\>p\_\{\\mathrm\{copy\}\}\. Dark bar:qcopy<pcopy/PAq\_\{\\mathrm\{copy\}\}<p\_\{\\mathrm\{copy\}\}/P\_\{A\}, where the logit gradient of exact KL points toward lowering the copy logit and that of the clipped objective toward raising it\. Light bar: both point toward lowering it\. \(c\) The same copied positions, plus observed stopping positions before maturity: mean−∂D/∂zcopy\-\\partial D/\\partial z\_\{\\mathrm\{copy\}\}\. Black and red: exact KL \(hypothetical\) and clipped KL, respectively, evaluated on the same positions and probabilities from the clipped run\. Positive values point toward raising the copy logit\. Phases as in Table[3](https://arxiv.org/html/2609.38995#S4.T3)\.
#### Copy token probability before maturity\.

We next examine the positions a repetition passes through before it becomes mature: how often the student places more probability than the teacher on the copy token, and whether each objective corrects this excess\.

For each 25\-update phaseϕ\\phi, let𝒢ϕ\\mathcal\{G\}\_\{\\phi\}denote the window of copied positions after the first copy and before maturity or an observed stop\. We record the fraction at which the student assigns greater probability to the copy token than the teacher:

Rϕ=∑t∈𝒢ϕ𝟏\{qt,copy\>pt,copy\}\|𝒢ϕ\|,R\_\{\\phi\}=\\frac\{\\sum\_\{t\\in\\mathcal\{G\}\_\{\\phi\}\}\\mathbf\{1\}\\\{q\_\{t,\\mathrm\{copy\}\}\>p\_\{t,\\mathrm\{copy\}\}\\\}\}\{\|\\mathcal\{G\}\_\{\\phi\}\|\},where𝟏​\{⋅\}\\mathbf\{1\}\\\{\\cdot\\\}is the indicator function and counts are pooled across repetitions and updates within the phase\. The result is in Figure[2](https://arxiv.org/html/2609.38995#S4.F2)b\. In the early phase, both runs of F, M, and L haveqcopy\>pcopyq\_\{\\mathrm\{copy\}\}\>p\_\{\\mathrm\{copy\}\}at 29 to 45 percent of these positions\. After ignition, this fraction falls to 4 to 6 percent in the exact KL runs but 32 to 77 percent in the clipped runs\. In the clipped runs,qcopyq\_\{\\mathrm\{copy\}\}lies specifically betweenpcopyp\_\{\\mathrm\{copy\}\}andpcopy/PAp\_\{\\mathrm\{copy\}\}/P\_\{A\}at 4 to 34 percent of positions\. The dark part of each bar in Figure[2](https://arxiv.org/html/2609.38995#S4.F2)b shows the fraction of positions at which the copy token’s logit gradient is reversed: the clipped objective points toward raising the copy token logit where exact KL points toward lowering it\.

We then measure the gradient of each objectiveD∈\{DFKL,Dclip\}D\\in\\\{D\_\{\\mathrm\{FKL\}\},D\_\{\\mathrm\{clip\}\}\\\}on the copy token\. For each phaseϕ\\phi, we record the mean of−∂D/∂zcopy\-\\partial D/\\partial z\_\{\\mathrm\{copy\}\}over the window𝒢ϕ\\mathcal\{G\}\_\{\\phi\}plus the observed stop positions, pooled across repetitions and updates\. We report the negative of the gradient so that positive values point toward raising the copy token logit\. We calculate∂DFKL/∂zcopy=qcopy−pcopy\\partial D\_\{\\mathrm\{FKL\}\}/\\partial z\_\{\\mathrm\{copy\}\}=q\_\{\\mathrm\{copy\}\}\-p\_\{\\mathrm\{copy\}\}and∂Dclip/∂zcopy\\partial D\_\{\\mathrm\{clip\}\}/\\partial z\_\{\\mathrm\{copy\}\}, given by Equation[5](https://arxiv.org/html/2609.38995#S4.E5)\. We evaluate both objectives on the positions and probabilities of the clipped run only, so that any difference between the two means comes from the objective alone\. Figure[2](https://arxiv.org/html/2609.38995#S4.F2)c shows the result: the net direction of each objective’s logit gradient on the copy token over all positions of the window\. In the early phase, both values are nearly zero in every family\. In the ignition and late phases, the mean of−∂DFKL/∂zcopy\-\\partial D\_\{\\mathrm\{FKL\}\}/\\partial z\_\{\\mathrm\{copy\}\}is negative, pointing toward lowering the copy token logit, whereas that ofDclipD\_\{\\mathrm\{clip\}\}is slightly positive in every family, pointing toward raising it\. At the positions that decide whether a repetition becomes mature, exact KL would thus correct the student’s excess probability on the copy token, and the clipped objective removes this correction\.

## 5Held\-Out Evaluation

We evaluate every checkpoint, saved at 25\-update intervals, on the AIME 2025 problems\. Figure[3](https://arxiv.org/html/2609.38995#S6.F3)reports accuracy and terminal loop rate at saved checkpoints\. It tests whether clipping increases looping and whether the additional loops accompany accuracy losses relative to the unclipped twin\.

No run significantly outperforms the base model \(Figure[3](https://arxiv.org/html/2609.38995#S6.F3)a\)\. K show only minimal accuracy changes, whereas F, L, and M degrade substantially, and most clipped runs perform worse than their unclipped twins\. M is the exception: both of its run collapse, and its twin reads zero only because it stops closing the thinking tag, so the strict scorer finds no answer\.

Terminal loop rate separates the runs more sharply \(Figure[3](https://arxiv.org/html/2609.38995#S6.F3)b\)\. Every clipped run ends in a terminal loop far more often than the base model, while every unclipped twin stays near the base rate\. Looped responses fail to reach a final answer and are therefore scored as incorrect\. Accuracy losses in the clipped runs follow the same ordering as their loop rates: smallest in K, intermediate in F, and largest in L and M\.

With 12 samples per problem, no clipped endpoint has lower pass@12 than its twin on either benchmark: every problem solved by the twin is also solved at least once by the clipped arm\. We interpret this pattern as evidence that clipped training reduces per\-sample reliability\. However, the unclipped F and L twins remain below the fixed\-base mean despite rarely looping\. Thus, clipping\-induced looping contributes to the observed degradation, but does not explain all accuracy loss relative to the base model\.

Appendix[D](https://arxiv.org/html/2609.38995#A4)gives avg@12, terminal loop rate, and pass@12 at every checkpoint on AIME 2024 and AIME 2025 benchmarks\.

## 6Conclusion

Pointwise clipping was introduced to stabilize training against dominant stylistic tokens, but it also changes the direction in which forward kl corrects the student\. In training this change appears as repetitions that continue to the end of the response, a loss of training stability\. Because this failure develops over updates, studies that train with the pointwise\-clipped objective could report accuracy and a text degeneration measure, across checkpoints, in addition to their best checkpoint\. We hope that understanding these clipping dynamics helps make on\-policy self\-distillation stable enough for its gains to persist over longer training\.

Figure 3:AIME 2025 across training\. \(a\) Avg@12, the mean correct ratio over 12 samples for each of 30 problems at the checkpoint saved after each 25 updates\. \(b\) The ratio of the 360 responses per checkpoint that end in a terminal loop\. Solid curves with circles are the clipped runs, dashed curves with squares their unclipped twins\. Each trained curve is one realized trajectory\. The black arrow on each vertical axis marks the base model, the fixed starting checkpoint averaged over eight evaluation\-sampling seeds\.
## References

- Agarwal et al\. \(2024\)Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem\.On\-policy distillation of language models: Learning from self\-generated mistakes\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=3zKtaqxLhW](https://openreview.net/forum?id=3zKtaqxLhW)\.
- Chen et al\. \(2026a\)Yunmeng Chen, Kunyu Wang, Peihan Li, Yi Wang, Shuyin Xia, Yi Liu, Xinyong Cheng, Dehui Wang, Xiangyong Zhai, Yanxing Liu, et al\.Scope\-opsd: Fisher\-conditioned privileged subspaces for on\-policy self\-distillation\.*arXiv preprint arXiv:2609\.12579*, 2026a\.
- Chen et al\. \(2026b\)Yutong Chen, Guangfu Guo, Zhichao Xu, and Kunpeng Liu\.Dualopsd: Adaptive privileged teachers for on\-policy self\-distillation\.*arXiv preprint arXiv:2608\.26019*, 2026b\.
- Feng et al\. \(2026\)Yangyang Feng, Zhuoyan Feng, and Junlan Chen\.Past: Privileged adaptation from complete student trajectories for on\-policy self\-distillation\.*arXiv preprint arXiv:2608\.08726*, 2026\.
- Gu et al\. \(2026\)Siyi Gu, Jialin Chen, Sophia Zhou, Arman Cohan, and Rex Ying\.Rethinking reward supervision: Rubric\-conditioned self\-distillation\.*arXiv preprint arXiv:2606\.19327*, 2026\.
- Guha et al\. \(2026\)Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, et al\.Openthoughts: Data recipes for reasoning models\.In*International Conference on Learning Representations*, volume 2026, pp\. 108059–108130, 2026\.
- Hou et al\. \(2026\)ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, et al\.Dash: Divergence\-adaptive supervision horizons for on\-policy self\-distillation of reasoning models\.*arXiv preprint arXiv:2608\.06243*, 2026\.
- Hübotter et al\. \(2026\)Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, et al\.Reinforcement learning via self\-distillation\.*arXiv preprint arXiv:2601\.20802*, 2026\.
- Ichihara et al\. \(2026\)Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, and Junpei Komiyama\.Privileged solutions or context\-induced teacher behavior? dissecting on\-policy self\-distillation\.*arXiv preprint arXiv:2608\.09228*, 2026\.
- \(10\)Hynek Kydlíček\.Math\-Verify: Math Verification Library\.URL[https://github\.com/huggingface/math\-verify](https://github.com/huggingface/math-verify)\.
- Li et al\. \(2026\)Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, and Nuno Vasconcelos\.On\-policy self\-distillation without any supervision\.*arXiv preprint arXiv:2608\.06296*, 2026\.
- Liang et al\. \(2026\)Zihan Liang, Yufei Ma, Ben Chen, Zhipeng Qian, Xuxin Zhang, Huangyu Dai, and Lingtao Mao\.Search\-e1: Self\-distillation drives self\-evolution in search\-augmented reasoning\.*arXiv preprint arXiv:2605\.22511*, 2026\.
- Liu et al\. \(2026\)Xiaogeng Liu, Xinyan Wang, Yingzi Ma, Yechao Zhang, and Chaowei Xiao\.When are teacher tokens reliable? position\-weighted on\-policy self\-distillation for reasoning\.*arXiv preprint arXiv:2605\.21606*, 2026\.
- Lu & Lab \(2025\)Kevin Lu and Thinking Machines Lab\.On\-policy distillation\.*Thinking Machines Lab: Connectionism*, 2025\.doi:10\.64434/tml\.20251026\.https://thinkingmachines\.ai/blog/on\-policy\-distillation\.
- Pan et al\. \(2026\)Leyi Pan, Shuchang Tao, Yunpeng Zhai, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, and Lijie Wen\.Rlcsd: Reinforcement learning with contrastive on\-policy self\-distillation\.*arXiv preprint arXiv:2606\.11709*, 2026\.
- Shenfeld et al\. \(2026\)Idan Shenfeld, Mehul Damani, Jonas Hübotter, and Pulkit Agrawal\.Self\-distillation enables continual learning\.In*Forty\-third International Conference on Machine Learning*, 2026\.URL[https://openreview\.net/forum?id=qA6FgH0nnZ](https://openreview.net/forum?id=qA6FgH0nnZ)\.
- Sheng et al\. \(2024\)Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu\.Hybridflow: A flexible and efficient rlhf framework\.*arXiv preprint arXiv: 2409\.19256*, 2024\.
- Shrestha & Tessier \(2026\)Samyak Shrestha and Alexander Tessier\.Rethinking privileged information in on\-policy self\-distillation\.*arXiv preprint arXiv:2608\.18271*, 2026\.
- Tan & Hong \(2026a\)Zhiquan Tan and Yinrong Hong\.Paint: Partial\-solution adaptive interpolated training for self\-distilled reasoners\.*arXiv preprint arXiv:2604\.26573*, 2026a\.
- Tan & Hong \(2026b\)Zhiquan Tan and Yinrong Hong\.Self\-supervised on\-policy distillation for reasoning language models\.*arXiv preprint arXiv:2605\.17497*, 2026b\.
- Wang et al\. \(2026\)Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Zheng Pan, Xin Li, and Lan\-Zhe Guo\.Trace: Distilling where it matters via token\-routed self on\-policy alignment\.*arXiv preprint arXiv:2605\.10194*, 2026\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Yang et al\. \(2026\)Yuxiao Yang, Xiaoyun Wang, and Weitong Zhang\.Ogls\-sd: On\-policy self\-distillation with outcome\-guided logit steering for llm reasoning\.*arXiv preprint arXiv:2605\.12400*, 2026\.
- Zhang et al\. \(2026\)XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, and Tat\-Seng Chua\.What does privileged information add to on\-policy self\-distillation?*arXiv preprint arXiv:2609\.20612*, 2026\.
- Zhang & Math\-AI \(2024\)Yifan Zhang and Team Math\-AI\.American invitational mathematics examination \(aime\) 2024, 2024\.
- Zhang & Math\-AI \(2025\)Yifan Zhang and Team Math\-AI\.American invitational mathematics examination \(aime\) 2025, 2025\.
- Zhao et al\. \(2026a\)Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover\.Self\-distilled reasoner: On\-policy self\-distillation for large language models\.In*Forty\-third International Conference on Machine Learning*, 2026a\.URL[https://openreview\.net/forum?id=Jpxfof0EaS](https://openreview.net/forum?id=Jpxfof0EaS)\.
- Zhao et al\. \(2026b\)Xuyang Zhao, Liting Zhang, Zichen Xu, Zhihu Wang, Xu Caiyue, Shiwan Zhao, and Qicheng Li\.Is more privileged information better? From solution traces to problem\-solving structure in self\-distilled reasoning\.*arXiv preprint arXiv:2608\.01589*, 2026b\.URL[https://arxiv\.org/abs/2608\.01589](https://arxiv.org/abs/2608.01589)\.
- Zheng et al\. \(2024\)Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody H Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al\.Sglang: Efficient execution of structured language model programs\.*Advances in neural information processing systems*, 37:62557–62583, 2024\.

## Appendix AProof of Proposition[1](https://arxiv.org/html/2609.38995#Thmproposition1)\.

Fix a response position and suppress its indextt\. Hold the finite support𝒮\\mathcal\{S\}and the teacher distributionppfixed\. By Equation[1](https://arxiv.org/html/2609.38995#S3.E1),pj,qj\>0p\_\{j\},q\_\{j\}\>0for everyj∈𝒮j\\in\\mathcal\{S\}, and both distributions sum to one on𝒮\\mathcal\{S\}\. For the temperature\-scaled student logitszj=bj/Tz\_\{j\}=b\_\{j\}/T, write

Z=∑j∈𝒮exp⁡\(zj\),qj=exp⁡\(zj\)Z\.Z=\\sum\_\{j\\in\\mathcal\{S\}\}\\exp\(z\_\{j\}\),\\qquad q\_\{j\}=\\frac\{\\exp\(z\_\{j\}\)\}\{Z\}\.
Consider any logit vectorz∘z^\{\\circ\}satisfyingdj≠τd\_\{j\}\\neq\\taufor allj∈𝒮j\\in\\mathcal\{S\}, and letAAandCCbe its active and clipped sets as defined in Section[4\.1](https://arxiv.org/html/2609.38995#S4.SS1)\. Eachdj=pj​log⁡\(pj/qj\)d\_\{j\}=p\_\{j\}\\log\(p\_\{j\}/q\_\{j\}\)is smooth inzz, sincepjp\_\{j\}is fixed andqjq\_\{j\}is a positive smooth function ofzz\. Because𝒮\\mathcal\{S\}is finite and nodjd\_\{j\}equalsτ\\tauatz∘z^\{\\circ\}, there is an open neighborhood ofz∘z^\{\\circ\}on whichAAandCC, and hencePA=∑j∈ApjP\_\{A\}=\\sum\_\{j\\in A\}p\_\{j\}, remain unchanged\. On this neighborhood, Equation[4](https://arxiv.org/html/2609.38995#S4.E4)gives

Dclip=∑j∈Adj\+\|C\|​τ,D\_\{\\mathrm\{clip\}\}=\\sum\_\{j\\in A\}d\_\{j\}\+\|C\|\\tau,withAAandCCfixed at their values atz∘z^\{\\circ\}\. This expression is smooth inzz, soDclipD\_\{\\mathrm\{clip\}\}is differentiable on the neighborhood and its gradient is obtained by differentiating the displayed expression\.

For anyi,j∈𝒮i,j\\in\\mathcal\{S\}, the softmax identitylog⁡qj=zj−log⁡Z\\log q\_\{j\}=z\_\{j\}\-\\log Zgives

dj=pj​log⁡pj−pj​zj\+pj​log⁡Z\.d\_\{j\}=p\_\{j\}\\log p\_\{j\}\-p\_\{j\}z\_\{j\}\+p\_\{j\}\\log Z\.Becauseppis fixed and

∂zj∂zi=𝟏\{j=i\},∂log⁡Z∂zi=1Z∂Z∂zi=exp⁡\(zi\)Z=qi,\\frac\{\\partial z\_\{j\}\}\{\\partial z\_\{i\}\}=\\mathbf\{1\}\\\{j=i\\\},\\qquad\\frac\{\\partial\\log Z\}\{\\partial z\_\{i\}\}=\\frac\{1\}\{Z\}\\frac\{\\partial Z\}\{\\partial z\_\{i\}\}=\\frac\{\\exp\(z\_\{i\}\)\}\{Z\}=q\_\{i\},it follows that

∂dj∂zi=−pj𝟏\{j=i\}\+pjqi\.\\frac\{\\partial d\_\{j\}\}\{\\partial z\_\{i\}\}=\-p\_\{j\}\\mathbf\{1\}\\\{j=i\\\}\+p\_\{j\}q\_\{i\}\.\(9\)
SinceDFKL=∑j∈𝒮djD\_\{\\mathrm\{FKL\}\}=\\sum\_\{j\\in\\mathcal\{S\}\}d\_\{j\}by Equation[3](https://arxiv.org/html/2609.38995#S3.E3), summing Equation[9](https://arxiv.org/html/2609.38995#A1.E9)over𝒮\\mathcal\{S\}gives

∂DFKL∂zi\\displaystyle\\frac\{\\partial D\_\{\\mathrm\{FKL\}\}\}\{\\partial z\_\{i\}\}=∑j∈𝒮\(−pj𝟏\{j=i\}\+pjqi\)\\displaystyle=\\sum\_\{j\\in\\mathcal\{S\}\}\\bigl\(\-p\_\{j\}\\mathbf\{1\}\\\{j=i\\\}\+p\_\{j\}q\_\{i\}\\bigr\)=−pi\+qi​∑j∈𝒮pj=qi−pi,\\displaystyle=\-p\_\{i\}\+q\_\{i\}\\sum\_\{j\\in\\mathcal\{S\}\}p\_\{j\}=q\_\{i\}\-p\_\{i\},at every logit vector, including points on the clipping boundaries\. For the clipped objective,\|C\|​τ\|C\|\\tauhas zero derivative on the neighborhood under consideration\. Summing only overAAtherefore gives

∂Dclip∂zi\\displaystyle\\frac\{\\partial D\_\{\\mathrm\{clip\}\}\}\{\\partial z\_\{i\}\}=∑j∈A\(−pj𝟏\{j=i\}\+pjqi\)\\displaystyle=\\sum\_\{j\\in A\}\\bigl\(\-p\_\{j\}\\mathbf\{1\}\\\{j=i\\\}\+p\_\{j\}q\_\{i\}\\bigr\)=−pi𝟏\{i∈A\}\+qi∑j∈Apj\\displaystyle=\-p\_\{i\}\\mathbf\{1\}\\\{i\\in A\\\}\+q\_\{i\}\\sum\_\{j\\in A\}p\_\{j\}=PAqi−pi𝟏\{i∈A\}\.\\displaystyle=P\_\{A\}q\_\{i\}\-p\_\{i\}\\mathbf\{1\}\\\{i\\in A\\\}\.\(10\)This establishes Equation[5](https://arxiv.org/html/2609.38995#S4.E5)\.

Now leti∈Ci\\in C\. Thendi\>τ\>0d\_\{i\}\>\\tau\>0, andpi\>0p\_\{i\}\>0implies

log⁡piqi\>0,hencepi\>qi\.\\log\\frac\{p\_\{i\}\}\{q\_\{i\}\}\>0,\\qquad\\text\{hence\}\\qquad p\_\{i\}\>q\_\{i\}\.Consequently,∂DFKL/∂zi=qi−pi<0\\partial D\_\{\\mathrm\{FKL\}\}/\\partial z\_\{i\}=q\_\{i\}\-p\_\{i\}<0\. Sincei∉Ai\\notin A, Equation[10](https://arxiv.org/html/2609.38995#A1.E10)reduces to∂Dclip/∂zi=PA​qi\\partial D\_\{\\mathrm\{clip\}\}/\\partial z\_\{i\}=P\_\{A\}q\_\{i\}\. To establish strict positivity, first note thatqi\>0q\_\{i\}\>0\. Moreover,∑j∈𝒮qj=∑j∈𝒮pj\\sum\_\{j\\in\\mathcal\{S\}\}q\_\{j\}=\\sum\_\{j\\in\\mathcal\{S\}\}p\_\{j\}implies the existence of a tokenjjwithqj≥pjq\_\{j\}\\geq p\_\{j\}\. For this token,

dj=pj​log⁡pjqj≤0<τ,d\_\{j\}=p\_\{j\}\\log\\frac\{p\_\{j\}\}\{q\_\{j\}\}\\leq 0<\\tau,soj∈Aj\\in AandPA≥pj\>0P\_\{A\}\\geq p\_\{j\}\>0\. Therefore,

∂DFKL∂zi=qi−pi<0,∂Dclip∂zi=PA​qi\>0,\\frac\{\\partial D\_\{\\mathrm\{FKL\}\}\}\{\\partial z\_\{i\}\}=q\_\{i\}\-p\_\{i\}<0,\\qquad\\frac\{\\partial D\_\{\\mathrm\{clip\}\}\}\{\\partial z\_\{i\}\}=P\_\{A\}q\_\{i\}\>0,which proves Equation[6](https://arxiv.org/html/2609.38995#S4.E6)\. Sincez∘z^\{\\circ\}was arbitrary away from the clipping boundaries, the result holds at every point specified in the proposition\. ∎

#### A simpler example of updating dynamics

To understand the direction of the updates on the logits as suggested by the previous proposition, the example below can give some intuitions\. In the most general case, decreasing certain logits on indexiidoes not automatically mean to decrease the corresponding probabilityqiq\_\{i\}–but in the case when otherzzvalues are fixed, this holds true\. We can construct a continuous\-time dynamical system to mimic the optimization process on the objective functions of the two model versions, defining the parameter updates as the negative gradient of the loss with respect to the logits \(d​zid​t=−∇ziD\\frac\{dz\_\{i\}\}\{dt\}=\-\\nabla\_\{z\_\{i\}\}D\)\. For the exact forward kl objectiveDF​K​LD\_\{FKL\}, this logit update is defined continuously as:

d​zid​t​\(t\)=pi−qi​\(t\)\\frac\{dz\_\{i\}\}\{dt\}\(t\)=p\_\{i\}\-q\_\{i\}\(t\)
For the clipped objective, the gradient space is instead split into a piecewise system governed by the local divergence thresholddjd\_\{j\}:

d​zid​t​\(t\)=\{pi−PA​qi​\(t\),if​dj≤τ−PA​qi​\(t\),if​dj\>τ\\frac\{dz\_\{i\}\}\{dt\}\(t\)=\\begin\{cases\}p\_\{i\}\-P\_\{A\}q\_\{i\}\(t\),&\\text\{if \}d\_\{j\}\\leq\\tau\\\\ \-P\_\{A\}q\_\{i\}\(t\),&\\text\{if \}d\_\{j\}\>\\tau\\end\{cases\}
To map these raw logit updatesd​zid​t\\frac\{dz\_\{i\}\}\{dt\}to the resulting probability trajectoriesd​qid​t\\frac\{dq\_\{i\}\}\{dt\}, we must account for the coupling effect of the softmax function, where the evolution of a single probabilityqiq\_\{i\}depends on the updates to all logits in the vocabulary\. However, to isolate the localized forces acting on a single prediction and visualize them within a simplified 2D direction field diagram, we introduce a specific framing assumption where all other logitszj,j≠iz\_\{j\},j\\neq iare temporarily held constant \(d​zjd​t=0\\frac\{dz\_\{j\}\}\{dt\}=0\)\. Under this isolated assumption, because the softmax function is strictly monotonically increasing with respect to its own logit,d​qid​t\\frac\{dq\_\{i\}\}\{dt\}acts as a positive monotonic scaling ofd​zid​t\\frac\{dz\_\{i\}\}\{dt\}\. This strict sign preservation guarantees that the direction of the probability update exactly mirrors the direction of the logit update, providing the mathematical justification to map the negative gradients directly onto a 2D phase portrait to evaluate their critical points and structural stability\. The direction field, null\-cline, contour for the clipping region and the diagonal \(for the correct alignment p=q\) is shown in figure[4](https://arxiv.org/html/2609.38995#A1.F4)\.

![Refer to caption](https://arxiv.org/html/2609.38995v1/figures/phase_direction_field.png)Figure 4:A simple example of updating direction of q under different losses\(Forward KL v\.s\. Clipped Loss Updates\.Applying this localized framework to the exact forward kl divergence \(left\) demonstrates that the gradient acts as a proportional restorative force across the entire probability space, driving the system toward a single, globally stable equilibrium along the critical boundary wherepi=qip\_\{i\}=q\_\{i\}\(the diagonal\)\. Because the directional derivative relies entirely on the unaltered residual, the vector field smoothly and pushes the predictions directly toward perfect calibration\. This ensures the parameter space remains entirely free of dead zones, structural barriers, or competing attractors\.

Conversely, the clipped objective \(right\) shatters the global stability of the forward kl\. In the active set \(unshaded area on bottom\-right corner\) where the divergence is low \(dj<τd\_\{j\}<\\tau\), scaling the predicted term byPAP\_\{A\}artificially shifts the fixed point away from true alignment, forcing the vectors toward a false attractor along the linepi=PA​qip\_\{i\}=P\_\{A\}q\_\{i\}\(the blue line with slope that is less steep\)\. When the local divergence exceeds theτ\\tauthreshold, the system crosses into the clipped set \(green shaded area on the top left\), completely erasing the target probabilityqiq\_\{i\}from the derivative calculation\. Inside this boundary, the dynamic devolves into pure probability decay\. Because the gradient equation yields a negative update \(−PA​qi\-P\_\{A\}q\_\{i\}\), it relentlessly pushes the predictions until00is reached, structurally preventing the system from ever recovering true calibration when initialized or driven too far from the target\. To sum up, the green shaded area along with the area between the red and blue lines on the right are where the arrows for the clipped objectives are in a different direction as compared with the ones for forward KL\. One might observe that if you trace a phase portrait alone the arrows, then if you start with a point in either clipped or active set, the trajectory never goes into the other–this may give some intuition for boundary check for the next proposition as well\.

## Appendix BProof of Proposition[2](https://arxiv.org/html/2609.38995#Thmproposition2)

#### Setup\.

Fix one response position\. Letppbe the teacher distribution on the finite candidate set𝒮\\mathcal\{S\}\. It is a softmax output, sopi\>0p\_\{i\}\>0for everyi∈𝒮i\\in\\mathcal\{S\}, and∑i∈𝒮pi=1\\sum\_\{i\\in\\mathcal\{S\}\}p\_\{i\}=1\. Letτ\>0\\tau\>0\. Student distributions are the points of the closed simplex,

qi≥0\(i∈𝒮\),∑i∈𝒮qi=1,q\_\{i\}\\geq 0\\quad\(i\\in\\mathcal\{S\}\),\\qquad\\sum\_\{i\\in\\mathcal\{S\}\}q\_\{i\}=1,so probabilities equal to zero are allowed\. We use the conventionpj​log⁡\(pj/0\)=\+∞p\_\{j\}\\log\(p\_\{j\}/0\)=\+\\infty\. The capped term of a zero\-probability token is thenmin⁡\{\+∞,τ\}=τ\\min\\\{\+\\infty,\\tau\\\}=\\tau, so

Dclip​\(q\)=∑j∈𝒮min⁡\{pj​log⁡pjqj,τ\}D\_\{\\mathrm\{clip\}\}\(q\)=\\sum\_\{j\\in\\mathcal\{S\}\}\\min\\Bigl\\\{p\_\{j\}\\log\\frac\{p\_\{j\}\}\{q\_\{j\}\},\\ \\tau\\Bigr\\\}is finite on the whole simplex\.

LetAAandCCbe fixed nonempty disjoint sets withA∪C=𝒮A\\cup C=\\mathcal\{S\}\. The*fixed clipping region*RRis the set of student distributionsqqwith

pi​log⁡piqi≤τ\(i∈A\),pj​log⁡pjqj\>τ\(j∈C\)\.p\_\{i\}\\log\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\leq\\tau\\quad\(i\\in A\),\\qquad p\_\{j\}\\log\\frac\{p\_\{j\}\}\{q\_\{j\}\}\>\\tau\\quad\(j\\in C\)\.WritePA=∑i∈ApiP\_\{A\}=\\sum\_\{i\\in A\}p\_\{i\}andQC=∑j∈CqjQ\_\{C\}=\\sum\_\{j\\in C\}q\_\{j\}\. SinceAAandCCare nonempty and everypi\>0p\_\{i\}\>0, we have0<PA<10<P\_\{A\}<1\. Forq∈Rq\\in R, every active term is at mostτ\\tau, so its minimum withτ\\tauis the term itself, and every clipped term exceedsτ\\tau, so its minimum withτ\\tauisτ\\tau\. Hence, forq∈Rq\\in R\(Equation[4](https://arxiv.org/html/2609.38995#S4.E4)\),

Dclip​\(q\)=∑i∈Api​log⁡piqi\+∑j∈Cτ=F⁡\(qA\)\+\|C\|​τ,F⁡\(qA\)=∑i∈Api​log⁡piqi,qA=\(qi\)i∈A\.D\_\{\\mathrm\{clip\}\}\(q\)=\\sum\_\{i\\in A\}p\_\{i\}\\log\\frac\{p\_\{i\}\}\{q\_\{i\}\}\+\\sum\_\{j\\in C\}\\tau=F\(q\_\{A\}\)\+\|C\|\\tau,\\qquad F\(q\_\{A\}\)=\\sum\_\{i\\in A\}p\_\{i\}\\log\\frac\{p\_\{i\}\}\{q\_\{i\}\},\\qquad q\_\{A\}=\(q\_\{i\}\)\_\{i\\in A\}\.This identity is valid onRR\. OutsideRRit fails in general, because a term that exceedsτ\\tauis capped\.

Claim\.The distributionq⋆q^\{\\star\}withqi⋆=pi/PAq^\{\\star\}\_\{i\}=p\_\{i\}/P\_\{A\}fori∈Ai\\in Aandqj⋆=0q^\{\\star\}\_\{j\}=0forj∈Cj\\in Cis the unique minimizer ofDclipD\_\{\\mathrm\{clip\}\}overRR, andDclip​\(q⋆\)=PA​log⁡PA\+\|C\|​τD\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)=P\_\{A\}\\log P\_\{A\}\+\|C\|\\tau\(Equation[7](https://arxiv.org/html/2609.38995#S4.E7)\)\.

#### Overview\.

The proof has two steps\.*Step 1*settles how much mass sits on the clipped set: any such mass can be moved onto an active token without leavingRR, and doing so strictly lowers the loss, so only distributions withQC=0Q\_\{C\}=0can compete\.*Step 2*settles how that mass is split within the active set\. Dropping the active\-token clipping inequalities leaves a larger, smooth problem: minimize the strictly convex functionFFover the positive allocations with∑i∈Aqi=1\\sum\_\{i\\in A\}q\_\{i\}=1\. Its unique global minimizer is the Lagrange pointqi=pi/PAq\_\{i\}=p\_\{i\}/P\_\{A\}, which exceedspip\_\{i\}becausePA<1P\_\{A\}<1\.*The crucial implication is that the unique global minimizer of the larger, smooth problem satisfies the original clipping constraints\.*It is therefore also the unique minimizer of the original, restricted problem\. The larger problem is only an auxiliary device: outsideRRthe formulaFFis not the clipped objective, and nothing is claimed there\.

#### Membership facts \(M1\)–\(M4\)\.

Steps 1 and 2 use four facts about when a token is active or clipped\. These are the only places where the threshold enters the proof\. Fix a tokeniiand regard its termpi​log⁡\(pi/x\)p\_\{i\}\\log\(p\_\{i\}/x\)as a function of its student probabilityx≥0x\\geq 0, with value\+∞\+\\inftyatx=0x=0\. Forx\>0x\>0,

dd​x​\[pi​log⁡pix\]=dd​x​\[pi​log⁡pi−pi​log⁡x\]=−pix<0,\\frac\{d\}\{dx\}\\Bigl\[p\_\{i\}\\log\\frac\{p\_\{i\}\}\{x\}\\Bigr\]=\\frac\{d\}\{dx\}\\bigl\[p\_\{i\}\\log p\_\{i\}\-p\_\{i\}\\log x\\bigr\]=\-\\frac\{p\_\{i\}\}\{x\}<0,so the term is strictly decreasing inxxon\(0,∞\)\(0,\\infty\), and atx=pix=p\_\{i\}it equalspi​log⁡1=0p\_\{i\}\\log 1=0\.

1. \(M1\)*Every active token hasqi\>0q\_\{i\}\>0\.*Ifqi=0q\_\{i\}=0, the term is\+∞\>τ\+\\infty\>\\tau, so tokeniiwould be clipped, not active\. ConsequentlyFFis finite onRR\.
2. \(M2\)*An active token that gains probability stays active\.*Ifqi′≥qi\>0q^\{\\prime\}\_\{i\}\\geq q\_\{i\}\>0, then, because the term is decreasing,pi​log⁡\(pi/qi′\)≤pi​log⁡\(pi/qi\)≤τp\_\{i\}\\log\(p\_\{i\}/q^\{\\prime\}\_\{i\}\)\\leq p\_\{i\}\\log\(p\_\{i\}/q\_\{i\}\)\\leq\\tau\.
3. \(M3\)*A clipped token that loses probability stays clipped, including at zero\.*Let0≤qj′≤qj0\\leq q^\{\\prime\}\_\{j\}\\leq q\_\{j\}\. Ifqj′=0q^\{\\prime\}\_\{j\}=0, the term is\+∞\>τ\+\\infty\>\\tau\. Ifqj′\>0q^\{\\prime\}\_\{j\}\>0, thenqj\>0q\_\{j\}\>0as well andpj​log⁡\(pj/qj′\)≥pj​log⁡\(pj/qj\)\>τp\_\{j\}\\log\(p\_\{j\}/q^\{\\prime\}\_\{j\}\)\\geq p\_\{j\}\\log\(p\_\{j\}/q\_\{j\}\)\>\\tau\.
4. \(M4\)*A token withqi≥piq\_\{i\}\\geq p\_\{i\}is active for everyτ\>0\\tau\>0\.*Because the term is decreasing,pi​log⁡\(pi/qi\)≤pi​log⁡\(pi/pi\)=0<τp\_\{i\}\\log\(p\_\{i\}/q\_\{i\}\)\\leq p\_\{i\}\\log\(p\_\{i\}/p\_\{i\}\)=0<\\tau\.

#### Step 1: mass on the clipped set prevents optimality\.

Letq∈Rq\\in RwithQC\>0Q\_\{C\}\>0\. Pick any active tokenk∈Ak\\in Aand move all clipped mass onto it:

qk′=qk\+QC,qj′=0\(j∈C\),qi′=qi\(i∈A,i≠k\)\.q^\{\\prime\}\_\{k\}=q\_\{k\}\+Q\_\{C\},\\qquad q^\{\\prime\}\_\{j\}=0\\ \\ \(j\\in C\),\\qquad q^\{\\prime\}\_\{i\}=q\_\{i\}\\ \\ \(i\\in A,\\ i\\neq k\)\.*Normalization\.*∑i∈𝒮qi′=∑i∈A,i≠kqi\+\(qk\+QC\)\+0=∑i∈Aqi\+QC=1\\sum\_\{i\\in\\mathcal\{S\}\}q^\{\\prime\}\_\{i\}=\\sum\_\{i\\in A,\\,i\\neq k\}q\_\{i\}\+\(q\_\{k\}\+Q\_\{C\}\)\+0=\\sum\_\{i\\in A\}q\_\{i\}\+Q\_\{C\}=1\.

*Same region\.*Tokenkkgains probability, so it stays active by \(M2\)\. Every clipped token drops to zero, so it stays clipped by \(M3\)\. The other active tokens are unchanged\. Henceq′∈Rq^\{\\prime\}\\in R, and the identityDclip=F\+\|C\|​τD\_\{\\mathrm\{clip\}\}=F\+\|C\|\\tauapplies to bothqqandq′q^\{\\prime\}\.

*Difference\.*The constants\|C\|​τ\|C\|\\taucancel, and so do all active terms withi≠ki\\neq k\. What remains is thekk\-th term:

Dclip​\(q′\)−Dclip​\(q\)\\displaystyle D\_\{\\mathrm\{clip\}\}\(q^\{\\prime\}\)\-D\_\{\\mathrm\{clip\}\}\(q\)=pk​log⁡pkqk\+QC−pk​log⁡pkqk\\displaystyle=p\_\{k\}\\log\\frac\{p\_\{k\}\}\{q\_\{k\}\+Q\_\{C\}\}\-p\_\{k\}\\log\\frac\{p\_\{k\}\}\{q\_\{k\}\}=pk​\[log⁡qk−log⁡\(qk\+QC\)\]=pk​log⁡qkqk\+QC\.\\displaystyle=p\_\{k\}\\bigl\[\\log q\_\{k\}\-\\log\(q\_\{k\}\+Q\_\{C\}\)\\bigr\]=p\_\{k\}\\log\\frac\{q\_\{k\}\}\{q\_\{k\}\+Q\_\{C\}\}\.Herepkp\_\{k\}andqkq\_\{k\}are the teacher and student probabilities of the single receiving tokenkk, not the set masses\. We havepk\>0p\_\{k\}\>0,qk\>0q\_\{k\}\>0by \(M1\), andQC\>0Q\_\{C\}\>0, so0<qk/\(qk\+QC\)<10<q\_\{k\}/\(q\_\{k\}\+Q\_\{C\}\)<1, the logarithm is negative, and

Dclip​\(q′\)<Dclip​\(q\)\.D\_\{\\mathrm\{clip\}\}\(q^\{\\prime\}\)<D\_\{\\mathrm\{clip\}\}\(q\)\.Hence every distribution inRRwithQC\>0Q\_\{C\}\>0has strictly larger loss than some distribution inRRwithQC=0Q\_\{C\}=0\.

#### Step 2: optimal allocation on the active set\.

Throughout this steppi\>0p\_\{i\}\>0for everyii, because the teacher is a softmax output, andqi\>0q\_\{i\}\>0for everyi∈Ai\\in Aby \(M1\)\. Neither is an extra assumption\. ConsequentlyFFis finite and smooth wherever it is evaluated, and every Hessian entrypi/qi2p\_\{i\}/q\_\{i\}^\{2\}below is strictly positive\.

*\(a\) The restricted problem\.*After Step 1 the competitors are the distributionsq∈Rq\\in RwithQC=0Q\_\{C\}=0\. For suchqqwe haveqj=0q\_\{j\}=0onCC,qi\>0q\_\{i\}\>0onAA,∑i∈Aqi=1\\sum\_\{i\\in A\}q\_\{i\}=1, andDclip​\(q\)=F⁡\(qA\)\+\|C\|​τD\_\{\\mathrm\{clip\}\}\(q\)=F\(q\_\{A\}\)\+\|C\|\\tau\. So we must minimizeFFover

Kτ=\{qA∈ℝA:qi\>0,∑i∈Aqi=1,pilogpiqi≤τ∀i∈A\}\.K\_\{\\tau\}=\\Bigl\\\{q\_\{A\}\\in\\mathbb\{R\}^\{A\}:\\ q\_\{i\}\>0,\\ \\ \\sum\_\{i\\in A\}q\_\{i\}=1,\\ \\ p\_\{i\}\\log\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\leq\\tau\\ \\ \\forall i\\in A\\Bigr\\\}\.Conversely, everyqA∈Kτq\_\{A\}\\in K\_\{\\tau\}, extended by zeros onCC, is a distribution inRRwithQC=0Q\_\{C\}=0\.

*\(b\) The relaxed problem\.*Drop the threshold inequalities and let

K=\{qA∈ℝA:qi\>0,∑i∈Aqi=1\}⊇Kτ\.K=\\Bigl\\\{q\_\{A\}\\in\\mathbb\{R\}^\{A\}:\\ q\_\{i\}\>0,\\ \\ \\sum\_\{i\\in A\}q\_\{i\}=1\\Bigr\\\}\\ \\supseteq\\ K\_\{\\tau\}\.We minimize the same formulaFFoverKK\. OnK∖KτK\\setminus K\_\{\\tau\}the formulaFFis no longer the clipped objective, because there some term exceedsτ\\tauand would be capped\. The relaxed problem is an auxiliary smooth problem\. It enters the argument only through the inclusionKτ⊆KK\_\{\\tau\}\\subseteq K, and nothing is claimed aboutDclipD\_\{\\mathrm\{clip\}\}outsideRR\.

*\(c\) Derivatives ofFF\.*We differentiate on the positive orthant\{qA:qi\>0\}\\\{q\_\{A\}:\\ q\_\{i\}\>0\\\}, which is open, convex, and containsKK\. Write

F⁡\(qA\)=∑i∈A\[pi​log⁡pi−pi​log⁡qi\]\.F\(q\_\{A\}\)=\\sum\_\{i\\in A\}\\bigl\[p\_\{i\}\\log p\_\{i\}\-p\_\{i\}\\log q\_\{i\}\\bigr\]\.Only theii\-th summand depends onqiq\_\{i\}, andpi​log⁡pip\_\{i\}\\log p\_\{i\}is a constant, so

∂F∂qi=∂∂qi​\[pi​log⁡pi−pi​log⁡qi\]=0−pi⋅1qi=−piqi\.\\frac\{\\partial F\}\{\\partial q\_\{i\}\}=\\frac\{\\partial\}\{\\partial q\_\{i\}\}\\bigl\[p\_\{i\}\\log p\_\{i\}\-p\_\{i\}\\log q\_\{i\}\\bigr\]=0\-p\_\{i\}\\cdot\\frac\{1\}\{q\_\{i\}\}=\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\.For the second derivatives, differentiate−pi/qi=−piqi−1\-p\_\{i\}/q\_\{i\}=\-p\_\{i\}\\,q\_\{i\}^\{\-1\}with respect toqmq\_\{m\}\. Ifm=im=i,

∂2F∂qi2=∂∂qi\[−piqi−1\]=−pi⋅\(−1\)qi−2=piqi2\.\\frac\{\\partial^\{2\}F\}\{\\partial q\_\{i\}^\{2\}\}=\\frac\{\\partial\}\{\\partial q\_\{i\}\}\\bigl\[\-p\_\{i\}\\,q\_\{i\}^\{\-1\}\\bigr\]=\-p\_\{i\}\\cdot\(\-1\)\\,q\_\{i\}^\{\-2\}=\\frac\{p\_\{i\}\}\{q\_\{i\}^\{2\}\}\.Ifm≠im\\neq i, the expression−pi/qi\-p\_\{i\}/q\_\{i\}does not depend onqmq\_\{m\}, so∂2F/∂qi​∂qm=0\\partial^\{2\}F/\\partial q\_\{i\}\\,\\partial q\_\{m\}=0\. The Hessian is therefore diagonal,∇2F​\(qA\)=diag​\(pi/qi2\)i∈A\\nabla^\{2\}F\(q\_\{A\}\)=\\mathrm\{diag\}\\bigl\(p\_\{i\}/q\_\{i\}^\{2\}\\bigr\)\_\{i\\in A\}, and for every vectorv∈ℝAv\\in\\mathbb\{R\}^\{A\},

v⊤​∇2F​\(qA\)​v=∑i∈A∑m∈Avi​∂2F∂qi​∂qm​vm=∑i∈Apiqi2​vi2\.v^\{\\top\}\\nabla^\{2\}F\(q\_\{A\}\)\\,v=\\sum\_\{i\\in A\}\\sum\_\{m\\in A\}v\_\{i\}\\,\\frac\{\\partial^\{2\}F\}\{\\partial q\_\{i\}\\,\\partial q\_\{m\}\}\\,v\_\{m\}=\\sum\_\{i\\in A\}\\frac\{p\_\{i\}\}\{q\_\{i\}^\{2\}\}\\,v\_\{i\}^\{2\}\.Every coefficientpi/qi2p\_\{i\}/q\_\{i\}^\{2\}is strictly positive, becausepi\>0p\_\{i\}\>0andqi\>0q\_\{i\}\>0\. Ifv≠0v\\neq 0, somevi≠0v\_\{i\}\\neq 0, so the sum is strictly positive: the Hessian is positive definite at every point of the orthant\.

*\(d\) Strict convexity\.*A functionffon a convex set is*strictly convex*iff⁡\(θ​x\+\(1−θ\)​y\)<θ​f​\(x\)\+\(1−θ\)​f​\(y\)f\(\\theta x\+\(1\-\\theta\)y\)<\\theta f\(x\)\+\(1\-\\theta\)f\(y\)for allx≠yx\\neq yin the set and allθ∈\(0,1\)\\theta\\in\(0,1\)\. We use two standard facts\. \(F1\) A twice differentiable function on an open convex set whose Hessian is positive definite at every point is strictly convex\. \(F2\) A differentiable strictly convex function lies strictly above each of its tangent planes:f\(y\)\>f\(x\)\+∇f\(x\)⊤\(y−x\)f\(y\)\>f\(x\)\+\\nabla f\(x\)^\{\\top\}\(y\-x\)for allx≠yx\\neq y\. By \(c\) and \(F1\),FFis strictly convex on the orthant, independently of normalization, and hence on its convex subsetKK\.

*\(e\) Lagrange point\.*Introduce a multiplierλ\\lambdafor the normalization constraint:

ℒ⁡\(qA,λ\)=F⁡\(qA\)\+λ⁡\(∑m∈Aqm−1\)\.\\mathcal\{L\}\(q\_\{A\},\\lambda\)=F\(q\_\{A\}\)\+\\lambda\\Bigl\(\\sum\_\{m\\in A\}q\_\{m\}\-1\\Bigr\)\.Using \(c\) and∂∂qi​\(∑m∈Aqm−1\)=1\\frac\{\\partial\}\{\\partial q\_\{i\}\}\\bigl\(\\sum\_\{m\\in A\}q\_\{m\}\-1\\bigr\)=1,

∂ℒ∂qi=−piqi\+λ=0⟹λ=piqi⟹qi=piλfor every​i∈A\.\\frac\{\\partial\\mathcal\{L\}\}\{\\partial q\_\{i\}\}=\-\\frac\{p\_\{i\}\}\{q\_\{i\}\}\+\\lambda=0\\quad\\Longrightarrow\\quad\\lambda=\\frac\{p\_\{i\}\}\{q\_\{i\}\}\\quad\\Longrightarrow\\quad q\_\{i\}=\\frac\{p\_\{i\}\}\{\\lambda\}\\qquad\\text\{for every \}i\\in A\.So the ratiopi/qip\_\{i\}/q\_\{i\}is the same for every active token\. Substituting into the normalization constraint∂ℒ/∂λ=∑i∈Aqi−1=0\\partial\\mathcal\{L\}/\\partial\\lambda=\\sum\_\{i\\in A\}q\_\{i\}\-1=0,

1=∑i∈Apiλ=1λ​∑i∈Api=PAλ⟹λ=PA⟹qi⋆=piPA\(i∈A\)\.1=\\sum\_\{i\\in A\}\\frac\{p\_\{i\}\}\{\\lambda\}=\\frac\{1\}\{\\lambda\}\\sum\_\{i\\in A\}p\_\{i\}=\\frac\{P\_\{A\}\}\{\\lambda\}\\quad\\Longrightarrow\\quad\\lambda=P\_\{A\}\\quad\\Longrightarrow\\quad q^\{\\star\}\_\{i\}=\\frac\{p\_\{i\}\}\{P\_\{A\}\}\\qquad\(i\\in A\)\.This is the only stationary point, since the equations forceqi=pi/λq\_\{i\}=p\_\{i\}/\\lambdaand thenλ=PA\\lambda=P\_\{A\}\. It lies inKK:qi⋆\>0q^\{\\star\}\_\{i\}\>0and∑i∈Aqi⋆=PA/PA=1\\sum\_\{i\\in A\}q^\{\\star\}\_\{i\}=P\_\{A\}/P\_\{A\}=1\.

*\(f\) Global minimality in the relaxed problem\.*The gradient at the Lagrange point is

∂F∂qi\|qA⋆=−piqi⋆=−pipi/PA=−PAfor every​i∈A,that is,∇F​\(qA⋆\)=−PA​𝟏\.\\frac\{\\partial F\}\{\\partial q\_\{i\}\}\\Big\|\_\{q^\{\\star\}\_\{A\}\}=\-\\frac\{p\_\{i\}\}\{q^\{\\star\}\_\{i\}\}=\-\\frac\{p\_\{i\}\}\{p\_\{i\}/P\_\{A\}\}=\-P\_\{A\}\\quad\\text\{for every \}i\\in A,\\qquad\\text\{that is,\}\\qquad\\nabla F\(q^\{\\star\}\_\{A\}\)=\-P\_\{A\}\\mathbf\{1\}\.LetqA∈Kq\_\{A\}\\in KwithqA≠qA⋆q\_\{A\}\\neq q^\{\\star\}\_\{A\}\. By \(F2\) withx=qA⋆x=q^\{\\star\}\_\{A\}andy=qAy=q\_\{A\},

F\(qA\)\>F\(qA⋆\)\+∇F\(qA⋆\)⊤\(qA−qA⋆\),F\(q\_\{A\}\)\>F\(q^\{\\star\}\_\{A\}\)\+\\nabla F\(q^\{\\star\}\_\{A\}\)^\{\\top\}\(q\_\{A\}\-q^\{\\star\}\_\{A\}\),and the last term vanishes because both allocations sum to one:

∇F\(qA⋆\)⊤\(qA−qA⋆\)=∑i∈A\(−PA\)\(qi−qi⋆\)=−PA\(∑i∈Aqi−∑i∈Aqi⋆\)=−PA\(1−1\)=0\.\\nabla F\(q^\{\\star\}\_\{A\}\)^\{\\top\}\(q\_\{A\}\-q^\{\\star\}\_\{A\}\)=\\sum\_\{i\\in A\}\(\-P\_\{A\}\)\(q\_\{i\}\-q^\{\\star\}\_\{i\}\)=\-P\_\{A\}\\Bigl\(\\sum\_\{i\\in A\}q\_\{i\}\-\\sum\_\{i\\in A\}q^\{\\star\}\_\{i\}\\Bigr\)=\-P\_\{A\}\(1\-1\)=0\.HenceF⁡\(qA\)\>F⁡\(qA⋆\)F\(q\_\{A\}\)\>F\(q^\{\\star\}\_\{A\}\)for everyqA∈Kq\_\{A\}\\in Kother thanqA⋆q^\{\\star\}\_\{A\}: the Lagrange point is the unique global minimizer of the relaxed problem\.

*\(g\) Membership of the candidate\.*Since0<PA<10<P\_\{A\}<1, we haveqi⋆=pi/PA\>piq^\{\\star\}\_\{i\}=p\_\{i\}/P\_\{A\}\>p\_\{i\}for everyi∈Ai\\in A, and

pi​log⁡piqi⋆=pi​log⁡pipi/PA=pi​log⁡PA<0<τ,p\_\{i\}\\log\\frac\{p\_\{i\}\}\{q^\{\\star\}\_\{i\}\}=p\_\{i\}\\log\\frac\{p\_\{i\}\}\{p\_\{i\}/P\_\{A\}\}=p\_\{i\}\\log P\_\{A\}<0<\\tau,in agreement with \(M4\)\. Every active term lies strictly below the threshold, for everyτ\>0\\tau\>0, so the candidate cannot cross into the clipped set\. ThusqA⋆∈Kτq^\{\\star\}\_\{A\}\\in K\_\{\\tau\}, and withqj⋆=0q^\{\\star\}\_\{j\}=0onCCwe getq⋆∈Rq^\{\\star\}\\in R\.

*\(h\) Back to the restricted problem\.*LetqA∈Kτq\_\{A\}\\in K\_\{\\tau\}withqA≠qA⋆q\_\{A\}\\neq q^\{\\star\}\_\{A\}\. SinceKτ⊆KK\_\{\\tau\}\\subseteq K, part \(f\) givesF⁡\(qA\)\>F⁡\(qA⋆\)F\(q\_\{A\}\)\>F\(q^\{\\star\}\_\{A\}\), andqA⋆∈Kτq^\{\\star\}\_\{A\}\\in K\_\{\\tau\}by \(g\)\. SoqA⋆q^\{\\star\}\_\{A\}is also the unique global minimizer of the restricted problem: a minimizer over the larger set that lies in the smaller set is the minimizer over the smaller set\. Its value is

F⁡\(qA⋆\)=∑i∈Api​log​pipi/PA=∑i∈Api​log​PA=\(∑i∈Api\)​log​PA=PA​log​PA\.F\(q^\{\\star\}\_\{A\}\)=\\sum\_\{i\\in A\}p\_\{i\}\\log\\frac\{p\_\{i\}\}\{p\_\{i\}/P\_\{A\}\}=\\sum\_\{i\\in A\}p\_\{i\}\\log P\_\{A\}=\\Bigl\(\\sum\_\{i\\in A\}p\_\{i\}\\Bigr\)\\log P\_\{A\}=P\_\{A\}\\log P\_\{A\}\.

#### Conclusion\.

Letq∈Rq\\in Rwithq≠q⋆q\\neq q^\{\\star\}\.*CaseQC\>0Q\_\{C\}\>0\.*Step 1 givesq′∈Rq^\{\\prime\}\\in RwithQC=0Q\_\{C\}=0andDclip​\(q\)\>Dclip​\(q′\)D\_\{\\mathrm\{clip\}\}\(q\)\>D\_\{\\mathrm\{clip\}\}\(q^\{\\prime\}\), and Step 2 givesDclip​\(q′\)=F⁡\(qA′\)\+\|C\|​τ≥F⁡\(qA⋆\)\+\|C\|τ=Dclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q^\{\\prime\}\)=F\(q^\{\\prime\}\_\{A\}\)\+\|C\|\\tau\\geq F\(q^\{\\star\}\_\{A\}\)\+\|C\|\\tau=D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)\. SoDclip​\(q\)\>Dclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q\)\>D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)\.*CaseQC=0Q\_\{C\}=0\.*Thenqj=0=qj⋆q\_\{j\}=0=q^\{\\star\}\_\{j\}onCC, soq≠q⋆q\\neq q^\{\\star\}meansqA≠qA⋆q\_\{A\}\\neq q^\{\\star\}\_\{A\}, and Step 2 givesDclip​\(q\)=F⁡\(qA\)\+\|C\|​τ\>F⁡\(qA⋆\)\+\|C\|τ=Dclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q\)=F\(q\_\{A\}\)\+\|C\|\\tau\>F\(q^\{\\star\}\_\{A\}\)\+\|C\|\\tau=D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)\.

In both casesDclip​\(q\)\>Dclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q\)\>D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)\. Henceq⋆q^\{\\star\}is the unique minimizer ofDclipD\_\{\\mathrm\{clip\}\}overRR, withDclip​\(q⋆\)=PA​log⁡PA\+\|C\|​τD\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)=P\_\{A\}\\log P\_\{A\}\+\|C\|\\tau\. Existence is not assumed: the minimizer is exhibited\. ∎

#### FixedQCQ\_\{C\}\(Equation[8](https://arxiv.org/html/2609.38995#S4.E8)\)\.

Letccbe a value ofQCQ\_\{C\}attained inRR\. Forq∈Rq\\in RwithQC=cQ\_\{C\}=c, the loss isF⁡\(qA\)\+\|C\|​τF\(q\_\{A\}\)\+\|C\|\\tauwithqi\>0q\_\{i\}\>0and∑i∈Aqi=1−c\\sum\_\{i\\in A\}q\_\{i\}=1\-c, and it does not depend on howccis split among the clipped tokens as long as it does not make any one of them active by going over thed\>τd\>\\taurestriction\. Repeat Step 2 with the constraint∑i∈Aqi=1−c\\sum\_\{i\\in A\}q\_\{i\}=1\-c\. Stationarity again givesqi=pi/λq\_\{i\}=p\_\{i\}/\\lambda, and now

1−c=∑i∈Apiλ=PAλ⟹λ=PA1−c⟹qi⋆​\(c\)=1−cPA​pi\.1\-c=\\sum\_\{i\\in A\}\\frac\{p\_\{i\}\}\{\\lambda\}=\\frac\{P\_\{A\}\}\{\\lambda\}\\quad\\Longrightarrow\\quad\\lambda=\\frac\{P\_\{A\}\}\{1\-c\}\\quad\\Longrightarrow\\quad q^\{\\star\}\_\{i\}\(c\)=\\frac\{1\-c\}\{P\_\{A\}\}\\,p\_\{i\}\.The gradient there is−pi/qi⋆\(c\)=−PA/\(1−c\)\-p\_\{i\}/q^\{\\star\}\_\{i\}\(c\)=\-P\_\{A\}/\(1\-c\)in every coordinate, so for any other positive allocation with the same sum,∇F\(qA⋆\(c\)\)⊤\(qA−qA⋆\(c\)\)=−PA1−c\(\(1−c\)−\(1−c\)\)=0\\nabla F\(q^\{\\star\}\_\{A\}\(c\)\)^\{\\top\}\(q\_\{A\}\-q^\{\\star\}\_\{A\}\(c\)\)=\-\\frac\{P\_\{A\}\}\{1\-c\}\\bigl\(\(1\-c\)\-\(1\-c\)\\bigr\)=0, and \(F2\) givesF⁡\(qA\)\>F⁡\(qA⋆​\(c\)\)F\(q\_\{A\}\)\>F\(q^\{\\star\}\_\{A\}\(c\)\): this is the unique global minimizer of the relaxed problem at levelcc\.

*Membership\.*Every clipped token has a term aboveτ\>0\\tau\>0, solog⁡\(pj/qj\)\>0\\log\(p\_\{j\}/q\_\{j\}\)\>0andqj<pjq\_\{j\}<p\_\{j\}\. Summing overCCgivesc<∑j∈Cpj=1−PAc<\\sum\_\{j\\in C\}p\_\{j\}=1\-P\_\{A\}, hence\(1−c\)/PA\>1\(1\-c\)/P\_\{A\}\>1andqi⋆​\(c\)\>piq^\{\\star\}\_\{i\}\(c\)\>p\_\{i\}\. By \(M4\) every active term is belowτ\\tau, so the relaxed minimizer satisfies the threshold inequalities and is also the minimizer of the restricted problem at levelcc\. Its loss is

Dclip⋆​\(c\)\\displaystyle D^\{\\star\}\_\{\\mathrm\{clip\}\}\(c\)=∑i∈Api​log⁡pi\(1−c\)​pi/PA\+\|C\|​τ\\displaystyle=\\sum\_\{i\\in A\}p\_\{i\}\\log\\frac\{p\_\{i\}\}\{\(1\-c\)\\,p\_\{i\}/P\_\{A\}\}\+\|C\|\\tau=PA​log⁡PA1−c\+\|C\|τ=PA​log⁡PA−PA​log⁡\(1−c\)\+\|C\|​τ,\\displaystyle=P\_\{A\}\\log\\frac\{P\_\{A\}\}\{1\-c\}\+\|C\|\\tau=P\_\{A\}\\log P\_\{A\}\-P\_\{A\}\\log\(1\-c\)\+\|C\|\\tau,and

d​Dclip⋆d​c=−PA⋅−11−c=PA1−c\>0\.\\frac\{dD^\{\\star\}\_\{\\mathrm\{clip\}\}\}\{dc\}=\-P\_\{A\}\\cdot\\frac\{\-1\}\{1\-c\}=\\frac\{P\_\{A\}\}\{1\-c\}\>0\.The level\-ccminimizer is unique in its active coordinates only\.

#### Remarks\.

*\(i\)FFis not a KL divergence\.*Its argumentspAp\_\{A\}andqAq\_\{A\}are sub\-probability vectors, andFFcan be negative:F⁡\(qA⋆\)=PA​log⁡PA<0F\(q^\{\\star\}\_\{A\}\)=P\_\{A\}\\log P\_\{A\}<0\. The proof uses only the strict convexity ofFF\.

*\(ii\) Scope\.*The statement concerns one fixed clipping region; no comparison across clipping patterns is claimed\. A softmax over finite logits assigns every token positive probability, so such a student never equalsq⋆q^\{\\star\}; by the fixed\-QCQ\_\{C\}statement the optimal loss at levelccdecreases toDclip​\(q⋆\)D\_\{\\mathrm\{clip\}\}\(q^\{\\star\}\)asc↓0c\\downarrow 0, soq⋆q^\{\\star\}is approached but not attained\.

## Appendix CTraining and Evaluation Configuration, Capture

#### Training configuration

In all eight runs, the student and the teacher are initialized from the same Qwen3\-4B checkpoint\. Each clipped/unclipped pair receives the same per\-update problems\. Table[4](https://arxiv.org/html/2609.38995#A3.T4)lists the detailed training configuration used in our experiment\.

Table 4:Training Configuration\.
#### Message templates\.

We give the user messages and the chat template used to serialize the student and teacher prompts\. The student’s user message is identical across all eight runs\. Its serialized prompt differs only by the empty thinking block described below\. Before chat serialization, the student’s single user message is:

> Problem: \{problem\} Please reason step by step, and put your final answer within \\boxed\{\}\.

The teacher’s single user message is:

> Problem: \{problem\} Here is a reference solution to this problem: === Reference Solution Begin === \{answer\} === Reference Solution End === After reading the reference solution above, make sure you truly understand the reasoning behind each step \- do not copy or paraphrase it\. Now, using your own words and independent reasoning, derive the same final answer to the problem above\. Please reason step by step, and put your final answer within \\boxed\{\}\.

The placeholder \{answer\} is filled with the full reference solution from the dataset, not the final answer alone\. Qwen serializes the one\-message prompt as

> <\|im\_start\|\>user \{content\}<\|im\_end\|\> <\|im\_start\|\>assistant

with no system or tool message\. When thinking is disabled, the serialization additionally appends an empty<think\>block followed by</think\>and a blank line\. The student’s sampled response token IDs are appended to the teacher prompt before teacher scoring\.

#### Capture setting\.

At every response position of all 200 updates of all eight runs, the trainer stores the token the student emitted, the IDs of the teacher’s top\-128 candidate tokens, the teacher’s log probabilities on those tokens, and the student’s log probabilities on the same tokens, gathered from its full\-vocabulary log\-softmax in the forward pass that computes the loss\.

## Appendix DComplete AIME2024 and AIME2025 Evaluation

#### Evaluation configuration\.

Checkpoints 25, 50, 75, 100, 125, 150, 175, and 200 are evaluated on AIME 2024 and AIME 2025\. Each benchmark contains 30 problems\.

Each saved checkpoint is evaluated with thinking enabled, drawing 12 samples per problem at temperature\.6\.6, top\-p=\.95p=\.95, top\-k=20k=20, and minimum\-p=0p=0, with a 38,912\-token output budget\. Scoring uses only the text after the last closing thinking tag: Math\-Verify\([Kydlíček,](https://arxiv.org/html/2609.38995#bib.bib10)\)compares it with the reference answer, and if that comparison fails, the last\\boxed\{\}expression is compared with the reference answer as a string or a number\. We evaluate the same base\-model checkpoint eight times for both benchmarks to measure variation from response sampling\.

Figure 5:AIME 2024 across training, in the layout of Figure[3](https://arxiv.org/html/2609.38995#S6.F3)\. \(a\) Avg@12, the mean correct ratio over 12 samples for each of 30 problems at the checkpoint saved after each 25 updates\. \(b\) The ratio of the 360 responses per checkpoint that end in a terminal periodic loop\. Solid curves with circles are the clipped runs, dashed curves with squares their unclipped twins\. Each trained curve is one realized trajectory\. The black arrow on each vertical axis marks the base model, averaged over eight evaluation\-sampling seeds\. The values appear in Table[6](https://arxiv.org/html/2609.38995#A4.T6)\.Table 5:Complete AIME 2025 evaluation at every checkpoint \(columns: training update\)\. \(a\) Avg@12, the mean correct ratio over 12 samples for each of 30 problems\. \(b\) Terminal loop rate, the ratio of the 360 responses that end in a terminal periodic loop\. \(c\) Pass@12, the ratio of the 30 problems solved by at least one of the 12 samples\. Each trained row is one realized trajectory\. The base row gives the mean and, in parentheses, the standard deviation over eight evaluation\-sampling seeds of the fixed starting checkpoint\.Table 6:Complete AIME 2024 evaluation at every checkpoint \(columns: training update\)\. \(a\) Avg@12, the mean correct ratio over 12 samples for each of 30 problems\. \(b\) Terminal loop rate, the ratio of the 360 responses that end in a terminal periodic loop\. \(c\) Pass@12, the ratio of the 30 problems solved by at least one of the 12 samples\. Each trained row is one realized trajectory\. The base row gives the mean and, in parentheses, the standard deviation over eight evaluation\-sampling seeds of the fixed starting checkpoint\.Figure[5](https://arxiv.org/html/2609.38995#A4.F5)visualize AIME 2024 evaluation result: \(a\) avg@12 and, \(b\) terminal loop rate\.

Table[6](https://arxiv.org/html/2609.38995#A4.T6)and[5](https://arxiv.org/html/2609.38995#A4.T5)shows the numerical value of per checkpoint evaluated on AIME 2024 and AIME 2025, respectively: \(a\) avg@12, \(b\) terminal loop rate and \(c\) pass@12\.

Similar Articles

On-Policy Distillation (5 minute read)

TLDR AI

This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

Hugging Face Daily Papers

The paper shows that Direct-On-Policy Distillation's token-level log-ratio reward can remain unchanged even when the teacher checkpoints' divergences vanish, motivating S²D-OPD, a method that masks supervision at low-divergence states using teacher-reference JSD. S²D-OPD improves held-out accuracy on AIME and HMMT benchmarks in 7 of 8 teacher-student settings across models from 1.7B to 8B parameters without extra forward passes.

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Hugging Face Daily Papers

This paper introduces β-OPSD, a generalization of on-policy self-distillation that frames it as a policy optimization family with a controllable KL penalty. The method uses distillation to approximate expensive policy optimization, improving stability and reasoning performance on mathematical benchmarks.

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

Hugging Face Daily Papers

This paper analyzes on-policy self-distillation (OPSD) for LLM reasoning and shows that unverified student scaffolds create an imitation gap that worsens with model scale, motivating OASIS, a method that supervises mostly verified by-label on-policy trajectories using model-generated attempts as teacher context, improving Qwen3 1.7B/4B/8B by 3.2-3.8 points on AIME/HMMT benchmarks.