Negative Self-Distillation: Learning to Reason by Avoiding Flaws

arXiv cs.CL Papers

Summary

This paper introduces Negative Self-Distillation (NSD), a framework for improving large language model reasoning by diverging from flawed reasoning instead of imitating privileged solutions, showing consistent gains over existing methods on mathematical benchmarks.

arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Original Article
View Cached Full Text

Cached at: 09/11/26, 08:33 AM

# Learning toReason by Avoiding Flaws
Source: [https://arxiv.org/html/2609.11699](https://arxiv.org/html/2609.11699)
## Negative Self\-Distillation: Learning to Reason by Avoiding Flaws

Rongcan PeiZhepei WeiAffiliation:Department of Computer Science, University of VirginiaEmail:[zhepei\.wei@virginia\.edu](mailto:)Shuyao XuAffiliation:Stanford UniversityEmail:[xinyuzhu@virginia\.edu](mailto:)Xinyu ZhuAffiliation:Department of Computer Science, University of VirginiaEmail:[wlchen@virginia\.edu](mailto:)Wei\-Lin ChenAffiliation:Department of Computer Science, University of VirginiaEmail:[yumeng5@virginia\.edu](mailto:)Yu MengEmail:[shuyao@stanford\.edu](mailto:)[GitHub](https://github.com/Prongcan/NSD)[Hugging Face](https://huggingface.co/collections/PassionPrc/nsd-negative-self-distillation)Affiliation:Department of Computer Science, University of Virginia

###### Abstract

On\-Policy Self\-Distillation \(OPSD\) has emerged as a popular paradigm for large language model \(LLM\) self\-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground\-truth solutions\. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self\-corrective behaviors required to solve challenging problems\. To address this, we introduce Negative Self\-Distillation \(NSD\), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions\. Instead of relying on ground\-truth answers or external supervision, NSD uses the model itself to generate a question\-specific negative condition \(e\.g\., acting as a “careless reasoner”\) and pushes the student’s distribution away from this self\-generated negative teacher\. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model’s foundational language capabilities\. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning\-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model’s linguistic priors\. Empirically, NSD consistently outperforms OPSD and other label\-free, self\-bootstrapping reinforcement learning \(RL\) baselines\. Across seven mathematical reasoning benchmarks \(AIME 24/25/26, HMMT, AMC, OlympiadBench, and MATH\), NSD achieves average gains of 2\.3%, 7\.5%, and 6\.0% for 1\.7B, 4B, and 8B models, respectively\. Further analyses show that NSD achieves higher training efficiency while preserving the self\-correction behaviors crucial for complex reasoning\.

![Refer to caption](https://arxiv.org/html/2609.11699v1/intro.png)Figure 1:Overview of the NSD framework\. \(Left\) We construct a negative teacher from the same base model via self\-generated negative conditioning\. \(Right\) The student model is optimized to diverge its distribution from that of the negative teacher\.## 1Introduction

Reinforcement Learning with Verifiable Rewards \(RLVR\)\([Shao et al\., 2024](https://arxiv.org/html/2609.11699#bib.bib1);[Yu et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib31);[Lambert et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib32)\)has emerged as an effective paradigm for enhancing the reasoning capabilities of large language models \(LLMs\)\. However, RLVR is often bottlenecked by computational inefficiency and training signal sparsity\. These challenges arise because \(1\) sampling multiple rollouts per query is expensive, and rollouts within a group frequently receive identical rewards on exceptionally easy or difficult problems, leading to advantage collapse and vanishing gradients\([Liao et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib8);[Xu et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib33);[Zhang et al\., 2025b](https://arxiv.org/html/2609.11699#bib.bib34)\); and \(2\) outcome\-based rewards are applied uniformly across the entire generated sequence, which obscures fine\-grained, token\-level credit assignment\. To mitigate these limitations, On\-Policy Distillation \(OPD\)\([Agarwal et al\., 2024](https://arxiv.org/html/2609.11699#bib.bib9);[Lu and Lab, 2025](https://arxiv.org/html/2609.11699#bib.bib19);[Song and Zheng, 2026](https://arxiv.org/html/2609.11699#bib.bib36)\)utilizes a stronger, external teacher model to provide dense token\-level supervision over the student model’s self\-sampled reasoning trajectories\. While this approach successfully yields richer feedback, it introduces a practical constraint: obtaining a strictly superior external teacher that is both sufficiently capable of providing accurate dense supervision and compatible with the student’s tokenizer is often impractical\.

To circumvent the reliance on external teacher models, On\-Policy Self\-Distillation \(OPSD\)\([Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10);[Shenfeld et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib11);[Hübotter et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib12)\)has been proposed as a scalable alternative\. In OPSD, the model acts as its own teacher by utilizing privileged information \(e\.g\., ground\-truth answers\) to generate dense supervision signals for the student’s self\-sampled trajectories\. However, because the OPSD teacher inherently knows the ground\-truth solution, it tends to produce artificially confident and highly linear reasoning trajectories\([Kim et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib7);[Harne et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib35)\)\. Consequently, forcing the student to minimize the divergence from this teacher distribution inadvertently suppresses high\-entropy exploration, expressions of uncertainty, and self\-corrective behaviors, which are essential for complex problem\-solving\.

Motivated by the observation that imitating a synthetically confident oracle can degrade natural reasoning processes, we explore an alternative training paradigm: optimizing the model to explicitly avoid flawed reasoning patterns\. We introduceNegative Self\-Distillation \(NSD\), a fully self\-bootstrapped framework that operates without external privileged data\. Instead of utilizing a teacher conditioned on the correct answer, the model is prompted to generate a question\-specific negative condition \(e\.g\., acting as a “careless reasoner”\) to instantiate a negative teacher\. The student is then optimized to move its token distribution away from the negative teacher, encouraging it to avoid premature conclusions and other flawed reasoning patterns\. Importantly, the negative signal is generated from the model itself and does not require ground\-truth solutions or external annotations\.

A central challenge, however, is that not every token assigned high likelihood by the negatively conditioned teacher corresponds to a reasoning error\. A naive divergence or unlikelihood objective\([Welleck et al\., 2020](https://arxiv.org/html/2609.11699#bib.bib2)\)can also penalize ordinary linguistic tokens, degrading the model’s pretrained linguistic priors\. NSD therefore introduces a dynamic token\-level gating mechanism that compares the negative teacher with a benign reference model and activates the negative objective only when the negative condition increases the likelihood of the sampled token\. We further stabilize these updates with a bounded unlikelihood formulation and a KL\-based regularization term, preventing excessive updates on high\-confidence structural tokens while retaining targeted supervision on reasoning\-critical tokens\. This design yields a training signal that is both selective and computationally efficient: NSD requires only a single student rollout per sample, avoids full\-vocabulary logit alignment, and can parallelize the reference and negative\-teacher computations\. Beyond accuracy, our analysis shows that NSD preserves and strengthens reflective self\-correction behavior rather than encouraging overly confident, linear reasoning\. Our main contributions are summarized as follows:

- •We propose Negative Self\-Distillation \(NSD\), a label\-free, fully self\-bootstrapped framework that enhances reasoning capabilities by optimizing the model to diverge from self\-generated flawed trajectories, eliminating the need for ground\-truth solutions or an external teacher\.
- •We introduce a token\-level gating mechanism together with a bounded unlikelihood objective, enabling targeted divergence from flawed reasoning while preserving foundational language priors\.
- •We demonstrate that NSD consistently outperforms existing training paradigms \(i\.e\., OPSD\([Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10)\), Intuitor\([Zhao et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib15)\)and TTRL\([Zuo et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib22)\)\) across 1\.7B, 4B, and 8B model sizes on seven reasoning tasks\. Furthermore, NSD achieves superior training efficiency, mitigates overconfidence, and preserves the model’s intrinsic reflection capabilities\.

## 2NSD: Negative Self\-Distillation

![Refer to caption](https://arxiv.org/html/2609.11699v1/method.png)Figure 2:Overview of Negative Self\-Distillation\. The student model generates negative conditions from the unlabeled training data \(Left\)\. We then compare the token distributions between benign and negative contexts, isolating the tokens whose probabilities are abnormally boosted by the negative condition \(Mid\)\. Finally, the model is penalized to suppress the probabilities of these isolated tokens, while the filtered benign tokens are regularized only by KL divergence \(Right\)\.We consider a label\-free training dataset denoted as𝒟raw=\{\(xi\)\}i=1N\\mathcal\{D\}\_\{\\text\{raw\}\}=\\\{\(x\_\{i\}\)\\\}\_\{i=1\}^\{N\}, wherexix\_\{i\}represents the problem statement\. Our method consists of two core components: self negative conditioning and NSD training\. We first prompt the student modelπθ\\pi\_\{\\theta\}to generate a negative condition promptnin\_\{i\}for each problem\. By conditioning the model on this promptnin\_\{i\}, we construct a negative teacherπneg\\pi\_\{\\text\{neg\}\}\. We then penalize the student’s alignment with the teacher under a simple gating mechanism to avoid applying penalty to reasoning\-irrelevant tokens\. The overview of NSD is shown in Figure[2](https://arxiv.org/html/2609.11699#S2.F2)\.

### 2\.1Negative Condition Prompt Generation

The objective of this module is to allocate a negative instructionnin\_\{i\}to each training sample designed to induce flawed reasoning patterns, thereby augmenting the original𝒟raw\\mathcal\{D\}\_\{\\text\{raw\}\}into a full negative\-conditioned dataset𝒟=\{\(xi,ni\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},n\_\{i\}\)\\\}\_\{i=1\}^\{N\}\. The negative condition generation strategy should follow theself\-generationoreasy\-to\-getprinciple, without utilizing any gold answer\. By default, we adopt an online generation strategy: For a given training problemxx, we first sample an initial solutionyinity\_\{\\text\{init\}\}from the student modelπθ\\pi\_\{\\theta\}\. Conditioned on both the problem and this initial response, we then prompt the student model to generate an adaptive negative conditionnnbased on its existing reasoning trace \(the complete prompt is provided in Appendix[C\.3](https://arxiv.org/html/2609.11699#A3.SS3)\.\):

yinit∼πθ\(⋅∣x\),n∼πθ\(⋅∣x,yinit\)y\_\{\\text\{init\}\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\),\\quad n\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{\\text\{init\}\}\)\(1\)Our framework can naturally accommodate alternative negative condition generation strategies \(discussed in Section[4\.3](https://arxiv.org/html/2609.11699#S4.SS3)\)\. We default to generating negative conditions on the fly during training as it provides the most stable and effective supervision signal\.

### 2\.2Background and Challenges in Unlikelihood Training

Our motivation of the NSD training objective is to move the student’s logits distribution away from the negative teacher model’s flawed reasoning behaviors through dense token\-level supervision\. A natural approach to achieve this is standard unlikelihood training \([Welleck et al\. \(2020\)](https://arxiv.org/html/2609.11699#bib.bib2)\), which minimizes the following objective to suppress the probability of undesirable tokens:

ℒunlikelihood=−log⁡\(1−πθ​\(yt∣xi,y<t\)\)\\mathcal\{L\}\_\{\\text\{unlikelihood\}\}=\-\\log\(1\-\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\)
However, directly optimizing the objective to distance the student model from the negative teacher’s distribution presents two critical challenges and research questions \(RQs\):

\(1\) Indiscriminately treating every highly probable token under the negatively conditioned teacher as a flaw and applying the unlikelihood training is problematic, as ordinary grammatical tokens can appear in both normal and flawed reasoning; unlearning them could easily lead to the degradation of fundamental reasoning capabilities\.RQ1: How to identify the tokens that represent genuine reasoning flaws?

\(2\) This unbounded unlikelihood objective is catastrophic for highly confident, trivial tokens \(e\.g\., punctuation or spaces\) — it triggers loss explosions and overly strong gradient that destabilize training and destroy the model’s inherent logic\. Specifically, asπθ→1\\pi\_\{\\theta\}\\to 1, theℒunlikelihood\\mathcal\{L\}\_\{\\text\{unlikelihood\}\}approaches∞\\inftyand the gradient approaches the maximum \(as detailed in Appendix[A](https://arxiv.org/html/2609.11699#A1)\)\. As a considerable number of tokens have a relatively high probability, this unbounded penalty triggers gradient explosions, also forcing the student to unlearn fixed fundamental linguistic priors \(e\.g\., how to use punctuations\) and rapidly update the model parameters in an unstable direction\.RQ2: How to formulate a penalty to avoid gradient and loss explosions for training stability?

### 2\.3The NSD Training Objective

#### Token\-level adaptive gating\.

To address RQ1, we propose the gating mechanism to filter out grammatical tokens\. During the training phase, the student model generates reasoning trajectoriesy=\(y1,…,yT\)∼πθy=\(y\_\{1\},\\dots,y\_\{T\}\)\\sim\\pi\_\{\\theta\}\. To construct the gating signals, we instantiate two frozen teacher models based on the same initial student model:

- •Reference model \(πref\\pi\_\{\\text\{ref\}\}\):Conditioned only on the original problemxix\_\{i\}, predicting the nominal probabilityπref​\(yt∣xi,y<t\)\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\.
- •Negative teacher \(πneg\\pi\_\{\\text\{neg\}\}\):Conditioned on both the problem and the generated negative promptnin\_\{i\}, predicting the negatively biased probabilityπneg​\(yt∣xi,ni,y<t\)\\pi\_\{\\text\{neg\}\}\(y\_\{t\}\\mid x\_\{i\},n\_\{i\},y\_\{<t\}\)\. Note thatπneg\\pi\_\{\\text\{neg\}\}shares the same model weights asπref\\pi\_\{\\text\{ref\}\}, differing only by the negative context\.

We introduce a simple gating function that compares the probabilities of both models to filter out ordinary linguistic tokens and identify the tokens sensitive to the negative injection\. For a given student\-generated tokenyt∼πθy\_\{t\}\\sim\\pi\_\{\\theta\}, the gate is defined as the adjusted positive divergence between the probability of negative and reference model on this token:

Gt=max⁡\(0,πneg​\(yt∣xi,ni,y<t\)−πref​\(yt∣xi,y<t\)\)G\_\{t\}=\\max\\Big\(0,\\pi\_\{\\text\{neg\}\}\(y\_\{t\}\\mid x\_\{i\},n\_\{i\},y\_\{<t\}\)\-\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\\Big\)\(2\)The gateGt∈\[0,1\]G\_\{t\}\\in\[0,1\]acts as an automatic noise filter\. Ifπref≥πneg\\pi\_\{\\text\{ref\}\}\\geq\\pi\_\{\\text\{neg\}\}, the token is not activated by a negative condition and naturally exempt from penalization, preserving the model’s original generative distribution\. Conversely, ifπneg\>πref\\pi\_\{\\text\{neg\}\}\>\\pi\_\{\\text\{ref\}\}, it indicates that the negative prompt has boosted the token’s likelihood, marking it as a critical target for suppression\. Crucially, the penalty weight scales proportionally to this positive gap: a larger divergence directly translates to a heavier penalization\.

#### Gated unlikelihood penalty\.

To formulate a mathematically sound penalty \(i\.e\., the second challenge\) and address RQ2, after filtering structural noise via the dynamic gateGtG\_\{t\}, we introduce a Sigmoid\-bounded unlikelihood penalty: We squash the penalty using a Sigmoid function, yielding12−πθ​\(yt∣xi,y<t\)\\frac\{1\}\{2\-\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\}\. The Gated Unlikelihood \(GU\) penalty is formulated as:

ℒGU\(t\)\\displaystyle\\mathcal\{L\}\_\{\\text\{GU\}\}^\{\(t\)\}=Gt⋅σ⁡\(−log⁡\(1−πθ​\(yt∣xi,y<t\)\)\)\\displaystyle=G\_\{t\}\\cdot\\sigma\\Big\(\-\\log\\big\(1\-\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\\big\)\\Big\)=Gt⋅12−πθ​\(yt∣xi,y<t\)\\displaystyle=G\_\{t\}\\cdot\\frac\{1\}\{2\-\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\}\(3\)
This bounded formulation actively repels the student from negative flaws while safely preserving essential structural tokens\. As shown in Figure[3](https://arxiv.org/html/2609.11699#S2.F3), our sigmoid formulation allocates the strongest unlearning signals to low\-to\-mid confidence tokens, thereby avoiding gradient explosion on high\-probability tokens\. We further discuss the GU objective in detail in Section[5\.3](https://arxiv.org/html/2609.11699#S5.SS3)\.

![Refer to caption](https://arxiv.org/html/2609.11699v1/case_study_punct_distribution.png)Figure 3:Left: After applying the sigmoid function, the gated unlikelihood \(GU\) values are reduced for basic tokens \(e\.g\., punctuations\), preventing gradient explosion\. Right: TheℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}value distribution over 4,096 tokens from 100 training samples, showing that the sigmoid objective avoids penalization spikes on high\-probability tokens, redistributing theℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}weights toward tokens with low\-to\-mid probabilities in the student model\.
#### Regularization and overall objective\.

Letπθ​\(yt∣xi,y<t\)\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)denote the current student model being optimized\. Our goal is to push the student’s distribution away from the identified vulnerabilities without destroying its fundamental linguistic priors\. To further regularize the objective, we introduce a point\-wise forward KL penalty evaluated on the sampled tokenyty\_\{t\}\. Instead of computing the full\-vocabulary KL divergence, which is computationally heavy during rollouts, we apply an empirical reference\-weighted anchor:

ℒKL\(t\)=πref​\(yt∣xi,y<t\)⋅log⁡πref​\(yt∣xi,y<t\)πθ​\(yt∣xi,y<t\)\\mathcal\{L\}\_\{\\text\{KL\}\}^\{\(t\)\}=\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\\cdot\\log\\frac\{\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\}\{\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\}\(4\)This is a single\-sample importance\-weighted estimator ofDKL\(πref∥πθ\)D\_\{\\text\{KL\}\}\(\\pi\_\{\\text\{ref\}\}\\\|\\pi\_\{\\theta\}\)evaluated on the sampled tokenyty\_\{t\}\. We finally formulate the NSD loss for a single tokenyty\_\{t\}as a composite objective:

ℒNSD\(t\)=ℒGU\(t\)\+α⋅ℒKL\(t\)\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}=\\mathcal\{L\}\_\{\\text\{GU\}\}^\{\(t\)\}\+\\alpha\\cdot\\mathcal\{L\}\_\{\\text\{KL\}\}^\{\(t\)\}\(5\)whereα\\alphais a hyperparameter\. The full NSD algorithm is shown in Algorithm[1](https://arxiv.org/html/2609.11699#alg1)\.

For every component in theℒNSD\\mathcal\{L\}\_\{\\text\{NSD\}\}, we validate its necessity and effectiveness through ablation studies in Section[5](https://arxiv.org/html/2609.11699#S5)\. The overall objective is calculated by aggregating the token\-level losses across the dataset:

𝒥⁡\(θ\)=𝔼\(x,n\)∼𝒟,y∼πθ​\[∑t=1\|y\|ℒNSD\(t\)\]\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{\(x,n\)\\sim\\mathcal\{D\},y\\sim\\pi\_\{\\theta\}\}\\left\[\\sum\_\{t=1\}^\{\|y\|\}\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}\\right\]\(6\)
Algorithm 1Negative Self\-Distillation \(NSD\) Training1:Unlabeled dataset

𝒟raw=\{xi\}i=1N\\mathcal\{D\}\_\{\\text\{raw\}\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}, Initial model

πθ0\\pi\_\{\\theta\_\{0\}\}, KL weight

α\\alpha
2:Optimized student model

πθ\\pi\_\{\\theta\}
3:Initialize student

πθ\\pi\_\{\\theta\}, and frozen teachers

πref,πneg←πθ0\\pi\_\{\\text\{ref\}\},\\pi\_\{\\text\{neg\}\}\\leftarrow\\pi\_\{\\theta\_\{0\}\}
4:for

xi∈𝒟rawx\_\{i\}\\in\\mathcal\{D\}\_\{\\text\{raw\}\}do

5:Sample reasoning trajectory

y=\(y1,…,yT\)∼πθ\(⋅∣xi\)y=\(y\_\{1\},\\ldots,y\_\{T\}\)\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{i\}\)and the negative prompt

ni∼πθ\(⋅∣xi,y\)n\_\{i\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{i\},y\)
6:Initialize loss

𝒥i←0\\mathcal\{J\}\_\{i\}\\leftarrow 0
7:for

t=1,…,Tt=1,\\ldots,Tdo

8:

pref←πref​\(yt∣xi,y<t\)p\_\{\\text\{ref\}\}\\leftarrow\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)
9:

pneg←πneg​\(yt∣xi,ni,y<t\)p\_\{\\text\{neg\}\}\\leftarrow\\pi\_\{\\text\{neg\}\}\(y\_\{t\}\\mid x\_\{i\},n\_\{i\},y\_\{<t\}\)
10:

pθ←πθ​\(yt∣xi,y<t\)p\_\{\\theta\}\\leftarrow\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)
11:

Gt←max⁡\(0,pneg−pref\)G\_\{t\}\\leftarrow\\max\(0,p\_\{\\text\{neg\}\}\-p\_\{\\text\{ref\}\}\)
12:

ℒNSD\(t\)←Gt2−pθ\+α​pref​log⁡prefpθ\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}\\leftarrow\\frac\{G\_\{t\}\}\{2\-p\_\{\\theta\}\}\+\\alpha p\_\{\\text\{ref\}\}\\log\\frac\{p\_\{\\text\{ref\}\}\}\{p\_\{\\theta\}\}
13:

𝒥i←𝒥i\+ℒNSD\(t\)\\mathcal\{J\}\_\{i\}\\leftarrow\\mathcal\{J\}\_\{i\}\+\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}
14:endfor

15:Update

θ\\thetausing gradient

∇θ𝒥i\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{i\}
16:endfor

17:return

πθ\\pi\_\{\\theta\}

## 3Experimental Setup

#### Training setup\.

We use the MATH\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.11699#bib.bib6)\)dataset as training dataset \(for NSD, Intuitor and TTRL training, we discard the gold labels\)\. We conduct training on the following models: Qwen3\-1\.7B, Qwen3\-4B, and Qwen3\-8B\([Team, 2025](https://arxiv.org/html/2609.11699#bib.bib4)\)\. All models are trained for a total of 2 epochs, which is enough to plateau in all baselines\. We setα=0\.01\\alpha=0\.01, top\-kk= 32, batch size = 32\. For NSD, we set the max generation length to 4096\.

#### Evaluation\.

We evaluate the math reasoning ability of all models on the following benchmarks: AIME 2024, AIME 2025, AIME 2026, HMMT 2025\([Dekoninck et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib17)\), MATH\-500, AMC 2023 and OlympiadBench\([He et al\., 2024](https://arxiv.org/html/2609.11699#bib.bib24)\)\. For OlympiadBench, we exclude the proof problems\. By default, we set hyperparameters according to the recommended setting in Qwen3 report\([Team, 2025](https://arxiv.org/html/2609.11699#bib.bib4)\): temperature = 0\.6; top\-pp= 0\.95; top\-kk= 20\. The output length is set to 32K\.

#### Baselines\.

We compare with the following methods representing three different training paradigms:OPSD\([Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10)\): A standard distillation framework that minimizes the full\-vocabulary KL divergence between the student and a teacher conditioned on the gold solution\.Intuitor\([Zhao et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib15)\): A representative RLIF \(Reinforcement Learning from Internal Feedback\) implementation, which is a variant of GRPO and utilizes average confidence \(self\-certainty\) as the intrinsic reward\.TTRL\([Zuo et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib22)\): A variant of GRPO that utilizes the majority\-voting consensus as pseudo\-gold labels\. While vanilla TTRL typically optimizes directly on the test set, we apply it to the training dataset to ensure a fair comparison with other baseline methods\.

A conceptual comparison of NSD with existing related methods, along with their implementation details and prompt templates, is provided in Appendix[B](https://arxiv.org/html/2609.11699#A2)and[C](https://arxiv.org/html/2609.11699#A3)\.

## 4Evaluation Results

In this section, we first present the main experimental results across multiple mathematical reasoning benchmarks \(§[4\.1](https://arxiv.org/html/2609.11699#S4.SS1)\)\. Then we empirically demonstrate NSD achieves better training efficiency and promotes reflection abilities compared to other baselines \(§[4\.2](https://arxiv.org/html/2609.11699#S4.SS2), §[4\.4](https://arxiv.org/html/2609.11699#S4.SS4)\)\. Finally, we show that employing simpler negative conditioning strategies in NSD can also yield comparable effectiveness \(§[4\.3](https://arxiv.org/html/2609.11699#S4.SS3)\)\.

### 4\.1Main Results

Table 1:Main evaluation results on mathematical reasoning benchmarks\. We report theAvg@8\(%\) performance under non\-thinking mode \(we report the performance under thinking mode in Appendix[D\.2](https://arxiv.org/html/2609.11699#A4.SS2)\)\.Δ\\DeltaAvg is the average absolute improvement over the same\-size baseline across all 7 benchmarks\. We report the best checkpoint on the validation set within 2 training epochs\.Boldmarks the best result in each model\-size group;underlinemarks the second best\.†\\daggerdenotes methods that require ground\-truth labels\. The last two columns report the 95% CI and one\-sidedpp\-value \(which measures the probability of observing an improvement at least as large as the really observed one if the improvement were due to chance\) forΔ\\DeltaAvg@8, respectively\. OlympiadBench is evaluated on the 675 open\-ended math problems \(excluding proof problems\) using the official judger with symbolic comparison\.MethodAIME2024AIME2025AIME2026HMMT2025 FebAMC2023Olympiad\-BenchMATH\-500𝚫\\mathbf\{\\Delta\}Avg95% CIpp1\.7B ModelsQwen3\-1\.7B9\.610\.09\.67\.144\.137\.162\.5———OPSD†15\.014\.28\.85\.844\.137\.262\.5\+1\.1\[−0\.3,\+2\.4\]\[\-0\.3,\+2\.4\]0\.060\.06Intuitor13\.88\.38\.36\.743\.435\.460\.6−\-0\.5\[−1\.8,\+0\.8\]\[\-1\.8,\+0\.8\]0\.290\.29TTRL11\.311\.39\.68\.341\.636\.863\.4\+0\.3\[−1\.0,\+1\.5\]\[\-1\.0,\+1\.5\]0\.290\.29NSD14\.217\.910\.07\.145\.938\.762\.6\+2\.3\[\+0\.7,\+4\.0\]\[\+0\.7,\+4\.0\]0\.0010\.0014B ModelsQwen3\-4B23\.820\.417\.910\.868\.847\.871\.2———OPSD†25\.422\.515\.815\.868\.847\.671\.7\+1\.0\[−0\.3,\+2\.4\]\[\-0\.3,\+2\.4\]0\.100\.10Intuitor24\.625\.818\.313\.870\.047\.769\.8\+1\.3\[−0\.5,\+3\.1\]\[\-0\.5,\+3\.1\]0\.050\.05TTRL25\.819\.618\.311\.768\.147\.171\.7\+0\.2\[−1\.2,\+1\.7\]\[\-1\.2,\+1\.7\]0\.330\.33NSD35\.831\.329\.216\.376\.351\.073\.1\+7\.5\[\+5\.4,\+9\.5\]\[\+5\.4,\+9\.5\]<10−4<10^\{\-4\}8B ModelsQwen3\-8B28\.819\.218\.311\.767\.248\.973\.1———OPSD†30\.021\.317\.112\.166\.948\.273\.5\+0\.3\[−1\.3,\+1\.9\]\[\-1\.3,\+1\.9\]0\.330\.33Intuitor34\.620\.418\.313\.870\.949\.572\.7\+1\.9\[\+0\.2,\+3\.4\]\[\+0\.2,\+3\.4\]0\.020\.02TTRL29\.218\.317\.110\.869\.149\.173\.0−\-0\.1\[−1\.4,\+1\.3\]\[\-1\.4,\+1\.3\]0\.570\.57NSD39\.626\.325\.017\.975\.650\.674\.1\+6\.0\[\+4\.0,\+7\.9\]\[\+4\.0,\+7\.9\]<10−4<10^\{\-4\}

The main results are shown in Table[1](https://arxiv.org/html/2609.11699#S4.T1)\. We highlight the following key observations:

#### NSD achieves the overall best performance on the models with different sizes\.

As shown in Table[1](https://arxiv.org/html/2609.11699#S4.T1), NSD consistently achieves the highest average improvements across all model scales, yieldingΔ\\DeltaAvg gains of\+2\.3%,\+7\.5%, and\+6\.0%on the three models respectively\. While baselines like OPSD†and RL excel narrowly on AIME 2024, our 4B and 8B models achieve broader generalization across diverse math tasks, maintaining peak AIME accuracies of 35\.8% and 39\.6%\. Notably, unlike other baselines where small improvements possibly partly stem from randomness, NSD guarantees stable performance gains, supported by a significantly lowpp\-value\.

#### NSD is more promising on larger model sizes due to self\-generated negative conditions\.

An observation from Table[1](https://arxiv.org/html/2609.11699#S4.T1)is that NSD exhibits stronger performance gains on larger models compared to the smaller 1\.7B variant\. This scaling behavior is tied to our online negative condition generation mechanism: NSD uses on the model itself to generate solution\-specific negative conditions\. By optimizing against these higher\-quality conditions, larger models receive a stronger contrastive training signal, which translates into substantial improvements on challenging reasoning tasks\.

#### Why does NSD outperform other baselines?

Compared to OPSD, label\-free training of NSD without the privileged information prevents bias \(e\.g\., reinforcing reasoning shortcuts due to the gold solution\) and reflection collapse caused by overconfidence\([Kim et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib7)\)\. We provide additional analysis in Section[4\.2](https://arxiv.org/html/2609.11699#S4.SS2)that further confirms NSD better promotes reflection behaviors than other methods\. Case studies in Appendix[D\.4](https://arxiv.org/html/2609.11699#A4.SS4)also illustrate how NSD\-trained models abandon the wrong reasoning trajectory and switch to the right one\. Compared to Intuitor and TTRL which use model confidence or majority voting to generate training signals, NSD removes the reliance on the model’s self\-judgement ability, which leads to possible incorrect training signals\. For example, weaker models hardly gain improvement from Intuitor \(\-0\.5% on Qwen3\-1\.7B\) because their high confidence does not necessarily equate to high accuracy\. Furthermore, those confidence\-based bootstrapping methods also degrade the reflection ability, as shown in Section[4\.2](https://arxiv.org/html/2609.11699#S4.SS2)\.

### 4\.2NSD Inspires Reflection

NSD prevents over\-confidence and preserves exploratory reflection\.We evaluate model reflection capabilities by measuring the average frequency of reflection tokens \(e\.g\., “Wait”\) across AIME and HMMT benchmarks \(Table[2](https://arxiv.org/html/2609.11699#S4.T2)\)\. The detailed definition of reflection tokens is shown in Appendix[E](https://arxiv.org/html/2609.11699#A5)\. We observe that OPSD and Intuitor severely suppress reflective behavior \(dropping to 2\.18 and 0\.75 per response, respectively\), as training on ground\-truth or unverified positive rollouts encourages overly direct, non\-verifying reasoning trajectories\.

Table 2:The average reflection token frequency per response on Qwen3\-4B\.MethodAIME 2024AIME 2025HMMT 2025AverageBaseline6\.82\.21\.73\.6OPSD2\.62\.11\.82\.2Intuitor0\.61\.00\.70\.8NSD6\.97\.58\.17\.5Conversely, NSD substantially enhances reflection frequency \(yielding up to 7\.5 per response\)\. By penalizing flawed reasoning paths, NSD avoids over\-confidence and enables the model to autonomously re\-evaluate potential errors during complex inference, which is also demonstrated by our case study in Appendix[D\.4](https://arxiv.org/html/2609.11699#A4.SS4)\.

### 4\.3Negative Condition Variants Study

In our main experiments, we default to an online self negative condition generation strategy \(denoted as online strategy briefly\)\. While intuitively well\-motivated, this approach incurs computational overhead from online rollouts\. To explore more efficient alternatives, we investigate the impact of simpler conditioning strategies, selecting the variants based on effective LLM negative conditioning paradigms identified in\([Chatziveroglou et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib5)\)\. In this section, we discuss the following offline generation strategies while our primary evaluations in the previous sections are conducted using the online paradigm:

Strategy 1: Solution\-aware negative conditioning\. Besides inputting the question, we let the student model rollout first, then prompt it to generate a negative condition based on the question and rollout\.

Strategy 2: Question\-only negative conditioning\. Input the training sample question to the frozen initial student and prompt it to generate a possible negative condition based on it\.

Strategy 3: Noise conditioning\. Simply add irrelevant Wikipedia articles as noise \(denoted as wiki\-irr strategy; “irr” stands for irrelevant\)\.

Figure 4:Evaluation results of NSD across different conditioning strategies\.Δ\\Deltadenotes the average absolute improvement over the base model across these 4 datasets\. The dashed line represents the referenceΔ\\Deltaachieved by the default online strategy\.We evaluate the three NSD conditioning strategies on the Qwen3\-4B model\. The experimental result of different conditioning strategies is shown in Figure[4](https://arxiv.org/html/2609.11699#S4.F4)\. Notably, the question\-only negative conditioning strategy achieves a 7\.3% average improvement, comparable to 7\.8% using our default online solution\-aware approach\. Furthermore, even the most lightweight offline strategy \(wiki\-irr\) also performs competitively with our default approach, highlighting NSD’s broad scalability to diverse and efficient negative conditions\. In contrast, the offline solution\-aware strategy exhibits relatively lower performance, primarily driven by its reliance on outdated offline\-generated solutions during conditioning\.

Figure 5:Average wall\-clock time per training step with 6 or 8 A100 GPUs\. For NSD and OPSD, the student model occupies 4 GPUs and the teacher occupies 2 GPUs\. For Intuitor, the generation stage is executed across all 8 GPUs\. Notably, the wiki\-irr strategy effectively reduces the latency compared to using the default online rollout in NSD\.
### 4\.4Efficiency of NSD

NSD exhibits superior training efficiency compared to OPSD and RLIF\.We focus our detailed latency analysis on the computational overhead of the rollout phase, which is the dominant source of training time discrepancy across different algorithms\.

In the rollout stage, the student model samplesbatch size×n\\text\{batch size\}\\times nrollouts, wherenndenotes the number of samples per prompt\. While GRPO\-based baselines \(Intuitor and TTRL\) demandn=8n=8, both NSD and OPSD require onlyn=1n=1\. Subsequently, OPSD and NSD perform additional forward passes on the generated sequences: OPSD prefills each concatenated prompt\-response pair to extract top\-kklog\-probabilities, wherek=32k=32in NSD and128128in OPSD; online NSD generates an online negative condition based on the student’s solution before running two forward passes to computeπref\\pi\_\{\\text\{ref\}\}andπneg\\pi\_\{\\text\{neg\}\}\.

The time consumption is shown in Figure[5](https://arxiv.org/html/2609.11699#S4.F5)\. Overall, NSD achieves superior training efficiency through three primary factors: \(1\)Minimal Rollout Overhead:Unlike multi\-sample GRPO\-style baselines, NSD requires only a single rollout per sample, reducing rollout time by∼\\sim60%\. Furthermore, static negative condition generation strategies \(e\.g\., wiki\-irr\) save online rollout time entirely, lowering the overall latency from 68s to 54s\. \(2\)Parallelized Forward Prefilling:Although computingπref\\pi\_\{\\text\{ref\}\}andπneg\\pi\_\{\\text\{neg\}\}involves two distinct prompts, both share the same model weights and can be prefilled concurrently in parallel\. \(3\)Scalar\-Only Loss Computation:NSD requires only three scalar token probabilities, avoiding full\-vocabulary logit projections\. The result shows that NSD trains faster overall than OPSD, proving that our parallelized prefilling costs substantially less than OPSD’s Top\-kklogit alignment\.

### 4\.5More Evaluations and Analyses

We conduct several supplementary evaluations provided in Appendix[D](https://arxiv.org/html/2609.11699#A4)\. First, we evaluate our models under the Pass@8 metric in Appendix[D\.1](https://arxiv.org/html/2609.11699#A4.SS1)\. Second, we report the performance under thinking mode in Appendix[D\.2](https://arxiv.org/html/2609.11699#A4.SS2)\. NSD continues to outperform all baselines under these settings\. Third, in Appendix[D\.3](https://arxiv.org/html/2609.11699#A4.SS3), we investigate an alternative objective formulation that treats the negative of the loss as an advantage signal for policy\-gradient optimization, demonstrating that the NSD framework is scalable to policy\-gradient\-style training paradigms\.

## 5Understanding NSD Training Objective

### 5\.1Adaptive Gating Analysis

The adaptive gating function is designed to filter out trivial tokens while retaining essential ones\. RLCSD\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18)\)rigorously conceptualizes this by categorizing tokens into style and task tokens, treating the former as noise\. A detailed definition is shown in Appendix[E](https://arxiv.org/html/2609.11699#A5)\. To evaluate how effectively the NSD gating function and existing weighting methods filter out style tokens, we sample a subset of 100 training queries and analyze the logit distributions across the initial 4,096 tokens and calculate the style\-task ratio \(𝒯\\mathcal\{T\}denotes the set of task tokens and𝒮\\mathcal\{S\}denotes the set of style tokens\):

R=\[1\|𝒮\|​∑t∈𝒮wt\]/\[1\|𝒯\|​∑t∈𝒯wt\]R=\{\\left\[\\dfrac\{1\}\{\|\\mathcal\{S\}\|\}\\displaystyle\\sum\_\{t\\in\\mathcal\{S\}\}w\_\{t\}\\right\]\}/\{\\left\[\\dfrac\{1\}\{\|\\mathcal\{T\}\|\}\\displaystyle\\sum\_\{t\\in\\mathcal\{T\}\}w\_\{t\}\\right\]\}\(7\)
We compare NSD adaptive gating with initial ratio, entropy\-based OPSD weighting\([Wang et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib14)\), and vanilla OPSD loss\([Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10)\):

NSD:wt=max⁡\(0,πneg​\(yt∣x,a,y<t\)−πref​\(yt∣x,y<t\)\)w\_\{t\}=\\max\\bigl\(0,\\ \\pi\_\{\\text\{neg\}\}\(y\_\{t\}\\mid x,a,y\_\{<t\}\)\-\\pi\_\{\\text\{ref\}\}\(y\_\{t\}\\mid x,y\_\{<t\}\)\\bigr\); Entropy\-OPSD:wt=−∑vπθ\(v∣x,y<t\)logπθ\(v∣x,y<t\)w\_\{t\}=\-\\sum\_\{v\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\log\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\); OPSD:wt=∑vπθ​\(v\)​log⁡πθ​\(v\)πgold​\(v\)w\_\{t\}=\\sum\_\{v\}\\pi\_\{\\theta\}\(v\)\\log\\frac\{\\pi\_\{\\theta\}\(v\)\}\{\\pi\_\{\\text\{gold\}\}\(v\)\}\.

A lowerRRinherently signifies a better approach\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18)\); it implies the method prioritizes task tokens, suppressing gradient generation on meaningless tokens — previous work\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18);[Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10)\)shows that in OPSD, the training signal might be dominated by style tokens, causing the student to imitate styles rather than learning reasoning ability\. As shown in Table[3](https://arxiv.org/html/2609.11699#S5.T3), the NSD gating mechanism alone filters style tokens more effectively than both entropy\-based and OPSD\-loss\-based weighting methods\.

Table 3:Comparison of style\-task ratio across different methods\. A lower value indicates a better approach\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18)\)\.MethodNSD gate\(wiki\)NSD gate\(solution\-aware\)NSD gate\(question\-only\)Entropy\-OPSDOPSDstyle\-task ratio2\.6×\\times3\.4×\\times3\.5×\\times3\.9×\\times5\.4×\\times
### 5\.2KL Ablation

We investigate the necessity of the KL divergence constraint within the NSD framework\. Figure[6](https://arxiv.org/html/2609.11699#S5.F6)illustrates this via an ablation study comparing the standard NSD against a variant without the KL anchor \(NSD\-noKL\)\.

Figure 6:Training log of NSD w/ and w/o KL constraint on Qwen3\-4B\. Left: Forward KL between the reference model and student model per training step; Right: The average activated gating \(Gt¯\\overline\{G\_\{t\}\}\) per training step\.Removing the KL constraint leads to a mid\-training collapse\. As shown in Figure[6](https://arxiv.org/html/2609.11699#S5.F6)\(left\), NSD\-noKL drastically shifts the student’s distribution, inducing a cycle of learning and forgetting, evidenced by sharp oscillations in the curve\. Furthermore, Figure[6](https://arxiv.org/html/2609.11699#S5.F6)\(right\) demonstrates that the gate activation ratio in NSD\-noKL initially increases but drops precipitously around the 120th step, coinciding precisely with the KL collapse\. We also observe that the meanℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}value decreases significantly, indicating that the gating mechanism activates spuriously and loses its effectiveness—a direct result of the model drifting excessively from the reference without KL regularization\.

### 5\.3Discussion on Gated Unlikelihood

A natural inherent consequence of the gating formulation is its sensitivity to minor probability fluctuations in highly predictable tokens \(whereπneg≈πref→1\\pi\_\{\\text\{neg\}\}\\approx\\pi\_\{\\text\{ref\}\}\\to 1, usually trivial tokens like punctuations\)\. Occasionally, inherent variance may cause the negative teacher model to assign a marginally higher probability than the reference model, bypassing the filter \(e\.g\., probabilities for space tokens often fluctuate slightly around 99%\)\. Nevertheless, this artifact is controlled: the minuscule divergence yields a near\-zero gate valueGtG\_\{t\}, ensuring that the overall gating on these high\-confidence tokens remains negligible\.

To further avoid distancing from these tokens, we bound the pure unlikelihood penalty via Sigmoid function \(Equation[3](https://arxiv.org/html/2609.11699#S2.E3)\), so that the model enjoys implicit gradient attenuation\. The Sigmoid penalty serves as a structural failsafe against gating imperfections\. As shown in Figure[3](https://arxiv.org/html/2609.11699#S2.F3),ℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}value on high\-probability tokens is significantly lower than the value of pure unlikelihood objectives\. Gradient analysis is mathematically discussed in Appendix[A](https://arxiv.org/html/2609.11699#A1)\.

## 6Related Work

On policy distillation\.The original OPD\([Agarwal et al\., 2024](https://arxiv.org/html/2609.11699#bib.bib9);[Lu and Lab, 2025](https://arxiv.org/html/2609.11699#bib.bib19);[Song and Zheng, 2026](https://arxiv.org/html/2609.11699#bib.bib36)\)relies on external reward models\. Self distillation\([Zhao et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib10);[Hübotter et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib12);[Shenfeld et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib11)\)removes external teachers by using ground\-truth solutions as hints, but suffers from solution bias and overconfidence\([Kim et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib7);[Harne et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib35);[Wang et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib30)\); Recent studies have increasingly optimized the distillation method across various dimensions, mainly including weak supervision\([He et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib13);[Li et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib38)\), credit assignment\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18);[Wang et al\., 2026d](https://arxiv.org/html/2609.11699#bib.bib37);[Xu et al\., 2026c](https://arxiv.org/html/2609.11699#bib.bib39);[Wang et al\., 2026c](https://arxiv.org/html/2609.11699#bib.bib40)\), agentic scenarios\([Wu et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib23);[Lu et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib41)\), and other better learning objectives\([Yang et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib42);[Heo et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib43);[Jiang et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib44);[Shen et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib20);[Kim et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib29)\)\.

Label\-free reinforcement learning\.Existing label\-free training methods primarily rely on substituting rewards with self\-generated ones \(usually based on confidence or entropy\) within RLVR frameworks\([Zhao et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib15);[Li et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib45);[Yuan et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib25);[Prabhudesai et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib46);[Huang et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib48);[Huang et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib49)\)or OPD frameworks\([Gkountouras et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib21);[Li et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib38)\), as well as generating gold labels by the model itself\([Zhang et al\., 2025a](https://arxiv.org/html/2609.11699#bib.bib47);[Zuo et al\., 2025](https://arxiv.org/html/2609.11699#bib.bib22)\)\.

Training with negative signals\.Unlikelihood objective\([Welleck et al\., 2020](https://arxiv.org/html/2609.11699#bib.bib2);[Li et al\., 2020](https://arxiv.org/html/2609.11699#bib.bib3)\)has been proposed to train earlier small language models\. Several studies incorporate both positive and negative trajectories into distillation or RLVR frameworks\([Xu et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib26);[Hamdan and Yuret, 2025](https://arxiv.org/html/2609.11699#bib.bib27);[Yang et al\., 2024](https://arxiv.org/html/2609.11699#bib.bib28)\)\. Notably, NSR\([Zhu et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib16)\)explores RLVR training driven exclusively by negative signals, demonstrating that it can preserve high\-confidence priors while mitigating overfitting\. Furthermore, in the context of self\-distillation, recent works introduce negative signals to alleviate student overconfidence\([Shen et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib20);[Kim et al\., 2026a](https://arxiv.org/html/2609.11699#bib.bib29)\)\.

## 7Conclusion

In this work, we propose Negative Self\-Distillation \(NSD\), a label\-free training framework comprising negative conditioning and gated unlikelihood training\. Empirical results demonstrate that NSD consistently outperforms existing baselines across seven mathematical reasoning benchmarks under various model sizes\. Comprehensive analyses and ablation studies show that our adaptive gating mechanism effectively isolates genuinely flawed tokens, while the sigmoid unlikelihood objective ensures smoother reasoning gradients\. Furthermore, NSD accommodates diverse negative conditioning strategies, establishing it as a highly scalable framework\. Its training efficiency is enhanced by bypassing full\-vocabulary computations and leveraging parallelized forward passes for the negative teacher and reference model\. Importantly, NSD inherently preserves and stimulates the model’s capacity for self\-reflection, highlighting its potential as a promising post\-training method for enhancing the reasoning capabilities of LLMs\.

## Limitations

NSD relies on the student model’s inherent capacity to generate negative conditions\. Consequently, this approach may be less effective for extremely small or weak models that struggle to produce meaningful negative contrasts for optimization\. However, given the rapid capability scaling of modern foundational models, this capacity bottleneck is expected to diminish naturally in future architectures or in stronger models\.

Under the online strategy, NSD requires negative\-condition generation and two forward passes through the two same frozen models\. Nevertheless, we explored alternative conditioning strategies, including efficient generation\-free methods like the wiki\-irr strategy, which can mitigate the rollout costs while maintaining competitive performance\. Moreover, executing the two forward passes in parallel at each training step effectively minimizes overall wall\-clock latency\.

## Acknowledgments

This research is partially funded by the NVIDIA Academic Grant and Amazon Research Award\. We thank Xinyu Wang and Yu Gu for their valuable feedback and suggestions, particularly for the experimental design\.

## References

- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. R\. Garea, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=3zKtaqxLhW)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Chatziveroglouet al\.\(2025\)G\. Chatziveroglou, R\. Yun, and M\. KelleherExploring LLM reasoning through controlled prompt variations\.External Links:2504\.02111,[Link](https://arxiv.org/abs/2504.02111)Cited by:[§4\.3](https://arxiv.org/html/2609.11699#S4.SS3.p1.1)\.
- Dekonincket al\.\(2026\)J\. Dekoninck, N\. Jovanović, T\. Gehrunger, K\. Rögnvaldsson, I\. Petrov, C\. Sun, and M\. VechevBeyond benchmarks: MathArena as an evaluation platform for mathematics with LLMs\.External Links:2605\.00674,[Link](https://arxiv.org/abs/2605.00674)Cited by:[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px2.p1.1)\.
- Ghimireet al\.\(2026\)M\. Ghimire, A\. Feng, L\. You, Y\. Luo, F\. Liu, and X\. ZhuPRISM: a unified framework for post\-training LLMs without verifiable rewards\.External Links:2601\.04700,[Link](https://arxiv.org/abs/2601.04700)Cited by:[§D\.2](https://arxiv.org/html/2609.11699#A4.SS2.p1.1)\.
- Gkountouraset al\.\(2026\)J\. Gkountouras, J\. Jukić, and I\. TitovConsensus as privileged context for label\-free self\-distillation\.External Links:2607\.13643,[Link](https://arxiv.org/abs/2607.13643)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Hamdan and Yuret \(2025\)S\. Hamdan and D\. YuretHow much do LLMs learn from negative examples?\.External Links:2503\.14391,[Link](https://arxiv.org/abs/2503.14391)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Harneet al\.\(2026\)S\. Harne, C\. Karkar, Y\. Pandya, A\. Awadallah, and A\. NambiPrivileged, but biased: how pi\-conditioned teachers break self\-distillation\.External Links:2608\.04794,[Link](https://arxiv.org/abs/2608.04794)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p2.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Heet al\.\(2024\)C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. L\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang, J\. Liu, L\. Qi, Z\. Liu, and M\. SunOlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.External Links:2402\.14008,[Link](https://arxiv.org/abs/2402.14008)Cited by:[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px2.p1.1)\.
- Heet al\.\(2026\)Y\. He, S\. Kaur, A\. Bhaskar, Y\. Yang, J\. Liu, N\. Ri, L\. Fowl, A\. Panigrahi, D\. Chen, and S\. AroraSelf\-Distillation Zero: self\-revision turns binary rewards into dense supervision\.External Links:2604\.12002,[Link](https://arxiv.org/abs/2604.12002)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the MATH dataset\.External Links:2103\.03874,[Link](https://arxiv.org/abs/2103.03874)Cited by:[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px1.p1.1)\.
- Heoet al\.\(2026\)B\. Heo, J\. Hwang, S\. Yun, and D\. HanOn\-policy delta distillation\.External Links:2607\.15161,[Link](https://arxiv.org/abs/2607.15161)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Huanget al\.\(2026a\)C\. Huang, H\. Liu, T\. Zheng, R\. Dai, L\. Huang, J\. Li, Z\. Li, Z\. Wei, Y\. Meng, and J\. HuangG\-Zero: self\-play for open\-ended generation from zero data\.arXiv preprint arXiv:2605\.09959\.Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Huanget al\.\(2026b\)C\. Huang, W\. Yu, X\. Wang, H\. Zhang, Z\. Li, R\. Li, J\. Huang, H\. Mi, and D\. YuR\-Zero: self\-evolving reasoning LLM from zero data\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=96apU6YzSO)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Hübotteret al\.\(2026\)J\. Hübotter, F\. Lübeck, L\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. K\. Buening, C\. Guestrin, and A\. KrauseReinforcement learning via self\-distillation\.External Links:2601\.20802,[Link](https://arxiv.org/abs/2601.20802)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p2.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Jianget al\.\(2026\)L\. Jiang, H\. Xu, Y\. Ding, and A\. ZhangTrajectory\-refined distillation\.External Links:2606\.08432,[Link](https://arxiv.org/abs/2606.08432)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Kimet al\.\(2026a\)J\. Kim, J\. Jeon, D\. Li, and Y\. YangRebellious student: reversing teacher signals for reasoning exploration with self\-distilled RLVR\.External Links:2605\.10781,[Link](https://arxiv.org/abs/2605.10781)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Kimet al\.\(2026b\)J\. Kim, X\. Luo, M\. Kim, S\. Lee, D\. Kim, J\. Jeon, D\. Li, and Y\. YangWhy does self\-distillation \(sometimes\) degrade the reasoning capability of LLMs?\.External Links:2603\.24472,[Link](https://arxiv.org/abs/2603.24472)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.11699#S4.SS1.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Lambertet al\.\(2025\)N\. Lambert, J\. Morrison, V\. Pyatkin, S\. Huang, H\. Ivison, F\. Brahman, L\. J\. V\. Miranda, A\. Liu, N\. Dziri, S\. Lyu, Y\. Gu, S\. Malik, V\. Graf, J\. D\. Hwang, J\. Yang, R\. L\. Bras, O\. Tafjord, C\. Wilhelm, L\. Soldaini, N\. A\. Smith, Y\. Wang, P\. Dasigi, and H\. HajishirziTulu 3: pushing frontiers in open language model post\-training\.External Links:2411\.15124,[Link](https://arxiv.org/abs/2411.15124)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Liet al\.\(2020\)M\. Li, S\. Roller, I\. Kulikov, S\. Welleck, Y\. Boureau, K\. Cho, and J\. WestonDon’t say that\! making inconsistent dialogue unlikely with unlikelihood training\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 4715–4728\.External Links:[Link](https://aclanthology.org/2020.acl-main.428/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.428)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Liet al\.\(2025\)P\. Li, M\. Skripkin, A\. Zubrey, A\. Kuznetsov, and I\. OseledetsConfidence is all you need: few\-shot RL fine\-tuning of language models\.External Links:2506\.06395,[Link](https://arxiv.org/abs/2506.06395)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Liet al\.\(2026\)Y\. Li, B\. Wang, Y\. Liang, Y\. Tian, D\. Fu, and N\. VasconcelosOn\-policy self\-distillation without any supervision\.External Links:2608\.06296,[Link](https://arxiv.org/abs/2608.06296)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Liaoet al\.\(2026\)B\. Liao, H\. Dong, X\. Xu, C\. Monz, and J\. BianSelf\-hinting language models enhance reinforcement learning\.External Links:2602\.03143,[Link](https://arxiv.org/abs/2602.03143)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Lu and Lab \(2025\)K\. Lu and T\. M\. LabOn\-policy distillation\.Thinking Machines Lab: Connectionism\.Note:https://thinkingmachines\.ai/blog/on\-policy\-distillationExternal Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[§D\.3](https://arxiv.org/html/2609.11699#A4.SS3.p1.1),[§1](https://arxiv.org/html/2609.11699#S1.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Luet al\.\(2026\)Z\. Lu, Z\. Yao, Z\. Han, Z\. Wang, J\. Wu, Q\. Gu, X\. Cai, W\. Lu, J\. Xiao, Y\. Zhuang, and Y\. ShenSelf\-distilled agentic reinforcement learning\.External Links:2605\.15155,[Link](https://arxiv.org/abs/2605.15155)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Panet al\.\(2026\)L\. Pan, S\. Tao, Y\. Zhai, L\. Zhang, Z\. Liu, B\. Ding, A\. Liu, and L\. WenRLCSD: reinforcement learning with contrastive on\-policy self\-distillation\.External Links:2606\.11709,[Link](https://arxiv.org/abs/2606.11709)Cited by:[Appendix E](https://arxiv.org/html/2609.11699#A5.p1.1),[§5\.1](https://arxiv.org/html/2609.11699#S5.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.11699#S5.SS1.p5.1),[Table 3](https://arxiv.org/html/2609.11699#S5.T3),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Prabhudesaiet al\.\(2026\)M\. Prabhudesai, L\. Chen, A\. Ippoliti, K\. Fragkiadaki, H\. Liu, and D\. PathakMaximizing confidence alone improves reasoning\.External Links:[Link](https://openreview.net/forum?id=Qhg479eBmo)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Shenet al\.\(2026\)G\. Shen, X\. Cheng, C\. Zhao, L\. Huang, J\. Li, D\. Zhao, and X\. YuAnti\-self\-distillation for reasoning RL via pointwise mutual information\.External Links:2605\.11609,[Link](https://arxiv.org/abs/2605.11609)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Shenfeldet al\.\(2026\)I\. Shenfeld, M\. Damani, J\. Hübotter, and P\. AgrawalSelf\-distillation enables continual learning\.External Links:2601\.19897,[Link](https://arxiv.org/abs/2601.19897)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p2.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Song and Zheng \(2026\)M\. Song and M\. ZhengA survey of on\-policy distillation for large language models\.arXiv preprint arXiv:2604\.00626\.Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Team \(2025\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2026a\)M\. Wang, H\. Zhao, W\. Liu, L\. Yang, G\. Liu, H\. Guo, G\. Xie, G\. Meng, H\. Liu, and F\. ZhuDenser≠\\neqbetter: limits of on\-policy self\-distillation for continual post\-training\.External Links:2607\.01763,[Link](https://arxiv.org/abs/2607.01763)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Wanget al\.\(2026b\)S\. Wang, L\. Yu, C\. Gao, C\. Zheng, S\. Liu, R\. Lu, K\. Dang, X\. Chen, J\. Yang, Z\. Zhang, Y\. Liu, A\. Yang, A\. Zhao, Y\. Yue, S\. Song, B\. Yu, G\. Huang, and J\. LinBeyond the 80/20 rule: high\-entropy minority tokens drive effective reinforcement learning for LLM reasoning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=yfcpdY4gMP)Cited by:[§5\.1](https://arxiv.org/html/2609.11699#S5.SS1.p3.1)\.
- Wanget al\.\(2026c\)Y\. Wang, S\. Lu, Y\. Gu, P\. Wang, Y\. Yang, Z\. Yan, C\. Xie, J\. Wu, and H\. YangNot all disagreement is learnable: token teachability in on\-policy distillation\.External Links:2605\.26844,[Link](https://arxiv.org/abs/2605.26844)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Wanget al\.\(2026d\)Z\. Wang, Z\. Lu, Z\. Yao, J\. Wu, J\. Wu, Z\. Cai, Y\. Sun, Z\. Ye, L\. Hao, Q\. Gu, X\. Cai, Y\. Shen, and Y\. YangAgentOPSD: recursive self\-distillation for agentic reinforcement learning\.External Links:2608\.05987,[Link](https://arxiv.org/abs/2608.05987)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Wellecket al\.\(2020\)S\. Welleck, I\. Kulikov, S\. Roller, E\. Dinan, K\. Cho, and J\. WestonNeural text generation with unlikelihood training\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SJeYe0NtvH)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p4.1),[§2\.2](https://arxiv.org/html/2609.11699#S2.SS2.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Wuet al\.\(2026\)J\. Wu, S\. Yang, Z\. Lu, F\. Zhang, Y\. Shen, L\. Feng, H\. Luo, Z\. Lian, S\. Zhang, Z\. Wen, and J\. TaoSEED: self\-evolving on\-policy distillation for agentic reinforcement learning\.External Links:2607\.14777,[Link](https://arxiv.org/abs/2607.14777)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Xuet al\.\(2026a\)H\. Xu, S\. Chen, R\. Qiu, Y\. Yan, C\. Luo, M\. X\. Cheng, J\. He, and H\. TongPrune as you generate: online rollout pruning for faster and better RLVR\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 13876–13893\.External Links:[Link](https://aclanthology.org/2026.acl-long.632/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.632),ISBN 979\-8\-89176\-390\-6Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Xuet al\.\(2026b\)S\. Xu, C\. Peng, J\. Long, W\. Xu, W\. Chu, and Y\. QiHarnessing negative signals: reinforcement distillation from teacher data for LLM reasoning\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 1618–1639\.External Links:[Link](https://aclanthology.org/2026.acl-long.74/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.74),ISBN 979\-8\-89176\-390\-6Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Xuet al\.\(2026c\)Y\. Xu, H\. Sang, Z\. Zhou, R\. He, Z\. Wang, and A\. GeramifardTIP: token importance in on\-policy distillation\.External Links:2604\.14084,[Link](https://arxiv.org/abs/2604.14084)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Yanget al\.\(2024\)K\. Yang, D\. Klein, A\. Celikyilmaz, N\. Peng, and Y\. TianRLCD: reinforcement learning from contrastive distillation for LM alignment\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=v3XXtxWKi6)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Yanget al\.\(2026\)S\. Yang, G\. Zhu, B\. Song, H\. Wang, M\. Xia, X\. Zheng, Y\. Ma, Z\. Chen, W\. Wang, J\. Zhao, and G\. ChenOPRD: on\-policy representation distillation\.External Links:2606\.06021,[Link](https://arxiv.org/abs/2606.06021)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Yuet al\.\(2025\)Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu, X\. Liu, H\. Lin, Z\. Lin, B\. Ma, G\. Sheng, Y\. Tong, C\. Zhang, M\. Zhang, W\. Zhang, H\. Zhu, J\. Zhu, J\. Chen, J\. Chen, C\. Wang, H\. Yu, Y\. Song, X\. Wei, H\. Zhou, J\. Liu, W\. Ma, Y\. Zhang, L\. Yan, M\. Qiao, Y\. Wu, and M\. WangDAPO: an open\-source LLM reinforcement learning system at scale\.External Links:2503\.14476,[Link](https://arxiv.org/abs/2503.14476)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Yuanet al\.\(2025\)W\. Yuan, R\. Y\. Pang, K\. Cho, X\. Li, S\. Sukhbaatar, J\. Xu, and J\. WestonSelf\-rewarding language models\.External Links:2401\.10020,[Link](https://arxiv.org/abs/2401.10020)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Zhanget al\.\(2025a\)K\. Zhang, Q\. YAO, S\. Liu, Y\. Wang, B\. Lai, J\. Ye, M\. Song, and D\. TaoConsistent paths lead to truth: self\-rewarding reinforcement learning for LLM reasoning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=ckW70ls93V)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Zhanget al\.\(2025b\)X\. Zhang, S\. Wen, W\. Wu, and L\. HuangEDGE\-GRPO: entropy\-driven GRPO with guided error correction for advantage diversity\.External Links:2507\.21848,[Link](https://arxiv.org/abs/2507.21848)Cited by:[§1](https://arxiv.org/html/2609.11699#S1.p1.1)\.
- Zhaoet al\.\(2026a\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-Distilled Reasoner: on\-policy self\-distillation for large language models\.External Links:2601\.18734,[Link](https://arxiv.org/abs/2601.18734)Cited by:[§D\.3](https://arxiv.org/html/2609.11699#A4.SS3.p1.1),[3rd item](https://arxiv.org/html/2609.11699#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2609.11699#S1.p2.1),[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px3.p1.1),[§5\.1](https://arxiv.org/html/2609.11699#S5.SS1.p3.1),[§5\.1](https://arxiv.org/html/2609.11699#S5.SS1.p5.1),[§6](https://arxiv.org/html/2609.11699#S6.p1.1)\.
- Zhaoet al\.\(2026b\)X\. Zhao, Z\. Kang, A\. Feng, S\. Levine, and D\. SongLearning to reason without external rewards\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=OU9nFEYR2M)Cited by:[§D\.2](https://arxiv.org/html/2609.11699#A4.SS2.p1.1),[3rd item](https://arxiv.org/html/2609.11699#S1.I1.i3.p1.1),[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.
- Zhuet al\.\(2026\)X\. Zhu, M\. Xia, Z\. Wei, W\. Chen, D\. Chen, and Y\. MengThe surprising effectiveness of negative reinforcement in LLM reasoning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=ftVlLG9cks)Cited by:[§6](https://arxiv.org/html/2609.11699#S6.p3.1)\.
- Zuoet al\.\(2025\)Y\. Zuo, K\. Zhang, L\. Sheng, S\. Qu, G\. Cui, X\. Zhu, H\. Li, Y\. Zhang, X\. Long, E\. Hua, B\. Qi, Y\. Sun, Z\. Ma, L\. Yuan, N\. Ding, and B\. ZhouTTRL: test\-time reinforcement learning\.External Links:2504\.16084,[Link](https://arxiv.org/abs/2504.16084)Cited by:[3rd item](https://arxiv.org/html/2609.11699#S1.I1.i3.p1.1),[§3](https://arxiv.org/html/2609.11699#S3.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.11699#S6.p2.1)\.

## Appendix AAnalysis of Candidate Gated Unlikelihood Gradients in NSD and OPSD Objectives

The NSD loss function is defined asℒ=α⋅DKL\(πref∥πθ\)\+ℒGU\\mathcal\{L\}=\\alpha\\cdot D\_\{\\text\{KL\}\}\(\\pi\_\{\\text\{ref\}\}\\parallel\\pi\_\{\\theta\}\)\+\\mathcal\{L\}\_\{\\text\{GU\}\}\. Specifically, we focus on isolating and analyzing theℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}term, which penalizes the student model on tokens vulnerable to negative conditions\. Letπc=πθ​\(c\)\\pi\_\{c\}=\\pi\_\{\\theta\}\(c\)denote the student’s predicted probability for the target sampled token, andG=max⁡\(0,πneg−πref\)G=\\max\(0,\\pi\_\{\\text\{neg\}\}\-\\pi\_\{\\text\{ref\}\}\)serve as the adaptive gate\. To demonstrate the necessity of ourℒGU\\mathcal\{L\}\_\{\\text\{GU\}\}item with Sigmoid design, we compare two distinct candidates for this penalty:

Standard Unlikelihood:GUstd=G⋅\[−log⁡\(1−πc\)\]\\displaystyle\\text\{GU\}\_\{\\text\{std\}\}=G\\cdot\[\-\\log\(1\-\\pi\_\{c\}\)\]\(8\)Ours:GUsig=G⋅σ⁡\(−log⁡\(1−πc\)\)=G⋅12−πc\\displaystyle\\text\{GU\}\_\{\\text\{sig\}\}=G\\cdot\\sigma\(\-\\log\(1\-\\pi\_\{c\}\)\)=G\\cdot\\frac\{1\}\{2\-\\pi\_\{c\}\}\(9\)The parameter update magnitude is driven by the gradient of the loss with respect to the pre\-softmax logit,∂GU∂zc\\frac\{\\partial\\text\{GU\}\}\{\\partial z\_\{c\}\}\. Given the logit\-probability Jacobian∂πc∂zc=πc​\(1−πc\)\\frac\{\\partial\\pi\_\{c\}\}\{\\partial z\_\{c\}\}=\\pi\_\{c\}\(1\-\\pi\_\{c\}\), we evaluate the optimization behavior of both formulations below\.

![Refer to caption](https://arxiv.org/html/2609.11699v1/gradient_functions_combined.png)Figure 7:\(Left and Mid\) Comparison of the variations of two types of gradients by probability\. \(Right\) The real gradient distribution in 100 training samples\. OPSD tends to assign larger gradients to high\-probability tokens, making the model more prone to drastic updates\. In contrast, compared to the vanilla unlikelihood loss, our GU objective further suppresses the gradients on high\-probability tokens\.### A\.1Gradient Hazard in Standard Unlikelihood

Applying the chain rule to the standard unbounded logarithmic penalty \(Equation[8](https://arxiv.org/html/2609.11699#A1.E8)\), the gradient with respect to the logit is:

∂GUstd∂zc=G⋅11−πc⋅πc​\(1−πc\)=G⋅πc\\frac\{\\partial\\text\{GU\}\_\{\\text\{std\}\}\}\{\\partial z\_\{c\}\}=G\\cdot\\frac\{1\}\{1\-\\pi\_\{c\}\}\\cdot\\pi\_\{c\}\(1\-\\pi\_\{c\}\)=G\\cdot\\pi\_\{c\}\(10\)
Mathematically, the gradient in standard unlikelihood scales strictly linearly with the student’s confidenceπc\\pi\_\{c\}\(as shown in the Figure[7](https://arxiv.org/html/2609.11699#A1.F7)\)\. This creates an optimization hazard: In causal language modeling, tokens with extreme confidence \(πc\>0\.9\\pi\_\{c\}\>0\.9\) are probably trivial structural tokens—such as fixed collocations, prepositions, and punctuation\. Under this formulation, whenever the gateGGis triggered whenπc→1\\pi\_\{c\}\\to 1, the optimizer delivers its almost maximum update magnitude to these hyper\-confident function words\. This aggressively penalizes the model’s fundamental linguistic priors, leading to a degradation in generation fluency, especially when the high\-probability token ratio is high per rollout\.

### A\.2Implicit Gradient Attenuation in Sigmoid\-Squashed GU

To construct a noise\-resilient supervision signal, our method utilizes the Sigmoid\-squashed penalty \(Equation[9](https://arxiv.org/html/2609.11699#A1.E9)\)\. Deriving the logit gradient for this formulation yields:

∂GUsig∂zc\\displaystyle\\frac\{\\partial\\text\{GU\}\_\{\\text\{sig\}\}\}\{\\partial z\_\{c\}\}=G⋅1\(2−πc\)2⋅πc​\(1−πc\)\\displaystyle=G\\cdot\\frac\{1\}\{\(2\-\\pi\_\{c\}\)^\{2\}\}\\cdot\\pi\_\{c\}\(1\-\\pi\_\{c\}\)=G⋅πc​\(1−πc\)\(2−πc\)2\\displaystyle=G\\cdot\\frac\{\\pi\_\{c\}\(1\-\\pi\_\{c\}\)\}\{\(2\-\\pi\_\{c\}\)^\{2\}\}\(11\)
This formulation introduces an elegant, parameter\-free implicit gradient attenuation mechanism\. The presence of the\(1−πc\)\(1\-\\pi\_\{c\}\)term in the numerator fundamentally alters the gradient landscape\. As the student model’s probability approaches11, the gradient magnitude decays toward zero:

limπc→1∂GUsig∂zc=0\\lim\_\{\\pi\_\{c\}\\to 1\}\\frac\{\\partial\\text\{GU\}\_\{\\text\{sig\}\}\}\{\\partial z\_\{c\}\}=0\(12\)
Sinceπc\>0\.9\\pi\_\{c\}\>0\.9predominantly corresponds to uninformative syntactic tokens, the Sigmoid function inherently protects the model’s structural fluency by silencing huge gradient on these tokens\. Instead, as shown in Figure[7](https://arxiv.org/html/2609.11699#A1.F7)it naturally concentrates the highest gradient magnitude on mid\-confidence tokens \(πc≈0\.6\\pi\_\{c\}\\approx 0\.6\), which are more likely to be the ambiguous, reasoning\-critical tokens where the student model requires the strongest corrective supervision\. Consequently, our squashed formulation guarantees that dense supervision remains targeted and stable\.

### A\.3OPSD Gradient Analysis

OPSD objective can be described by:

ℒOPSD\(t\)=DKL\(πθ\(⋅∣xi,si,y<t\)∥πθ\(⋅∣xi,y<t\)\)\\mathcal\{L\}\_\{\\text\{OPSD\}\}^\{\(t\)\}=D\_\{\\text\{KL\}\}\\Big\(\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{i\},s\_\{i\},y\_\{<t\}\)\\parallel\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{i\},y\_\{<t\}\)\\Big\)\(13\)
sis\_\{i\}denotes the gold solution inii\-th training sample\. The gradient visualization is shown in Figure[7](https://arxiv.org/html/2609.11699#A1.F7)\. This figure shows that, compared with the NSD objective, whose gradient generally decreases as the token probability increases, the OPSD objective exhibits an increasing trend\. This indicates that OPSD encourages the model to learn more from high\-probability tokens, which may lead to certain forms of reward hacking, such as overlearning style tokens\. In contrast, the candidate objectives in NSD exhibit relatively stable gradient patterns, while the sigmoid\-based GU can more effectively suppress gradients on high\-probability tokens\.

## Appendix BComparison of NSD with Other Methods

Table 4:Comparison of NSD with other methods\. We conceptually compare them across the following dimensions:Samplingdenotes whether the training relies on trajectories generated by the model itself;Source of reward signalindicates the core component driving the training loss function;Teacherspecifies whether the approach depends on an external teacher model;Gold labelrefers to whether ground\-truth answers are required; andMonitor signal qualityrepresents whether the method actively filters training signals \(e\.g\., unconsciously or intentionally\) rather than indiscriminately optimizing over all tokens\. Note that our NSD is a label\-free approach, which is not directly comparable to baselines that rely on additional or external supervision\. Consequently, our main experiments mostly focus on comparable methods, with OPSD as a representative label\-dependent method for reference\.MethodSamplingSource ofreward signalTeacherGold labelMonitor signalqualitySFT/Off\-PolicyDistillation☹off\-policyexternal teacher☹external☺no☹noRLVR \(GRPO\)☺on\-policygold label☺no☹needed☺noise gradients are☺counteractedOPD☺on\-policyexternal reward☹external☺no☹noOPSD/SDPO☺on\-policygold label☺self☹needed☹noSD\-Zero☺on\-policytrained reviser☺self☹needed☹noRLIF☺on\-policyinternal metric☺no☺no☹noTTRL/U\-OPSD☺on\-policymajority\-voting☺no☺no☹noNSD☺on\-policynegative condition☺self☺no☺noise is filtered bygating

## Appendix CExperiment Details

### C\.1Hyperparameters

Table[5](https://arxiv.org/html/2609.11699#A3.T5)lists the training hyperparameters for all methods\. All experiments are conducted on a single node equipped with 8 NVIDIA A100 \(80GB\) GPUs\. Unless otherwise specified, we adopt the default hyperparameters from the respective official implementations, with the following controlled adjustments for fair comparison: For OPSD, we evaluate configurations both with and without LoRA and report the best\-performing variant \(where LoRA achieves superior results on the 1\.7B and 4B models\)\. For Intuitor, we standardize the training batch size to 128, deviating from their scale\-dependent defaults \(64 for smaller models and 128 for larger models\)\. For TTRL, as majority voting relies on complete final solutions, we extend the maximum generation length to 8192 tokens to prevent output truncation\.

Table 5:Training hyperparameters for all methods\. “—” means not applicable\.HyperparameterNSD \(Online, Solution\-aware\)OPSDIntuitorTTRLGPUs4 for actor \+ 2 for teacher4 for actor \+ 2 for teacher88Train batch size32321288PPO mini\-batch size32321281Max prompt length512512512512Max response length4096409630728192Actor learning rate1×10−61\\times 10^\{\-6\}5×10−65\\times 10^\{\-6\}3×10−63\\times 10^\{\-6\}5×10−75\\times 10^\{\-7\}LR warmup ratio0\.10\.10\.10\.03Rollout per samplenn1188Top\-kklogits32\-1——KL coefficient0\.01—0\.0050\.00Total epochs2222OPSD LoRA target modules: all\-linear, with LoRA rank = 64 and alpha = 128\.

### C\.2Templates

We use the Qwen3 instruct chat template throughout\. All training are conducted innon\-thinking mode: the chat template is invoked withenable\_thinking=False, which causes the model to emit an empty<think\>block and proceed directly to the answer\. This applies uniformly to the student rollout, the teacher log\-probability computation, and all downstream evaluations\.

The template for a single\-turn exchange takes the following form:

Qwen3 Chat Template \(non\-thinking, enable\_thinking=False\)``` <|im_start|>user {user message} <|im_end|> <|im_start|>assistant <think> </think> {model response} ```

The empty<think\>…</think\>block is prepended automatically by the template whenenable\_thinking=Falseandadd\_generation\_prompt=True\. The model then generates its response after the second blank line\.

### C\.3Prompts

The student always receives the plain problem prompt below\. During NSD training the teacher receives either the same prompt \(reference pass\) or a negative prompt \(negative pass\), depending on the variant\. All prompts are wrapped in the chat template described in Appendix[C\.2](https://arxiv.org/html/2609.11699#A3.SS2)\.

Student / reference teacher prompt \(all methods\)\.

Student PromptProblem: \{problem\} Let’s think step by step and output the final answer within \\boxed\{\}\.

#### NSD negative condition prompt generator \(Question\-only\)\.

The following meta\-prompt is sent to a helper LLM to produce the per\-sample negative condition promptnin\_\{i\}used in the question\-only offline variant\. The generated prompt replaces the system context seen by the teacher model\.

Meta\-Prompt: Question\-only Negative Condition GenerationYou are an expert Math Educator and AI Prompt Engineer\. Your task is to analyze the following math problem and generate a “Generalized Attack Prompt” that will force an LLM to make a highly plausible, human\-like cognitive error\.Anatomy of a Universal Attack Prompt:1\.Persona:Must start exactly with*“You are a student who…”*\. Describe a specific bad habit relevant to*this*problem\.2\.Trigger:Abstract the problem’s mathematical class\.*Never*use specific numbers or variables from the current problem\.3\.Flawed Execution:Instruct a naive heuristic or impulsive shortcut that would give a wrong answer\.4\.Fatal Omission:Explicitly forbid the critical verification step\.Now, perform this task for the following problem: Problem:\{problem\}Output*only*the “Generalized Attack Prompt”\. Start your response with*“You are a student who…”*\. Keep it concise \(2–3 sentences\)\.

#### NSD negative condition prompt generator \(Solution\-aware\)\.

When the model’s own rollout is available, the meta\-prompt is augmented with the student’s solution to produce a more targeted negative condition\.

Meta\-Prompt: Solution\-aware Negative Condition GenerationYou are an expert Math Educator and AI Prompt Engineer\. Your task is to analyze the following math problem and a student’s existing solution, then generate a “Targeted Attack Prompt” that exploits the exact reasoning steps the student used to cause a highly plausible cognitive error\.The student’s solution reveals*how*they solved this problem—use that to craft an attack targeting their specific reasoning steps\.Anatomy of a Targeted Attack Prompt:1\.Persona:Must start exactly with*“You are a student who…”*\. Describe a specific bad habit that would corrupt the*exact*step where this student’s reasoning is most fragile\.2\.Trigger:Reference the*type*of reasoning the student used \(not specific numbers or variables from this problem\)\.3\.Flawed Execution:Instruct a shortcut that mirrors the student’s approach but introduces a subtle error\.4\.Fatal Omission:Forbid the specific verification the student performed correctly\.Problem:\{problem\} Student’s Existing Solution: \{solution\}Output*only*the “Targeted Attack Prompt”\. Start your response with*“You are a student who…”*\. Keep it concise \(2–3 sentences\)\.

#### NSD teacher prompt \(wiki\-irr variant\)\.

In the wiki\-irr variant no meta\-prompt generator is used\. Instead, each training sample is paired with a randomly sampled Wikipedia passage that is concatenated as spurious “context”\. The teacher sees the following prompt while the student still receives the plain student prompt above\.

Teacher Prompt: Wiki Irrelevant Negative ConditionProblem: \{problem\} Below is some context you may find useful to answering the question above: \{wikipedia\_passage\} Let’s think step by step and output the final answer within \\boxed\{\}\.

#### NSD teacher prompt\.

For the question\-only and solution\-aware offline variants, the teacher receives the following prompt, where\{negative\_condition\}is the output of the meta\-prompt generator above\.

Teacher Prompt: Question\-only / Solution\-aware Negative ConditionProblem: \{problem\} \{negative\_condition\} Now solve the problem following this instruction: Let’s think step by step and output the final answer within \\boxed\{\}\.

## Appendix DAdditional Experimental Results

### D\.1Pass@8 Performance

We report the performance of pass@8 in Table[6](https://arxiv.org/html/2609.11699#A4.T6)\.

Table 6:Main evaluation results reported aspass@8\(%\): at least one of 8 sampled solutions is correct\. Same evaluation setting as Table[1](https://arxiv.org/html/2609.11699#S4.T1)\.Δ\\DeltaAvg is the average absolute improvement over the same\-size baseline across all 7 benchmarks\.Boldmarks the best result in each model\-size group;underlinemarks the second best\.†\\daggerdenotes methods that require ground\-truth labels\.MethodAIME2024AIME2025AIME2026HMMT2025 FebAMC2023Olympiad\-BenchMATH\-500𝚫\\mathbf\{\\Delta\}Avg1\.7B ModelsQwen3\-1\.7B16\.723\.313\.316\.770\.057\.877\.6—OPSD†40\.023\.313\.316\.772\.559\.677\.6\+3\.9Intuitor30\.023\.316\.713\.375\.057\.076\.2\+2\.3TTRL30\.026\.723\.313\.377\.558\.777\.0\+4\.4NSD33\.336\.723\.316\.777\.560\.977\.6\+7\.24B ModelsQwen3\-4B50\.040\.040\.020\.095\.067\.081\.6—OPSD†40\.046\.736\.730\.092\.566\.481\.6\+0\.0Intuitor50\.053\.336\.726\.795\.066\.180\.0\+2\.0TTRL63\.340\.046\.723\.390\.064\.381\.6\+2\.2NSD60\.063\.353\.330\.095\.068\.382\.0\+8\.38B ModelsQwen3\-8B56\.730\.043\.323\.392\.566\.882\.0—OPSD†46\.743\.340\.023\.387\.567\.681\.8−\-0\.6Intuitor60\.040\.036\.726\.795\.069\.081\.8\+2\.1TTRL56\.733\.336\.720\.092\.566\.882\.0−\-1\.0NSD70\.046\.760\.040\.095\.069\.081\.8\+9\.7

### D\.2Performance on Thinking Mode

We evaluate the performance of all methods under the thinking mode on the Qwen3\-4B model, reporting the results of the best\-performing checkpoints evaluated under the thinking mode\. As shown in Table[7](https://arxiv.org/html/2609.11699#A4.T7), NSD also outperforms all other baselines overall\. Notably, both OPSD and NSD achieve more improvements on challenging datasets such as AIME and Olympiad Bench, aligning with the observations from the non\-thinking setting\. On datasets with limited headroom for improvement \(e\.g\., AMC\), all methods perform comparably to the base model\. Furthermore, we observe that the Intuitor\-trained model tends to over\-think, causing many responses to exceed the maximum generation length limit \(even after extending it to 38k tokens\), which leads to a severe degradation in accuracy\. This phenomenon is also discussed in the previous works\([Zhao et al\., 2026b](https://arxiv.org/html/2609.11699#bib.bib15);[Ghimire et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib50)\)\.

Table 7:Thinking mode evaluation results on 4B models reported as avg@8 \(%\)\. Same evaluation setting as Table 1\.ModelAIME2024AIME2025AIME2026HMMT2025AMC2023MATH\-500OlympiadBenchΔ\\DeltaAvgQwen3\-4B75\.869\.167\.546\.097\.279\.845\.9\-OPSD76\.269\.867\.246\.296\.680\.046\.7\+0\.2Intuitor52\.945\.851\.239\.690\.978\.143\.9\-11\.3TTRL72\.564\.365\.146\.096\.679\.244\.4\-1\.9NSD77\.373\.367\.748\.497\.879\.957\.9\+3\.0
### D\.3Alternative Objective: Policy Gradient Optimization

While the NSD lossℒNSD\(t\)\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}can be directly backpropagated as a supervised objective, we find it also fits a sampled\-token advantage policy\-gradient framework, following the spirit of[Zhao et al\. \(2026a\)](https://arxiv.org/html/2609.11699#bib.bib10)and[Lu and Lab \(2025\)](https://arxiv.org/html/2609.11699#bib.bib19)\. For each tokenyty\_\{t\}in a student rollouty∼πθ\(⋅∣xi\)y\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{i\}\), we define a token\-level advantage as the NSD loss:

At=−ℒNSD\(t\)A\_\{t\}=\-\\mathcal\{L\}\_\{\\text\{NSD\}\}^\{\(t\)\}\(14\)Intuitively, a token with high NSD loss receives a strongly negative advantage, signaling the policy to reduce its probability\. Conversely, tokens with low NSD loss receive near\-zero or positive advantage, leaving their probabilities unchanged\. We treatAtA\_\{t\}as a constant with respect toθ\\thetaand optimize the student via the standard policy gradient surrogate objective:

𝒥PG​\(θ\)=𝔼\(x,a\)∼𝒟,y∼πθ​\[∑tAt​log⁡πθ​\(yt∣xi,y<t\)\]\\mathcal\{J\}\_\{\\text\{PG\}\}\(\\theta\)=\\mathbb\{E\}\_\{\(x,a\)\\sim\\mathcal\{D\},y\\sim\\pi\_\{\\theta\}\}\\left\[\\sum\_\{t\}A\_\{t\}\\log\\pi\_\{\\theta\}\(y\_\{t\}\\mid x\_\{i\},y\_\{<t\}\)\\right\]\(15\)
We evaluate the NSD based on the alternative objective under the same setting as our main experiment\. The result is shown in Table[8](https://arxiv.org/html/2609.11699#A4.T8)\. Compared to models optimized with the𝒥\\mathcal\{J\}objective \(Eq\.[6](https://arxiv.org/html/2609.11699#S2.E6)\), NSD trained under𝒥PG\\mathcal\{J\}\_\{\\text\{PG\}\}\(Eq\.[15](https://arxiv.org/html/2609.11699#A4.E15)\) achieves superior performance on the 1\.7B model \(\+5\.0% on average\)\. However, on the 4B and 8B models, the𝒥\\mathcal\{J\}\-objective NSD yields better overall results\. Notably, the wiki\-irr strategy consistently performs best under the policy gradient setting, while the online solution\-aware strategy emerges as the second best\. This discrepancy arises because the gradients are truncated by the advantage function, decreasing the capture of richer gradient signals\. In contrast, wiki\-irr utilizes noise to introduce more generalized interference \(causing an overall degradation of the model’s reasoning capabilities in long contexts\), which ultimately makes it a more effective strategy in this regime\.

Table 8:Evaluation results of NSD with different negative conditioning strategies on mathematical reasoning benchmarks\. The models are trained based on the NSD policy gradient objective in Eq\.[15](https://arxiv.org/html/2609.11699#A4.E15)\.Method / VariantAIME 2024AIME 2025HMMT Feb 2025MATH\-500𝚫\\mathbf\{\\Delta\}Avg1\.7B ModelsNSD \(Solution\-aware, Offline\)15\.814\.25\.463\.2\+2\.4NSD \(Solution\-aware, Online\)15\.812\.57\.563\.3\+2\.5NSD \(Wiki\-irr, Offline\)20\.015\.09\.664\.4\+5\.04B ModelsNSD \(Question\-only, Offline\)30\.424\.616\.272\.7\+4\.4NSD \(Solution\-aware, Offline\)33\.822\.514\.672\.1\+4\.2NSD \(Solution\-aware, Online\)31\.325\.016\.374\.0\+5\.1NSD \(Wiki\-irr, Offline\)33\.825\.417\.173\.5\+5\.98B ModelsNSD \(Solution\-aware, Offline\)30\.823\.312\.573\.6\+1\.9NSD \(Solution\-aware, Online\)35\.021\.713\.873\.6\+2\.8NSD \(Wiki\-irr, Offline\)34\.228\.819\.273\.9\+5\.8

### D\.4Case Study

A case study is shown in Table[9](https://arxiv.org/html/2609.11699#A4.T9)\. These results indicate that NSD\-trained models more readily explore novel and correct solutions that are entirely absent from the outputs of both the base and OPSD\-trained models\. Furthermore, by prompting an external LLM \(Sonnet\) to analyze the reasoning traces of each response, we observe that the NSD\-trained model engages in several reflection steps, successfully circumventing erroneous trajectories that commonly trap the base model\.

Table 9:Case study on AIME 2025 II \#12 comparing Qwen3\-4B baseline, OPSD, and NSD\. The table shows a summary from Sonnet\. Numbers denote correct samples out of 8 independent draws\.✓correct;✗incorrect\.ProblemBaselineOPSDNSD \(Ours\)AIME 2025 II \#12 — Geometry\.LetA1A2⋯A11A\_\{1\}A\_\{2\}\\cdots A\_\{11\}be a non\-convex simple 11\-gon satisfying: \(1\)\[Ai​A1​Ai\+1\]=1\[A\_\{i\}A\_\{1\}A\_\{i\+1\}\]=1for2≤i≤102\\leq i\\leq 10; \(2\)cos⁡\(∠​Ai​A1​Ai\+1\)=1213\\cos\(\\angle A\_\{i\}A\_\{1\}A\_\{i\+1\}\)=\\tfrac\{12\}\{13\}for2≤i≤102\\leq i\\leq 10; \(3\) perimeter=20=20\. ExpressA1​A2\+A1​A11=m​n−pqA\_\{1\}A\_\{2\}\+A\_\{1\}A\_\{11\}=\\tfrac\{m\\sqrt\{n\}\-p\}\{q\}\(nnsquarefree, no prime divides all ofm,p,qm,p,q\); findm\+n\+p\+qm\+n\+p\+q\.Pass@80/8✗0/8✗4/8✓Key reasoningFromcos⁡θ=1213\\cos\\theta=\\tfrac\{12\}\{13\}derivessin⁡θ=513\\sin\\theta=\\tfrac\{5\}\{13\}, hence\|A1​Ai\|⋅\|A1​Ai\+1\|=265\|A\_\{1\}A\_\{i\}\|\\cdot\|A\_\{1\}A\_\{i\+1\}\|=\\tfrac\{26\}\{5\}\. The product constraint gives an alternating sequencea2=x,a3=265​x,a4=x,…a\_\{2\}=x,\\;a\_\{3\}=\\tfrac\{26\}\{5x\},\\;a\_\{4\}=x,\\;\\ldotsAttempts to use the perimeter butconflates the sum of radii fromA1A\_\{1\}with the polygon perimeter, obtaining5​x\+26x=205x\+\\tfrac\{26\}\{x\}=20which has no clean closed form\. After extensive numerical trials,guesses the symmetric solutionx=135x=\\tfrac\{13\}\{\\sqrt\{5\}\}\(i\.e\.a2=a10a\_\{2\}=a\_\{10\}\), givingA1​A2\+A1​A11=135\+2​5=23​55A\_\{1\}A\_\{2\}\+A\_\{1\}A\_\{11\}=\\tfrac\{13\}\{\\sqrt\{5\}\}\+2\\sqrt\{5\}=\\tfrac\{23\\sqrt\{5\}\}\{5\}, som=23,n=5,p=0,q=5m=23,n=5,p=0,q=5\. Final:𝟑𝟑\\mathbf\{33\}✗Same product relation and alternating sequence\. Applies Law of Cosines: sincexi​xi\+1=265x\_\{i\}x\_\{i\+1\}=\\tfrac\{26\}\{5\}, each inner\-polygon side satisfiesd2=xi2\+xi\+12−485d^\{2\}=x\_\{i\}^\{2\}\+x\_\{i\+1\}^\{2\}\-\\tfrac\{48\}\{5\}, so all 9 inner sides are equal\. Correctly writes the perimeter equationa\+9​d\+265​a=20a\+9d\+\\tfrac\{26\}\{5a\}=20and setsS=a\+265​aS=a\+\\tfrac\{26\}\{5a\}\.Instead of solving forSS, minimisesSSvia AM–GM:min⁡\(a\+265​a\)=2​265=2​1305\\min\\\!\\bigl\(a\+\\tfrac\{26\}\{5a\}\\bigr\)=2\\sqrt\{\\tfrac\{26\}\{5\}\}=\\tfrac\{2\\sqrt\{130\}\}\{5\}, and incorrectlytreats this minimum as the answer, concludingA1​A2\+A1​A11=2​1305A\_\{1\}A\_\{2\}\+A\_\{1\}A\_\{11\}=\\tfrac\{2\\sqrt\{130\}\}\{5\}, som=2,n=130,p=0,q=5m=2,n=130,p=0,q=5\. Final:𝟏𝟑𝟕\\mathbf\{137\}✗Same product relation and alternating sequence\. Applies Law of Cosines; all 9 inner sides equaldd\. Writes the perimeter equationa\+9​d\+265​a=20a\+9d\+\\tfrac\{26\}\{5a\}=20\. Thenattempts a symmetric\-guess approach: testsx=26/5x=\\sqrt\{26/5\},x=2x=2,x=13/5x=13/5,x=13/5x=13/\\sqrt\{5\}in turn, each time verifying numerically that the two expressions ford2d^\{2\}do*not*agree\. SetsS=a\+265​aS=a\+\\tfrac\{26\}\{5a\}, so the perimeter equation givesd=20−S9d=\\tfrac\{20\-S\}\{9\}\. Rewritesd2d^\{2\}via Law of Cosines:d2=a2\+67625​a2−485=\(a\+265​a\)2−525−485=S2−20d^\{2\}=a^\{2\}\+\\tfrac\{676\}\{25a^\{2\}\}\-\\tfrac\{48\}\{5\}=\\bigl\(a\+\\tfrac\{26\}\{5a\}\\bigr\)^\{2\}\-\\tfrac\{52\}\{5\}\-\\tfrac\{48\}\{5\}=S^\{2\}\-20\. Substitutingd=20−S9d=\\tfrac\{20\-S\}\{9\}yields\(20−S9\)2=S2−20\\bigl\(\\tfrac\{20\-S\}\{9\}\\bigr\)^\{2\}=S^\{2\}\-20, which expands to4​S2\+2​S−101=04S^\{2\}\+2S\-101=0\. Quadratic formula:S=−2±16208=9​5−14S=\\tfrac\{\-2\\pm\\sqrt\{1620\}\}\{8\}=\\tfrac\{9\\sqrt\{5\}\-1\}\{4\}\(positive root\), som=9,n=5,p=1,q=4m=9,n=5,p=1,q=4\. Final:𝟏𝟗\\mathbf\{19\}✓ReflectionNo self\-correction\.After the perimeter approach yields no clean form, the modelcommits to a guess\(x=135x=\\tfrac\{13\}\{\\sqrt\{5\}\}\) without checking whether it satisfies the original constraints, and submits the result directly\.No self\-correction\.The model sets up the correct equation structure butreplaces the constraint with its relaxation: once the AM–GM bound is computed it is treated as the solution, with no attempt to verify that the minimum is actually attained\.After the guessing strategy fails on multiple candidates \(x=26/5x=\\\!\\sqrt\{26/5\},22,135\\tfrac\{13\}\{5\},135\\tfrac\{13\}\{\\sqrt\{5\}\}\), the modelexplicitly abandons the approach\(“Hmm\. Maybe my approach is not working\. Alternative idea:”\) andreframes the problemaround the aggregate variableS=a\+265​aS=a\+\\tfrac\{26\}\{5a\}, turning an intractable system into a single quadratic4​S2\+2​S−101=04S^\{2\}\+2S\-101=0\.

## Appendix EDefinition of Task, Style, and Reflection Tokens

Considering the similar task setting with RLCSD\([Pan et al\., 2026](https://arxiv.org/html/2609.11699#bib.bib18)\), we follow theie definition of both task and style tokens:

1. \(1\)empty or whitespace\-only→\\rightarrowstyle;
2. \(2\)matches any of the math regexes \(a digit\\d; an arithmetic operator in\+−=∗/<\>×÷≤≥≠\+\-=\*/<\>\\times\\div\\leq\\geq\\neq; a LaTeX command\\\[A\-Za\-z\]\+; a double backslash; or one of$`^``\_`\)→\\rightarrowtask;
3. \(3\)the normalized form is in the math wordlist \{mod, prime, factor, gcd, lcm, log, ln, sin, cos, tan, exp, integral, sqrt, boxed, frac, sum, prod, pi, alpha, beta, gamma, theta, delta, lambda, mu, sigma, infty, leq, geq, neq, cdot, times, div\}→\\rightarrowtask;
4. \(4\)pure punctuation or a literal newline token \(\\n,\\\\n\)→\\rightarrowstyle;
5. \(5\)the normalized form is in the discourse wordlist \(connectivestherefore, so, thus, hence, then, because, since; hedgeswait, maybe, perhaps, seems, okay, ok, well, now, first, next, finally, actually, alternatively, however; scaffoldingstep, answer, let, lets; closed\-class function wordsis, are, us, we, the, a, an, of, to, for, in, on, by, at, as, and, or, but, if, yes, no, this, that, these, those, it, its, be, been, being, have, has, had, do, does, did, will, would, should, could, can, may\)→\\rightarrowstyle;
6. \(6\)otherwise→\\rightarrowneutral\.

The following tokens are considered as reflection tokens:

wait, actually, hmm, let me reconsider, let me rethink, i made an error, i made a mistake, that’s wrong, that is wrong, incorrect, reconsider, rethink, re\-examine, let me check, let me verify, double check, double\-check, going back, revisit, on second thought

Similar Articles

Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

Hugging Face Daily Papers

Proposes Anti-Self-Distillation (AntiSD) which reverses the knowledge transfer direction in self-distillation to improve math reasoning efficiency and accuracy, achieving GRPO baseline accuracy in 2-10x fewer steps and up to 11.5 points higher final accuracy across models from 4B to 30B parameters.

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

arXiv cs.CL

The paper identifies 'Thinking Collapse' in on-policy self-distillation for large language models, characterized by a decline in intermediate reasoning steps, and proposes AD-OPSD, a control framework that mitigates this collapse by anchoring high-suppression-risk tokens to a reference prior. The method achieves up to +4.1% absolute average accuracy improvement on mathematical benchmarks.