基于位置选择性自蒸馏,从语言反馈中训练 LLM 评判模型

arXiv cs.CL 论文

摘要

该论文提出了一种位置选择性自蒸馏方法,用于利用自然语言反馈训练 LLM 评判模型。该方法通过逐位置的熵变化来屏蔽容易记忆的词元,在主观任务上将分布外泛化能力提升了 2–9 个点,优于 GRPO 等结果监督强化学习方法。

arXiv:2609.38792v1 Announce Type: new Abstract: We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:45

# Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Source: [https://arxiv.org/html/2609.38792](https://arxiv.org/html/2609.38792)
Changlong YuAffiliation:AmazonZhenghao XuAffiliation:Georgia Institute of TechnologyXin LiuAffiliation:AmazonYuwei ZhangAffiliation:UC San DiegoQin LuAffiliation:AmazonBing YinAffiliation:AmazonTuo ZhaoAffiliation:Amazon

###### Abstract

We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them\. The dominant approach, outcome\-supervised RL \(e\.g\., GRPO\), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion\-choice tokens and ignoring the rich language feedback \(e\.g\., preference rationales\) that naturally accompanies preference labels\. Self\-Distillation \(SD\) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position\-level supervision\. However, not all positions carry equally useful signal\. Using the per\-position entropy shift between teacher and student, we identify two regimes:*context sharpening*, where the teacher concentrates probability on a particular feedback\-aligned criterion expression, and*context spreading*, where the teacher distributes probability across multiple feedback\-aligned alternatives\. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives\. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy\-shift distribution\. Experiments show that masking higher\-entropy\-shift positions improves out\-of\-distribution generalization over naive SD\. The resulting self\-distilled judges outperform judges trained with outcome\-supervised RL by 2–9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones\.

††footnotetext:†Work done during internship at Amazon\.005050100100150150200200828284848686Training stepRM\-BenchAcc\. \(%\)Qwen3\-4B\-Instruct005050100100150150200200848486868888Training stepRM\-BenchAcc\. \(%\)Qwen3\-30B\-A3B\-Instruct

Training LLM Judges from Language Feedback via Position\-Selective Self\-Distillation

Figure 1:Out\-of\-distributionRM\-Benchaccuracy during training\. Curves show the debiased EMA \(β=0\.6\\beta=0\.6\) of accuracy evaluated every 20 training steps\. All methods start from a judge prompt template tuned for best base\-model accuracy\. SD\+mask outperforms naive SD and Dr\. GRPO\.
## 1Introduction

Training LLMs to*judge*responses, both as standalone evaluators and as reward models for downstream training, is a core building block of modern post\-training\. A judge’s verdict varies along two distinct axes: 1\)*criterion choice*, i\.e\., which evaluation criteria the model invokes and how it weighs them, and 2\)*application rigor*, i\.e\., how rigorously those criteria are applied to evaluate responses\. Judgment tasks fall into two regimes by which of these two axes decides the verdict:*objective*tasks, where the decisive criterion is clear\-cut \(e\.g\., correctness\), so the verdict depends on application rigor \(e\.g\., math, coding\), and*subjective*tasks, where the decisive criterion is subtle and multifaceted \(e\.g\., what counts as helpful for this chat\), so the choice and weighting of criteria can drive the verdict\. Outcome\-supervised RL \(e\.g\., GRPO\[[Shao et al\., 2024](https://arxiv.org/html/2609.38792#bib.bib7)\], Dr\. GRPO\[[Liu et al\., 2025d](https://arxiv.org/html/2609.38792#bib.bib9)\], DAPO\[[Yu and others, 2025](https://arxiv.org/html/2609.38792#bib.bib8)\]\) with a single verdict\-correctness reward is the dominant approach to train LLM judges\[[Whitehouse et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib11),[Hong et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib18),[Guo et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib19),[Chen et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib12),[Wang et al\., 2026a](https://arxiv.org/html/2609.38792#bib.bib3),[Xu et al\., 2026a](https://arxiv.org/html/2609.38792#bib.bib20)\], and works well on objective tasks by sharpening reasoning over a clear\-cut criterion\.

For subjective tasks, however, outcome\-supervised RL provides limited explicit guidance on criterion choice\. It credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict; it shapes criterion selection and weighting indirectly, through whether a given choice produces a correct verdict, without explicitly distinguishing the contributions of criterion choice and criterion application\. Even with high application rigor, a model that applies the*wrong*criterion \(or weighs competing criteria poorly\) still produces a wrong verdict\. Compounding the problem, outcome\-supervised RL is known to drive policy entropy downward over training\[[Cui and others, 2025](https://arxiv.org/html/2609.38792#bib.bib10),[Yu and others, 2025](https://arxiv.org/html/2609.38792#bib.bib8)\], which further narrows the pool of criteria the judge explores\.

Many preference datasets naturally contain per\-example natural language feedback alongside the preference label, since obtaining the label typically requires annotators to articulate why one response is preferred over the other: preference rationales that explicitly name the decisive criterion\[[Wang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib13),[Liu et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib14)\], a signal that the outcome\-only objective leaves unused\. Self\-Distillation \(SD\)\[[Hübotter et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib1),[Zhao et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib6)\]is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher and provides dense distributional guidance at every position of the student’s rollouts\. Unlike outcome\-supervised RL, this dense per\-position supervision now provides separate credit at the criterion\-choice tokens, since the teacher’s distributional signal at those positions is directly shaped by the language feedback that names the decisive criterion\.

Not every position in the distillation loss carries equally useful signal\. To characterize this heterogeneity, we define the per\-position*entropy shift*as the entropy reduction in the teacher’s next\-token distribution relative to the student’s, induced by conditioning the teacher on the language feedback, and use it to identify two regimes\. At positions with a large positive entropy shift \(*context sharpening*\), the teacher concentrates probability on a particular feedback\-aligned criterion expression\. At positions with a large negative entropy shift \(*context spreading*\), the teacher distributes probability across multiple feedback\-aligned alternatives\. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives\. Figures[3](https://arxiv.org/html/2609.38792#S3.F3),[4](https://arxiv.org/html/2609.38792#S3.F4),[6](https://arxiv.org/html/2609.38792#A1.F6), and[7](https://arxiv.org/html/2609.38792#A1.F7)illustrate concrete examples of these two regimes\. Motivated by this asymmetry, we introduce*position masking based on entropy shift*: retain the lower tail of each generation’s entropy\-shift distribution for the distillation loss, favoring context\-spreading positions while removing the highest\-shift positions\.

Experiments show that self\-distilled judges outperform judges trained with outcome\-supervised RL \(Dr\. GRPO\) by22–99percentage points on the evaluated subjective subcategories, while Dr\. GRPO remains competitive on objective subcategories where application rigor matters most\. Further, masking higher\-entropy\-shift positions improves out\-of\-distribution \(OOD\) generalization over naive SD\. As Figure[1](https://arxiv.org/html/2609.38792#S0.F1)shows, SD\+mask leads both naive SD and Dr\. GRPO onRM\-Bench\[[Liu et al\., 2025c](https://arxiv.org/html/2609.38792#bib.bib17)\]throughout training for Qwen3\-4B\-Instruct and Qwen3\-30B\-A3B\-Instruct\[[Yang and others, 2025](https://arxiv.org/html/2609.38792#bib.bib29)\], surpassing strong reasoning judges including DeepSeek\-R1\[[DeepSeek\-AI, 2025](https://arxiv.org/html/2609.38792#bib.bib21)\]and Claude\-Sonnet\-4\[[Anthropic, 2025](https://arxiv.org/html/2609.38792#bib.bib22)\]at the 30B scale\.

## 2Related Work

##### Learning from natural language feedback\.

Prior work uses natural language feedback primarily through refinement or critique pipelines\. At inference time, models are prompted to iteratively revise outputs from self\- or external\-model\-generated language feedback\[[Madaan et al\., 2023](https://arxiv.org/html/2609.38792#bib.bib49),[Wadhwa et al\., 2024](https://arxiv.org/html/2609.38792#bib.bib50),[Lee et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib53)\]\. At training time, models are fine\-tuned on refinements that incorporate language feedback\[[Scheurer et al\., 2023](https://arxiv.org/html/2609.38792#bib.bib48)\], on self\-critiques and revisions generated from natural language principles\[[Bai et al\., 2022](https://arxiv.org/html/2609.38792#bib.bib51)\], while other methods augment GRPO with additional refinement rollouts conditioned on a critique of the initial response\[[Zhang et al\., 2025a](https://arxiv.org/html/2609.38792#bib.bib54)\]\. Text2Grad\[[Wang et al\., 2026b](https://arxiv.org/html/2609.38792#bib.bib5)\]instead aligns critique phrases with response spans and converts these alignments into per\-span differentiable reward signals that drive gradient updates on the offending tokens\. A more recent line of on\-policy SD uses the same model, additionally conditioned on language feedback unavailable to the student at inference, as the teacher\. This framework incorporates diverse types of language feedback, including reference solutions, environmental feedback, successful rollouts, expert demonstrations, and dynamically summarized skills\[[Zhao et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib6),[Hübotter et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib1),[Shenfeld et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib55),[Wang et al\., 2026c](https://arxiv.org/html/2609.38792#bib.bib52)\]\. Our work uses one\- or two\-sentence annotator rationales as language feedback for judge training\. These rationales expose the evaluation criteria behind each preference label, which is especially useful for subjective tasks where generalization depends on selecting and weighting the right criteria\.

##### Token selection methods for post\-training\.

Recent work has begun to replace uniform token\-level supervision with selective updates during LLM post\-training\. In supervised fine\-tuning and preference optimization, several methods filter or reweight tokens based on influence\-based quality, counterfactual importance, per\-token KL, or preference\-derived importance scores\[[Pang et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib36),[Ruan et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib37),[Zeng et al\., 2024](https://arxiv.org/html/2609.38792#bib.bib38),[Liu et al\., 2025a](https://arxiv.org/html/2609.38792#bib.bib39),[Yang et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib40)\]\. In RLVR, high\-entropy token selection identifies a small set of uncertain “forking” tokens that dominate policy\-gradient learning, while polarity–entropy decomposition and gradient\-magnitude selection further refine token\-level credit assignment\[[Wang et al\., 2025a](https://arxiv.org/html/2609.38792#bib.bib4),[He et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib44),[Lv et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib41)\]\. Closest to our setting, on\-policy distillation methods select or reweight token losses using teacher entropy, student entropy and teacher–student divergence, log\-probability gaps with LLM\-judged relevance, training\-trajectory dynamics, position\-based teacher reliability, or asymmetric updates in non\-positive\-advantage regions\[[Feng and Vaid, 2026](https://arxiv.org/html/2609.38792#bib.bib42),[Jin et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib2),[Xu et al\., 2026b](https://arxiv.org/html/2609.38792#bib.bib43),[Shen et al\., 2026b](https://arxiv.org/html/2609.38792#bib.bib45),[Liu et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib46),[Jia et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib47)\]\. In contrast, our method studies position selection in full\-logit on\-policy SD for judge training via the entropy shift between the student and feedback\-conditioned teacher distributions\.

## 3Method

### 3\.1Preliminary: Outcome\-Supervised RL for LLM Judges

A pairwise LLM judge is trained on examples of the form\(x,yA,yB,c⋆\)\(x,y\_\{A\},y\_\{B\},c^\{\\star\}\): a promptxx, two candidate responsesyA,yBy\_\{A\},y\_\{B\}, and a gold preference labelc⋆∈𝒞c^\{\\star\}\\in\\mathcal\{C\}where𝒞\\mathcal\{C\}is a finite set of possible verdicts\. The verdict set𝒞\\mathcal\{C\}could be simply binary \(yA≻yBy\_\{A\}\\succ y\_\{B\}oryA≺yBy\_\{A\}\\prec y\_\{B\}\) or multiclass\[[Hong et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib18),[Wang et al\., 2026a](https://arxiv.org/html/2609.38792#bib.bib3)\], for example includingtie,unknown/unclear, or a multi\-way ordinal preference such asyA≻≻yBy\_\{A\}\\succ\\succ y\_\{B\},yA≻yBy\_\{A\}\\succ y\_\{B\},yA∼yBy\_\{A\}\\sim y\_\{B\}\. In this paper, we consider the binary setting with𝒞=\{A,B\}\\mathcal\{C\}=\\\{A,B\\\}\. Given\(x,yA,yB\)\(x,y\_\{A\},y\_\{B\}\), the judge generates a token sequenceτ=\(a1,a2,…,aT\)\\tau=\(a\_\{1\},a\_\{2\},\\ldots,a\_\{T\}\)one token at a time,at∼π\(⋅∣st\)a\_\{t\}\\sim\\pi\(\\cdot\\mid s\_\{t\}\), wherest=\(x,yA,yB,a1,…,at−1\)s\_\{t\}=\(x,y\_\{A\},y\_\{B\},a\_\{1\},\\ldots,a\_\{t\-1\}\)is the prefix at positiontt\. The sequence consists of an intermediate trace \(criterion choice and application\) and a final verdictc≡aT∈𝒞c\\equiv a\_\{T\}\\in\\mathcal\{C\}\.

The dominant training algorithm is outcome\-supervised RL\. Each rollout is scored by whether its final verdict matches the gold preference label:

R\(τ\)=\[c=c⋆\]\.R\(\\tau\)\\;=\\;\\mathbf\{1\}\\\!\\left\[c=c^\{\\star\}\\right\]\.\(1\)The judge is then optimized against this reward using GRPO\-style policy\-optimization algorithms\[[Shao et al\., 2024](https://arxiv.org/html/2609.38792#bib.bib7),[Liu et al\., 2025d](https://arxiv.org/html/2609.38792#bib.bib9),[Yu and others, 2025](https://arxiv.org/html/2609.38792#bib.bib8)\]\. Even when a training example contains language feedbackzz, such as a preference rationale that articulates why one response is preferred over the other and identifies the decisive criterion, this outcome\-only reward leaveszzunused\. It also assigns the same outcome\-based credit to every token position, with no direct supervision where the judge chooses and weighs evaluation criteria\.

### 3\.2Language Feedback Self\-Distillation for LLM Judges

To use the language feedback left unused by the outcome\-only objective, we apply self\-distillation \(SD\) to provide dense, position\-specific supervision for the judge’s intermediate trace\.

Each training example carries a piece of language feedbackzzunderlying its preference labelc⋆c^\{\\star\}\. We use a single modelπ\\piin two roles: as the*student*, conditioned on the prompt alone,π\(⋅∣st\)\\pi\(\\cdot\\mid s\_\{t\}\); and as the*teacher*, which additionally conditions onzz,π\(⋅∣st,z\)\\pi\(\\cdot\\mid s\_\{t\},z\)\. The context is*privileged*in the sense that the student is never givenzz, either during training or at deployment\. We adopt the on\-policy reverse\-KL SD objective from\[[Hübotter et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib1),[Zhao et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib6)\]as the underlying loss\. Over a training set𝒟\\mathcal\{D\}of judge examples\(x,yA,yB,z\)\(x,y\_\{A\},y\_\{B\},z\)with on\-policy rolloutsτ=\(a1,…,aT\)∼π\\tau=\(a\_\{1\},\\ldots,a\_\{T\}\)\\sim\\pidrawn per example,

ℒSD\(π\)=𝔼\(x,yA,yB,z\)∼𝒟,τ∼π\[1T∑t=1TKL\(π\(⋅∣st\)∥sg\[π\(⋅∣st,z\)\]\)\],\\mathcal\{L\}\_\{\\text\{SD\}\}\(\\pi\)=\\mathbb\{E\}\_\{\(x,y\_\{A\},y\_\{B\},z\)\\sim\\mathcal\{D\},\\;\\tau\\sim\\pi\}\\left\[\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathrm\{KL\}\\\!\\left\(\\pi\(\\cdot\\mid s\_\{t\}\)\\;\\Big\\\|\\;\\mathrm\{sg\}\\\!\\left\[\\pi\(\\cdot\\mid s\_\{t\},z\)\\right\]\\right\)\\right\],\(2\)wheresg⁡\[⋅\]\\mathrm\{sg\}\[\\cdot\]is stop\-gradient and the inner KL is a full\-vocabulary sum at each position\. Unlike a verdict\-level reward, this objective transfers the feedback\-conditioned teacher distribution at every token position, including positions where the judge selects evaluation criteria\.

Figure 2:BothΔ​Ht\\Delta H\_\{t\}tails account for most of the distillation loss\.\(a\)Three\-way per\-sequence percentile partition \(top\-30%Δ​Ht\\Delta H\_\{t\}/ middle 40% / bottom\-30%\): both tails carry∼\\sim20–28×\\timesthe per\-position KL of the neutral middle\.\(b\)Binned mean reverse KL alongΔ​Ht\\Delta H\_\{t\}using 20 equal\-count bins \(95% x\-range shown\)\.![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/revkl_by_ig_tier_a.png)\(a\)Mean reverse KL byΔ​Ht\\Delta H\_\{t\}tier\.
![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/revkl_by_ig_tier_b.png)\(b\)Mean reverse KL as a function ofΔ​Ht\\Delta H\_\{t\}\.

### 3\.3Analysis of Per\-Position Self\-Distillation Signal

Eq\.[2](https://arxiv.org/html/2609.38792#S3.E2)treats every response position as an equally valid imitation target, but the language feedbackzzdoes not affect the teacher uniformly across positions\. We characterize this heterogeneity through the per\-position entropy shift, then examine its connection to criterion choice\. As discussed above, language feedback identifies the decisive criteria behind a preference label, allowing self\-distillation to provide direct supervision at criterion\-choice tokens\. We therefore study how this feedback changes the teacher’s next\-token distribution at positions where the judge names and defines evaluation criteria\.

##### Per\-position entropy shift\.

To measure how the language feedbackzzchanges next\-token uncertainty at positiontt, we define

ΔH\(st,z\)=H\(π\(⋅∣st\)\)−H\(π\(⋅∣st,z\)\),\\Delta H\(s\_\{t\},z\)\\;=\\;H\\\!\\left\(\\pi\(\\cdot\\mid s\_\{t\}\)\\right\)\\;\-\\;H\\\!\\left\(\\pi\(\\cdot\\mid s\_\{t\},z\)\\right\),\(3\)the entropy of the model’s next\-token distribution withoutzzminus the entropy withzz, both evaluated at positiontt\. When the language feedback is fixed for an example, we abbreviate this asΔ​Ht\\Delta H\_\{t\}\.*Positive*Δ​Ht\\Delta H\_\{t\}means the teacher conditioned onzzhas lower entropy than the student;zzhas sharpened the teacher’s distribution\.*Negative*Δ​Ht\\Delta H\_\{t\}means the teacher has higher entropy than the student;zzhas spread the teacher’s distribution\.*Near\-zero*Δ​Ht\\Delta H\_\{t\}means little change in entropy, though not necessarily little change in the next\-token distribution\.

##### BothΔ​Ht\\Delta H\_\{t\}tails carry substantial distillation loss\.

We analyze 104,046 response positions from 102HelpSteer3\-Preference\[[Wang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib13)\]validation rollouts generated by Qwen3\-30B\-A3B\-Instruct\-2507\[[Yang and others, 2025](https://arxiv.org/html/2609.38792#bib.bib29)\], with 34 examples each from code, general, and STEM\. Empirically, per\-position reverse KL exhibits a U\-shaped relationship withΔ​Ht\\Delta H\_\{t\}: positions at either tail of the per\-sequenceΔ​Ht\\Delta H\_\{t\}distribution carry substantially more per\-position KL than positions in the near\-zero middle \(Figure[2](https://arxiv.org/html/2609.38792#S3.F2)\)\. Partitioning each rollout into the bottom 30%, middle 40%, and top 30% byΔ​Ht\\Delta H\_\{t\}, both tails carry approximately 20–28×\\timesthe per\-position KL of the middle\. Thus, substantial distillation loss occurs both where the language feedback sharpens the teacher’s distribution and where it spreads it\.

![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/30b_examples/30b_sharpening_main.png)Figure 3:Context sharpening on twoHelpSteer3\-Preferencerollouts \(Qwen3\-30B\-A3B\-Instruct\-2507\)\. Each panel shows the user task, the rationale \(anchor wordred\), the student rollout excerpt with the annotated position⟨\\langlePosNN⟩\\ranglemarked inblue, and student vs\. teacher top\-5 next\-token distributions at that position\. All positions sit at criterion\-naming slots in the rollout’s opening evaluation\-criteria list\. The teacher concentrates probability on a token semantically equivalent to the rationale’s anchor while the student is uncertain across multiple plausible criteria\. Two further rollouts are shown in Figure[6](https://arxiv.org/html/2609.38792#A1.F6)\.
##### Criterion\-choice tokens account for disproportionate distillation loss\.

Criterion choice determines which evaluation criteria the judge invokes and how it weighs them, and is particularly important for subjective tasks\. Positions where the judge names and defines these criteria make this choice explicit in the intermediate trace\. Since the preference rationale identifies the decisive criteria, these positions provide a natural location to examine how language feedback shapes criterion choice\.

We analyze the same 102 rollouts using GPT\-5\.5\[[OpenAI, 2026](https://arxiv.org/html/2609.38792#bib.bib23)\], with the generation and rationale annotated separately\. From the generation alone, the annotator marks*criterion\-selection spans*, comprising tokens that name and define evaluation criteria, and*criterion\-name positions*, marking each criterion’s head noun\. These annotations are mapped to model\-token positions\. From the rationale alone, the annotator extracts the decisive criteria, which serve as the reference for the semantic analysis below\.

Criterion\-selection spans comprise only 16\.7% of response tokens but account for 44\.4% of total reverse\-KL mass, corresponding to approximately 4\.0×\\timesthe per\-token KL of other positions\. Criterion\-name positions alone comprise 0\.4% of tokens but account for 11\.1% of total reverse\-KL mass\. They also exhibit 5\.7×\\timeslarger mean\|Δ​Ht\|\|\\Delta H\_\{t\}\|than other response positions, and 94\.8% fall within the two extreme 30% tails, which together contain 60% of response positions\. Thus, criterion\-choice tokens account for disproportionate distillation loss, and criterion\-name positions concentrate in both entropy\-shift tails\. This motivates examining how the two tails differ in the criterion choices favored by the teacher\.

##### Two regimes of language\-feedback influence on criterion choice\.

The sign ofΔ​Ht\\Delta H\_\{t\}distinguishes sharpening from spreading, but entropy alone does not reveal which criteria the teacher favors\. We therefore examine whether its candidate criterion names align with the decisive criteria extracted from the preference rationale\.

Among the annotated criterion\-name positions, we select those with\|Δ​Ht\|≥0\.50\|\\Delta H\_\{t\}\|\\geq 0\.50, yielding 87 sharpening positions from 56 rollouts and 71 spreading positions from 51 rollouts\. At each position, we take the teacher’s top\-10 next\-token candidates as candidate criterion\-name tokens\. We force each token and greedily complete it into a criterion name, truncating at the head noun\. Using Qwen3\-Embedding\-8B\[[Zhang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib28)\], we score each completed name by its maximum cosine similarity to the extracted decisive criteria\. For sharpening, we score the name obtained from the top\-1 token, measuring alignment of the teacher’s most probable criterion\. For spreading, we average across all ten candidates, measuring alignment across the broader set of criteria\.

As a control, we replace the teacher’s language feedback with another example’s rationale and repeat the candidate selection and completion procedure\. The student prefix, evaluation position, and regime assignment remain fixed as determined under the matched condition\. Both conditions are scored against the same decisive criteria from the original rationale\. We report paired mean differences, with 95% confidence intervals obtained from 2,000 bootstrap resamples at the rollout level\.

![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/30b_examples/30b_spreading_main.png)Figure 4:Context spreading on twoHelpSteer3\-Preferencerollouts \(Qwen3\-30B\-A3B\-Instruct\-2507\)\. Same layout as Figure[3](https://arxiv.org/html/2609.38792#S3.F3)\. Each panel features a position where the student commits to a criterion that is non\-decisive in the rationale \(top\-1 probability≥0\.79\\geq 0\.79\), while the teacher reopens with a spread of alternatives that include semantic equivalents of the rationale’s decisive criterion\. All positions sit at criterion\-naming slots in the rollout’s opening evaluation\-criteria list\. Two further rollouts are shown in Figure[7](https://arxiv.org/html/2609.38792#A1.F7)\.*Context sharpening\.*At positions with large positiveΔ​Ht\\Delta H\_\{t\}, the language feedback makes the teacher more confident than the student\. Figure[3](https://arxiv.org/html/2609.38792#S3.F3)illustrates how this concentrates probability on a specific decisive criterion\. In the TypeScript\-explanation rollout, the rationale phrase*“correctly includes”*causes the teacher to concentrate onCorrectness\(top\-10\.970\.97\), while the student is uncertain across several plausible criteria\. In the legal\-term explanation rollout,*“difficult to read”*causes the teacher to concentrate onReadability\(top\-11\.001\.00\)\.

Across the annotated sharpening positions, the teacher’s top\-1 criterion name has mean similarity 0\.713 to the decisive criteria under the matched rationale, compared with 0\.647 under the control\. The paired difference is\+0\.066\+0\.066, with a 95% confidence interval of\[0\.033,0\.100\]\[0\.033,0\.100\]\. Together with the lower teacher entropy, this supports the interpretation that context sharpening concentrates probability on a particular feedback\-aligned criterion expression\. We interpret this concentrated supervision as encouraging memorization of a particular criterion expression rather than understanding of the underlying criterion\.

*Context spreading\.*At positions with large negativeΔ​Ht\\Delta H\_\{t\}, the language feedback makes the teacher less certain than the student\. Figure[4](https://arxiv.org/html/2609.38792#S3.F4)illustrates how this reopens a criterion choice by spreading probability across alternatives\. In the story\-writing rollout, the student commits toOriginality\(top\-10\.830\.83\), while the rationale’s emphasis on*“factually verifiable”*details shifts the teacher toward alternatives includingCorrectness,Error, andFidelity\. In the project\-description rollout, the student commits toCompleteness, while the*“one or two sentences”*constraint shifts the teacher toward alternatives includingPrecision,Focus, andAdherence\.

Across the annotated spreading positions, mean similarity over the teacher’s top\-10 criterion names is 0\.646 under the matched rationale, compared with 0\.610 under the control\. The paired difference is\+0\.037\+0\.037, with a 95% confidence interval of\[0\.022,0\.051\]\[0\.022,0\.051\]\. The result is consistent when using the top\-3 or top\-5 candidates\. Together with the higher teacher entropy, this supports the interpretation that context spreading distributes probability across multiple feedback\-aligned alternatives\. We interpret this supervision as promoting semantic understanding of the underlying criterion by preserving multiple feedback\-aligned alternatives\.

### 3\.4Position Masking Based on Entropy Shift

Section[3\.3](https://arxiv.org/html/2609.38792#S3.SS3)shows that context sharpening concentrates probability on a particular feedback\-aligned criterion expression, while context spreading distributes probability across multiple feedback\-aligned alternatives\. Motivated by the possibility that preserving these alternatives promotes semantic understanding, we propose a per\-generation mask that drops the upper tail and retains the lower tail of the entropy\-shift distribution\. Concretely, we rank positions within each generation byΔ​Ht\\Delta H\_\{t\}, mask the topρ\\rhofraction with the largest values, and retain the remaining bottom1−ρ1\-\\rhofraction:

ℒSD\(ρ\)​\(π\)\\displaystyle\\mathcal\{L\}\_\{\\text\{SD\}\}^\{\(\\rho\)\}\(\\pi\)=𝔼\(x,yA,yB,z\)∼𝒟,τ∼π\[1∑tmt∑t=1Tmt⋅KL\(π\(⋅∣st\)∥sg\[π\(⋅∣st,z\)\]\)\],\\displaystyle=\\;\\mathbb\{E\}\_\{\(x,y\_\{A\},y\_\{B\},z\)\\sim\\mathcal\{D\},\\;\\tau\\sim\\pi\}\\left\[\\frac\{1\}\{\\sum\_\{t\}m\_\{t\}\}\\sum\_\{t=1\}^\{T\}m\_\{t\}\\cdot\\mathrm\{KL\}\\\!\\left\(\\pi\(\\cdot\\mid s\_\{t\}\)\\;\\Big\\\|\\;\\mathrm\{sg\}\\bigl\[\\pi\(\\cdot\\mid s\_\{t\},z\)\\bigr\]\\right\)\\right\],\(4\)mt\\displaystyle m\_\{t\}=\[ΔHt≤Q1−ρ\],\\displaystyle=\\;\\mathbf\{1\}\\\!\\left\[\\Delta H\_\{t\}\\leq Q\_\{1\-\\rho\}\\right\],whereQ1−ρQ\_\{1\-\\rho\}is the\(1−ρ\)\(1\-\\rho\)\-quantile of\{Δ​H1,…,Δ​HT\}\\\{\\Delta H\_\{1\},\\dots,\\Delta H\_\{T\}\\\}for that generation\. The mask is detached from the computation graph\. Settingρ=0\\rho=0recovers naive SD\.

##### Entropy\-shift masking is associated with broader criterion diversity\.

![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/criteria_pool_sweep.png)Figure 5:Distinct semantic clusters of evaluation criteria invoked by trained 30B judges across a sweep of complete\-link cosine clustering thresholds\. Naive SD invokes a markedly smaller pool than SD\+mask at every threshold; the gap widens with stricter clustering\.Section[3\.3](https://arxiv.org/html/2609.38792#S3.SS3)shows that, at the analyzed criterion\-name positions, context spreading distributes probability across multiple feedback\-aligned alternatives, whereas context sharpening concentrates probability on a particular criterion expression\. This suggests that prioritizing lower\-entropy\-shift positions during distillation helps preserve a broader repertoire of evaluation criteria\. We examine this possibility by measuring the diversity of criteria invoked by the trained judges at inference time\. For each judgment a trained judge produces \(across 300 samples each fromRM\-Bench\[[Liu et al\., 2025c](https://arxiv.org/html/2609.38792#bib.bib17)\],RewardBench v2\[[Malik et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib16)\], andHelpSteer3\-Preferencevalidation\[[Wang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib13)\]\), we extract the evaluation criteria the model proposes and uses \(e\.g\.,“adherence to user intent”,“factual accuracy”,“relevance to the request”\), apply text normalization \(lowercase, strip punctuation, sort tokens to merge order\-variants\), and embed each criterion string with Qwen3\-Embedding\-8B\[[Zhang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib28)\]\. We then cluster the resulting 4096\-dimensional vectors with*complete\-link agglomerative clustering*at cosine thresholdtt: two criterion strings belong to the same cluster only if*every pair*within the cluster has cosine similarity≥t\\geq t\. This procedure groups semantically similar criterion expressions \(e\.g\.,“factual accuracy,”“factual correctness,”and“accuracy of facts”\), reducing sensitivity to differences in wording\. Largerttenforces tighter clusters; lowerttallows looser merging\.

Figure[5](https://arxiv.org/html/2609.38792#S3.F5)shows that SD\+mask invokes more distinct criterion clusters than naive SD at every evaluated clustering threshold, with the difference increasing from 4 clusters att=0\.65t=0\.65to 26 att=0\.85t=0\.85\. This pattern is consistent with the interpretation suggested by §[3\.3](https://arxiv.org/html/2609.38792#S3.SS3): supervision that preserves multiple feedback\-aligned alternatives helps maintain a broader criterion repertoire after training\. Different prompts call for different evaluation criteria, making criterion diversity a relevant property of a general\-purpose judge\.

## 4Experiments

### 4\.1Setup

##### Base models\.

##### Training data\.

All methods are trained on a cleaned version ofHelpSteer3\-Preference\[[Wang et al\., 2025b](https://arxiv.org/html/2609.38792#bib.bib13)\]with three modifications applied in order: \(1\) filter out rows withdomain == "multilingual"oroverall\_preference == 0\(tie\); \(2\) within\-split deduplication by content hash \(sha1\(context, response1, response2\)\), keeping the first occurrence; upstreamHelpSteer3\-Preferencecontains∼\\sim35% byte\-identical duplicate rows after step \(1\); \(3\) cross\-split deduplication: drop validation rows whose content hash also appears in train\-split; upstreamHelpSteer3\-Preferencehas∼\\sim918 of∼\\sim1,553 filtered\-deduped validation rows that are byte\-identical to a train\-split row and would otherwise cause train/val leakage\. Each retained example contains a prompt, two candidate responses, a gold preference label, and a one\- or two\-sentence human\-written rationale explaining the label\.

##### Language feedback types\.

We mainly use the preference rationale as the source of language feedbackzz\. Preference rationales are the most naturally available and easiest\-to\-collect form of textual feedback during preference\-data construction\. To assign a binary preference label, an annotator must already compare the two responses and identify why one is better\. To test sensitivity to a substantially different language\-feedback format, we also generated a per\-example rubric for everyHelpSteer3\-Preferencesample using Claude Opus 4\.8\[[Anthropic, 2026a](https://arxiv.org/html/2609.38792#bib.bib24)\], following the iterative rubric\-generation procedures of\[[Shen et al\., 2026a](https://arxiv.org/html/2609.38792#bib.bib15)\]\. Generating rubrics requires a separate generation or annotation process and careful design to ensure that the criteria are discriminative, non\-redundant, and aligned with the response pair and preference direction\. The generated rubrics average 2,062 characters, making them approximately 7\.7 times longer than the original preference rationales\. Results using these rubrics as language feedback are reported in Appendix[B](https://arxiv.org/html/2609.38792#A2)\.

##### Evaluation benchmarks\.

We evaluate out\-of\-distribution \(OOD\) judge accuracy onRM\-Bench\[[Liu et al\., 2025c](https://arxiv.org/html/2609.38792#bib.bib17)\]andRewardBench v2\[[Malik et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib16)\]\.

##### Baselines\.

We compare three on\-policy training methods\. We use Dr\. GRPO as the outcome\-supervised RL baseline\. Naive SD \(ρ=0\\rho=0\) is the objective in Eq\.[2](https://arxiv.org/html/2609.38792#S3.E2)with all positions included\. SD\+mask usesρ=0\.7\\rho=0\.7, masking the top 70% of positions byΔ​Ht\\Delta H\_\{t\}and keeping the 30% of positions with the lowest entropy shifts\. Implementation details for all three are in Appendix[D](https://arxiv.org/html/2609.38792#A4)\. We additionally compare against strong prompting baselines \(DeepSeek\-R1\[[DeepSeek\-AI, 2025](https://arxiv.org/html/2609.38792#bib.bib21)\], Claude\-Sonnet\-4\[[Anthropic, 2025](https://arxiv.org/html/2609.38792#bib.bib22)\]\) and trained reward\-model baselines \(Think\-RM\-8B\[[Hong et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib18)\], RM\-R1\-DeepSeek\-Distilled\-Qwen\-7B and RM\-R1\-DeepSeek\-Distilled\-Qwen\-32B\[[Chen et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib12)\], Llama\-3\.3\-Nemotron\-Super\-49B\-GenRM\[[NVIDIA, 2025](https://arxiv.org/html/2609.38792#bib.bib31)\]\) evaluated from their authors’ released checkpoints, and RationaleRM\-30B\[[Wang et al\., 2026a](https://arxiv.org/html/2609.38792#bib.bib3)\]for which we cite the authors’ reported numbers\.

Table 1:Judge accuracy \(%\) onRM\-BenchandRewardBench v2\. Dark gray \(in bold\) and light gray highlight the best and second\-best performance per column, respectively\. RationaleRM\-30B numbers are taken from[Wang et al\. \[2026a\]](https://arxiv.org/html/2609.38792#bib.bib3)as the model has not been released;RewardBench v2cells are left empty\.

### 4\.2Main Results

Table[1](https://arxiv.org/html/2609.38792#S4.T1)reports judge accuracy onRM\-BenchandRewardBench v2for the Qwen3\-4B\-Instruct and Qwen3\-30B\-A3B\-Instruct base models, using preference rationales as language feedback, alongside prompting\-based and trained reward\-model baselines\.

##### Main findings\.

Per\-category results in Table[1](https://arxiv.org/html/2609.38792#S4.T1)can be split along the objective/subjective axis\.

*Outcome\-supervised RL vs\. self\-distillation\.*On the objective subcategories \(e\.g\., math, coding\), where the decisive criterion is clear\-cut, Dr\. GRPO sharpens application rigor against that criterion and is competitive with both naive SD and SD\+mask: at 30B it matches them on RM\-Bench Math \(95\.59 \(Dr\. GRPO\) vs 94\.45 \(SD\) / 95\.63 \(SD\+mask\)\) and beats them onRewardBench v2Math \(87\.43 vs 81\.42 / 84\.15\)\. On the subjective subcategories \(e\.g\., chat helpfulness; factuality, where the judge must prioritize factual accuracy over otherwise persuasive presentation; focus, which tests detection of high\-quality, on\-topic answers to general user queries\), the gap flips: both SD variants gain over Dr\. GRPO by 2–9 percentage points \(30B Chat 74\.07 vs 76\.74 / 80\.53; 30B Factuality 71\.79 vs 78\.53 / 79\.16; 30B Focus 80\.00 vs 86\.26 / 85\.66; 4B Chat 68\.91 vs 75\.71 / 77\.95; 4B Factuality 62\.74 vs 68\.42 / 70\.74; 4B Focus 75\.76 vs 80\.20 / 77\.78\)\.

*Naive SD vs\. SD\+mask\.*SD\+mask consistently improves over naive SD on overall accuracy at both scales \(Total Avg:77\.70→79\.0477\.70\\to 79\.04at 4B;80\.17→82\.5080\.17\\to 82\.50at 30B\), with the largest per\-subcategory gains spread across different subcategories \(30B RM\-Bench Code76\.56→80\.0776\.56\\to 80\.07; 30B Chat76\.74→80\.5376\.74\\to 80\.53; 30BRewardBench v2Precise IF40\.62→48\.7540\.62\\to 48\.75\)\. Because both benchmarks are OOD with respect to the training data, these gains are consistent with the proposed mechanism: masking higher\-entropy\-shift positions reduces overreliance on particular criterion expressions, while the broader inference\-time criterion vocabulary in Figure[5](https://arxiv.org/html/2609.38792#S3.F5)provides additional support for this interpretation\.

Our 30B SD\+mask judge outperforms leading judge baselines in the literature, including RM\-R1\-DeepSeek\-Distilled\-Qwen\-32B, Llama\-3\.3\-Nemotron\-Super\-49B\-GenRM, and RationaleRM\-30B, on overall benchmark accuracy, and remains comparable to Claude\-Sonnet\-4\.

### 4\.3Downstream Utility for Policy Optimization

The benchmark results above evaluate judges directly, but a reward model is ultimately used to determine which on\-policy outputs receive higher reward and advantage during policy optimization\. The strongest downstream validation would train separate policies with GRPO using each learned judge as the reward model\. Because this requires multiple full RLHF runs, we instead evaluate the core pairwise\-selection operation used by judge\-guided GRPO through a controlled multi\-round best\-of\-eight tournament\.

##### Setup\.

For each subjective creative\-writing prompt from Arena\-Hard\-v2\[[Li et al\., 2024](https://arxiv.org/html/2609.38792#bib.bib27)\], GLM\-4\.5\-Air\[[Zeng and others, 2025](https://arxiv.org/html/2609.38792#bib.bib26)\]generates eight responses with temperature1\.01\.0, top\-pp1\.01\.0, and a maximum of 8K new tokens\. We randomly pair the eight responses into four matchups and use a judge to select the winner of each pair\. We repeat this procedure with the four winners, and then with the two remaining responses, until one final response is selected\. We run the tournament separately using Qwen3\-30B\-A3B\-Instruct judges trained with Dr\. GRPO, naive SD, and SD\+mask \(ρ=0\.7\\rho=0\.7\)\. Claude Sonnet 5\[[Anthropic, 2026b](https://arxiv.org/html/2609.38792#bib.bib25)\]evaluates the selected responses against the Arena\-Hard\-v2 reference responses to compute the creative\-writing win rate\.

Table 2:Creative\-writing win rate of responses selected through multi\-round best\-of\-eight tournaments\. Each tournament uses a different Qwen3\-30B\-A3B\-Instruct judge as its pairwise selector; selected responses are evaluated against the Arena\-Hard\-v2 reference responses by Claude Sonnet 5\.SD\+mask selects the strongest downstream responses, improving the creative\-writing win rate over the Dr\. GRPO\-trained judge by7\.67\.6percentage points and over naive SD by2\.02\.0percentage points\. Although this tournament does not include the subsequent gradient\-based policy updates of a full GRPO run, it directly tests the repeated pairwise reward comparisons that determine which on\-policy outputs would be preferentially reinforced\. The result therefore provides evidence that the gains from SD\+mask are not confined to standalone judge benchmarks: when used as a reward selector, it more reliably favors responses preferred under the downstream evaluation\.

### 4\.4Ablations

#### 4\.4\.1Mask\-fraction \(ρ\\rho\) sweep

Table 3:ρ\\rho\-sweep on Qwen3\-30B\-A3B\-Instruct, all masking the top\-ρ\\rhofraction byΔ​Ht\\Delta H\_\{t\}and training on the remaining bottom\(1−ρ\)\(1\-\\rho\)\.We sweep the mask fractionρ∈\{0,0\.3,0\.5,0\.7\}\\rho\\in\\\{0,0\.3,0\.5,0\.7\\\}on Qwen3\-30B\-A3B\-Instruct \(Table[3](https://arxiv.org/html/2609.38792#S4.T3)\) to characterize how aggressive the mask should be; the corresponding Qwen3\-4B\-Instruct results are reported in Appendix[C](https://arxiv.org/html/2609.38792#A3)\.ρ=0\\rho=0recovers naive SD \(no positions masked\)\. Overall accuracy rises as we mask more of the upperΔ​Ht\\Delta H\_\{t\}tail \(Total Avg:80\.17→80\.63→80\.68→82\.5080\.17\\to 80\.63\\to 80\.68\\to 82\.50acrossρ∈\{0,0\.3,0\.5,0\.7\}\\rho\\in\\\{0,0\.3,0\.5,0\.7\\\}\), withρ=0\.7\\rho=0\.7achieving the highest overall accuracy among the tested mask fractions\. The same pattern holds at 4B, whereρ=0\.7\\rho=0\.7also obtains the highest total average \(Table[6](https://arxiv.org/html/2609.38792#A3.T6)\)\.

#### 4\.4\.2Selector ablation

Table 4:Selector ablation on Qwen3\-30B\-A3B\-Instruct, all at matched mask fractionρ=0\.7\\rho=0\.7\.Is the OOD gain specific to the entropy\-shift criterion, or does any sensible per\-position selector at matchedρ\\rhogive the same benefit? We compare our top\-Δ​Ht\\Delta H\_\{t\}\-masked selector against five alternatives at matchedρ=0\.7\\rho=0\.7on Qwen3\-30B\-A3B\-Instruct \(Table[4](https://arxiv.org/html/2609.38792#S4.T4)\):

- •Δ​Ht\\Delta H\_\{t\}, mask bottomρ\\rhofraction \(*direction flip*\): drop the positions with the lowest entropy shifts and retain the highest\-shift positions instead\. Tests whether the gain depends on the masking direction\.
- •\|Δ​Ht\|\|\\Delta H\_\{t\}\|, mask bottomρ\\rhofraction \(*both tails*\): drop positions with near\-zero entropy shifts and retain positions with large absolute entropy shifts, where the teacher and student entropies differ substantially\. Tests whether retaining positions with large absolute entropy shifts is sufficient, regardless of whether the shift is positive or negative\.
- •Random masking: mask a uniformly sampledρ\\rhofraction of response positions\. Tests whether the gain follows simply from reducing the number of supervised positions\.
- •Student\-entropy masking: mask the bottomρ\\rhofraction of positions by student next\-token entropy\. Tests whether retaining positions where the unconditioned student is uncertain is sufficient\.
- •Teacher\-entropy masking: mask the bottomρ\\rhofraction of positions by teacher next\-token entropy\. Tests whether retaining positions where the feedback\-conditioned teacher is uncertain is sufficient\.

Our top\-Δ​Ht\\Delta H\_\{t\}\-masked selector outperforms all five controls on overallRM\-BenchandRewardBench v2accuracy\. The gain is therefore not explained by the masking fraction alone, by either distribution’s entropy in isolation, by reversing the masking direction, or by retaining high\-\|Δ​Ht\|\|\\Delta H\_\{t\}\|positions of either sign\. Together, these ablations support selecting positions by signed entropy shift, with lower\-shift selection outperforming the tested alternatives\. The corresponding Qwen3\-4B\-Instruct results are reported in Appendix[C](https://arxiv.org/html/2609.38792#A3)\.

## 5Conclusion

We studied on\-policy dense supervision for training LLM judges on subjective tasks, where the verdict hinges on which evaluation criteria the judge invokes and how it weighs them\. Outcome\-supervised RL credits every token by the final verdict and does not provide explicit guidance on criterion choice; SD converts the per\-example language feedback that names the decisive criterion into per\-position supervision, but not every position carries equally useful signal\. Using the per\-position entropy shift between teacher and student, we identified two regimes, context sharpening and context spreading, and proposed a simple mask that retains the lower tail of the entropy\-shift distribution\. Across model sizes, SD outperforms outcome\-supervised RL on the evaluated subjective tasks, and our mask further improves generalization over naive SD\.

## 6Limitations and Scope

Our method has two main limitations\. First, the method depends on feedback quality: entropy\-shift masking does not guarantee useful supervision when the feedback fails to identify a decisive criterion or provides only generic guidance\. Second, our method does not explicitly address known limitations of self\-distillation, including hallucination and training instability in long\-chain\-of\-thought reasoning models\[[Kim et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib35)\]\. We focus on instruction\-tuned models to support efficient judge inference; extending the method to long\-chain\-of\-thought reasoning models remains future work\.

## References

- Anthropic \(2025\)AnthropicIntroducing Claude 4\.Note:[https://www\.anthropic\.com/news/claude\-4](https://www.anthropic.com/news/claude-4)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p5.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1)\.
- Anthropic \(2026a\)AnthropicIntroducing Claude Opus 4\.8\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-8](https://www.anthropic.com/news/claude-opus-4-8)Cited by:[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px3.p1.1)\.
- Anthropic \(2026b\)AnthropicIntroducing Claude Sonnet 5\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-5](https://www.anthropic.com/news/claude-sonnet-5)Cited by:[§4\.3](https://arxiv.org/html/2609.38792#S4.SS3.SSS0.Px1.p1.1)\.
- Baiet al\.\(2022\)Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.Constitutional AI: harmlessness from AI feedback\.arXiv preprint arXiv:2212\.08073\.External Links:[Link](https://arxiv.org/abs/2212.08073)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2026\)X\. Chen, G\. Li, Z\. Wang, B\. Jin, C\. Qian, Y\. Wang, H\. Wang, Y\. Zhang, D\. Zhang, T\. Zhang, H\. Tong, and H\. JiRM\-R1: reward modeling as reasoning\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2505.02387)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1)\.
- Cuiet al\.\(2025\)G\. Cuiet al\.The entropy mechanism of reinforcement learning for reasoning language models\.arXiv preprint arXiv:2505\.22617\.External Links:[Link](https://arxiv.org/abs/2505.22617)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p2.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-AIDeepSeek\-R1: incentivizing reasoning capability in LLMs via reinforcement learning\.Nature645,pp\. 633–638\.External Links:[Link](https://arxiv.org/abs/2501.12948)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p5.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1)\.
- Feng and Vaid \(2026\)R\. Feng and P\. VaidBringing capabilities in distribution via relevance\-masked self\-distillation\.Note:Applied Compute ResearchExternal Links:[Link](https://www.appliedcompute.com/research/relevance-masked-self-distillation)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Guoet al\.\(2025\)J\. Guo, Z\. Chi, L\. Dong, Q\. Dong, X\. Wu, S\. Huang, and F\. WeiReward reasoning model\.arXiv preprint arXiv:2505\.14674\.External Links:[Link](https://arxiv.org/abs/2505.14674)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1)\.
- Heet al\.\(2026\)Y\. He, H\. Wu, S\. Liu, H\. Ge, H\. Zhou, K\. Wu, Z\. Zheng, Q\. Lin, Z\. Zhong, and Y\. ZhangRethinking token\-level credit assignment in RLVR: a polarity\-entropy analysis\.arXiv preprint arXiv:2604\.11056\.External Links:[Link](https://arxiv.org/abs/2604.11056)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Honget al\.\(2025\)I\. Hong, C\. Yu, L\. Qiu, W\. Yan, Z\. Xu, H\. Jiang, Q\. Zhang, Q\. Lu, X\. Liu, C\. Zhang, and T\. ZhaoThink\-RM: enabling long\-horizon reasoning in generative reward models\.arXiv preprint arXiv:2505\.16265\.External Links:[Link](https://arxiv.org/abs/2505.16265)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.38792#S3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1)\.
- Hübotteret al\.\(2026\)J\. Hübotter, F\. Lübeck, L\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. Kleine Buening, C\. Guestrin, and A\. KrauseReinforcement learning via self\-distillation\.arXiv preprint arXiv:2601\.20802\.External Links:[Link](https://arxiv.org/abs/2601.20802)Cited by:[Appendix D](https://arxiv.org/html/2609.38792#A4.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.38792#S1.p3.1),[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2609.38792#S3.SS2.p2.1)\.
- Jiaet al\.\(2026\)N\. Jia, H\. Yang, X\. Ma, J\. Lian, S\. Zhang, W\. Zhang, K\. Zeng, X\. Cai, and Z\. SunAsymmetric on\-policy distillation: bridging exploitation and imitation at the token level\.arXiv preprint arXiv:2605\.06387\.External Links:[Link](https://arxiv.org/abs/2605.06387)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Jinet al\.\(2026\)W\. Jin, T\. Min, Y\. Yang, S\. R\. Kadhe, Y\. Zhou, D\. Wei, N\. Baracaldo, and K\. LeeEntropy\-aware on\-policy distillation of language models\.arXiv preprint arXiv:2603\.07079\.External Links:[Link](https://arxiv.org/abs/2603.07079)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2026\)J\. Kim, X\. Luo, M\. Kim, S\. Lee, D\. Kim, J\. Jeon, D\. Li, and Y\. YangWhy does self\-distillation \(sometimes\) degrade the reasoning capability of LLMs?\.arXiv preprint arXiv:2603\.24472\.External Links:[Link](https://arxiv.org/abs/2603.24472)Cited by:[§6](https://arxiv.org/html/2609.38792#S6.p1.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with PagedAttention\.InACM Symposium on Operating Systems Principles \(SOSP\),External Links:[Link](https://arxiv.org/abs/2309.06180)Cited by:[Appendix D](https://arxiv.org/html/2609.38792#A4.SS0.SSS0.Px1.p1.1)\.
- Leeet al\.\(2025\)Y\. Lee, J\. Boen, and C\. FinnFeedback descent: open\-ended text optimization via pairwise comparison\.arXiv preprint arXiv:2511\.07919\.External Links:[Link](https://arxiv.org/abs/2511.07919)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2024\)T\. Li, W\. Chiang, E\. Frick, L\. Dunlap, T\. Wu, B\. Zhu, J\. E\. Gonzalez, and I\. StoicaFrom crowdsourced data to high\-quality benchmarks: Arena\-Hard and BenchBuilder pipeline\.arXiv preprint arXiv:2406\.11939\.External Links:[Link](https://arxiv.org/abs/2406.11939)Cited by:[§4\.3](https://arxiv.org/html/2609.38792#S4.SS3.SSS0.Px1.p1.1)\.
- Liuet al\.\(2025a\)A\. Liu, H\. Bai, Z\. Lu, Y\. Sun, X\. Kong, S\. Wang, J\. Shan, A\. M\. Jose, X\. Liu, L\. Wen, P\. S\. Yu, and M\. CaoTIS\-DPO: token\-level importance sampling for direct preference optimization with estimated weights\.arXiv preprint arXiv:2410\.04350\.External Links:[Link](https://arxiv.org/abs/2410.04350)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025b\)T\. Liu, R\. Xu, T\. Yu, I\. Hong, C\. Yang, T\. Zhao, and H\. WangOpenRubrics: towards scalable synthetic rubric generation for reward modeling and LLM alignment\.arXiv preprint arXiv:2510\.07743\.External Links:[Link](https://arxiv.org/abs/2510.07743)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p3.1)\.
- Liuet al\.\(2026\)X\. Liu, X\. Wang, Y\. Ma, Y\. Zhang, and C\. XiaoWhen are teacher tokens reliable? Position\-Weighted on\-policy self\-distillation for reasoning\.arXiv preprint arXiv:2605\.21606\.External Links:[Link](https://arxiv.org/abs/2605.21606)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2025c\)Y\. Liu, Z\. Yao, R\. Min, Y\. Cao, L\. Hou, and J\. LiRM\-Bench: benchmarking reward models of language models with subtlety and style\.InInternational Conference on Learning Representations \(ICLR\),Note:OralExternal Links:[Link](https://arxiv.org/abs/2410.16184)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p5.1),[§3\.4](https://arxiv.org/html/2609.38792#S3.SS4.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px4.p1.1)\.
- Liuet al\.\(2025d\)Z\. Liu, C\. Chen, W\. Li, P\. Qi, T\. Pang, C\. Du, W\. S\. Lee, and M\. LinUnderstanding R1\-zero\-like training: a critical perspective\.arXiv preprint arXiv:2503\.20783\.External Links:[Link](https://arxiv.org/abs/2503.20783)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.38792#S3.SS1.p2.2)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/1711.05101)Cited by:[Appendix D](https://arxiv.org/html/2609.38792#A4.SS0.SSS0.Px2.p1.1)\.
- Lvet al\.\(2026\)O\. Lv, Y\. Zhang, and X\. ZhangGMTS: gradient magnitude\-based token selection improves RLVR training for LLM reasoning\.InOpenReview preprint,External Links:[Link](https://openreview.net/forum?id=JGvOicAo3g)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Welleck, B\. P\. Majumder, S\. Gupta, A\. Yazdanbakhsh, and P\. ClarkSelf\-refine: iterative refinement with self\-feedback\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:[Link](https://arxiv.org/abs/2303.17651)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Maliket al\.\(2025\)S\. Malik, V\. Pyatkin, S\. Land, J\. Morrison, N\. A\. Smith, H\. Hajishirzi, and N\. LambertRewardBench 2: advancing reward model evaluation\.arXiv preprint arXiv:2506\.01937\.External Links:[Link](https://arxiv.org/abs/2506.01937)Cited by:[§3\.4](https://arxiv.org/html/2609.38792#S3.SS4.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px4.p1.1)\.
- NVIDIA \(2025\)NVIDIALlama\-3\.3\-Nemotron\-Super\-49B\-GenRM\.Note:[https://huggingface\.co/nvidia/Llama\-3\_3\-Nemotron\-Super\-49B\-GenRM](https://huggingface.co/nvidia/Llama-3_3-Nemotron-Super-49B-GenRM)Cited by:[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1)\.
- OpenAI \(2026\)OpenAIIntroducing GPT\-5\.5\.Note:[https://openai\.com/index/introducing\-gpt\-5\-5/](https://openai.com/index/introducing-gpt-5-5/)Cited by:[§3\.3](https://arxiv.org/html/2609.38792#S3.SS3.SSS0.Px3.p2.1)\.
- Panget al\.\(2025\)J\. Pang, N\. Di, Z\. Zhu, J\. Wei, H\. Cheng, C\. Qian, and Y\. LiuToken cleaning: fine\-grained data selection for LLM supervised fine\-tuning\.InInternational Conference on Machine Learning \(ICML\),External Links:[Link](https://arxiv.org/abs/2502.01968)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Ruanet al\.\(2025\)Z\. Ruan, Y\. Li, H\. Zhu, Y\. Chen, P\. Li, Y\. Liu, and G\. ChenEnhancing large language model reasoning via selective critical token fine\-tuning\.arXiv preprint arXiv:2510\.10974\.External Links:[Link](https://arxiv.org/abs/2510.10974)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Scheureret al\.\(2023\)J\. Scheurer, J\. A\. Campos, T\. Korbak, J\. S\. Chan, A\. Chen, K\. Cho, and E\. PerezTraining language models with language feedback at scale\.InarXiv preprint arXiv:2303\.16755,External Links:[Link](https://arxiv.org/abs/2303.16755)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.External Links:[Link](https://arxiv.org/abs/2402.03300)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.38792#S3.SS1.p2.2)\.
- Shenet al\.\(2026a\)W\. F\. Shen, X\. Qiu, C\. Whitehouse, L\. Alazraki, S\. Goel, F\. Barbieri, T\. Willi, A\. Mathur, and I\. LeontiadisRethinking rubric generation for improving LLM judge and reward modeling for open\-ended tasks\.arXiv preprint arXiv:2602\.05125\.External Links:[Link](https://arxiv.org/abs/2602.05125)Cited by:[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px3.p1.1)\.
- Shenet al\.\(2026b\)Z\. Shen, J\. Hu, Z\. Qin, H\. Chen, W\. Ye, Z\. Huang, Y\. Zhuang, G\. Lu, J\. Zhou, and J\. ZhaoTraining\-trajectory\-aware token selection\.arXiv preprint arXiv:2601\.10348\.External Links:[Link](https://arxiv.org/abs/2601.10348)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Shenfeldet al\.\(2026\)I\. Shenfeld, M\. Damani, J\. Hübotter, and P\. AgrawalSelf\-distillation enables continual learning\.arXiv preprint arXiv:2601\.19897\.External Links:[Link](https://arxiv.org/abs/2601.19897)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Shenget al\.\(2025\)G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu, W\. Zhang, R\. Zhang, Y\. Peng, H\. Lin, and C\. WuHybridFlow: a flexible and efficient RLHF framework\.InEuropean Conference on Computer Systems \(EuroSys\),External Links:[Link](https://arxiv.org/abs/2409.19256)Cited by:[Appendix D](https://arxiv.org/html/2609.38792#A4.SS0.SSS0.Px1.p1.1)\.
- Wadhwaet al\.\(2024\)M\. Wadhwa, X\. Zhao, J\. J\. Li, and G\. DurrettLearning to refine with fine\-grained natural language feedback\.InFindings of the Association for Computational Linguistics: EMNLP,External Links:[Link](https://arxiv.org/abs/2407.02397)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2026a\)B\. Wang, Y\. Liu, Y\. Liu, T\. Tang, S\. Wang, C\. Gao, C\. Zheng, Y\. Zhang, L\. Yu, S\. Liu, T\. Gui, Q\. Zhang, X\. Huang, B\. Yu, F\. Huang, and J\. LinOutcome accuracy is not enough: aligning the reasoning process of reward models\.arXiv preprint arXiv:2602\.04649\.External Links:[Link](https://arxiv.org/abs/2602.04649)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.38792#S3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px5.p1.1),[Table 1](https://arxiv.org/html/2609.38792#S4.T1),[Table 1](https://arxiv.org/html/2609.38792#S4.T1.7)\.
- Wanget al\.\(2026b\)H\. Wang, L\. Wang, C\. Zhang, T\. Mao, S\. Qin, Q\. Lin, S\. Rajmohan, and D\. ZhangText2Grad: reinforcement learning from natural language feedback\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2505.22338)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2026c\)H\. Wang, G\. Wang, H\. Xiao, Y\. Zhou, Y\. Pan, J\. Wang, K\. Xu, Y\. Wen, X\. Ruan, X\. Chen, and H\. QiSkill\-SD: skill\-conditioned self\-distillation for multi\-turn LLM agents\.arXiv preprint arXiv:2604\.10674\.External Links:[Link](https://arxiv.org/abs/2604.10674)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025a\)S\. Wang, L\. Yu, C\. Gao, C\. Zheng, S\. Liu, R\. Lu, K\. Dang, X\. Chen, J\. Yang, Z\. Zhang, Y\. Liu, A\. Yang, A\. Zhao, Y\. Yue, S\. Song, B\. Yu, G\. Huang, and J\. LinBeyond the 80/20 rule: high\-entropy minority tokens drive effective RL for LLM reasoning\.arXiv preprint arXiv:2506\.01939\.External Links:[Link](https://arxiv.org/abs/2506.01939)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2025b\)Z\. Wang, J\. Zeng, O\. Delalleau, H\. Shin, F\. Soares, A\. Bukharin, E\. Evans, Y\. Dong, and O\. KuchaievHelpSteer3\-preference: open human\-annotated preference data across diverse tasks and languages\.arXiv preprint arXiv:2505\.11475\.External Links:[Link](https://arxiv.org/abs/2505.11475)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p3.1),[§3\.3](https://arxiv.org/html/2609.38792#S3.SS3.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2609.38792#S3.SS4.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.38792#S4.SS1.SSS0.Px2.p1.1)\.
- Whitehouseet al\.\(2025\)C\. Whitehouse, T\. Wang, P\. Yu, X\. Li, J\. Weston, I\. Kulikov, and S\. SahaJ1: incentivizing thinking in LLM\-as\-a\-judge via reinforcement learning\.arXiv preprint arXiv:2505\.10320\.External Links:[Link](https://arxiv.org/abs/2505.10320)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1)\.
- Xuet al\.\(2026a\)R\. Xu, T\. Liu, Z\. Dong, T\. Yu, I\. Hong, C\. Yang, L\. Zhang, T\. Zhao, and H\. WangAlternating reinforcement learning for rubric\-based reward modeling in non\-verifiable LLM post\-training\.arXiv preprint arXiv:2602\.01511\.External Links:[Link](https://arxiv.org/abs/2602.01511)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1)\.
- Xuet al\.\(2026b\)Y\. Xu, H\. Sang, Z\. Zhou, R\. He, Z\. Wang, and A\. GeramifardTIP: token importance in on\-policy distillation\.arXiv preprint arXiv:2604\.14084\.External Links:[Link](https://arxiv.org/abs/2604.14084)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2025\)A\. Yanget al\.Qwen3 Technical Report\.arXiv preprint arXiv:2505\.09388\.External Links:[Link](https://arxiv.org/abs/2505.09388)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p5.1),[§3\.3](https://arxiv.org/html/2609.38792#S3.SS3.SSS0.Px2.p1.1)\.
- Yanget al\.\(2026\)N\. Yang, H\. Lin, Y\. Liu, B\. Tian, G\. Liu, and H\. ZhangToken\-importance guided direct preference optimization\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2505.19653)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Yuet al\.\(2025\)Q\. Yuet al\.DAPO: an open\-source LLM reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.External Links:[Link](https://arxiv.org/abs/2503.14476)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p1.1),[§1](https://arxiv.org/html/2609.38792#S1.p2.1),[§3\.1](https://arxiv.org/html/2609.38792#S3.SS1.p2.2)\.
- Zenget al\.\(2025\)A\. Zenget al\.GLM\-4\.5: agentic, reasoning, and coding \(ARC\) foundation models\.arXiv preprint arXiv:2508\.06471\.External Links:[Link](https://arxiv.org/abs/2508.06471)Cited by:[§4\.3](https://arxiv.org/html/2609.38792#S4.SS3.SSS0.Px1.p1.1)\.
- Zenget al\.\(2024\)Y\. Zeng, G\. Liu, W\. Ma, N\. Yang, H\. Zhang, and J\. WangToken\-level direct preference optimization\.InInternational Conference on Machine Learning \(ICML\),External Links:[Link](https://arxiv.org/abs/2404.11999)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025a\)X\. Zhang, Y\. Zhang, H\. Sun, K\. Feng, C\. Lu, C\. Yang, and H\. MengCritique\-GRPO: advancing LLM reasoning with natural language and numerical feedback\.arXiv preprint arXiv:2506\.03106\.External Links:[Link](https://arxiv.org/abs/2506.03106)Cited by:[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025b\)Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. ZhouQwen3 Embedding: advancing text embedding and reranking through foundation models\.arXiv preprint arXiv:2506\.05176\.External Links:[Link](https://arxiv.org/abs/2506.05176)Cited by:[§3\.3](https://arxiv.org/html/2609.38792#S3.SS3.SSS0.Px4.p2.1),[§3\.4](https://arxiv.org/html/2609.38792#S3.SS4.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2026\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-distilled reasoner: on\-policy self\-distillation for large language models\.arXiv preprint arXiv:2601\.18734\.External Links:[Link](https://arxiv.org/abs/2601.18734)Cited by:[§1](https://arxiv.org/html/2609.38792#S1.p3.1),[§2](https://arxiv.org/html/2609.38792#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2609.38792#S3.SS2.p2.1)\.
- Zhaoet al\.\(2023\)Y\. Zhao, A\. Gu, R\. Varma, L\. Luo, C\. Huang, M\. Xu, L\. Wright, H\. Shojanazeri, M\. Ott, S\. Shleifer, A\. Desmaison, C\. Balioglu, P\. Damania, B\. Nguyen, G\. Chauhan, Y\. Hao, A\. Mathews, and S\. LiPyTorch FSDP: experiences on scaling fully sharded data parallel\.Proceedings of the VLDB Endowment16\(12\),pp\. 3848–3860\.External Links:[Link](https://arxiv.org/abs/2304.11277)Cited by:[Appendix D](https://arxiv.org/html/2609.38792#A4.SS0.SSS0.Px1.p1.1)\.

## Appendix AAdditional Context\-Sharpening and Context\-Spreading Examples

Figures[6](https://arxiv.org/html/2609.38792#A1.F6)and[7](https://arxiv.org/html/2609.38792#A1.F7)present two further rollouts each for context sharpening and context spreading, supplementing the two featured examples in the main body \(Figures[3](https://arxiv.org/html/2609.38792#S3.F3)and[4](https://arxiv.org/html/2609.38792#S3.F4)\)\. The plot layout and construction are identical to the main\-body figures\.

![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/30b_examples/30b_sharpening_appx.png)Figure 6:Additional context\-sharpening examples on twoHelpSteer3\-Preferencerollouts \(Qwen3\-30B\-A3B\-Instruct\-2507\)\.![Refer to caption](https://arxiv.org/html/2609.38792v1/figures/30b_examples/30b_spreading_appx.png)Figure 7:Additional context\-spreading examples on twoHelpSteer3\-Preferencerollouts \(Qwen3\-30B\-A3B\-Instruct\-2507\)\.
## Appendix BAdditional Experiments with Rubric Feedback

We repeat the Qwen3\-4B\-Instruct self\-distillation experiments using the generated per\-example rubrics instead of the original preference rationales as language feedbackzz\. Table[5](https://arxiv.org/html/2609.38792#A2.T5)compares naive SD and SD\+mask under both feedback formats, along with the same Dr\. GRPO outcome\-supervised baseline reported in the main results\.

Table 5:Judge accuracy \(%\) for Qwen3\-4B\-Instruct using preference rationales or generated per\-example rubrics as language feedback\. Dark gray \(in bold\) and light gray highlight the best and second\-best performance per column, respectively\.Entropy\-shift masking improves the total average under both feedback formats: from 75\.61 to 77\.65 with generated rubrics and from 77\.70 to 79\.04 with preference rationales\. This suggests that the benefit of entropy\-shift masking is not specific to one language\-feedback format\.

Second, the rationale\-based configuration achieves stronger overall performance than the rubric\-based configuration, despite using substantially shorter feedback\. One possible explanation is that the original rationales focus directly on the decisive reasons for the observed preferences, whereas generated rubrics introduce additional criteria that are less relevant to those preferences\.

These results support our practical choice of preference rationales: they are the most naturally available and easiest\-to\-collect textual feedback during preference\-data construction, require little additional annotation beyond the preference judgment itself, and outperform the substantially longer, separately generated rubrics in our experiments\.

## Appendix CAdditional Ablation Results

Table[6](https://arxiv.org/html/2609.38792#A3.T6)reports the full mask\-fraction sweep for Qwen3\-4B\-Instruct, complementing the Qwen3\-30B\-A3B\-Instruct results in Table[3](https://arxiv.org/html/2609.38792#S4.T3)\.

Table 6:ρ\\rho\-sweep on Qwen3\-4B\-Instruct, all masking the top\-ρ\\rhofraction byΔ​Ht\\Delta H\_\{t\}and training on the remaining bottom\(1−ρ\)\(1\-\\rho\)\.At 4B, the total average changes from 77\.70 with naive SD to 77\.53, 78\.52, and 79\.04 asρ\\rhoincreases to0\.30\.3,0\.50\.5, and0\.70\.7, respectively\. Thus, the canonicalρ=0\.7\\rho=0\.7setting is also the strongest configuration at the smaller model scale\.

Table[7](https://arxiv.org/html/2609.38792#A3.T7)reports the corresponding selector ablation at 4B\. As in the 30B results, every selector uses the same mask fraction,ρ=0\.7\\rho=0\.7\.

Table 7:Selector ablation on Qwen3\-4B\-Instruct, all at matched mask fractionρ=0\.7\\rho=0\.7\.At 4B, top\-Δ​Ht\\Delta H\_\{t\}masking obtains a total average of 79\.04, compared with 78\.35 for the direction\-flipped selector and at most 77\.80 for the\|Δ​Ht\|\|\\Delta H\_\{t\}\|, random, student\-entropy, and teacher\-entropy controls\. As at 30B, these results support selecting positions by signed entropy shift, with lower\-shift selection outperforming the tested alternatives\.

## Appendix DImplementation Details

##### Training framework\.

All runs use theverl\[[Sheng et al\., 2025](https://arxiv.org/html/2609.38792#bib.bib32)\]222Apache\-2\.0 license\.on\-policy RL framework with FSDP\[[Zhao et al\., 2023](https://arxiv.org/html/2609.38792#bib.bib30)\]for parameter/optimizer sharding and vLLM\[[Kwon et al\., 2023](https://arxiv.org/html/2609.38792#bib.bib33)\]333Apache\-2\.0 license\.for rollouts\.

##### Optimizer and schedule\.

We use AdamW\[[Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.38792#bib.bib34)\]with a constant learning\-rate schedule \(no warmup\)\.

##### Common training hyperparameters\.

The following are shared across all six training runs \(Qwen3\-4B\-Instruct and Qwen3\-30B\-A3B\-Instruct, each with Dr\. GRPO, naive SD, and SD\+mask\): train batch size256256\(mini\-batch and PPO mini\-batch both equal to256256\);33training epochs; same order of training samples across runs; max prompt length40964096, max response length81928192; per\-epoch A/B response\-order swap enabled\.

##### Per\-method hyperparameters\.

Dr\. GRPO uses learning rate1​e−61\\mathrm\{e\}\{\-\}6\(4B\) /2​e−62\\mathrm\{e\}\{\-\}6\(30B\) and88rollouts per prompt\. Naive SD and SD\+mask both use learning rate5​e−65\\mathrm\{e\}\{\-\}6\(4B\) /1​e−51\\mathrm\{e\}\{\-\}5\(30B\),44rollouts per prompt, andk=100k=100\. To reduce computational cost, we approximate the full\-vocabulary reverse KL using the teacher’s top\-kktokens and a tail\-remainder bucket\[[Hübotter et al\., 2026](https://arxiv.org/html/2609.38792#bib.bib1)\]\. Student and teacher entropies used to computeΔ​Ht\\Delta H\_\{t\}are calculated from their full\-vocabulary distributions\. The teacher parameters are maintained as an exponential moving average \(EMA\) of the student parameters, with an update rate of0\.010\.01, to stabilize training\.

##### Inference / evaluation decoding\.

At evaluation we sample one rollout per prompt with temperature0\.70\.7, top\-pp0\.80\.8, and top\-kk2020, following the official Qwen3\-Instruct model cards\. We use the prompt template in Appendix[E](https://arxiv.org/html/2609.38792#A5)\.

## Appendix EPrompt Templates

##### Input format\.

We send a single user message of the form shown in the templates below; no system prompt is used\. The judge expects each turn of the dialog and each candidate response to be wrapped with<user\>\.\.\.</user\>and<assistant\>\.\.\.</assistant\>tags\.

The placeholder\{context\}is the user\-side input: for a single\-turn query it is just one<user\>\.\.\.</user\>block; for a multi\-turn conversation it is the full alternating dialog \(which must alternate user, assistant, user,…\\ldotsand end on a<user\>turn\)\. The placeholders\{response\_a\}and\{response\_b\}are the two candidate replies, each wrapped in a single<assistant\>block\. The teacher prompt inserts the language feedback at\{feedback\}\. This field is omitted from the student and evaluation prompts\.

##### Rationale language\-feedback template\.

> You are an impartial judge tasked with determining which of two assistant responses is better for the given context\. Below is a context \(a user query or a conversation between the user and an assistant\) and two assistant responses to that context\. \[Start of Context\] \{context\} \[End of Context\] \[Start of Assistant A’s Response\] \{response\_a\} \[End of Assistant A’s Response\] \[Start of Assistant B’s Response\] \{response\_b\} \[End of Assistant B’s Response\] \{feedback\} Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions\. When correctness matters, solve the problem yourself and check each response for any errors\. After your analysis, determine which response is better overall and provide your final verdict \(A or B only\) in <verdict\>\.\.\.</verdict\>\.

##### Rubric language\-feedback template\.

The rubric\-feedback experiments use the same prompt, except for the final instruction paragraph:

> You are an impartial judge tasked with determining which of two assistant responses is better for the given context\. Below is a context \(a user query or a conversation between the user and an assistant\) and two assistant responses to that context\. \[Start of Context\] \{context\} \[End of Context\] \[Start of Assistant A’s Response\] \{response\_a\} \[End of Assistant A’s Response\] \[Start of Assistant B’s Response\] \{response\_b\} \[End of Assistant B’s Response\] \{feedback\} Identify the rubric that matters most for this specific task: the hard requirements the response must satisfy, ranked by importance, and the discriminative criteria that most decisively separate a better response from a worse one, ranked most decisive first, stating for each what makes a response better versus worse\. Then evaluate and compare the two assistant responses step by step against that rubric\. When correctness matters, solve the problem yourself and check each response for any errors\. After your analysis, determine which response is better overall and provide your final verdict \(A or B only\) in <verdict\>\.\.\.</verdict\>\.

##### Single\-turn example\.

> \[Start of Context\] <user\> What is the capital of France? </user\> \[End of Context\] \[Start of Assistant A’s Response\] <assistant\> The capital of France is Paris\. </assistant\> \[End of Assistant A’s Response\] \[Start of Assistant B’s Response\] <assistant\> Lyon\. </assistant\> \[End of Assistant B’s Response\]

##### Multi\-turn example\.

> \[Start of Context\] <user\> I’m planning a 3\-day trip to Tokyo next month\. Any recommendations? </user\> <assistant\> Sure, what kind of activities are you interested in \(food, history, nightlife, shopping\)? </assistant\> <user\> Mostly food and history\. </user\> \[End of Context\] \[Start of Assistant A’s Response\] <assistant\> Day 1: Tsukiji outer market for breakfast…\\ldots </assistant\> \[End of Assistant A’s Response\] \[Start of Assistant B’s Response\] <assistant\> Just go to Shibuya and figure it out when you get there\. </assistant\> \[End of Assistant B’s Response\]

相似文章

面向Lean定理证明的LLM反馈蒸馏

arXiv cs.AI

提出反馈蒸馏(Feedback Distillation),一种利用来自LLM的token级监督来改进复杂推理的训练方法,在Lean 4定理证明上进行了评估。该方法比GRPO更好地保持了多样性,且两种方法互补。