Weak-to-Strong On-Policy Distillation

arXiv cs.LG Papers

Summary

Introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), a framework that improves a strong language model by distilling from multiple weaker models using contrast pairs in logit space, consistently outperforming standard on-policy distillation on math and code benchmarks.

arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger teacher exists, or consolidate multiple domain experts trained from a shared base, which requires costly training at the student's scale. We introduce Weak-to-Strong On-Policy Distillation (W2S-OPD), a simple yet effective OPD framework that improves the strong student by distilling from multiple weak models. W2S-OPD constructs a proxy teacher in logit space from a contrast pair of a positive and a negative model, both smaller than the student and cheap to obtain. Their logit difference isolates the capability direction, which is added to the student's own base model, yielding a proxy teacher that couples this direction while staying distributionally adjacent to the student. The student then distills it by minimizing the per-token reverse KL on its own rollouts. We instantiate the contrast pair as i) a post-RL expert against its pre-RL initialization, isolating the skill RL instills, ii) a larger against a smaller base model, isolating the capability from scale, and iii) a small base model with correct versus wrong hints, isolating the instance-level direction toward the solution. Across four math and three code benchmarks, W2S-OPD outperforms OPD, enables the student to surpass the domain teacher, and keeps improving the student even when every supervision source is weaker. Analysis shows different contrasts yield distinct signals: the post-RL and hint contrasts emphasize reasoning frameworks, while the scale contrast emphasizes the solving procedure. Our code will be available at https://github.com/Yu-Fangxu/W2S-OPD.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:56 AM

# Weak-to-Strong On-Policy Distillation
Source: [https://arxiv.org/html/2607.26246](https://arxiv.org/html/2607.26246)
Fangxu Yu1, Zinan Lin2,Xiaodong Liu2,Weijia Xu2,Michael Xu2, Tianyi Zhou3,Jianfeng Gao2 1University of Maryland, College Park,2Microsoft Research,3MBZUAI

###### Abstract

On\-policy distillation \(OPD\), which aligns a student with the teacher’s token\-level distribution on the student’s own rollouts, has become an effective paradigm for transferring capabilities across large language models \(LLMs\)\. Prevailing approaches assume a teacher at least as capable as the student, and either distill a larger model into a smaller one, which fails at the frontier when no larger teacher exists, or train multiple domain experts from a shared base and consolidate them into one student, which requires costly training at the student’s scale\. To tackle these challenges, we introduce Weak\-to\-Strong On\-Policy Distillation \(W2S\-OPD\), a simple yet effective OPD framework that improves the strong student by distilling from multiple weak models\. Specifically, W2S\-OPD constructs a proxy teacher in logit space from a contrast pair of a positive and a negative model, both smaller than the student and cheap to obtain\. Their logit difference isolates the capability direction, which is then added to the student’s own base model\. The resulting proxy teacher thus couples this direction while staying distributionally adjacent to the student\. The student then distills it by minimizing the per\-token reverse KL on its own rollouts\. We instantiate the contrast pair as i\) a post\-RL expert against its pre\-RL initialization, isolating the skill RL instills, ii\) a larger against a smaller base model, isolating the capability from scale, and iii\) a small base model with correct and wrong hints, isolating the instance\-level direction toward the solution\. Across four math and three code benchmarks, W2S\-OPD consistently outperforms OPD and even enables the student to surpass the domain teacher and continues to improve the student when every supervision source is weaker\. Further analysis shows that different contrasts yield distinct learning signals: the post\-RL and hint contrast emphasizes reasoning frameworks, while the scale contrast emphasizes the solving procedure\. Our code will be available at[https://github\.com/Yu\-Fangxu/W2S\-OPD](https://github.com/Yu-Fangxu/W2S-OPD)\.

## 1Introduction

![Refer to caption](https://arxiv.org/html/2607.26246v1/x1.png)Figure 1:W2S\-OPD improves Qwen3\-8B using 4B models as teachers\.Accuracy is averaged over 4 math reasoning and 3 code generation benchmarks\. \(a\) With a post\-RL Qwen3\-4B expert, W2S\-OPD beats OPD\. \(b\) From 2 off\-the\-shelf base models \(Qwen3\-4B and 0\.6B\), both weaker than the student and used without training, W2S\-OPD still lifts the student above itself and both sources\.Reinforcement learning with verifiable rewards \(RLVR\)\(Guoet al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib6); Shaoet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib128); Yuet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib121)\)and knowledge distillation\(Guet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib71); Xiaoet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib67)\)are two dominant paradigms for improving the reasoning ability of large language models \(LLMs\)\. RLVR scales optimization directly from verifiable outcomes, yet its sparse reward provides limited fine\-grained supervision\. In contrast, knowledge distillation offers dense token\-level supervision from a teacher, but the learning from teacher\-generated off\-policy trajectories suffers from exposure bias\. On\-policy distillation \(OPD\)\(Agarwalet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib27); Lu and Lab,[2025](https://arxiv.org/html/2607.26246#bib.bib73)\)combines the complementary strengths of both, which supervises the student with a teacher’s token\-level fine\-grained supervision, yet on trajectories sampled from the student’s own policy\. This approach delivers dense credit assignment that alleviates exposure bias\. However, the premise is a teacher at least as capable as the student\.

This premise breaks down in the two ways current practice makes concrete\. The strong\-to\-weak paradigm distills a larger model into a smaller one\. For instance, Qwen3\(Yanget al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib26)\)distills a large\-scale model \(e\.g\.,235B\) into smaller variants \(e\.g\.,8B, 14B\)\. As models approach the frontier, however, no larger teacher exists to distill from, capping the performance ceiling\. A second line instead performs same\-origin consolidation: rather than assuming a single superior teacher, it trains multiple domain experts in parallel from a shared base model and distills their capabilities back into one student\. MOPD\(Maet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib57)\)instantiates this by applying per\-domain RL to obtain a set of domain teachers and merging them in policy space, while DeepSeek\-V4\(Xuet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib59)\)scales the same specialize\-then\-consolidate recipe to the trillion\-parameter regime\. Yet training multiple domain experts at the student’s scale is computationally expensive, inflating the cost of OPD\. We therefore move to the weak\-to\-strong paradigm\(Burnset al\.,[2023](https://arxiv.org/html/2607.26246#bib.bib28); Yanget al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib60)\), which induces supervisory signals from weaker models that already exist or can be trained inexpensively\. This enables a strong model to keep improving when no stronger teacher is available at a lower cost\.

However, directly applying OPD to weak\-to\-strong learning is challenging\. First, the weak teacher’s distribution is distinct from the student’s, and such a mismatch weakens the distillation signals\(Koet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib32); Liet al\.,[2026c](https://arxiv.org/html/2607.26246#bib.bib36)\)\. Second, since the teacher is weaker than the student, distillation toward it bounds the student at the teacher’s level and risks eroding the general capability it already possesses\.

To address these challenges, we propose Weak\-to\-Strong On\-Policy Distillation \(W2S\-OPD\), a simple yet effective OPD framework that distills the strong student from multiple weak models\. Figure[2](https://arxiv.org/html/2607.26246#S3.F2)provides an overview\. Specifically, W2S\-OPD constructs a proxy teacher in logit space from a contrast pair of a positive model and a negative model, both substantially smaller than the student and cheap to obtain\. Their logit difference isolates the capability direction, and W2S\-OPD adds this direction to the student’s own base model\. The proxy teacher therefore inherits the isolated direction while remaining distributionally close to the student\. The student then distills it by minimizing the per\-token reverse KL on its own rollouts\.

We propose three ways to construct the proxy teacher from weak models: i\)pre\-RL and post\-RL models, where the positive model is a domain expert trained by RL and the negative model is its pre\-RL initialization, whose difference isolates the skill that RL instills; since the RL need only be run at the small model’s scale, this transfers an expensive capability to the large student without ever running RL at the student’s scale; ii\)larger and smaller base models, two off\-the\-shelf base models of different sizes, whose difference isolates the capability that emerges from scale; this signal comes for free from models that already exist, requiring no new data, reward design, or training; and iii\)a single base model conditioned on correct and wrong hints, which extracts the instance\-level direction toward the solution from single model and provides token\-level supervision more efficiently\. We evaluate all three settings on four math reasoning and three code generation benchmarks\. Our main findings are threefold:

- •Surpassing the teacher it learns from\.In the pre\-RL / post\-RL setting, W2S\-OPD outperforms OPD by 11\.4% and 12\.0% relative on math reasoning under single\- and multi\-teacher distillation, and lifts the student above the domain expert itself\.
- •Improving from purely weaker models\.In the smaller/larger and contrastive\-hints settings, W2S\-OPD still improves the student even though every supervision source is weaker than it\.
- •Different contrasts reinforce complementary reasoning patterns\.The post\-RL and hint contrast places more emphasis on the tokens for reasoning framework \(e\.g\.,planning and monitoring the solving process\), whereas the scale contrast emphasizes the solving procedure\.

More broadly, W2S\-OPD reframes weak\-to\-strong learning as isolating and transferring a capability direction rather than imitating a weak supervisor, suggesting how frontier models can keep improving from abundant existing weak signals rather than waiting for a stronger teacher to be built\.

## 2Preliminary: On\-Policy Distillation

On\-policy distillation \(OPD\), formalized by the Generalized Knowledge Distillation \(GKD\) framework\(Agarwalet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib27)\), bridges reinforcement learning and imitation learning by providing dense, token\-level supervision on the student’s own trajectories\. Unlike off\-policy distillation that learns from a fixed teacher\-generated dataset, OPD lets the student policyπS\\pi\_\{S\}generate rollouts on\-policy, and a powerful frozen teacherπT\\pi\_\{T\}then scores every token of these self\-generated trajectories, against which the student minimizes a per\-token divergence:

ℒOPD\(πS\)=𝔼x∼𝒟,y∼πS\(⋅∣x\)\[∑t=1\|y\|D\(πS\(⋅∣st\)∥πT\(⋅∣st\)\)\],\\mathcal\{L\}\_\{\\mathrm\{OPD\}\}\(\\pi\_\{S\}\)\\;=\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\;y\\sim\\pi\_\{S\}\(\\cdot\\mid x\)\}\\bigg\[\\sum\_\{t=1\}^\{\|y\|\}D\\big\(\\pi\_\{S\}\(\\cdot\\mid s\_\{t\}\)\\,\\big\\\|\\,\\pi\_\{T\}\(\\cdot\\mid s\_\{t\}\)\\big\)\\bigg\],\(1\)wherest=\(x,y<t\)s\_\{t\}=\(x,y\_\{<t\}\)is the student\-generated prefix andDDis a per\-token divergence\. Because states are drawn from the student’s own rollouts, the teacher corrects the student precisely on the states the student actually visits, thereby avoiding the exposure bias of off\-policy distillation\. We instantiateDDas the reverse KL divergenceKL​\(πS∥πT\)\\mathrm\{KL\}\(\\pi\_\{S\}\\,\\\|\\,\\pi\_\{T\}\)\. In practice, we employ Top\-kkOPD for computational efficiency; see more details in Appendix[A](https://arxiv.org/html/2607.26246#A1)\.

Two prominent uses of OPD are strong\-to\-weak distillation and same\-origin consolidation\. The former compresses a larger, more capable teacher into a smaller student\(Yanget al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib26)\), while the latter distills several RL\-trained domain experts back into a single model, with all experts initialized from the student’s base so that their distributions stay close\(Xiaoet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib67); Maet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib57)\)\. Both presuppose a teacher at least as capable as the student, which becomes ineffective when training the frontier model, since no larger teacher exists to distill from, and training qualified experts at the student’s scale is computationally expensive\. These limitations motivate us to instead leverage small, cheaply obtained models to supervise a stronger student\.

## 3Weak\-to\-Strong On\-Policy Distillation

![Refer to caption](https://arxiv.org/html/2607.26246v1/x2.png)Figure 2:Overview of W2S\-OPD\. W2S\-OPD synthesizes a proxy teacher and distills it into a student\. A positive and a negative model form a contrast pair whose logit difference isolates a capability direction\. The studentπS\\pi\_\{S\}, initialized from the same base model, generates on\-policy rollouts and minimizes the per\-token reverse KL toward this proxy teacher\. The contrast pair can be instantiated as a post\-RL expert against its pre\-RL initialization, a larger base model against a smaller one, or a single model conditioned on correct against wrong hints\.Improving a strong student with supervision from weaker models presents two central challenges: i\) the distributional mismatch between the weak models and the student renders direct distillation ineffective\(Koet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib32); Liet al\.,[2026c](https://arxiv.org/html/2607.26246#bib.bib36)\); and ii\) since the supervision sources are weaker than the student, direct distillation constrains the student to a lower\-capability policy and degrades the general ability it already possesses\. To address both challenges, we adopt a directional perspective, as illustrated in Figure[2](https://arxiv.org/html/2607.26246#S3.F2)\. Instead of imitating a weak model directly, we isolate the capability direction that separates a stronger positive model from a weaker negative model, a signal that is largely disentangled from model scale and hence transferable across scales, and re\-anchor this direction onto the student’s base model\. The resulting proxy teacher is simultaneously equipped with the target domain capability and distributionally adjacent to the student, which stabilizes optimization and preserves the student’s general ability\. We detail the proxy teacher construction and distillation objective in §[3\.1](https://arxiv.org/html/2607.26246#S3.SS1), the instantiations of the contrast pair in §[3\.2](https://arxiv.org/html/2607.26246#S3.SS2), and its extension to multi\-teacher distillation in §[3\.3](https://arxiv.org/html/2607.26246#S3.SS3)\.

### 3\.1Proxy Teacher Synthesis and Distillation

Suppose the studentπS\\pi\_\{S\}is initialized from a strong base model, while the available weak models form a contrast pair: a positive modelm\+m^\{\+\}and a weaker negative modelm−m^\{\-\}\. Specifically,m\+m^\{\+\}is obtained either by applying domain\-specific RL to its initializationm−m^\{\-\}, or taken as a stronger base model thanm−m^\{\-\}\(we detail these instantiations in §[3\.2](https://arxiv.org/html/2607.26246#S3.SS2)\)\. W2S\-OPD constructs the teacher directly from these three frozen models, requiring only their output logits\.

The key observation behind W2S\-OPD is that, although bothm\+m^\{\+\}andm−m^\{\-\}are weaker than the strong base model, their difference still encodes a transferable capability direction\. What the two weak models agree on reflects their shared, limited ability, so subtracting their logits cancels this common component and retains precisely the direction along whichm\+m^\{\+\}improves overm−m^\{\-\}\. Adding this capability direction onto the strong base model therefore yields a proxy teacher that combines it with the student’s general strength while staying distributionally close to the student\. We instantiate this construction following decoding\-time experts\(Liuet al\.,[2021](https://arxiv.org/html/2607.26246#bib.bib62);[2024](https://arxiv.org/html/2607.26246#bib.bib63)\), wherem\+m^\{\+\}acts as the positive model whose logits are additively combined,m−m^\{\-\}acts as the negative model whose logits are negatively combined, and the student’s base model serves as the anchor\. Formally, at each positiontt, we condition the three models on the prefixsts\_\{t\}to obtain the logit scoreszbasez\_\{\\mathrm\{base\}\},z\+z^\{\+\}, andz−z^\{\-\}, and synthesize the proxy teacher as:

πT,α\(⋅∣st\)=softmax\(zbase\(st\)\+α\(z\+\(st\)−z−\(st\)\)\),\\pi\_\{T,\\alpha\}\(\\cdot\\mid s\_\{t\}\)\\;=\\;\\mathrm\{softmax\}\\Big\(\\,z\_\{\\mathrm\{base\}\}\(s\_\{t\}\)\\,\+\\,\\alpha\\big\(z^\{\+\}\(s\_\{t\}\)\-z^\{\-\}\(s\_\{t\}\)\\big\)\\Big\),\(2\)whereα≥0\\alpha\\geq 0is an amplification coefficient controlling the strength of the injected capability direction, and thus trades off signal strength against distributional proximity\. A smallα\\alphainjects little of the direction and keeps the proxy teacher close to the base model, yielding supervision too weak to drive improvement, whereas a largeα\\alphainjects a stronger signal but risks distorting the distribution and pushing the teacher away from the student\. Given the constructed proxy teacher, W2S\-OPD instantiates the OPD objective in Eq\.[1](https://arxiv.org/html/2607.26246#S2.E1)withπT=πT,α\\pi\_\{T\}=\\pi\_\{T,\\alpha\}:

ℒW2S\-OPD\(πS\)=𝔼x∼𝒟,y∼πS\(⋅∣x\)\[∑t=1\|y\|KL\(πS\(⋅∣st\)∥πT,α\(⋅∣st\)\)\],\\mathcal\{L\}\_\{\\text\{W2S\-OPD\}\}\(\\pi\_\{S\}\)\\;=\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\;y\\sim\\pi\_\{S\}\(\\cdot\\mid x\)\}\\bigg\[\\sum\_\{t=1\}^\{\|y\|\}\\mathrm\{KL\}\\big\(\\pi\_\{S\}\(\\cdot\\mid s\_\{t\}\)\\,\\big\\\|\\,\\pi\_\{T,\\alpha\}\(\\cdot\\mid s\_\{t\}\)\\big\)\\bigg\],\(3\)Taken together, W2S\-OPD address both challenges raised at the beginning of this section\. Anchoring the proxy teacher at the student’s own base model keeps the supervision target within a distribution the student already realizes \(withα\\alphabounding how far it moves\), and since the weak models enter only through their logit difference rather than their absolute level, distillation transfers their capability direction without pulling the student down toward them, which alleviates the capability ceiling\. Appendix[B](https://arxiv.org/html/2607.26246#A2)gives an alternative view of this objective as reward maximization under a KL trust region\.

### 3\.2Instantiations of the Positive and Negative Models

W2S\-OPD requires a contrast pair\(m\+,m−\)\(m^\{\+\},m^\{\-\}\)to construct the proxy teacher\. We propose three instantiations and summarize in Table[1](https://arxiv.org/html/2607.26246#S3.T1):

i\)Pre\-RL and post\-RL: the positive modelm\+m^\{\+\}is a domain expert obtained by applying RL to a small base model, and the negative modelm−m^\{\-\}is its pre\-RL initialization, whose difference isolates the domain skill acquired through RL\. This is a cheaper way of reducing the cost of training domain experts at the student’s own scale\.

ii\)Smaller and larger: the positive and negative models are two off\-the\-shelf base models of different sizes \(e\.g\.,Qwen3\-4B and Qwen3\-0\.6B\)\. Their difference isolates the capability that emerges purely from scale and requires no additional training\. This signal is already available in released models and can be used for free to further improve the capability of frontier models\.

iii\)Correct and wrong hints: the positive and negative models are a single base model conditioned on a correct and a wrong solution hint of the identical format, respectively\. Since the hint\-conditioning is shared, their difference cancels the style shift it induces\(Panet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib44)\)and isolates the instance\-level direction toward the correct solution\. Needing only one small model and a reference solution, which token\-level supervision efficiently\.

These instantiations indicate that W2S\-OPD can support various origins of the contrast, in which the pair exhibiting a separable capability gap can serve as the supervision source, since the construction of the proxy teacher requires only the output logits of the pair\.

### 3\.3Unification of Single\- and Multi\- Proxy Teacher Distillation

W2S\-OPD can naturally extend to multiple proxy teacher distillation scenarios\. GivenKKpositive models\{mk\+\}k=1K\\\{m^\{\+\}\_\{k\}\\\}\_\{k=1\}^\{K\}, each paired with a weaker negative modelmk−m^\{\-\}\_\{k\}, W2S\-OPD composes their capability directions on the shared base model:

πT,𝜶\(⋅∣st\)=softmax\(zbase\(st\)\+∑k=1Kαk\(zk\+\(st\)−zk−\(st\)\)\),\\pi\_\{T,\\bm\{\\alpha\}\}\(\\cdot\\mid s\_\{t\}\)\\;=\\;\\mathrm\{softmax\}\\Big\(\\,z\_\{\\mathrm\{base\}\}\(s\_\{t\}\)\\,\+\\,\\sum\_\{k=1\}^\{K\}\\alpha\_\{k\}\\big\(z^\{\+\}\_\{k\}\(s\_\{t\}\)\-z^\{\-\}\_\{k\}\(s\_\{t\}\)\\big\)\\Big\),\(4\)whereαk\\alpha\_\{k\}controls the strength of capability directionkk\. Since each subtraction isolates its own capability direction, the summation injects multiple domain skills into a single proxy teacher, and one distillation run with Eq\.[3](https://arxiv.org/html/2607.26246#S3.E3)merges them into a unified student\. For example,αk\\alpha\_\{k\}can be set to\{0,1\}\\\{0,1\\\}\. In this way, each query is routed to its most suitable positive and negative model pair for distillation\.

Table 1:The three instantiations of the contrast pair in W2S\-OPD\. All models are from Qwen3 series\.

## 4Experiments

We evaluate W2S\-OPD under the three contrast\-pair settings\. §[4\.2](https://arxiv.org/html/2607.26246#S4.SS2)reports the main results for each, §[4\.3](https://arxiv.org/html/2607.26246#S4.SS3)examines the factors behind its effectiveness, and §[4\.4](https://arxiv.org/html/2607.26246#S4.SS4)characterizes the learning signals the contrast provides\.

### 4\.1Experimental Setups

Benchmarks\.For math reasoning, we evaluate on four competition\-level benchmarks: AIME24\(AI\-MO,[2024](https://arxiv.org/html/2607.26246#bib.bib77)\), AIME25\(Zhang and Math\-AI,[2025](https://arxiv.org/html/2607.26246#bib.bib75)\), HMMT25 \(Feb\.\), and HMMT25 \(Nov\.\)\(Balunovicet al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib80)\)\. For code generation, we adopt HumanEval\+, MBPP\+\(Liuet al\.,[2023](https://arxiv.org/html/2607.26246#bib.bib81)\), and LiveCodeBench\-V6\(Jainet al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib79)\)\. To assess out\-of\-domain generalization, we further evaluate on GPQA\-Diamond\(Reinet al\.,[2023](https://arxiv.org/html/2607.26246#bib.bib61)\)for scientific reasoning and IFBench\(Pyatkinet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib56)\)for instruction following\.

Baselines and Evaluation Metrics\.We compare W2S\-OPD against OPD, which distills the student directly from the positive model under an identical training configuration\. As reference points, we additionally report the positive and negative models that form the contrast pair, as well as the student itself\. In all evaluations, we set the temperature to 1\.0 and top\-ppto 1\.0\. On each math reasoning benchmark, we sample 32 solutions per problem, whereas on each code generation benchmark, we sample 4 solutions per problem, and report the average accuracy over all samples on each benchmark for a robust evaluation\. We adopt Math\-Verify to validate answer correctness for math reasoning\.

Training details\.We use Qwen3\-8B\(Yanget al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib26)\)in non\-thinking mode as the studentπS\\pi\_\{S\}; a frozen copy of the same checkpoint serves as the anchor\. Both W2S\-OPD and OPD are trained for 100 steps under an identical configuration\. See Appendix[C](https://arxiv.org/html/2607.26246#A3)for more implementation details\.

Contrast pairs\.Table[1](https://arxiv.org/html/2607.26246#S3.T1)summarizes the contrast pair of each setting\. In the pre\-RL / post\-RL setting, the positive modelm\+m^\{\+\}is a Qwen3\-4B domain expert obtained by applying GRPO\(Shaoet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib128)\)to the Qwen3\-4B base, for 500 steps on the level\-6 subset of DeepMath\-103K\(Heet al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib76)\)for math reasoning and for 300 steps on Eurus\-RL\-Code\(Cuiet al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib74)\)for code generation, and the negative modelm−m^\{\-\}is its pre\-RL initialization,i\.e\.,the original Qwen3\-4B model\. In the smaller/larger setting, the positive and negative models are the off\-the\-shelf Qwen3\-4B and Qwen3\-0\.6B base models, requiring no additional training\. In the contrastive\-hint setting, the positive and negative models share the Qwen3\-4B base model, conditioned on a correct and a wrong solution hint of the identical format, respectively\.

### 4\.2Main Results

Table 2:Results for Pre\-RL / Post\-RL contrast setting\. W2S\-OPD beats OPD and even surpasses the 4B\-expert on math\.Improv\.reports the absolute gain over OPD\.Table 3:Results for contrastive hints setting\. W2S\-OPD improves the 8B student above its own base with only a single weaker and smaller model\.Improv\.reports the absolute gain of W2S\-OPD over the student base model\.i\) Distillation from Pre\-RL / Post\-RL\.Table[2](https://arxiv.org/html/2607.26246#S4.T2)compares W2S\-OPD against OPD under single\-teacher distillation, which learns from a single domain expert, and multi\-teacher distillation, which merges multiple domain experts \(here a math and a code expert\) into one student\. We have the following observations:\(a\) A small post\-RL expert substantially improves the large student\.W2S\-OPD consistently outperforms SFT and OPD across all benchmarks in the single\-teacher setting, with 11\.4% and 3\.7% average relative improvements on math and code, respectively\. Notably, on math reasoning W2S\-OPD surpasses the domain teacher itself, whereas OPD remains below it\. On code generation, W2S\-OPD likewise narrows the gap to the teacher\.\(b\) W2S\-OPD outperforms OPD when distilling from multiple domain teachers\.We conduct experiments in the multi\-teacher setting, where we aim to merge the capabilities from different domain teachers into a single student through OPD, which is obtained by applying domain\-specific RL to the same base model\. Specifically, we mix the math and coding training data and route each sample to the corresponding domain teacher through their domain labels\. The results in Table[2](https://arxiv.org/html/2607.26246#S4.T2)demonstrate that W2S\-OPD consistently leads to better performance than OPD on all benchmarks\.\(c\) W2S\-OPD learns faster and more stably\.Figure[3](https://arxiv.org/html/2607.26246#S4.F3)tracks benchmark performance over training and shows that W2S\-OPD outperforms OPD throughout the training\.

Table 4:Results for Smaller and Larger contrast setting\. W2S\-OPD improves the 8B student above its own base even though both source models are weaker than it\.Improv\.reports the absolute gain of W2S\-OPD over the student base model\.![Refer to caption](https://arxiv.org/html/2607.26246v1/x3.png)Figure 3:Performance on the math and code benchmarks over training steps\. W2S\-OPD improves faster and outperforms OPD\.ii\) Distillation from Smaller / Larger Base Models\.In this setting, we find thatTwo weak base models can improve the stronger student\.As shown in Table[4](https://arxiv.org/html/2607.26246#S4.T4), when the proxy teacher is constructed from Qwen3\-4B and Qwen3\-0\.6B, both weaker than the student, W2S\-OPD still improves the student by an absolute 6\.0% on math reasoning and 1\.2% on code generation on average\. This confirms that the inherent gap between two off\-the\-shelf weak models also encodes a transferable improving direction to distill from\.

iii\) Distillation from a Single Model with Correct / Wrong Hints\.As shown in Table[3](https://arxiv.org/html/2607.26246#S4.T3), a contrastive hint direction from a single weak model improves the student\. W2S\-OPD improves the student by an absolute 1\.4% on math reasoning and 1\.1% on code generation on average, despite the hint model being a 4B model weaker than the student\. This construction distills the privileged information and shows that a meaningful learning signal can also be extracted from a difference in context without requiring two different models\.

### 4\.3Training Analysis

Using the Pre\-RL / Post\-RL setting, we conduct two further analyses: the impact of the amplification coefficientα\\alphaon distillation, and out\-of\-domain generalization compared with OPD\.

#### 4\.3\.1Effect of the Amplification Coefficientα\\alpha

To investigate the impact of the amplification coefficient, we varyα\\alphawith all other configurations fixed and report the results in Figure[4](https://arxiv.org/html/2607.26246#S4.F4)\. Whenα\\alphais small,πT,α\\pi\_\{T,\\alpha\}stays close to the student’s own base model, where the capability direction is barely injected, the distillation signal is weak, and the distilled student falls behind OPD\. Asα\\alphaincreases, more domain skill is injected into the proxy teacher, and the student’s performance improves accordingly\. When theα\\alphacontinues increasing, performance starts to decline on most benchmarks, likely because an overly largeα\\alphadrives the proxy teacher further away from the student’s distribution, making the supervision increasingly hard for the student to learn\. Consequently, a moderateα\\alphastrikes the best balance between the strength of the injected signal and the adjacency of the teacher\.

![Refer to caption](https://arxiv.org/html/2607.26246v1/x4.png)Figure 4:Performance on the math and code benchmarks with differentα\\alpha\. OPD is included for reference, denoted by the gray line\. W2S\-OPD outperforms OPD over a wide range ofα\\alpha\.
#### 4\.3\.2Generalization to Out\-of\-Domain Tasks

Table 5:Results for OOD generalization on GPQA\-Diamond and IFBench\. Both distillation methods are trained only on the math task, W2S\-OPD transfers out of domain and improves general ability, whereas OPD can degrade it below the base\.Improv\.indicates the absolute gain over OPD\.A crucial concern for weak\-to\-strong distillation is whether learning from a small domain expert erodes the strong student’s general capability\. To investigate this, we evaluate the student trained in the Pre\-RL / Post\-RL setting on two out\-of\-domain benchmarks\. As shown in Table[5](https://arxiv.org/html/2607.26246#S4.T5), the distilled reasoning skill transfers well beyond the training domain\. W2S\-OPD lifts the student from 38\.9 to 56\.5 on GPQA\-Diamond, outperforming OPD by an absolute 2\.1 %\. The two models diverge on IFBench, where OPD degrades the student below its initialization, consistent with the student absorbing the small expert’s limitations, yet W2S\-OPD instead improves compared to the base model and outperforms OPD by 1\.1%\. These results indicate that W2S\-OPD transfers domain skill without sacrificing and even enhancing the student’s general ability\.

### 4\.4What Kinds of Tokens Are Most Strengthened?

Since all settings improve the student, we raise the question of what each capability direction actually provides\. We quantitatively analyze this question by recording each token’s offset between the positive and negative models on a shared solution generated by the student model\. At every positionttalong this trace, the capability direction in logit space isΔ​zt=z\+​\(st\)−z−​\(st\)\\Delta z\_\{t\}=z^\{\+\}\(s\_\{t\}\)\-z^\{\-\}\(s\_\{t\}\)\. To measure how strongly it reinforces the realized tokenxtx\_\{t\}, we take the log\-probability each model assigns toxtx\_\{t\}and define:

Δt=log⁡π\+​\(xt∣st\)−log⁡π−​\(xt∣st\),π±=softmax​\(z±\)\.\\Delta\_\{t\}\\;=\\;\\log\\pi^\{\+\}\(x\_\{t\}\\mid s\_\{t\}\)\\;\-\\;\\log\\pi^\{\-\}\(x\_\{t\}\\mid s\_\{t\}\),\\qquad\\pi^\{\\pm\}=\\mathrm\{softmax\}\(z^\{\\pm\}\)\.\(5\)
Tokens with largeΔt\\Delta\_\{t\}are those most strongly reinforced by the capability direction\. To identify which reasoning steps each contrast reinforces, we collect the top\-1% highest\-Δ\\Deltatokens across 200 math\-reasoning traces and label every token with one of eight problem\-solving episodes\(Schoenfeld,[2014](https://arxiv.org/html/2607.26246#bib.bib46)\)using ThinkARM\(Liet al\.,[2026b](https://arxiv.org/html/2607.26246#bib.bib47)\), an automatic episode classifier\. Table[6](https://arxiv.org/html/2607.26246#S4.T6)shows the token distribution\. Relative to the scale contrast, the post\-RL and hint contrasts place more weight on the reasoning*framework*such as thePlanandMonitorepisodes that structure and track the solution, whereas the scale contrast keeps more weight on the core solving steps,AnalyzeandImplement\. The hint contrast additionally concentrates on the finalAnswertokens, which is a consequence of conditioning on a correct versus a wrong answer\. Figure[5](https://arxiv.org/html/2607.26246#S4.F5)illustrates the same division on a single trace:the three contrasts reinforce different kinds of tokens\.These emphases are complementary, resulting in a student learning to improve the reasoning by distilling different patterns\.

Table 6:Distribution of the top\-1% highest\-Δ\\Deltatokens over the eight Schoenfeld episodes\. The first row gives the distribution of each episodes of all tokens\.Question:Calculate the line integral of1z\\frac\{1\}\{z\}over a contour that consists of a square and a circle, both centered at the origin, oriented counterclockwise\.\(answer4​π​i4\\pi i; wrong\-answer hint2​π​i2\\pi i\) We areintegrating∮C1z​𝑑z\\oint\_\{C\}\\frac\{1\}\{z\}dz⋯\\cdotsStep 1: Understand the function⋯\\cdots1/z1/zhas asingularityatz=0z\{=\}0, inside both⋯\\cdotsStep 2: Cauchy’s Theorem⋯\\cdotsHowever, it applies only to holomorphicff⋯\\cdotsButwe can use the IntegralFormula⋯\\cdotsStep 3: known result⋯\\cdotspositivelyoriented⇒2​π​i\\Rightarrow 2\\pi i⋯\\cdotsso the total=2​π​i\+2​π​i=4​π​i=2\\pi i\{\+\}2\\pi i=4\\pi i⋯\\cdotsFinal Answer:4​π​i4\\pi i\.\(a\) Pre\-RL / Post\-RLWe are integrating∮C1z​𝑑z\\oint\_\{C\}\\frac\{1\}\{z\}dz⋯\\cdotsStep 1: Understand thefunction⋯\\cdots1/z1/zhas a singularity atz=0z\{=\}0, insideboth⋯\\cdots Step 2: Cauchy’s Theorem⋯\\cdotsHowever, it applies only toholomorphicff⋯\\cdotsBut we canusethe Integral Formula⋯\\cdotsStep 3: known result⋯\\cdotspositively oriented⇒2​π​i\\Rightarrow 2\\pi i⋯\\cdotsso thetotal==2​π​i\+2​π​i2\\pi i\{\+\}2\\pi i=4​π​i=4\\pi i⋯\\cdotsFinal Answer:4​π​i4\\pi i\.\(b\) Smaller / Larger ModelsWe are integrating∮C1z​𝑑z\\oint\_\{C\}\\frac\{1\}\{z\}dz⋯\\cdotsStep 1: Understand the function⋯\\cdots1/z1/zhas a singularity atz=0z\{=\}0, inside both⋯\\cdotsStep 2: Cauchy’s Theorem⋯\\cdotsHowever, it applies only to holomorphicff⋯\\cdotsBut we can use the Integral Formula⋯\\cdotsStep 3: known result⋯\\cdotspositively oriented⇒2​π​i\\Rightarrow 2\\pi i⋯\\cdotsso thetotal=2​π​i\+2​π​i==2\\pi i\{\+\}2\\pi i=44π​i\\pi i⋯\\cdotsFinal Answer:4​π​i4\\pi i\.\(c\) Correct / Wrong Hints

Figure 5:A case study that the three contrasts strengthen different reasoning tokens\. The shade of color represents the relative magnitude ofΔt\\Delta\_\{t\}\.

## 5Related Work

On\-Policy Distillation\.OPD\(Lu and Lab,[2025](https://arxiv.org/html/2607.26246#bib.bib73); Song and Zheng,[2026](https://arxiv.org/html/2607.26246#bib.bib65); Agarwalet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib27)\)supervises the student on its own rollouts with a superior teacher\. It has become a key post\-training paradigm for strong\-to\-weak distillation\(Zenget al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib58); Yanget al\.,[2025](https://arxiv.org/html/2607.26246#bib.bib26)\)and for merging multi\-domain experts into one model\(Xiaoet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib67); Xuet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib59); Chenet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib38); Yanget al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib70)\), with follow\-ups relaxing its requirements through self\-distillation from privileged information\(Zhaoet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib40); Shenfeldet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib37); Hübotteret al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib39); Yeet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib35)\), stabilized optimization against the student–teacher gap\(Jinet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib31); Koet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib32); Xinget al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib34); Janget al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib33)\), and multimodal extensions\(Liet al\.,[2026a](https://arxiv.org/html/2607.26246#bib.bib24); Liuet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib30); Yoonet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib29)\)\. These works assume a teacher at least as capable as the student\. In contrast, W2S\-OPD works weak\-to\-strong, synthesizing a teacher from weak models rather than requiring a strong superior model\. Direct\-OPD\(Fenget al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib43)\)is a concurrent work that transfers a weak teacher’s pre\-/post\-RL log\-ratio as a dense reward\. However, its contrast is confined to RL\-trained teachers, without incorporating the directions from multiple contrastive sources and only evaluated on math reasoning tasks\.

Weak\-to\-Strong Generalization\.Weak\-to\-Strong elicits the capabilities of a stronger model with the supervision of a weak model\(Burnset al\.,[2023](https://arxiv.org/html/2607.26246#bib.bib28)\)\. This is critical when stronger model is hard to obtain\(Christianoet al\.,[2018](https://arxiv.org/html/2607.26246#bib.bib25)\)\. Recent work extends it to LLM reasoning and alignment\(Yaoet al\.,[2025b](https://arxiv.org/html/2607.26246#bib.bib72);[a](https://arxiv.org/html/2607.26246#bib.bib69); Yuanet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib68); Yanget al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib60); Zhaoet al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib66)\)\. Another line casts learning by teaching where a strong model improves by instructing weaker students and turning their comprehension into a training signal\(Ninget al\.,[2024](https://arxiv.org/html/2607.26246#bib.bib49); Cetinet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib48)\)\. However, such supervision primarily provides trajectory\-level signals, while W2S\-OPD instead distills a synthesized proxy teacher with token\-level dense supervision\.

Controllable Text Generation\.Controlling the outputs of an LLM has been widely studied\(Prabhumoyeet al\.,[2020](https://arxiv.org/html/2607.26246#bib.bib55)\)along two directions: training\-based methods finetune the model to elicit desired properties\(Keskaret al\.,[2019](https://arxiv.org/html/2607.26246#bib.bib54); Chanet al\.,[2020](https://arxiv.org/html/2607.26246#bib.bib53)\)but are computationally costly, whereas decoding\-time methods steer generation by shifting or composing output distributions at inference\(Liuet al\.,[2021](https://arxiv.org/html/2607.26246#bib.bib62); Krauseet al\.,[2021](https://arxiv.org/html/2607.26246#bib.bib50); Qinet al\.,[2022](https://arxiv.org/html/2607.26246#bib.bib52)\)\. W2S\-OPD adopts this decoding\-time composition to synthesize its proxy teacher, but turns it into a training signal that transfers capability from weak models to a strong student\.

## 6Conclusion

We introduce W2S\-OPD, a weak\-to\-strong on\-policy distillation framework that improves a strong student using only smaller or weaker models\. By isolating the capability direction between a contrast pair of weak models and re\-anchoring it onto the student’s base model, W2S\-OPD synthesizes a proxy teacher that couples the isolated capability while staying distributionally adjacent to the student\. Across math reasoning and code generation benchmarks, W2S\-OPD substantially outperforms OPD, surpasses the weak teacher itself, and remains effective when the contrast pair comes from RL training, model scale, or contrastive hints, while preserving out\-of\-domain capabilities\.

Discussion\.As models approach the frontier, the assumption of an ever\-stronger teacher no longer holds, and where the supervision for OPD should come from becomes an open question\. Weak supervision sources are abundant, yet how far they can push a stronger student before weak\-to\-strong supervision saturates, and how to elicit more informative signals from them, remain to be explored\. Future work may develop weak\-to\-strong learning into a sustained paradigm for OPD and post\-training more broadly, in which frontier models continue to improve from supervision sources cheaper and weaker than themselves\.

## References

- R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. Ramos Garea, M\. Geist, and O\. Bachem \(2024\)On\-policy distillation of language models: learning from self\-generated mistakes\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 21246–21263\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1),[§2](https://arxiv.org/html/2607.26246#S2.p1.2),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- AI\-MO \(2024\)AIMO validation AIME\.Note:[https://huggingface\.co/datasets/AI\-MO/aimo\-validation\-aime](https://huggingface.co/datasets/AI-MO/aimo-validation-aime)90 problems from AIME 2022–2024, extracted from the AoPS wikiCited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- M\. Balunovic, J\. Dekoninck, I\. Petrov, N\. Jovanovic, and M\. Vechev \(2025\)Matharena: evaluating llms on uncontaminated math competitions, february 2025\.URL https://matharena\. ai8\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- C\. Burns, P\. Izmailov, J\. H\. Kirchner, B\. Baker, L\. Gao, L\. Aschenbrenner, Y\. Chen, A\. Ecoffet, M\. Joglekar, J\. Leike,et al\.\(2023\)Weak\-to\-strong generalization: eliciting strong capabilities with weak supervision\.arXiv preprint arXiv:2312\.09390\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p2.1),[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- E\. Cetin, T\. Zhao, and Y\. Tang \(2026\)Reinforcement learning teachers of test time scaling\.Advances in Neural Information Processing Systems38,pp\. 107533–107567\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- A\. Chan, Y\. Ong, B\. Pung, A\. Zhang, and J\. Fu \(2020\)Cocon: a self\-supervised approach for controlled text generation\.arXiv preprint arXiv:2006\.03535\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- T\. Chen, J\. Ou, Z\. Liu, R\. Tang, J\. Liang, and H\. Li \(2026\)Counteraction\-aware multi\-teacher on\-policy distillation for general capability recovery with domain preservation\.arXiv preprint arXiv:2605\.27115\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- P\. Christiano, B\. Shlegeris, and D\. Amodei \(2018\)Supervising strong learners by amplifying weak experts\.arXiv preprint arXiv:1810\.08575\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- G\. Cui, L\. Yuan, Z\. Wang, H\. Wang, Y\. Zhang, J\. Chen, W\. Li, B\. He, Y\. Fan, T\. Yu,et al\.\(2025\)Process reinforcement through implicit rewards\.arXiv preprint arXiv:2502\.01456\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p4.2)\.
- S\. Feng, H\. Gao, H\. Chi, H\. Wu, Z\. Zhang, Z\. Jiang, B\. He, W\. Ma, Y\. Zhang, and H\. Zhou \(2026\)Weak\-to\-strong generalization via direct on\-policy distillation\.arXiv preprint arXiv:2607\.05394\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- Y\. Gu, L\. Dong, F\. Wei, and M\. Huang \(2024\)Minillm: knowledge distillation of large language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 32694–32717\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1)\.
- Z\. He, T\. Liang, J\. Xu, Q\. Liu, X\. Chen, Y\. Wang, L\. Song, D\. Yu, Z\. Liang, W\. Wang,et al\.\(2025\)Deepmath\-103k: a large\-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning\.arXiv preprint arXiv:2504\.11456\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p4.2)\.
- J\. Hübotter, F\. Lübeck, L\. Behric, A\. Baumann, M\. Bagatella, D\. Marta, I\. Hakimi, I\. Shenfeld, T\. K\. Buening, C\. Guestrin,et al\.\(2026\)Reinforcement learning via self\-distillation\.arXiv preprint arXiv:2601\.20802\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- N\. Jain, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. Stoica \(2025\)Livecodebench: holistic and contamination free evaluation of large language models for code\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 58791–58831\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- I\. Jang, J\. Yeom, J\. Yeo, H\. Lim, and T\. Kim \(2026\)Stable on\-policy distillation through adaptive target reformulation\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 42217–42227\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- W\. Jin, T\. Min, Y\. Yang, S\. R\. Kadhe, Y\. Zhou, D\. Wei, N\. Baracaldo, and K\. Lee \(2026\)Entropy\-aware on\-policy distillation of language models\.arXiv preprint arXiv:2603\.07079\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- N\. S\. Keskar, B\. McCann, L\. R\. Varshney, C\. Xiong, and R\. Socher \(2019\)Ctrl: a conditional transformer language model for controllable generation\.arXiv preprint arXiv:1909\.05858\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- J\. Ko, S\. Abdali, Y\. J\. Kim, T\. Chen, and P\. Cameron \(2026\)Scaling reasoning efficiently via relaxed on\-policy distillation\.arXiv preprint arXiv:2603\.11137\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p3.1),[§3](https://arxiv.org/html/2607.26246#S3.p1.1),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- B\. Krause, A\. D\. Gotmare, B\. McCann, N\. S\. Keskar, S\. Joty, R\. Socher, and N\. F\. Rajani \(2021\)Gedi: generative discriminator guided sequence generation\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 4929–4952\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- J\. Li, H\. Yin, H\. Xu, B\. Xu, W\. Tan, Z\. He, J\. Ju, Z\. Luo, and J\. Luan \(2026a\)Video\-opd: efficient post\-training of multimodal large language models for temporal video grounding via on\-policy distillation\.arXiv preprint arXiv:2602\.02994\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- M\. Li, C\. Fan, Y\. Cheng, S\. Feizi, and T\. Zhou \(2026b\)Schoenfeld’s anatomy of mathematical reasoning by language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 32773–32802\.Cited by:[§4\.4](https://arxiv.org/html/2607.26246#S4.SS4.p2.2)\.
- Y\. Li, Y\. Zuo, B\. He, J\. Zhang, C\. Xiao, C\. Qian, T\. Yu, H\. Gao, W\. Yang, Z\. Liu,et al\.\(2026c\)Rethinking on\-policy distillation of large language models: phenomenology, mechanism, and recipe\.arXiv preprint arXiv:2604\.13016\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p3.1),[§3](https://arxiv.org/html/2607.26246#S3.p1.1)\.
- A\. Liu, X\. Han, Y\. Wang, Y\. Tsvetkov, Y\. Choi, and N\. A\. Smith \(2024\)Tuning language models by proxy\.arXiv preprint arXiv:2401\.08565\.Cited by:[§3\.1](https://arxiv.org/html/2607.26246#S3.SS1.p2.11)\.
- A\. Liu, M\. Sap, X\. Lu, S\. Swayamdipta, C\. Bhagavatula, N\. A\. Smith, and Y\. Choi \(2021\)DExperts: decoding\-time controlled text generation with experts and anti\-experts\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 6691–6706\.Cited by:[§3\.1](https://arxiv.org/html/2607.26246#S3.SS1.p2.11),[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- J\. Liu, C\. S\. Xia, Y\. Wang, and L\. Zhang \(2023\)Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation\.Advances in neural information processing systems36,pp\. 21558–21572\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- R\. Liu, X\. Lv, G\. Li, X\. Zhu, Z\. Wang, Z\. Zhang, J\. Chen, Z\. Li, B\. Li, J\. Gao,et al\.\(2026\)Visual\-advantage on\-policy distillation for vision\-language models\.arXiv preprint arXiv:2605\.21924\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- K\. Lu and T\. M\. Lab \(2025\)On\-policy distillation\.Thinking Machines Lab: Connectionism\.Note:https://thinkingmachines\.ai/blog/on\-policy\-distillationExternal Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[Appendix A](https://arxiv.org/html/2607.26246#A1.p1.4),[§1](https://arxiv.org/html/2607.26246#S1.p1.1),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- W\. Ma, J\. Wei, L\. Zhao, H\. Zhang, B\. Xiao, L\. Li, Q\. Yang, B\. Gao, Y\. Wang, R\. Li,et al\.\(2026\)MOPD: multi\-teacher on\-policy distillation for capability integration in llm post\-training\.arXiv preprint arXiv:2606\.30406\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p2.1),[§2](https://arxiv.org/html/2607.26246#S2.p2.1)\.
- X\. Ning, Z\. Wang, S\. Li, Z\. Lin, P\. Yao, T\. Fu, M\. B\. Blaschko, G\. Dai, H\. Yang, and Y\. Wang \(2024\)Can llms learn by teaching for better reasoning? a preliminary study\.Advances in Neural Information Processing Systems37,pp\. 71188–71239\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- L\. Pan, S\. Tao, Y\. Zhai, L\. Zhang, Z\. Liu, B\. Ding, A\. Liu, and L\. Wen \(2026\)RLCSD: reinforcement learning with contrastive on\-policy self\-distillation\.arXiv preprint arXiv:2606\.11709\.Cited by:[§3\.2](https://arxiv.org/html/2607.26246#S3.SS2.p4.1)\.
- S\. Prabhumoye, A\. W\. Black, and R\. Salakhutdinov \(2020\)Exploring controllable text generation techniques\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 1–14\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- V\. Pyatkin, S\. Malik, V\. Graf, H\. Ivison, S\. Huang, P\. Dasigi, N\. Lambert, and H\. Hajishirzi \(2026\)Generalizing verifiable instruction following\.Advances in Neural Information Processing Systems38\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- L\. Qin, S\. Welleck, D\. Khashabi, and Y\. Choi \(2022\)Cold decoding: energy\-based constrained text generation with langevin dynamics\.Advances in Neural Information Processing Systems35,pp\. 9538–9551\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p3.1)\.
- D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman \(2023\)Gpqa: a graduate\-level google\-proof q&a benchmark\.arXiv preprint arXiv:2311\.12022\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- A\. H\. Schoenfeld \(2014\)Mathematical problem solving\.Elsevier\.Cited by:[§4\.4](https://arxiv.org/html/2607.26246#S4.SS4.p2.2)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo \(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.CoRRabs/2402\.03300\.External Links:[Link](https://doi.org/10.48550/arXiv.2402.03300),[Document](https://dx.doi.org/10.48550/ARXIV.2402.03300),2402\.03300Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1),[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p4.2)\.
- I\. Shenfeld, M\. Damani, J\. Hübotter, and P\. Agrawal \(2026\)Self\-distillation enables continual learning\.arXiv preprint arXiv:2601\.19897\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- M\. Song and M\. Zheng \(2026\)A survey of on\-policy distillation for large language models\.arXiv preprint arXiv:2604\.00626\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- B\. Xiao, B\. Xia, B\. Yang, B\. Gao, B\. Shen, C\. Zhang, C\. He, C\. Lou, F\. Luo, G\. Wang,et al\.\(2026\)Mimo\-v2\-flash technical report\.arXiv preprint arXiv:2601\.02780\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1),[§2](https://arxiv.org/html/2607.26246#S2.p2.1),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- X\. Xing, H\. Wang, B\. Gao, Z\. Li, and Y\. Tang \(2026\)Trust region on\-policy distillation\.arXiv preprint arXiv:2606\.01249\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.\(2026\)Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p2.1),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p2.1),[§2](https://arxiv.org/html/2607.26246#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p3.1),[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- S\. Yang, G\. Zhu, B\. Song, H\. Wang, M\. Xia, X\. Zheng, Y\. Ma, Z\. Chen, W\. Wang, and G\. Chen \(2026\)OPRD: on\-policy representation distillation\.arXiv preprint arXiv:2606\.06021\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- Y\. Yang, Y\. Ma, and P\. Liu \(2024\)Weak\-to\-strong reasoning\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 8350–8367\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p2.1),[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- W\. Yao, G\. Xu, H\. Tang, W\. Yang, D\. Di, Z\. Wang, and Y\. Liu \(2025a\)On weak\-to\-strong generalization and f\-divergence\.arXiv preprint arXiv:2506\.03109\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- W\. Yao, W\. Yang, Z\. Wang, Y\. Lin, and Y\. Liu \(2025b\)Revisiting weak\-to\-strong generalization in theory and practice: reverse kl vs\. forward kl\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 2860–2888\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- T\. Ye, L\. Dong, X\. Wu, S\. Huang, and F\. Wei \(2026\)On\-policy context distillation for language models\.arXiv preprint arXiv:2602\.12275\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- H\. S\. Yoon, E\. Yoon, J\. Jang, S\. Eom, J\. W\. Hong, M\. Hasegawa\-Johnson, Q\. Dai, C\. Luo, and C\. D\. Yoo \(2026\)Decomposed on\-policy distillation for vision\-language reasoning: steering gradients for visual grounding\.arXiv preprint arXiv:2606\.00564\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- F\. Yu, L\. Jiang, H\. Kang, S\. Hao, and L\. Qin \(2024\)Flow of reasoning: training llms for divergent reasoning with minimal examples\.arXiv preprint arXiv:2406\.05673\.Cited by:[§1](https://arxiv.org/html/2607.26246#S1.p1.1)\.
- Y\. Yuan, T\. Xiao, S\. Tao, X\. Wang, J\. Gao, B\. Ding, and B\. Xu \(2026\)Incentivizing strong reasoning from weak supervision\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7138–7156\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- A\. Zeng, X\. Lv, Z\. Hou, Z\. Du, Q\. Zheng, B\. Chen, D\. Yin, C\. Ge, C\. Huang, C\. Xie,et al\.\(2026\)Glm\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- Y\. Zhang and T\. Math\-AI \(2025\)American invitational mathematics examination \(aime\) 2025\.Cited by:[§4\.1](https://arxiv.org/html/2607.26246#S4.SS1.p1.1)\.
- S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. Grover \(2026\)Self\-distilled reasoner: on\-policy self\-distillation for large language models\.arXiv preprint arXiv:2601\.18734\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p1.1)\.
- X\. Zhao, X\. Yang, T\. Pang, C\. Du, L\. Li, Y\. Wang, and W\. Y\. Wang \(2024\)Weak\-to\-strong jailbreaking on large language models\.arXiv preprint arXiv:2401\.17256\.Cited by:[§5](https://arxiv.org/html/2607.26246#S5.p2.1)\.
- W\. Zhu, R\. Xie, R\. Wang, and P\. Liu \(2026\)Hybrid policy distillation for llms\.arXiv preprint arXiv:2604\.20244\.Cited by:[Appendix A](https://arxiv.org/html/2607.26246#A1.p1.4)\.

## Appendix AKL\-Divergence Estimation

Single\-token OPD\. An alternative estimator, adopted by several recent OPD implementations\(Lu and Lab,[2025](https://arxiv.org/html/2607.26246#bib.bib73); Zhuet al\.,[2026](https://arxiv.org/html/2607.26246#bib.bib41)\), extracts a single scalar of teacher information per sampled token: the log\-ratio on the sampled token is treated as a token\-level advantage,

At=log⁡πT​\(yt∣st\)−log⁡πSold​\(yt∣st\),A\_\{t\}\\;=\\;\\log\\pi\_\{T\}\(y\_\{t\}\\mid s\_\{t\}\)\\;\-\\;\\log\\pi\_\{S\}^\{\\mathrm\{old\}\}\(y\_\{t\}\\mid s\_\{t\}\),\(6\)and plugged into a clipped policy\-gradient update, whereπSold\\pi\_\{S\}^\{\\mathrm\{old\}\}denotes the policy that generated the rollouts\. This variant reuses the RL infrastructure, but conveys only 1 scalar per position and thus exhibits higher variance than the dense top\-KKobjective, which backpropagates the teacher’s distribution at every position through a directly differentiable divergence and requires no advantage estimation or clipping\. We therefore adopt the dense top\-KKobjective for both distillation arms in all experiments\.

Top\-KKOPD\. Materializing the teacher’s full distribution requires storing and transferring\|𝒱\|\|\\mathcal\{V\}\|logits \(roughly 150K for the Qwen3 family\) at every response position, which is prohibitively expensive to communicate across training workers\. Since the teacher distribution is highly peaked, its top\-KKtokens already carry the majority of probability mass at nearly all positions\. The teacher worker therefore ships only the top\-KKtoken indices and log\-probabilities per position, and the divergence in Eq\.[1](https://arxiv.org/html/2607.26246#S2.E1)is estimated on this truncated support:

ℒOPDtop\-​K​\(πS\)=𝔼x∼𝒟,y∼πS\(⋅∣x\)​\[∑t=1\|y\|KL​\(π¯S𝒦t∥π¯T𝒦t\)\],\\mathcal\{L\}^\{\\text\{top\-\}K\}\_\{\\mathrm\{OPD\}\}\(\\pi\_\{S\}\)\\;=\\;\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\;y\\sim\\pi\_\{S\}\(\\cdot\\mid x\)\}\\bigg\[\\sum\_\{t=1\}^\{\|y\|\}\\mathrm\{KL\}\\big\(\\bar\{\\pi\}\_\{S\}^\{\\mathcal\{K\}\_\{t\}\}\\,\\big\\\|\\,\\bar\{\\pi\}\_\{T\}^\{\\mathcal\{K\}\_\{t\}\}\\big\)\\bigg\],\(7\)where𝒦t\\mathcal\{K\}\_\{t\}denotes the top\-KKsupport of the teacher distribution at positiontt, andπ¯𝒦t\\bar\{\\pi\}^\{\\mathcal\{K\}\_\{t\}\}denotes the corresponding distribution restricted to𝒦t\\mathcal\{K\}\_\{t\}and renormalized\. The truncation error is bounded by the tiny probability mass that falls outside the teacher’s top\-KKset, so the estimation preserves the supervision almost losslessly while reducing the cost\.

## Appendix BFrom logit arithmetic to a KL\-constrained objective

Letr​\(s,a\):=log⁡π\+​\(a∣s\)−log⁡π−​\(a∣s\)r\(s,a\):=\\log\\pi^\{\+\}\(a\\mid s\)\-\\log\\pi^\{\-\}\(a\\mid s\)denote the*capability reward*carried by the contrast pair: a token\-level, relative quantity measuring how much more the positive model favorsaathan the negative one does, in which the absolute competence of either source never appears\. Sinceez\+​\(a\)−z−​\(a\)e^\{z^\{\+\}\(a\)\-z^\{\-\}\(a\)\}equalsπ\+​\(a∣s\)/π−​\(a∣s\)\\pi^\{\+\}\(a\\mid s\)/\\pi^\{\-\}\(a\\mid s\)up to a factor independent ofaa, which the softmax normalizer absorbs, Eq\.[2](https://arxiv.org/html/2607.26246#S3.E2)is an exponential tilting of the student’s own base model,

πT,α​\(a∣s\)=πbase​\(a∣s\)​eα​r​\(s,a\)Zα​\(s\),Zα​\(s\)=𝔼a∼πbase\(⋅∣s\)​\[eα​r​\(s,a\)\]\.\\pi\_\{T,\\alpha\}\(a\\mid s\)=\\frac\{\\pi\_\{\\mathrm\{base\}\}\(a\\mid s\)\\,e^\{\\alpha r\(s,a\)\}\}\{Z\_\{\\alpha\}\(s\)\},\\qquad Z\_\{\\alpha\}\(s\)=\\mathbb\{E\}\_\{a\\sim\\pi\_\{\\mathrm\{base\}\}\(\\cdot\\mid s\)\}\\big\[e^\{\\alpha r\(s,a\)\}\\big\]\.\(8\)Substituting Eq\.[8](https://arxiv.org/html/2607.26246#A2.E8)intoKL​\(q∥πT,α\)\\mathrm\{KL\}\(q\\,\\\|\\,\\pi\_\{T,\\alpha\}\)and rearranging gives, for everyq∈Δ​\(𝒱\)q\\in\\Delta\(\\mathcal\{V\}\),

𝔼a∼q\[r\(s,a\)\]−1αKL\(q∥πbase\(⋅∣s\)\)⏟𝒥α​\(q;s\)=1αlogZα\(s\)−1αKL\(q∥πT,α\(⋅∣s\)\)\.\\underbrace\{\\mathbb\{E\}\_\{a\\sim q\}\\big\[r\(s,a\)\\big\]\-\\tfrac\{1\}\{\\alpha\}\\mathrm\{KL\}\\big\(q\\,\\\|\\,\\pi\_\{\\mathrm\{base\}\}\(\\cdot\\mid s\)\\big\)\}\_\{\\textstyle\\mathcal\{J\}\_\{\\alpha\}\(q\\,;s\)\}\\;=\\;\\tfrac\{1\}\{\\alpha\}\\log Z\_\{\\alpha\}\(s\)\\;\-\\;\\tfrac\{1\}\{\\alpha\}\\mathrm\{KL\}\\big\(q\\,\\\|\\,\\pi\_\{T,\\alpha\}\(\\cdot\\mid s\)\\big\)\.\(9\)The first term on the right does not depend onqqand the second is nonnegative and vanishes only atq=πT,αq=\\pi\_\{T,\\alpha\}\. The proxy teacher is therefore the unique closed\-form maximizer

πT,α\(⋅∣s\)=argmaxq∈Δ​\(𝒱\)\{𝔼a∼q​\[r​\(s,a\)\]⏟inject capability−1αKL\(q∥πbase\(⋅∣s\)\)⏟stay adjacent to the student\},\\pi\_\{T,\\alpha\}\(\\cdot\\mid s\)\\;=\\;\\arg\\max\_\{q\\in\\Delta\(\\mathcal\{V\}\)\}\\Big\\\{\\underbrace\{\\mathbb\{E\}\_\{a\\sim q\}\\big\[r\(s,a\)\\big\]\}\_\{\\text\{inject capability\}\}\\;\-\\;\\tfrac\{1\}\{\\alpha\}\\underbrace\{\\mathrm\{KL\}\\big\(q\\,\\\|\\,\\pi\_\{\\mathrm\{base\}\}\(\\cdot\\mid s\)\\big\)\}\_\{\\text\{stay adjacent to the student\}\}\\Big\\\},\(10\)where1/α1/\\alphais the strength of a trust region centered at the student rather than at any weak model\.

A compositional view\.Eq\.[8](https://arxiv.org/html/2607.26246#A2.E8)can be stated without reference to logits at all,

πT,α​\(a∣s\)∝πbase​\(a∣s\)​\(π\+​\(a∣s\)π−​\(a∣s\)\)α,\\pi\_\{T,\\alpha\}\(a\\mid s\)\\;\\propto\\;\\pi\_\{\\mathrm\{base\}\}\(a\\mid s\)\\,\\Big\(\\frac\{\\pi^\{\+\}\(a\\mid s\)\}\{\\pi^\{\-\}\(a\\mid s\)\}\\Big\)^\{\\alpha\},\(11\)which exhibits the proxy teacher as a product of experts over the shared vocabulary: the student’s base model is the anchor,π\+\\pi^\{\+\}enters as a conjoined expert andπ−\\pi^\{\-\}as a negated one, andα\\alphais the exponent controlling how sharply the pair is applied\. The capability direction is thus a*negation*operation, which divides out what the two weak models agree on and the trust region in Eq\.[10](https://arxiv.org/html/2607.26246#A2.E10)is precisely the price of keeping the resulting product near the anchor, with1α​log⁡Zα​\(s\)\\tfrac\{1\}\{\\alpha\}\\log Z\_\{\\alpha\}\(s\)its free energy atss\. WithKKcontrast pairs the same reading gives:

πT,𝜶​\(a∣s\)∝πbase​\(a∣s\)​∏k=1K\(πk\+​\(a∣s\)πk−​\(a∣s\)\)αk,\\pi\_\{T,\\bm\{\\alpha\}\}\(a\\mid s\)\\;\\propto\\;\\pi\_\{\\mathrm\{base\}\}\(a\\mid s\)\\prod\_\{k=1\}^\{K\}\\Big\(\\frac\{\\pi^\{\+\}\_\{k\}\(a\\mid s\)\}\{\\pi^\{\-\}\_\{k\}\(a\\mid s\)\}\\Big\)^\{\\alpha\_\{k\}\},\(12\)so merging capability directions is multiplication of experts\. In particular the result is invariant to the order in which the pairs are introduced, unlike sequentially fine\-tuning or distilling one teacher after another, where the outcome depends on the schedule and earlier skills may be overwritten\.

Three design choices follow: i\)*Adjacency*: the reference measure in Eq\.[10](https://arxiv.org/html/2607.26246#A2.E10)is the student’s own base model, so the proxy teacher is a reweighting of a distribution the student already realizes, andα\\alphasets how far it may move;α→0\\alpha\\\!\\to\\\!0recoversπbase\\pi\_\{\\mathrm\{base\}\}and supplies no supervision, whereasα→∞\\alpha\\\!\\to\\\!\\inftyremoves the trust region and collapses ontoarg⁡maxa⁡r​\(s,a\)\\arg\\max\_\{a\}r\(s,a\), so a moderateα\\alphais optimal \(Figure[4](https://arxiv.org/html/2607.26246#S4.F4)\)\. ii\)*Contrast*: the weak models enter only through the ratiorr, so a pair with no capability gap leaves the anchor untouched and the student is never ceilinged at the sources, unlike direct OPD whose optimum isπ\+\\pi^\{\+\}itself\. iii\)*Composition*: rewards add where policies do not, so Eq\.[4](https://arxiv.org/html/2607.26246#S3.E4)is the same maximizer with composite reward∑kαk​rk\\sum\_\{k\}\\alpha\_\{k\}r\_\{k\}, andKKcapability directions are merged by a single distillation run\.

## Appendix CImplementation Details

All experiments are implemented on top of the verl framework, and each distillation run is conducted on 2 NVIDIA B200 GPUs\. Table[8](https://arxiv.org/html/2607.26246#A3.T8)lists the training hyperparameters shared by W2S\-OPD and OPD\. For the proxy teacher, the divergence is estimated on the teacher’s top\-KKsupport withK=32K=32\(Appendix[A](https://arxiv.org/html/2607.26246#A1)\)\. Table[8](https://arxiv.org/html/2607.26246#A3.T8)reports the GRPO recipe used to obtain domain experts in the pre\-RL / post\-RL setting, and Table[9](https://arxiv.org/html/2607.26246#A3.T9)shows the rollout instructions used during OPD training\.

Table 7:Training hyperparameters of W2S\-OPD\. Paired entries denote math / code\.
Table 8:Training hyperparameters of the domain experts \(RL\-trained positive models\)\.

Table 9:On\-policy rollout prompts used during OPD training\.
## Appendix DAdditional Experimental Results

### D\.1Runtime Analysis

Table 10:Average wall\-clock time per training step \(s\) in the pre\-RL / post\-RL setting; W2S\-OPD adds only 20% over OPD\.Training efficiency matters for practical deployment\. To quantify the overhead W2S\-OPD introduces, we record the average wall\-clock time per training step in the pre\-RL / post\-RL setting and report it in Table[10](https://arxiv.org/html/2607.26246#A4.T10)\. W2S\-OPD incurs only a 20% increase in per\-step time over OPD, despite forwarding 3 frozen models \(the 8B anchor and the 4B contrast pair\) instead of a single 4B teacher\. The overhead stays modest because the per\-step cost is dominated by on\-policy rollout generation, which the two methods share, while the additional teacher\-side scoring is forward\-only\. These scoring passes are also independent of one another and can be further parallelized across devices\.

Table 11:Average math reasoning accuracy of W2S\-OPD with base\-model contrast pairs of different capability gaps\. The student is Qwen3\-8B\.
### D\.2Performance with Different Scales of Base Models

To investigate how the choice of base models affects the Smaller / Larger setting, we fix the Qwen3\-8B student and the Qwen3\-4B positive model, and vary the negative model between Qwen3\-0\.6B and Qwen3\-1\.7B\. As shown in Table[11](https://arxiv.org/html/2607.26246#A4.T11), both pairs improve the student, and the gain grows with the capability gap within the pair: the wider 4B−\-0\.6B pair yields the larger improvement of 6\.0 % on math reasoning, whereas the narrower 4B−\-1\.7B pair yields 4\.7\. We attribute this to the amount of capability the subtraction isolates, where a larger gap spans a longer stretch of the direction along which capability grows with scale, so the difference extracts a stronger and more informative signal, and re\-anchoring it onto the student base provides supervision that pushes the student further\. These results suggest instantiating W2S\-OPD with the most separated base\-model pair available\.

## Appendix EExample Generations

We provide qualitative examples shown in Figures[E](https://arxiv.org/html/2607.26246#A5)–[E](https://arxiv.org/html/2607.26246#A5)\. Figures[E](https://arxiv.org/html/2607.26246#A5)and[E](https://arxiv.org/html/2607.26246#A5)come from the pre\-RL / post\-RL setting and compare W2S\-OPD against OPD; Figures[E](https://arxiv.org/html/2607.26246#A5)and[E](https://arxiv.org/html/2607.26246#A5)come from the smaller/larger setting and compare the Qwen3\-8B student before and after W2S\-OPD distillation\. In each case the W2S\-OPD student succeeds while the baseline fails, with the baseline’s errors marked in red\.

Case Study: Math Reasoning \(Pre\-RL / Post\-RL Setting\)Question:Letppbe the least prime number for which there exists a positive integernnsuch thatn4\+1n^\{4\}\+1is divisible byp2p^\{2\}\. Find the least positive integermmsuch thatm4\+1m^\{4\}\+1is divisible byp2p^\{2\}\.Correct Answer:110Direct OPD:We need to find the least primeppsuch that there exists a positive integernnwithn4\+1n^\{4\}\+1divisible byp2p^\{2\}\. Then, once we find thatpp, we need to find the smallest positive integermmsuch thatm4\+1m^\{4\}\+1is divisible byp2p^\{2\}\.Step 1: Understand the problem\.We are looking for the smallest primeppsuch that there exists an integernnwheren4\+1≡0\(modp2\)n^\{4\}\+1\\equiv 0\\pmod\{p^\{2\}\},i\.e\.,n4≡−1\(modp2\)n^\{4\}\\equiv\-1\\pmod\{p^\{2\}\}\. Then, once we find suchpp, we need to find the minimal positive integermmsatisfying the same congruence\.Step 2: Find the minimal primepp\.To solve this congruence, we need to know when−1\-1is a quartic residue\. Let’s try small primes:Tryp=2p=2:we needn4\+1n^\{4\}\+1divisible by 4\.n=1n=1:22;n=2n=2:1717;n=3n=3:8282;n=4n=4:257≡1\(mod4\)257\\equiv 1\\pmod\{4\}— none divisible by 4\. No solution\.Tryp=3p=3:we needn4≡−1≡8\(mod9\)n^\{4\}\\equiv\-1\\equiv 8\\pmod\{9\}\. Checkingn=0,…,8n=0,\\dots,8:0,1,7,0,4,4,0,7,10,1,7,0,4,4,0,7,1\(e\.g\.,52=25≡75^\{2\}=25\\equiv 7, so54≡72=49≡45^\{4\}\\equiv 7^\{2\}=49\\equiv 4;8≡−18\\equiv\-1, so84≡18^\{4\}\\equiv 1\)\. None give 8\. No solution\.Tryp=5p=5:check mod 5 first:x4mod5x^\{4\}\\bmod 5forx=0,…,4x=0,\\dots,4gives0,1,1,1,10,1,1,1,1, so allx4≡0x^\{4\}\\equiv 0or1\(mod5\)1\\pmod\{5\}, and−1≡4\-1\\equiv 4is not a quartic residue\. Not possible\.Tryp=7p=7:x4mod7x^\{4\}\\bmod 7takes values\{0,1,2,4\}\\\{0,1,2,4\\\}\(24=16≡22^\{4\}=16\\equiv 2;34=81≡43^\{4\}=81\\equiv 4;54≡25^\{4\}\\equiv 2;6≡−16\\equiv\-1gives 1\)\.−1≡6\-1\\equiv 6is not among them\. No solution\.Tryp=11p=11:computingx2x^\{2\}thenx4x^\{4\}mod 11 forx=1,…,10x=1,\\dots,10gives quartic residues\{1,3,4,5,9\}\\\{1,3,4,5,9\\\};−1≡10\-1\\equiv 10is not among them\. No solution\.Tryp=13p=13:−1≡12\-1\\equiv 12\. Checking:34≡33^\{4\}\\equiv 3;44≡94^\{4\}\\equiv 9;52=25≡125^\{2\}=25\\equiv 12,54≡122=144≡15^\{4\}\\equiv 12^\{2\}=144\\equiv 1;64≡96^\{4\}\\equiv 9;74≡97^\{4\}\\equiv 9;82=64≡128^\{2\}=64\\equiv 12,84≡18^\{4\}\\equiv 1;94≡99^\{4\}\\equiv 9;104≡310^\{4\}\\equiv 3;114≡311^\{4\}\\equiv 3;124≡112^\{4\}\\equiv 1\. Noxxwithx4≡12x^\{4\}\\equiv 12\. No solution\.Maybe we can think about orders: in the cyclic group modpp, the equationx4≡−1x^\{4\}\\equiv\-1has solutions iff−1\-1lies in the image of the 4th\-power map…this is getting complex\. Let me try concrete examples\.Tryp=17p=17:tryx=2x=2:24=16≡−1\(mod17\)2^\{4\}=16\\equiv\-1\\pmod\{17\}\. Yes\! So24≡−1\(mod17\)2^\{4\}\\equiv\-1\\pmod\{17\}, andp=17p=17is such a prime\. Since we checkedp=2,3,5,7,11,13p=2,3,5,7,11,13and found no solutions, and there is no prime between 13 and 17,p=17p=17is the smallest such prime\.Step 3: Find the leastmmwithm4\+1≡0\(mod289\)m^\{4\}\+1\\equiv 0\\pmod\{289\}\.We know24≡−1\(mod17\)2^\{4\}\\equiv\-1\\pmod\{17\}, but24\+1=172^\{4\}\+1=17is divisible by 17, not by17217^\{2\}\. So we lift with Hensel’s lemma\. Writem=2\+17​km=2\+17k; modulo17217^\{2\}, terms with\(17​k\)2\(17k\)^\{2\}vanish, so\(2\+17​k\)4≡24\+4⋅23⋅\(17​k\)=16\+32⋅17​k\(mod289\)\.\(2\+17k\)^\{4\}\\equiv 2^\{4\}\+4\\cdot 2^\{3\}\\cdot\(17k\)=16\+32\\cdot 17k\\pmod\{289\}\.We want16\+544​k≡−1\(mod289\)16\+544k\\equiv\-1\\pmod\{289\}\. Since544≡255\(mod289\)544\\equiv 255\\pmod\{289\}, this is255​k≡−17\(mod289\)255k\\equiv\-17\\pmod\{289\}, and255≡−34255\\equiv\-34, so−34​k≡−17\-34k\\equiv\-17,i\.e\.,34​k≡17\(mod289\)34k\\equiv 17\\pmod\{289\}\. Dividing by 17:2​k≡1\(mod17\)2k\\equiv 1\\pmod\{17\}, sok≡9\(mod17\)k\\equiv 9\\pmod\{17\}\(since2×9=18≡12\\times 9=18\\equiv 1\)\. Thereforem=2\+17×9=155m=2\+17\\times 9=155\.Is there a smallermm?If we considerm≡2\(mod17\)m\\equiv 2\\pmod\{17\}, the possiblemm’s are2,19,36,53,70,87,104,121,138,155,…2,19,36,53,70,87,104,121,138,155,\\dotsCheckm=2m=2:24\+1=172^\{4\}\+1=17, not divisible by 289\. Checkm=19m=19:192=361≡72\(mod289\)19^\{2\}=361\\equiv 72\\pmod\{289\},194≡722=5184≡271\(mod289\)19^\{4\}\\equiv 72^\{2\}=5184\\equiv 271\\pmod\{289\}\(since289×17=4913289\\times 17=4913and5184−4913=2715184\-4913=271\), so194\+1≡272≠019^\{4\}\+1\\equiv 272\\neq 0\. Not divisible\. Checkingm=36m=36would be tedious\.Alternatively, note that the solution is unique in the congruence class, so the minimalmmis 155\.Since we used Hensel’s lemma and the lifting worked, we can trust that this is the correct answer\. Thus, the least positive integermmis 155\.Final Answer:155\\boxed\{155\}\(←\\leftarrowthe lift itself is valid,1554≡−1mod289155^\{4\}\\equiv\-1\\bmod 289, but the trace only examines the rootm≡2\(mod17\)m\\equiv 2\\pmod\{17\}; the congruencex4≡−1\(mod17\)x^\{4\}\\equiv\-1\\pmod\{17\}has 4 roots \(2,8,9,152,8,9,15\), and lifting the root88gives110<155110<155\.\)W2S\-OPD:We are given a problem involving prime numbers and divisibility\. Let me break it down step by step\.Step 1: Understand the problem\.We need two things: \(1\) find the smallest primeppsuch thatp2p^\{2\}dividesn4\+1n^\{4\}\+1for some positive integernn; \(2\) for thatpp, find the smallestmmsuch thatp2p^\{2\}dividesm4\+1m^\{4\}\+1\. So first, we need to determine what primesppsatisfy the condition that there exists somennwheren4≡−1\(modp2\)n^\{4\}\\equiv\-1\\pmod\{p^\{2\}\}\.Step 2: Analyze the equationn4≡−1\(modp2\)n^\{4\}\\equiv\-1\\pmod\{p^\{2\}\}\.This is equivalent ton8≡1\(modp2\)n^\{8\}\\equiv 1\\pmod\{p^\{2\}\}, since squaring both sides gives\(n4\)2≡\(−1\)2=1\(modp2\)\(n^\{4\}\)^\{2\}\\equiv\(\-1\)^\{2\}=1\\pmod\{p^\{2\}\}\. So the order ofnnmodulop2p^\{2\}must divide 8, but not divide 4 \(sincen4≡−1≠1n^\{4\}\\equiv\-1\\neq 1\)\. Let me start checking small primes and see which ones satisfy this condition\.Step 3: Check small primes\.Forp=2p=2:n=1n=1gives1\+1=21\+1=2, not divisible by22=42^\{2\}=4;n=2n=2gives16\+1=1716\+1=17, not divisible by 4\. No solution\. Forp=3p=3:n=1,…,6n=1,\\dots,6given4\+1≡2,8,1,5,5,1\(mod9\)n^\{4\}\+1\\equiv 2,8,1,5,5,1\\pmod\{9\}, none zero\. No solution\. Forp=5p=5: checkingn=1,…,10n=1,\\dots,10againstp2=25p^\{2\}=25, the valuesn4\+1mod25n^\{4\}\+1\\bmod 25are2,17,7,7,1,22,2,22,12,12,17,7,7,1,22,2,22,12,1, none zero\. No solution\.For larger primes, let me think algebraically\. Ifppdividesn4\+1n^\{4\}\+1, thenn4≡−1\(modp\)n^\{4\}\\equiv\-1\\pmod\{p\}, hencen8≡1\(modp\)n^\{8\}\\equiv 1\\pmod\{p\}, so the order ofnnmoduloppis exactly 8\. Since the multiplicative group modppis cyclic of orderp−1p\-1, we need8∣p−18\\mid p\-1,i\.e\.,p≡1\(mod8\)p\\equiv 1\\pmod\{8\}\. This rules outp=7,13p=7,13\(verified directly: nonnwithn4≡−1n^\{4\}\\equiv\-1exists for either\)\. The smallest prime withp≡1\(mod8\)p\\equiv 1\\pmod\{8\}isp=17p=17\.Step 4: Verifyp=17p=17\.The group mod 17 has order 16, so elements of order 8 exist\. Computingn4mod17n^\{4\}\\bmod 17for smallnn: atn=8n=8,82=64≡138^\{2\}=64\\equiv 13, so84≡132=169≡−1\(mod17\)8^\{4\}\\equiv 13^\{2\}=169\\equiv\-1\\pmod\{17\}\. Son=8n=8satisfiesn4\+1≡0\(mod17\)n^\{4\}\+1\\equiv 0\\pmod\{17\}\.But we needp2∣n4\+1p^\{2\}\\mid n^\{4\}\+1\. Forn=8n=8:84\+1=40978^\{4\}\+1=4097, and4097=14×289\+514097=14\\times 289\+51, so4097≡51\(mod289\)4097\\equiv 51\\pmod\{289\}— divisible by 17 but not by17217^\{2\}\. So we must lift the solution\.Step 5: Hensel lifting\.Writen=8\+17​kn=8\+17kand expand modulo17217^\{2\}, dropping terms with\(17​k\)2\(17k\)^\{2\}:\(8\+17​k\)4\+1≡84\+4⋅83⋅17​k\+1\(mod289\)\.\(8\+17k\)^\{4\}\+1\\;\\equiv\\;8^\{4\}\+4\\cdot 8^\{3\}\\cdot 17k\+1\\pmod\{289\}\.Since84=4096≡50\(mod289\)8^\{4\}=4096\\equiv 50\\pmod\{289\}, the condition becomes51\+4⋅83⋅17​k≡0\(mod289\)51\+4\\cdot 8^\{3\}\\cdot 17k\\equiv 0\\pmod\{289\}\. Note51=3×1751=3\\times 17, so dividing through by 17 reduces the condition to4⋅83⋅k≡−3\(mod17\)\.4\\cdot 8^\{3\}\\cdot k\\;\\equiv\\;\-3\\pmod\{17\}\.Compute82≡138^\{2\}\\equiv 13,83≡8⋅13=104≡2\(mod17\)8^\{3\}\\equiv 8\\cdot 13=104\\equiv 2\\pmod\{17\}, so4⋅83≡8\(mod17\)4\\cdot 8^\{3\}\\equiv 8\\pmod\{17\}and the condition is8​k≡−3≡14\(mod17\)8k\\equiv\-3\\equiv 14\\pmod\{17\}\. The inverse of 8 mod 17 is 15 \(since8×15=120≡18\\times 15=120\\equiv 1\), givingk≡14×15=210≡6\(mod17\)k\\equiv 14\\times 15=210\\equiv 6\\pmod\{17\}\. Thereforen=8\+17×6=110n=8\+17\\times 6=110satisfiesn4\+1≡0\(mod289\)n^\{4\}\+1\\equiv 0\\pmod\{289\}\.Step 6: Check that no smallermmworks\.Any solution must satisfym4≡−1\(mod17\)m^\{4\}\\equiv\-1\\pmod\{17\}, som≡8m\\equiv 8orm≡9\(mod17\)m\\equiv 9\\pmod\{17\}\(among the order\-8 residues\)\. For the rootm≡9\(mod17\)m\\equiv 9\\pmod\{17\}:94=6561≡203\(mod289\)9^\{4\}=6561\\equiv 203\\pmod\{289\}, and repeating the lifting with93≡159^\{3\}\\equiv 15,4⋅93≡9\(mod17\)4\\cdot 9^\{3\}\\equiv 9\\pmod\{17\}gives9​k≡−12≡5\(mod17\)9k\\equiv\-12\\equiv 5\\pmod\{17\}; the inverse of 9 is 2, sok≡10\(mod17\)k\\equiv 10\\pmod\{17\}and the minimal solution on this branch ism=9\+170=179\>110m=9\+170=179\>110\.For the branchm≡8\(mod17\)m\\equiv 8\\pmod\{17\}, the candidates below 110 are8,25,42,59,76,938,25,42,59,76,93, corresponding tok=0,1,…,5k=0,1,\\dots,5; since the lifting requiresk≡6\(mod17\)k\\equiv 6\\pmod\{17\}, none of them works \(checked explicitly form=8m=8:4097≡514097\\equiv 51;m=25m=25:252=625≡4725^\{2\}=625\\equiv 47,472=2209≡18647^\{2\}=2209\\equiv 186,186\+1=187≠0\(mod289\)186\+1=187\\neq 0\\pmod\{289\}\)\. Thereforem=110m=110is the least positive integer with172∣m4\+117^\{2\}\\mid m^\{4\}\+1\.Final Answer:110\\boxed\{110\}

Figure 6:Math case study under Pre\-RL/ Post\-RL setting, with the errors of direct OPD marked in red\.Case Study: Code Generation \(Pre\-RL/Post\-RL Setting\)Question:Write a Python function to find the last digit of a given number\.Direct OPD:To find thelast digitof a given number, we can use themodulo 10operation\.Reasoning:For any numbern, the last digit is the remainder whennis divided by 10\. For example:123 % 10 = 3→\\tolast digit is 3;456 % 10 = 6→\\tolast digit is 6;0 % 10 = 0→\\tolast digit is 0;\-123 % 10 = 7→\\tobecause−123\-123is equivalent to−123\+130=7\-123\+130=7\(in mod 10\)\.But wait\! When the number is negative, the modulo operation in Python returns a non\-negative result\.So for negative numbers, the last digit is still the correct one\.So, the function can be written as:def last\_Digit\(n\): return n % 10Test Case:last\_Digit\(123\)→\\to3;last\_Digit\(456\)→\\to6;last\_Digit\(\-123\)→\\to7\(because\-123 % 10is7\);last\_Digit\(0\)→\\to0\. So this function works for all integer inputs\.Final Code: def last\_Digit\(n\): return n % 10 \# Test the function assert last\_Digit\(123\) == 3\(←\\leftarrowthe intended last digit of−123\-123is33, not77; the trace convinces itself that Python’s non\-negative modulo is already the desired behavior and never takes the absolute value\.\)W2S\-OPD:Let’s firstthinkabout the problem:Problem:We want a function that returns thelast digitof a given number\. The last digit of a number is simply the remainder when the number is divided by 10\. For example:123 % 10gives3;4567 % 10gives7;0 % 10gives0\.So, the strategy is straightforward: take the given numbernand returnn % 10\.But we must consider the possibility ofnegative numbers\. For instance:\-123 % 10gives7in Python — this is not the intended behavior for this problem\. We want theabsolute value’s last digit\. So, we should take the absolute value ofnbefore taking the modulo 10\.So the correct formula is: last digit = abs\(n\) % 10\.Let’s test with sample input:last\_Digit\(123\)→\\toabs\(123\) % 10→\\to3✓;last\_Digit\(\-456\)→\\toabs\(\-456\) % 10→\\to6✓;last\_Digit\(0\)→\\to0 % 10→\\to0✓;last\_Digit\(9\)→\\to9 % 10→\\to9✓\.Now let’s write the Python function\.Python Code: def last\_Digit\(n\): return abs\(n\) % 10Now, test it with the assert statement: assert last\_Digit\(123\) == 3 This should pass without error\.Let me present the final code as you requested\.✓Final Answer: def last\_Digit\(n\): return abs\(n\) % 10

Figure 7:Code case study under the pre\-RL/post\-RL setting with the errors of direct OPD marked in red\.Case Study: Math Reasoning \(Smaller & Larger Setting\)Question:LetA​B​C​DABCDandC​E​F​GCEFGbe squares such thatCClies on segmentD​G¯\\overline\{DG\}andEElies on segmentB​C¯\\overline\{BC\}\. LetOObe the circumcenter of triangleA​E​GAEG\. Given thatA,D,A,D,andOOare collinear andA​B=1AB=1, computeF​GFG\.Correct Answer:3−1\\sqrt\{3\}\-1Qwen3\-8B \(student\):We are given two squares,A​B​C​DABCDandC​E​F​GCEFG, such thatA​B=1AB=1;CClies on segmentD​G¯\\overline\{DG\};EElies on segmentB​C¯\\overline\{BC\};OOis thecircumcenterof triangleA​E​GAEG; andA,D,A,D,andOOarecollinear\. We are to find the length ofF​GFG\.Step 1: Assign coordinates\.Place squareA​B​C​DABCDwithA=\(0,0\)A=\(0,0\),B=\(1,0\)B=\(1,0\),C=\(1,1\)C=\(1,1\),D=\(0,1\)D=\(0,1\), so squareA​B​C​DABCDhas side length 1\. SinceEElies onB​C¯\\overline\{BC\}, a vertical segment from\(1,1\)\(1,1\)to\(1,0\)\(1,0\), letE=\(1,t\)E=\(1,t\)with0≤t≤10\\leq t\\leq 1\. The direction fromCCtoEEis downward, so the squareC​E​F​GCEFGextends down and left from pointCC\. The vector fromCCtoEEis\(0,t−1\)\(0,t\-1\), and the vector fromEEtoFFshould be orthogonal of the same length,so rotate the vector\(0,t−1\)\(0,t\-1\)90∘90^\{\\circ\}clockwise: the rotation of\(x,y\)\(x,y\)clockwise is\(y,−x\)\(y,\-x\), so the vector fromEEtoFFis\(t−1,0\)\(t\-1,0\), givingF=\(1\+t−1,t\)=\(t,t\)F=\(1\+t\-1,\\,t\)=\(t,t\)\.Then fromFF, move by the same vector fromCCtoEE:G=F\+\(0,t−1\)=\(t,2​t−1\)G=F\+\(0,t\-1\)=\(t,\\,2t\-1\)\. Therefore the squareC​E​F​GCEFGhas verticesC=\(1,1\)C=\(1,1\),E=\(1,t\)E=\(1,t\),F=\(t,t\)F=\(t,t\),G=\(t,2​t−1\)G=\(t,2t\-1\)\.Step 2: TriangleA​E​GAEG\.We haveA=\(0,0\)A=\(0,0\),E=\(1,t\)E=\(1,t\),G=\(t,2​t−1\)G=\(t,2t\-1\)\. We will find the circumcenterOOand then enforce thatAA,DD,OOare collinear\.Step 3: Find the circumcenter of triangleA​E​GAEG\.Using perpendicular bisectors\. For sideA​EAE: midpointM1=\(12,t2\)M\_\{1\}=\(\\tfrac\{1\}\{2\},\\tfrac\{t\}\{2\}\), slopett, so the perpendicular bisector isy=−1t​x\+1\+t22​t\.y=\-\\frac\{1\}\{t\}x\+\\frac\{1\+t^\{2\}\}\{2t\}\.For sideA​GAG: midpointM2=\(t2,2​t−12\)M\_\{2\}=\(\\tfrac\{t\}\{2\},\\tfrac\{2t\-1\}\{2\}\), slope2​t−1t\\tfrac\{2t\-1\}\{t\}, perpendicular slope−t2​t−1\-\\tfrac\{t\}\{2t\-1\}, soy=−t2​t−1​x\+t22​t−1\+2​t−12=−t2​t−1​x\+6​t2−4​t\+12​\(2​t−1\)\.y=\-\\frac\{t\}\{2t\-1\}x\+\\frac\{t^\{2\}\}\{2t\-1\}\+\\frac\{2t\-1\}\{2\}=\-\\frac\{t\}\{2t\-1\}x\+\\frac\{6t^\{2\}\-4t\+1\}\{2\(2t\-1\)\}\.Solving \(1\) and \(2\) directly is complicated, but we will use a key condition:AA,DD, andOOare collinear\. SinceA=\(0,0\)A=\(0,0\)andD=\(0,1\)D=\(0,1\), lineA​DADis vertical,i\.e\.,x=0x=0\. So thexx\-coordinate ofOOmust be 0, and from \(1\):O=\(0,1\+t22​t\)O=\\big\(0,\\frac\{1\+t^\{2\}\}\{2t\}\\big\)\.Step 4: Enforce thatOOlies on both perpendicular bisectors\.Pluggingx=0x=0into \(2\) givesy=6​t2−4​t\+12​\(2​t−1\)y=\\frac\{6t^\{2\}\-4t\+1\}\{2\(2t\-1\)\}, so for consistency we require1\+t22​t=6​t2−4​t\+12​\(2​t−1\)⟹\(1\+t2\)​\(2​t−1\)=t​\(6​t2−4​t\+1\)\.\\frac\{1\+t^\{2\}\}\{2t\}=\\frac\{6t^\{2\}\-4t\+1\}\{2\(2t\-1\)\}\\;\\Longrightarrow\\;\(1\+t^\{2\}\)\(2t\-1\)=t\(6t^\{2\}\-4t\+1\)\.Expanding:2​t3−t2\+2​t−1=6​t3−4​t2\+t2t^\{3\}\-t^\{2\}\+2t\-1=6t^\{3\}\-4t^\{2\}\+t, which gives the cubic−4​t3\+3​t2\+t−1=0\.\-4t^\{3\}\+3t^\{2\}\+t\-1=0\.Try rational roots\.t=1t=1:−4\+3\+1−1=−1≠0\-4\+3\+1\-1=\-1\\neq 0\.t=12t=\\tfrac\{1\}\{2\}:−12\+34\+12−1=−14≠0\-\\tfrac\{1\}\{2\}\+\\tfrac\{3\}\{4\}\+\\tfrac\{1\}\{2\}\-1=\-\\tfrac\{1\}\{4\}\\neq 0\.t=14t=\\tfrac\{1\}\{4\}:−116\+316\+416−1616=−1116≠0\-\\tfrac\{1\}\{16\}\+\\tfrac\{3\}\{16\}\+\\tfrac\{4\}\{16\}\-\\tfrac\{16\}\{16\}=\-\\tfrac\{11\}\{16\}\\neq 0\.t=23t=\\tfrac\{2\}\{3\}: messy\. Let’s try an interval method:t=1t=1gives−1\-1;t=32t=\\tfrac\{3\}\{2\}: messy, but try a better approach\.Wait, this is very algebra\-heavy\. But we already found thatO=\(0,1\+t22​t\)O=\\big\(0,\\frac\{1\+t^\{2\}\}\{2t\}\\big\), and from the earlier step, whent=12t=\\tfrac\{1\}\{2\}, this works? Pluggingt=12t=\\tfrac\{1\}\{2\}:1\+1/41=54\\frac\{1\+1/4\}\{1\}=\\tfrac\{5\}\{4\}, theyy\-coordinate ofOO; also, from tentative earlier steps, this works\. So let’s assumet=12t=\\tfrac\{1\}\{2\}\.Step 5: ComputeF​GFGwitht=12t=\\tfrac\{1\}\{2\}\.RecallF=\(t,t\)=\(12,12\)F=\(t,t\)=\(\\tfrac\{1\}\{2\},\\tfrac\{1\}\{2\}\)andG=\(t,2​t−1\)=\(12,0\)G=\(t,2t\-1\)=\(\\tfrac\{1\}\{2\},0\)\. So segmentF​GFGis vertical with length\|12−0\|=12\|\\tfrac\{1\}\{2\}\-0\|=\\tfrac\{1\}\{2\}\.Final Answer:12\\boxed\{\\tfrac\{1\}\{2\}\}\(←\\leftarrowtwo failures: the square is built by rotating clockwise, placingGGat\(t,2​t−1\)\(t,2t\-1\), for whichCCnever lies onD​G¯\\overline\{DG\}, and this constraint is never checked; then, unable to solve the resulting cubic, the trace assumest=12t=\\tfrac\{1\}\{2\}immediately after computing thatt=12t=\\tfrac\{1\}\{2\}does not satisfy the equation\.\)W2S\-OPD:We are given two squares:A​B​C​DABCDandC​E​F​GCEFG, with the following constraints: pointCClies on segmentD​G¯\\overline\{DG\}; pointEElies on segmentB​C¯\\overline\{BC\};OOis thecircumcenterof triangleA​E​GAEG; pointsAA,DD, andOOarecollinear;A​B=1AB=1\. We are to compute the lengthF​GFG\. We’ll solve thisstep by stepusing coordinate geometry\.Step 1: Set up a coordinate system\.Let’s place squareA​B​C​DABCDin the coordinate plane for simplicity\. SinceA​B=1AB=1andA​B​C​DABCDis a square, we can assignA=\(0,0\)A=\(0,0\),B=\(1,0\)B=\(1,0\),C=\(1,1\)C=\(1,1\),D=\(0,1\)D=\(0,1\)\. So squareA​B​C​DABCDis in the first quadrant, with sides aligned with the axes\.Step 2: Place squareC​E​F​GCEFG\.We need to place squareC​E​F​GCEFGsuch thatC=\(1,1\)C=\(1,1\),EElies onB​C¯\\overline\{BC\},CClies onD​G¯\\overline\{DG\}, andOOis the circumcenter ofA​E​GAEGwithAA,DD,OOcollinear\. Let’s try to find the coordinates ofEE,FF, andGG\.Step 3: Place pointEEonB​C¯\\overline\{BC\}\.SegmentB​C¯\\overline\{BC\}goes fromB=\(1,0\)B=\(1,0\)toC=\(1,1\)C=\(1,1\), so it is avertical segmentatx=1x=1, fromy=0y=0toy=1y=1\. Let’s sayE=\(1,e\)E=\(1,e\), where0<e<10<e<1\.Step 4: Place squareC​E​F​GCEFG\.SideC​E→\\overrightarrow\{CE\}isvertical, going fromC=\(1,1\)C=\(1,1\)toE=\(1,e\)E=\(1,e\), so it is of length\|1−e\|\|1\-e\|\. The square can be constructed in two possible orientations, either to the left or to the right ofC​ECE, and we must choose the one consistent with the constraint thatCClies onD​G¯\\overline\{DG\}\. Trycounterclockwisefirst:C​E→=\(0,e−1\)\\vec\{CE\}=\(0,e\-1\); rotatingC​E→\\vec\{CE\}90∘90^\{\\circ\}counterclockwise gives\(1−e,0\)\(1\-e,0\); so fromE=\(1,e\)E=\(1,e\), moving in direction\(1−e,0\)\(1\-e,0\)givesF=\(2−e,e\)F=\(2\-e,\\,e\); then fromFF, moving opposite toC​E→\\vec\{CE\}by\(0,1−e\)\(0,1\-e\)givesG=\(2−e,1\)G=\(2\-e,\\,1\)\.Let’s verify this shape is a square:C​E→=\(0,e−1\)\\vec\{CE\}=\(0,e\-1\),E​F→=\(1−e,0\)\\vec\{EF\}=\(1\-e,0\),F​G→=\(0,1−e\)\\vec\{FG\}=\(0,1\-e\),G​C→=\(e−1,0\)\\vec\{GC\}=\(e\-1,0\)\. All sides are of length\|e−1\|\|e\-1\|, and adjacent sides are perpendicular→\\tothis is indeed a square\.Now, also note that segmentD​G¯\\overline\{DG\}goes fromD=\(0,1\)D=\(0,1\)toG=\(2−e,1\)G=\(2\-e,1\), soD​G¯\\overline\{DG\}is ahorizontal segment aty=1y=1, fromx=0x=0tox=2−ex=2\-e\. So pointC=\(1,1\)C=\(1,1\)lies onD​G¯\\overline\{DG\}if and only if0≤1≤2−e⇒e≤10\\leq 1\\leq 2\-e\\Rightarrow e\\leq 1, which is true since0<e<10<e<1\. ✓ So this configuration satisfies the condition thatCClies on segmentD​G¯\\overline\{DG\}\.Step 5: Coordinates of all relevant points\.Recall:A=\(0,0\)A=\(0,0\),D=\(0,1\)D=\(0,1\),E=\(1,e\)E=\(1,e\),G=\(2−e,1\)G=\(2\-e,1\)\. We need to find thecircumcenterOOof triangleA​E​GAEG, and use the condition thatAA,DD, andOOare collinear\.Step 6: Find the circumcenterOOof triangleA​E​GAEG\.DenoteO=\(h,k\)O=\(h,k\)\. SinceOOis the circumcenter, it is equidistant from all three vertices:h2\+k2=\(h−1\)2\+\(k−e\)2=\(h−\(2−e\)\)2\+\(k−1\)2\.h^\{2\}\+k^\{2\}=\(h\-1\)^\{2\}\+\(k\-e\)^\{2\}=\(h\-\(2\-e\)\)^\{2\}\+\(k\-1\)^\{2\}\.Equation 1:O​A2=O​E2OA^\{2\}=OE^\{2\}\.Expanding the right side:\(h−1\)2\+\(k−e\)2=h2−2​h\+1\+k2−2​e​k\+e2\(h\-1\)^\{2\}\+\(k\-e\)^\{2\}=h^\{2\}\-2h\+1\+k^\{2\}\-2ek\+e^\{2\}\. Setting it equal toh2\+k2h^\{2\}\+k^\{2\}and cancelingh2,k2h^\{2\},k^\{2\}from both sides:0=−2h\+1−2ek\+e2⟹2h\+2ek=1\+e2\.\(Equation A\)0=\-2h\+1\-2ek\+e^\{2\}\\;\\Longrightarrow\\;2h\+2ek=1\+e^\{2\}\.\\quad\\text\{\(Equation A\)\}Equation 2:O​E2=O​G2OE^\{2\}=OG^\{2\}\.The left side ish2−2​h\+1\+k2−2​e​k\+e2h^\{2\}\-2h\+1\+k^\{2\}\-2ek\+e^\{2\}; the right side ish2−2​\(2−e\)​h\+\(2−e\)2\+k2−2​k\+1h^\{2\}\-2\(2\-e\)h\+\(2\-e\)^\{2\}\+k^\{2\}\-2k\+1\. Cancelingh2h^\{2\},k2k^\{2\},11and bringing everything to the left:−2​h−2​e​k\+e2\+4​h−2​e​h\+2​k−\(4−4​e\+e2\)=0,\-2h\-2ek\+e^\{2\}\+4h\-2eh\+2k\-\(4\-4e\+e^\{2\}\)=0,which simplifies term by term \(−2​h\+4​h=2​h\-2h\+4h=2h;−2​e​k−2​e​h=−2​e​\(k\+h\)\-2ek\-2eh=\-2e\(k\+h\);e2−e2=0e^\{2\}\-e^\{2\}=0\) to2​h−2​e​\(h\+k\)\+2​k−4\+4​e=0\.2h\-2e\(h\+k\)\+2k\-4\+4e=0\.Dividing by 2 and grouping like terms:\(h\+k\)​\(1−e\)\+2​\(e−1\)=0\(h\+k\)\(1\-e\)\+2\(e\-1\)=0\. Sincee−1=−\(1−e\)e\-1=\-\(1\-e\), this factors as\(1−e\)​\[h\+k−2\]=0\.\(1\-e\)\\,\[\\,h\+k\-2\\,\]=0\.So either1−e=0⇒e=11\-e=0\\Rightarrow e=1,butthis isnot allowedsince thenE=CE=C, which would collapse the square; orh\+k=2h\+k=2\.\(Equation B\)Step 7: Combine equations \(A\) and \(B\)\.From \(B\),k=2−hk=2\-h\. Substituting into \(A\):2​h\+2​e​\(2−h\)=1\+e2⟹2​h\+4​e−2​e​h=1\+e2⟹2​h​\(1−e\)=1\+e2−4​e,2h\+2e\(2\-h\)=1\+e^\{2\}\\;\\Longrightarrow\\;2h\+4e\-2eh=1\+e^\{2\}\\;\\Longrightarrow\\;2h\(1\-e\)=1\+e^\{2\}\-4e,soh=1\+e2−4​e2​\(1−e\)\.h=\\frac\{1\+e^\{2\}\-4e\}\{2\(1\-e\)\}\.Step 8: Use collinearity ofA=\(0,0\)A=\(0,0\),D=\(0,1\)D=\(0,1\),O=\(h,k\)O=\(h,k\)\.PointsAAandDDare on thevertical linex=0x=0, so forOOto be on lineA​DADit must also satisfyh=0h=0\. Setting the numerator to zero:1\+e2−4​e=0⟹e2−4​e\+1=0⟹e=4±16−42=2±3\.1\+e^\{2\}\-4e=0\\;\\Longrightarrow\\;e^\{2\}\-4e\+1=0\\;\\Longrightarrow\\;e=\\frac\{4\\pm\\sqrt\{16\-4\}\}\{2\}=2\\pm\\sqrt\{3\}\.Now0<e<10<e<1, so check values:2\+3≈3\.7322\+\\sqrt\{3\}\\approx 3\.732is too big;2−3≈0\.26792\-\\sqrt\{3\}\\approx 0\.2679is valid\. Soe=2−3e=2\-\\sqrt\{3\}\.Step 9: FindF​GFG\.RecallF=\(2−e,e\)F=\(2\-e,\\,e\)andG=\(2−e,1\)G=\(2\-e,\\,1\)\. SegmentF​GFGisvertical, samexx\-coordinate, difference inyy\-coordinates:F​G=\|1−e\|=1−e=1−\(2−3\)=3−1\.FG=\|1\-e\|=1\-e=1\-\(2\-\\sqrt\{3\}\)=\\sqrt\{3\}\-1\.✓Final Answer:3−1\\boxed\{\\sqrt\{3\}\-1\}

Figure 8:Math case study from under the smaller/larger setting, with the errors of the Qwen3\-8B student marked in red\.Case Study: Code Generation \(Smaller & Larger Setting\)Question:Given two listsoperatorandoperand, where the first list contains basic algebra operations and the second is a list of integers, use the two lists to build an algebraic expression and return its evaluation,e\.g\.,operator = \[’\+’, ’\*’, ’\-’\]andoperand = \[2, 3, 4, 5\]represent2\+3×4−5=92\+3\\times 4\-5=9\.Qwen3\-8B \(student\):To solve this problem, we need toapply the given list of operators to the corresponding elements of the operand list in a sequential manner\. Given that the length of theoperatorlist is one less than the length of theoperandlist, we can construct the expression step\-by\-step and evaluate it\.Here’s the step\-by\-step plan: 1\.Initialize the resultwith the first operand\.2\.Iterateover the operators and operands in pairs\. 3\.Apply each operatorto the current result and the next operand\. 4\.Return the final resultafter all operations are completed\.Here is the Python code that implements the above logic:def do\_algebra\(operator, operand\): result = operand\[0\] for i in range\(len\(operator\)\): op = operator\[i\] num = operand\[i \+ 1\] if op == ’\+’: result \+= num elif op == ’\-’: result \-= num elif op == ’\*’: result \*= num elif op == ’//’: result //= num elif op == ’\*\*’: result \*\*= num return resultExample Usage: print\(do\_algebra\(\[’\+’, ’\*’, ’\-’\], \[2, 3, 4, 5\]\)\) \# Output: 9This code handles all the required operations and follows the constraints specified in the problem statement\.\(←\\leftarrowevaluates strictly left to right, ignoring operator precedence: for\[’\+’, ’\*’, ’\-’\]and\[2, 3, 4, 5\]the code computes\(\(2\+3\)×4\)−5=15\(\(2\+3\)\\times 4\)\-5=15, not the claimed 9 \(=2\+3×4−5=2\+3\\times 4\-5\), so even its own example comment is inconsistent with the code\.\)W2S\-OPD:To solve this problem, we need to construct an algebraic expression using the givenoperatorandoperandlists, then evaluate it\.Plan:1\. Theoperatorlist has one fewer element than theoperandlist\.2\. We can iterate through theoperatorlist and apply each operator to consecutive elements from theoperandlist\.3\. Use Python’seval\(\)function to safely evaluate the constructed expression string\.4\. Make sure the expression string is properly formatted with spaces to avoid any ambiguity or syntax errors\.Important Notes:We assume that the input is valid as per the problem constraints\. We useeval\(\)with caution, but in this context it’s safe since the inputs are controlled and only contain allowed operations and integers\. We format the expression string with spaces to ensure correct parsing\.Python Code:def do\_algebra\(operator, operand\): \# Build the expression string expression = str\(operand\[0\]\) for i in range\(len\(operator\)\): expression \+= f" \{operator\[i\]\} \{operand\[i \+ 1\]\}" \# Evaluate the expression return eval\(expression\)Example Usage: operator = \[’\+’, ’\*’, ’\-’\] operand = \[2, 3, 4, 5\] print\(do\_algebra\(operator, operand\)\) \# Output: 9Explanation:The expression string becomes"2 \+ 3 \* 4 \- 5"\. Evaluating this gives2\+3×4−5=2\+12−5=92\+3\\times 4\-5=2\+12\-5=9\. This solution is concise, readable, and leverages Python’s built\-in capabilities to achieve the desired result\.

Figure 9:Code case study under the smaller/larger setting, with the errors of the Qwen3\-8B student marked in red\.

Similar Articles

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Hugging Face Daily Papers

The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.

On-policy distillation: one of the hottest terms on PapersWithCode [R]

Reddit r/MachineLearning

Hugging Face's Niels introduces On-policy Distillation (OPD), a key post-training technique used in models like Qwen 3.6/3.7, GLM-5.1, and DeepSeek-V4, now featured on PapersWithCode with a linked whiteboard explanation by Sasha Rush and Dwarkesh Patel.