Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers Papers

Summary

This paper proposes Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD) to address unbalanced feedback when merging specialist language models, showing performance improvements on benchmarks like mathematics and instruction-following.

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Original Article
View Cached Full Text

Cached at: 09/29/26, 08:15 AM

Paper page - Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Source: https://huggingface.co/papers/2609.35347

Abstract

Reinforcementlearningcanturnonelanguagemodelintoseveralspecialists,eachexcellentatasingleskillsuchasmathematics,codingorfollowinginstructions,butusersneedonemodelwithalloftheseskills.Multi-teacheron-policydistillation(MOPD)mergesthembylettingthespecialiststeachonestudent:thestudentanswerseachprompt,andthespecialistforthatprompt’sdomaingivesfeedbackoneverytoken.Thisroutingdecideswhichspecialistteaches,butnothowstronglyitsfeedbackmovesthesharedstudent.InQwen3.5modelsatthreesizes,wefindthatMOPD’sstudentdoesnotbeatonetaughtbythebestsinglespecialistandgainslittleofthemathematicsspecialist’sadvantage.Thefeedbackisunbalanced:instruction-followingfeedbackisseveraltimesmorespreadoutthanmathematicsfeedbackanddominatesthestudent’supdates.WeproposeDomain-NormalizedMOPD(DN-MOPD),whichkeepstheroutingandrescaleseachdomain’sfeedbackbyitsmeasuredspread.Onsixpublicbenchmarks,DN-MOPDimprovestheaveragescoreoverMOPDateverysize,acrossthreerandomseedsandundertwoanswer-lengthlimits,andrecoversmostofthelostmathematicsgain.Controlswithfixeddomainweightsshowthatthegaincomesmainlyfromturningdowninstruction-followingfeedbackratherthanturningupmathematicsalone,andthatfixedweightsclosetothoseDN-MOPDmeasuresperformcomparably.Combiningspecialiststhereforerequiresdecidingnotonlywhichoneteaches,butalsohowstronglyitsfeedbackcounts.

View arXiv pageView PDFProject pageGitHub5Add to collection

Get this paper in your agent:

hf papers read 2609\.35347

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35347 in a model README.md to link it from this page.

Datasets citing this paper1

#### XINLI1997/DN-MOPD-Data Viewer• Updatedabout 2 hours ago • 26.7k

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35347 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Hugging Face Daily Papers

MOPD proposes a multi-teacher on-policy distillation paradigm for LLM post-training, enabling efficient integration of multiple domain capabilities by distilling specialized RL teachers into a student model using its own rollouts. It outperforms existing methods like Mix-RL and Cascade RL, and has been deployed in industrial-scale models.

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers

MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.

Multi-Rollout On-Policy Distillation via Peer Successes and Failures

arXiv cs.LG

Introduces Multi-Rollout On-Policy Distillation (MOPD), a method that conditions the teacher on both successful and failed peer rollouts to provide denser token-level supervision for language model post-training, improving performance across multiple benchmarks.

DOPD: Dual On-policy Distillation

Hugging Face Daily Papers

DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.