Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Summary
This paper proposes Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD) to address unbalanced feedback when merging specialist language models, showing performance improvements on benchmarks like mathematics and instruction-following.
View Cached Full Text
Cached at: 09/29/26, 08:15 AM
Paper page - Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Source: https://huggingface.co/papers/2609.35347
Abstract
Reinforcementlearningcanturnonelanguagemodelintoseveralspecialists,eachexcellentatasingleskillsuchasmathematics,codingorfollowinginstructions,butusersneedonemodelwithalloftheseskills.Multi-teacheron-policydistillation(MOPD)mergesthembylettingthespecialiststeachonestudent:thestudentanswerseachprompt,andthespecialistforthatprompt’sdomaingivesfeedbackoneverytoken.Thisroutingdecideswhichspecialistteaches,butnothowstronglyitsfeedbackmovesthesharedstudent.InQwen3.5modelsatthreesizes,wefindthatMOPD’sstudentdoesnotbeatonetaughtbythebestsinglespecialistandgainslittleofthemathematicsspecialist’sadvantage.Thefeedbackisunbalanced:instruction-followingfeedbackisseveraltimesmorespreadoutthanmathematicsfeedbackanddominatesthestudent’supdates.WeproposeDomain-NormalizedMOPD(DN-MOPD),whichkeepstheroutingandrescaleseachdomain’sfeedbackbyitsmeasuredspread.Onsixpublicbenchmarks,DN-MOPDimprovestheaveragescoreoverMOPDateverysize,acrossthreerandomseedsandundertwoanswer-lengthlimits,andrecoversmostofthelostmathematicsgain.Controlswithfixeddomainweightsshowthatthegaincomesmainlyfromturningdowninstruction-followingfeedbackratherthanturningupmathematicsalone,andthatfixedweightsclosetothoseDN-MOPDmeasuresperformcomparably.Combiningspecialiststhereforerequiresdecidingnotonlywhichoneteaches,butalsohowstronglyitsfeedbackcounts.
View arXiv pageView PDFProject pageGitHub5Add to collection
Get this paper in your agent:
hf papers read 2609\.35347
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35347 in a model README.md to link it from this page.
Datasets citing this paper1
#### XINLI1997/DN-MOPD-Data Viewer• Updatedabout 2 hours ago • 26.7k
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35347 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
The paper introduces Open-MOPD, a framework that diagnoses and fixes capability imbalance in multi-teacher on-policy distillation by balancing token-level budgets, improving headroom recovery from 35.6% to 83.4% through dynamic allocation and reward refresh.
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
MOPD proposes a multi-teacher on-policy distillation paradigm for LLM post-training, enabling efficient integration of multiple domain capabilities by distilling specialized RL teachers into a student model using its own rollouts. It outperforms existing methods like Mix-RL and Cascade RL, and has been deployed in industrial-scale models.
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.
Multi-Rollout On-Policy Distillation via Peer Successes and Failures
Introduces Multi-Rollout On-Policy Distillation (MOPD), a method that conditions the teacher on both successful and failed peer rollouts to provide denser token-level supervision for language model post-training, improving performance across multiple benchmarks.
DOPD: Dual On-policy Distillation
DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.