PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Summary
PMOPD 提出一种几何感知的多教师在线策略蒸馏方法,通过构建任务子空间记忆并对梯度和优化器更新做投影来缓解能力跷跷板效应,并加入冲突探测与任务循环策略,在 Code/Reason/Math 任务上于 Qwen2.5-7B 和 Llama-3.1-8B 上分别平均提升 2.54 和 2.09 分。
View Cached Full Text
Cached at: 09/30/26, 04:17 PM
Paper page - PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Source: https://huggingface.co/papers/2609.34605
Abstract
Multi-teacheron-policydistillation(MOPD)hasemergedasapopularpost-trainingparadigmforintegratingspecializedcapabilitiesinfrontierlanguagemodels.ExistingOPDresearchhasprimarilyfocusedonoptimizingsingle-taskdistillationthroughobjectivedesign,distillationscope,andteachersignalconstruction,whereasMOPDmustaggregatemultiplecapabilitiesinsharedparametersandaddresstheresultingcapabilityseesaw,inwhichimprovingonedomainsuppressescapabilitiesacquiredfromanother.InspiredbythedistinctiveupdategeometryofOPD,wefindthatparameterupdatesfromdifferenttasksrapidlyconcentrateintheirrespectivelow-dimensionalsubspacesduringMOPD,providingadirectgeometricbasisforidentifyingandcontrollingcross-taskinterference.WethereforeproposePMOPD(Projection-basedMulti-TeacherOn-PolicyDistillation),whichconstructssubspacememoriesfromthecumulativeparameterdisplacementsofdifferenttasksandprojectsbothgradientsandoptimizerupdatestoremovecomponentsthatinterferewithprotectedtaskdirections.Wefurtherdevelopalightweightconflictprobetocharacterizetaskinteractionsandguidetaskordering,togetherwithacyclingstrategythatbalancessubspaceestimationandtimelytaskrevisitation.ExperimentsonrepresentativeCode,Reason,andMathtasksshowthatPMOPDimproveseveryevaluatedcapabilityoverMOPD,raisingtheaveragescoreacrossthethreetasksby2.54pointsonQwen2.5-7Band2.09pointsonLlama-3.1-8B.Theseconsistentgainsestablishgeometry-awareoptimizationasaneffectiveandtransferableapproachtobalancedmulti-teacherdistillation.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.34605
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34605 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34605 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34605 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
The paper introduces Open-MOPD, a framework that diagnoses and fixes capability imbalance in multi-teacher on-policy distillation by balancing token-level budgets, improving headroom recovery from 35.6% to 83.4% through dynamic allocation and reward refresh.
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
MOPD proposes a multi-teacher on-policy distillation paradigm for LLM post-training, enabling efficient integration of multiple domain capabilities by distilling specialized RL teachers into a student model using its own rollouts. It outperforms existing methods like Mix-RL and Cascade RL, and has been deployed in industrial-scale models.
Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
Poly-OPD is a framework for distilling complementary strengths from heterogeneous text-to-image flow models into a single compact flow-matching student, using pixel bridges and gradient-compatible adapters. It improves GenEval and DrawBench scores while consolidating multiple teacher capabilities.
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
This paper proposes Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD) to address unbalanced feedback when merging specialist language models, showing performance improvements on benchmarks like mathematics and instruction-following.