PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers Papers

Summary

PMOPD 提出一种几何感知的多教师在线策略蒸馏方法,通过构建任务子空间记忆并对梯度和优化器更新做投影来缓解能力跷跷板效应,并加入冲突探测与任务循环策略,在 Code/Reason/Math 任务上于 Qwen2.5-7B 和 Llama-3.1-8B 上分别平均提升 2.54 和 2.09 分。

Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:17 PM

Paper page - PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Source: https://huggingface.co/papers/2609.34605

Abstract

Multi-teacheron-policydistillation(MOPD)hasemergedasapopularpost-trainingparadigmforintegratingspecializedcapabilitiesinfrontierlanguagemodels.ExistingOPDresearchhasprimarilyfocusedonoptimizingsingle-taskdistillationthroughobjectivedesign,distillationscope,andteachersignalconstruction,whereasMOPDmustaggregatemultiplecapabilitiesinsharedparametersandaddresstheresultingcapabilityseesaw,inwhichimprovingonedomainsuppressescapabilitiesacquiredfromanother.InspiredbythedistinctiveupdategeometryofOPD,wefindthatparameterupdatesfromdifferenttasksrapidlyconcentrateintheirrespectivelow-dimensionalsubspacesduringMOPD,providingadirectgeometricbasisforidentifyingandcontrollingcross-taskinterference.WethereforeproposePMOPD(Projection-basedMulti-TeacherOn-PolicyDistillation),whichconstructssubspacememoriesfromthecumulativeparameterdisplacementsofdifferenttasksandprojectsbothgradientsandoptimizerupdatestoremovecomponentsthatinterferewithprotectedtaskdirections.Wefurtherdevelopalightweightconflictprobetocharacterizetaskinteractionsandguidetaskordering,togetherwithacyclingstrategythatbalancessubspaceestimationandtimelytaskrevisitation.ExperimentsonrepresentativeCode,Reason,andMathtasksshowthatPMOPDimproveseveryevaluatedcapabilityoverMOPD,raisingtheaveragescoreacrossthethreetasksby2.54pointsonQwen2.5-7Band2.09pointsonLlama-3.1-8B.Theseconsistentgainsestablishgeometry-awareoptimizationasaneffectiveandtransferableapproachtobalancedmulti-teacherdistillation.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.34605

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34605 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34605 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34605 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers

MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Hugging Face Daily Papers

MOPD proposes a multi-teacher on-policy distillation paradigm for LLM post-training, enabling efficient integration of multiple domain capabilities by distilling specialized RL teachers into a student model using its own rollouts. It outperforms existing methods like Mix-RL and Cascade RL, and has been deployed in industrial-scale models.