Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Summary
The paper introduces Latent-MOPD, the first representation-level multi-teacher on-policy distillation method for LLMs, which integrates multiple specialist teachers through both their output distributions and hidden states without extra teacher training, outperforming token-only, representation-only, and averaging baselines across nine math, code, and logic benchmarks.
View Cached Full Text
Cached at: 10/05/26, 10:03 AM
# Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Source: [https://arxiv.org/html/2610.02381](https://arxiv.org/html/2610.02381)
###### Abstract
On\-policy distillation \(OPD\) trains a student on the responses it generates\. Existing LLM multi\-teacher OPD transfers*what*specialists predict through their output distributions\. We introduce Latent\-MOPD, to our knowledge the first representation\-level multi\-teacher OPD method for LLMs\. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training\. To coordinate representation supervision from multiple specialists, we select late\-layer targets according to the teacher–student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain\. Each teacher’s supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist\. In our main same\-family setting, Latent\-MOPD outperforms the token\-only, representation\-only and uniform\-averaging baselines on all nine benchmarks across math, code and logic\. With the same parameter count as each teacher, the student also surpasses the per\-benchmark best teacher on a majority of these benchmarks\. With larger, separately developed cross\-family teachers, Latent\-MOPD outperforms both single\-channel baselines on all benchmarks\. A same\-family all\-layer representation\-only control remains stable with domain\-pure updates but collapses when teacher domains are interleaved within an update\. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations\.
## 1Introduction
Reinforcement learning \(RL\) has become an important tool for improving large language models \(LLMs\)\([Ouyang et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib31);[Shao et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib34)\), and specialized RL pipelines now produce models with complementary strengths in domains such as mathematical reasoning, coding, and logic\. While each of these models excels in its own domain, a broader goal is a single model that combines their strengths by reusing the existing specialists \(Figure[1](https://arxiv.org/html/2610.02381#S1.F1)\)\. Multi\-teacher on\-policy distillation \(MOPD\)\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)extends on\-policy distillation \(OPD\)\([Agarwal et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib1)\)to several specialists, routing each student\-generated response to its domain’s teacher\. The selected teacher scores the student’s tokens, and its next\-token distribution provides dense token\-level guidance along the student’s own rollout\. In this way, the student learns what each specialist predicts, with supervision defined entirely at the output level\.
Figure 1:Same\-family gains beyond individual teachers, with a cross\-family extension\.\(a\)Same\-family;\(b\)cross\-family \(last 1, Linear\)\. Both panels use shared per\-axis scales, with each outer vertex fixed at the same\-family Latent\-MOPD score\. Solid polygons show Latent\-MOPD; the dotted ring in\(b\)repeats the same\-family reference\. Representation\-only denotes OPRD\-style \(all layers\) in\(a\)and Rep\-only \(last 1\) in\(b\)\. Labels give Latent\-MOPD scores;boldwith \(\>\>Teacher\) marks those above the strongest teacher\. Table[1](https://arxiv.org/html/2610.02381#S3.T1)reports all nine same\-family benchmarks\.These predictions, however, are the end result of internal computations\. Interpretability studies show that hidden representations carry intermediate reasoning steps and causally influence later computation, even when these steps are not verbalized\([Lindsey et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib25);[Gurnee et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib12)\)\. This suggests that internal representations carry information about*how*a teacher computes its predictions, not only*what*it predicts\. Representation\-level OPD methods such as LastOPD and OPRD show that, with suitable layer pairing and training schedules, hidden\-state supervision can improve transfer over token\-only OPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48);[Yang et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib49)\)\. However, these approaches study learning from a single teacher\. Learning from multiple specialists poses a different challenge: their distinct representation targets must jointly train the same student backbone\. Routing determines whose supervision a prompt receives, but leaves a central question:*how should multi\-teacher OPD select representation targets and organize their supervision so that one student integrates the capabilities of several specialists?*
Choosing representation targets requires more than matching layer indices\.Before training, centered kernel alignment \(CKA\)\([Kornblith et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib18)\)between each same\-family specialist and the base student remains high through most layers, with larger differences near the output \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)a\)\. Across model families, this correspondence is uneven, with a sharp drop in similarity at intermediate depths \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)b\)\. These patterns suggest different considerations for alignment\. Within a shared lineage, the common representational structure provides a basis for learning from the specialists’ late\-layer differences; across lineages, weak correspondence at intermediate depths instead motivates targets with a shared computational role, such as the states directly used for next\-token prediction\. The choice of representation targets, therefore, depends on both the relationship between the models and the role of the states being aligned\.
Teacher routing does not determine how representation supervision should be organized over training\.In a same\-family all\-layer representation\-only control, mixing samples supervised by different teachers within an optimizer update collapses training within twenty updates, whereas domain\-pure batches remain stable \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c; Section[5](https://arxiv.org/html/2610.02381#S5.SS0.SSS0.Px3)\)\. Each sample receives supervision from the same specialist in both settings; the difference is whether several teachers contribute to a single parameter update\. Sustained representation\-only supervision is also problematic: in the cross\-family setting, it falls below the student’s initial performance in aggregate, so how latent supervision is scheduled over training matters as much as how it is batched\.
To answer this question, we introduce Latent\-MOPD\. It selects late\-layer targets according to the teacher–student relationship and organizes their supervision throughdomain\-pure updatesand aper\-teacher crossfade\(Figure[2](https://arxiv.org/html/2610.02381#S1.F2); Section[3](https://arxiv.org/html/2610.02381#S3)\)\. Each rollout’s specialist supplies both token predictions and hidden\-state targets\. We align hidden states directly within the same family and use a shared trainable map when widths differ\. The crossfade shifts each specialist’s supervision from latent states to token predictions on its own update clock, adapting the transient latent signal studied by LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)to domain\-routed multi\-teacher training\. All teachers remain frozen, and the student needs neither the teachers nor the map at inference\.
- 1Representation\-level multi\-teacher OPD\.We introduce, to our knowledge, the first representation\-level multi\-teacher OPD method for LLMs\. We identify effective layer choices for different teacher panels and show that grouping routed targets by domain stabilizes all\-layer representation\-only training\.
- 2Same\-family gains beyond individual teachers\.One1\.51\.5B student outperforms token\-only, representation\-only, and uniform\-averaging baselines on all nine benchmarks across math, code, and logic, and surpasses the per\-benchmark best teacher on five of them at the same model size \(Figure[1](https://arxiv.org/html/2610.02381#S1.F1); Section[4\.2](https://arxiv.org/html/2610.02381#S4.SS2)\)\. With the other settings fixed, crossfade improves eight of nine benchmarks over constant joint weighting\.
- 3Transfer across teacher panels and initializations\.With separately developed77B Qwen\-based teachers, Latent\-MOPD outperforms both single\-channel baselines on all six benchmarks \(Section[4\.3](https://arxiv.org/html/2610.02381#S4.SS3)\)\. Starting from a parameter\-merged initialization,6262distillation steps improve performance on five of six benchmarks across all three domains \(Section[4\.4](https://arxiv.org/html/2610.02381#S4.SS4)\)\. Further analyses examine layer selection, training dynamics, batch organization, and generalization \(Section[5](https://arxiv.org/html/2610.02381#S5)\)\.
Code:[https://github\.com/fangzy96/Latent\-MOPD](https://github.com/fangzy96/Latent-MOPD)\.
Figure 2:Three supervision designs on shared student\-generated prefixes\.\(a\)Uniform averaging gives each teacher’s token probabilities equal weight\.\(b\)MOPD\-style routes token supervision to one frozen teacher\.\(c\)Latent\-MOPD adds a representation target from the same teacher, matchinggψ\(zS\)g\_\{\\psi\}\(z^\{S\}\)to its hidden state\. PG denotes the policy\-gradient token update from outputs after the LM heads; bars are schematic distributions\. Routed columns illustrate a math prompt\. One layer pair is drawn: same\-family aligns the last three using the identity, cross\-family the last using a shared trainable linear map\. Both use crossfade \(Figure[4](https://arxiv.org/html/2610.02381#S3.F4)\)\.
## 2Related work
#### On\-policy and representation\-level distillation\.
OPD provides teacher supervision on student\-generated responses\([Agarwal et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib1);[Gu et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib10)\)\. Representation matching predates this setting: FitNets introduced intermediate hints, and subsequent methods aligned hidden states or attention statistics\([Romero et al\., 2015](https://arxiv.org/html/2610.02381#bib.bib32);[Sun et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib39);[Jiao et al\., 2020](https://arxiv.org/html/2610.02381#bib.bib17);[Wang et al\., 2020](https://arxiv.org/html/2610.02381#bib.bib43)\)\. OPRD\([Yang et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib49)\)brings hidden\-state alignment onto the student’s own rollouts, pairing layers by relative depth\. LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)studies latent collapse across sizes and proposes last\-layer supervision that crossfades to token\-only OPD\. PR\-OPD\([Li et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib21)\)aligns an agent’s hidden states with its own skill\-conditioned copy\. These studies focus on supervision from one teacher\. We study how to select and organize representation targets from several specialists\.
#### Multi\-teacher distillation\.
Off\-policy methods combine teacher outputs or align several teachers’ features on fixed training data\([You et al\., 2017](https://arxiv.org/html/2610.02381#bib.bib50);[Yuan et al\., 2021](https://arxiv.org/html/2610.02381#bib.bib52);[Formont et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib6)\); FuseLLM combines language models through their output distributions\([Wan et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib42)\)\. MOPD’s reported pipeline\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)trains domain specialists from a shared SFT checkpoint, then routes student rollouts to the frozen teachers for token\-level distillation\. Our experiments directly reuse existing specialist checkpoints and add hidden\-state targets from the same routed teacher\. MT\-SDPO verifies teacher candidates against answers\([He et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib13)\), while Open\-MOPD examines capability imbalance under domain routing\([Gao et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib9)\)\.
#### Model merging and composition\.
Model soups average compatible model weights\([Wortsman et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib45)\), while task arithmetic combines task\-specific parameter changes\([Ilharco et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib15)\)\. We also use a parameter merge to initialize the student for6262steps of distillation, testing whether distillation improves over the merged initialization while retaining a single deployable student\. Appendix[I](https://arxiv.org/html/2610.02381#A9)extends the comparison and discusses representation diagnostics\.
## 3Latent\-MOPD: Representation alignment from multiple teachers
Latent\-MOPD distillsNNfrozen teachers\{πTi\}i=1N\\\{\\pi\_\{T\_\{i\}\}\\\}\_\{i=1\}^\{N\}into one studentπθ\\pi\_\{\\theta\}\. Each promptxxcarries a domain labeld∈\{1,…,N\}d\\in\\\{1,\\dots,N\\\}selecting its specialist: math, code or logic here\. The student samples a response, the selected teacher evaluates it, and the student learns from both that teacher’s token predictions and selected hidden states\. The same teacher supplies token targets throughout the response and latent targets at the positions specified for each regime \(Figure[2](https://arxiv.org/html/2610.02381#S1.F2)\)\.
Lety^∼πθ\(⋅∣x\)\\hat\{y\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)denote a student response andMMthe valid response positions in a loss calculation\. At each prefixst=\(x,y^<t\)s\_\{t\}=\(x,\\hat\{y\}\_\{<t\}\),pθ,tp\_\{\\theta,t\}andpTi,tp\_\{T\_\{i\},t\}are next\-token distributions over the full vocabulary\. We writep¯t\\bar\{p\}\_\{t\}for the fixed student scoring distribution andVk,tV\_\{k,t\}for its top\-kksupport, withk=16k=16\. For a selected layer pair,ztSz^\{S\}\_\{t\}andztTiz^\{T\_\{i\}\}\_\{t\}are the hidden states; the final pair is taken after final normalization and before the LM heads\. We writeν\(z\)=z/max\(∥z∥2,ϵ\)\\nu\(z\)=z/\\max\(\\lVert z\\rVert\_\{2\},\\epsilon\)forℓ2\\ell\_\{2\}normalization andsg\\mathrm\{sg\}for stop\-gradient\. Appendix[C](https://arxiv.org/html/2610.02381#A3)gives the corresponding single\-teacher objectives\. Representation matching is used only during training; inference uses the student alone\.
Figure 3:Representation targets and update organization\.\(a,b\)Native linear CKA between matched block outputs on fixed base\-student responses before training\. Same\-family differences increase near the output \(magnified scale\); cross\-family similarities show a middle\-layer trough\. Hatching marks the selected depths; bands give one prompt\-bootstrap standard deviation \(Appendix[E](https://arxiv.org/html/2610.02381#A5)\)\.\(c\)Same\-family all\-layer representation\-only supervision remains stable with domain\-pure updates; interleaved updates collapse and finish far below the base \(dotted line\)\. Both use MATH\-500 with88samples per problem \(Appendix[F](https://arxiv.org/html/2610.02381#A6)\)\.#### Domain\-routed token channel \(MOPD\-style\)\.
Averaging teacher distributions \(Figure[2](https://arxiv.org/html/2610.02381#S1.F2)a\) can dilute a specialist’s confident prediction when the other teachers assign it less probability\. MOPD\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)instead routes each prompt to its domain’s teacherπTd\\pi\_\{T\_\{d\}\}\. This assignment stays fixed throughout the response\. We use domain\-pure batches of3232\(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c; Section[5](https://arxiv.org/html/2610.02381#S5.SS0.SSS0.Px3)\)\. At each response position, we score the student’s top\-kkcandidates using teacher\-to\-student log\-probability ratios\. Candidate weights renormalizep¯t\\bar\{p\}\_\{t\}withinVk,tV\_\{k,t\}, while both log probabilities retain full\-vocabulary normalization\. Scaling these rewards by their domain’s masked standard deviation gives fixedadvantagesAt,vA\_\{t,v\}\(Appendix[C](https://arxiv.org/html/2610.02381#A3)\)\. The token update is
ℒOPDroute=−1\|M\|∑t∈M∑v∈Vk,tsg\[At,v\]logpθ,t\(v\)\.\\mathcal\{L\}\_\{\\mathrm\{OPD\}\}^\{\\mathrm\{route\}\}\\;=\\;\-\\frac\{1\}\{\|M\|\}\\sum\_\{t\\in M\}\\sum\_\{v\\in V\_\{k,t\}\}\\mathrm\{sg\}\[A\_\{t,v\}\]\\,\\log p\_\{\\theta,t\}\(v\)\.\(1\)Each position receives one specialist’s token targets\. The representation channel additionally supervises hidden states upstream of the head\.
Figure 4:Crossfade across rotating teachers\.Domain\-pure updates rotate through math, code and logic\. Two selected cycles illustrate token weights rising and latent weights fading on each teacher’s own clock \(Equation[3](https://arxiv.org/html/2610.02381#S3.E3)\); bar lengths are schematic\. Both regimes use this schedule, with different windows\.
#### Selective late\-layer alignment\.
Same\-family teachers share the student’s architecture and initialization lineage\. Their native CKA remains high through most blocks and decreases near the output \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)a\)\. We align the final three layers, supported by the layer ablations in Section[5](https://arxiv.org/html/2610.02381#S5)\. Across scales and training lineages, the middle\-layer trough \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)b\) cautions against assuming correspondence at every depth\. We align the final state at each selected token: it integrates prediction\-relevant information from preceding computation and directly feeds the next\-token prediction head in each model, giving the target a common functional role\. Within each panel, the teachers share a hidden widthdTd\_\{T\}\. The mapgψ:ℝdS→ℝdTg\_\{\\psi\}:\\mathbb\{R\}^\{d\_\{S\}\}\\\!\\to\\mathbb\{R\}^\{d\_\{T\}\}is the identity in the same\-family setting and, by default, a trainable linear map shared across routed teachers and selected layers in the cross\-family setting\. The linear map is initialized by a ridge fit to domain\-routed student–teacher state pairs \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\. Selecting33of2828layers requires89\.3%89\.3\\%less selected\-state storage at fixed token count and precision \(Appendix[A](https://arxiv.org/html/2610.02381#A1.SS0.SSS0.Px12)\)\. For16,38416\{,\}384positions and BF16 width\-1,5361\{,\}536states, the student and teacher tensor copies total0\.280\.28GiB for three layers, versus2\.632\.63GiB for all2828\. For one selected layer pair, the weighted penalty at positionttis
ℓrep,troute=λdT∥ν\(gψ\(ztS\)\)−sg\[ν\(ztTd\)\]∥22\.\\ell^\{\\mathrm\{route\}\}\_\{\\mathrm\{rep\},t\}\\;=\\;\\frac\{\\lambda\}\{d\_\{T\}\}\\big\\lVert\\nu\\\!\\left\(g\_\{\\psi\}\(z^\{S\}\_\{t\}\)\\right\)\-\\mathrm\{sg\}\\\!\\left\[\\nu\(z^\{T\_\{d\}\}\_\{t\}\)\\right\]\\big\\rVert\_\{2\}^\{2\}\.\(2\)Hard domain routing selects only the prompt’s specialist: math responses receive math\-teacher states and code responses receive code\-teacher states\. The lossℒreproute\\mathcal\{L\}\_\{\\mathrm\{rep\}\}^\{\\mathrm\{route\}\}aggregates these penalties over all response positions at the last three depth\-matched layers in the same\-family setting, and over the final2,0002\{,\}000response positions at the last layer in the cross\-family setting \(all positions for shorter responses\)\. For each selected layer, we average over selected response\-token positions across participating responses, then average equally over layers\. Thus selecting three layers does not triple the representation weight\. Gradients update the student and, cross\-family, the shared map; every teacher stays frozen\. We use a fixed coefficientλ\\lambdawithin each training configuration, with no running loss\-share normalization \(reductions in Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\.
#### One schedule: a crossfade on a per\-teacher clock\.
Both regimes use a transient representation term, adopting the crossfade form of LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)\. The token channel ramps in as the representation channel fades out\. At global optimizer stepss, letuucount completed updates for the active teacher andTwT\_\{w\}denote its crossfade window:
ℒ\(s\)=α\(u\)ℒOPDroute\+β\(u\)ℒreproute,α\(u\)=min\(1,uTw\),β\(u\)=max\(0,1−uTw\)\.\\mathcal\{L\}^\{\(s\)\}\\;=\\;\\alpha\(u\)\\,\\mathcal\{L\}\_\{\\mathrm\{OPD\}\}^\{\\mathrm\{route\}\}\+\\beta\(u\)\\,\\mathcal\{L\}\_\{\\mathrm\{rep\}\}^\{\\mathrm\{route\}\},\\qquad\\alpha\(u\)=\\min\\\!\\Big\(1,\\tfrac\{u\}\{T\_\{w\}\}\\Big\),\\qquad\\beta\(u\)=\\max\\\!\\Big\(0,\\,1\-\\tfrac\{u\}\{T\_\{w\}\}\\Big\)\.\(3\)Each domain\-pure optimizer update uses one pair of weights; only the active teacher’s counter advances afterward\. Each teacher receives the same sequence of weights before switching to token\-only updates\. We useTw=10T\_\{w\}=10per\-teacher steps same\-family andTw=7T\_\{w\}=7cross\-family\. Under the three\-teacher rotation, these windows span3030and2121of the6262global steps, respectively \(Figure[4](https://arxiv.org/html/2610.02381#S3.F4)\)\. MOPD\-style retains a constant token\-loss weight throughout training; representation\-only baselines retain the latent objective for the full run without a token term\.
#### Initialization from a parameter merge\.
When the teachers share the student’s lineage, their weights can also be averaged into a zero\-training*parameter merge*\(a “model soup”,[Wortsman et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib45)\), combining the experts’ weights before distillation begins\. We use this merge to initialize the student for6262steps with the same\-family crossfade recipe \(Section[4\.4](https://arxiv.org/html/2610.02381#S4.SS4)\)\.
Table 1:Same\-family1\.51\.5B teachers, distilled into the1\.51\.5B student for6262steps\. Gray rows are teachers; blue rows are Latent\-MOPD, with darker blue marking the default last\-three\-layer setting\.Boldandunderlinemark the best and second\-best non\-teacher scores;†\\daggerexceeds the best teacher\.Norm: base=0=0, best\-teacher envelope=1=1\. Appendix[A](https://arxiv.org/html/2610.02381#A1)specifies evaluation, Norm and baseline implementation\.
## 4Experiments
### 4\.1Setup
#### Models, regimes and protocol\.
We distill existing specialist checkpoints into DeepSeek\-R1\-Distill\-Qwen\-1\.51\.5B\([Guo et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib11)\), without additional teacher training\. Our mainsame\-familyregime uses three RL\-tuned1\.51\.5B teachers from the same model lineage as the student, one per domain\.Cross\-familyuses three separately developed77B teachers with hidden width3,5843\{,\}584\(student:1,5361\{,\}536\)\. Here “cross\-family” denotes different scale and post\-training; all four models have2828blocks and use a Qwen backbone\.*Merge initialization*starts the student from the same\-family teachers’ parameter merge \(Table[3](https://arxiv.org/html/2610.02381#S4.T3)\)\. Training prompts come from DAPO\-Math\-17k, OpenCodeReasoning and Reasoning Gym\([Yu et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib51);[Ahmad et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib2);[Stojanovski et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib37)\)\. The pool contains17,85617\{,\}856prompts \(5,9525\{,\}952per domain\)\. Training and evaluation share no problem instances, as checked before training\. Domain\-pure blocks of3232rotate without shuffling:6262updates process1,9841\{,\}984prompt presentations and generate7,9367\{,\}936responses \(44per prompt\), without verifier rewards\. The main tables evaluate each trained student at its final checkpoint \(step6262\)\.
#### Evaluation and baselines\.
The main tables cover*math*,*code*and*logic*; Figure[5](https://arxiv.org/html/2610.02381#S4.F5)b tests general ability on suites excluded from distillation\. GYM is a frozen925925\-problem set with a1616k\-token budget \(scoring and datasets in Appendices[A](https://arxiv.org/html/2610.02381#A1)and[B](https://arxiv.org/html/2610.02381#A2)\)\. Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)compare the base student, teachers,*MOPD\-style \(token\-only\)*and representation\-only baselines\. The MOPD\-style baseline retains domain\-based teacher routing\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)\. Its token estimator matches Latent\-MOPD: student\-selected top\-kksupport with full\-vocabulary log\-probability ratios and student\-normalized candidate weights\.*OPRD\-style \(all layers\)*adapts[Yang et al\. \(2026b\)](https://arxiv.org/html/2610.02381#bib.bib49)to routed multi\-teacher supervision of all2828layers;*Rep\-only \(last 1\)*and*\(last 3\)*use the final one or three\. These baselines retain representation supervision throughout training without a token loss\.*Latent\-MOPD \(OPRD\-style\)*adds the routed token channel to all\-layer supervision, with both weights constant\. The routed methods share prompts, batches and optimizer, with the reported differences in teachers, initialization, objective and schedule\. Appendix[A](https://arxiv.org/html/2610.02381#A1)distinguishes our MOPD\-style implementation from the published system and specifies the*uniform averaging*protocol\. Table[3](https://arxiv.org/html/2610.02381#S4.T3)compares a zero\-training*parameter merge*with students distilled from it\. Norm measures aggregate improvement over the base divided by the summed base\-to\-best\-teacher gaps: base=0=0and best\-teacher envelope=1=1\.
### 4\.2Same\-family teachers: surpassing individual specialists
With three existing same\-family1\.51\.5B specialists, the default last\-three\-layer student surpasses the per\-benchmark best teacher on five of nine benchmarks: BBH, MuSR, MBPP, MBPP\+, and Minerva \(Table[1](https://arxiv.org/html/2610.02381#S3.T1)\)\. These gains span logic, code and math in one student with the same parameter count as each teacher\. Norm reaches1\.051\.05, above the best\-teacher envelope at11, compared with0\.900\.90for MOPD\-style,0\.640\.64–0\.730\.73for representation\-only baselines, and0\.940\.94for Latent\-MOPD \(OPRD\-style\)\.
Latent\-MOPD also improves over the token\-only, representation\-only and averaging baselines on all nine benchmarks\. Relative to MOPD\-style, Minerva rises from33\.533\.5to36\.336\.3, LiveCodeBench\-easy from68\.368\.3to73\.173\.1, and MuSR from50\.750\.7to52\.452\.4\. These gains use the same prompts and number of updates as the routed single\-channel baselines\. Among joint methods, crossfade improves eight of nine benchmarks over the matched constant\-weight, last\-three\-layer control \(Norm1\.051\.05vs\.0\.940\.94\)\. Our complete default configuration also outperforms Latent\-MOPD \(OPRD\-style\) on eight of nine benchmarks\.
### 4\.3Extension to cross\-family teachers
Table 2:Cross\-family77B teachers\.Skywork\-OR1\-Math, AceReason\-Nemotron and R1\-Distill\-7B, each∼2\.3×\\sim\\\!2\.3\\timesthe student’s hidden width, distilled into the same1\.51\.5B student\. Darker blue highlights the MLP variant with the highest Norm\. Score markings and Norm follow Table[1](https://arxiv.org/html/2610.02381#S3.T1)\.To test transfer beyond the same\-family setting, we use three separately developed77B teachers\. The default last\-layer Latent\-MOPD with a linear map outperforms token\-only and representation\-only baselines on all six benchmarks across math, code and logic \(Table[2](https://arxiv.org/html/2610.02381#S4.T2)\)\. Relative to MOPD\-style, AIME24 improves from36\.536\.5to39\.439\.4and MBPP from55\.455\.4to57\.957\.9\. Logic sees the largest gains: GYM rises from29\.229\.2to34\.134\.1and BBH from45\.845\.8to49\.549\.5\. The Norm score rises from0\.160\.16for MOPD\-style to0\.260\.26for Latent\-MOPD\. A two\-layer MLP also outperforms both single\-channel baselines on all six benchmarks \(Norm0\.270\.27\); the two projector configurations split benchmark wins three to three\.
### 4\.4Exploring a parameter\-merged initialization
We also explore distillation from a parameter\-merged initialization \(Table[3](https://arxiv.org/html/2610.02381#S4.T3)\)\. After6262steps, Latent\-MOPD improves five of six benchmarks across math, code and logic\. GYM rises from46\.346\.3to52\.252\.2, BBH from60\.660\.6to63\.463\.4, and AIME24 from48\.148\.1to53\.853\.8\. Norm increases from0\.870\.87to1\.031\.03, above token\-only \(0\.970\.97\) and representation\-only \(0\.910\.91\) training from the same merge\. The student also exceeds the best same\-family teacher on BBH, MBPP and MBPP\+\.
Table 3:Parameter\-merge initialization\.Uniform merging and6262\-step distillation on the same six benchmarks as Table[2](https://arxiv.org/html/2610.02381#S4.T2)\. Gray rows are teachers; blue marks Latent\-MOPD\. Bold and underline mark the best and second\-best non\-teacher scores, respectively\.Figure 5:Training dynamics, general retention and representation alignment\.\(a\)Default same\-family \(last\-three\-layer\) and cross\-family \(last\-layer\) Latent\-MOPD on MATH\-500, with88samples per problem\.\(b\)Same\-family general retention under00\-shot likelihood scoring\.\(c\)Same\-family raw representation loss, normalized to each domain’s first logged value and excludingβ\\beta\. Crossfade lasts1010updates per domain \(shaded\), followed by token\-only supervision\. Appendix[A](https://arxiv.org/html/2610.02381#A1)specifies scoring and loss measurement\.
## 5Ablations and analysis
Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)compare token\-only, representation\-only and joint supervision; Table[1](https://arxiv.org/html/2610.02381#S3.T1)also tests crossfade against constant weights\. We examine layer selection, update organization, training dynamics and teacher matching \(Figures[3](https://arxiv.org/html/2610.02381#S3.F3),[5](https://arxiv.org/html/2610.02381#S4.F5)and[6](https://arxiv.org/html/2610.02381#S5.F6)\); Appendix[H](https://arxiv.org/html/2610.02381#A8)compares solutions\. For the last\-layer and last\-three\-layer Latent\-MOPD comparisons, coefficient, response positions and crossfade are fixed within each regime\.
#### Layer selection depends on the teacher panel\.
In the same\-family setting, three\-layer supervision improves eight of nine benchmarks, raising Norm from0\.990\.99to1\.051\.05\(Table[1](https://arxiv.org/html/2610.02381#S3.T1)\)\. The representation\-only controls also favor three layers: they outperform the last\-layer and all\-layer alternatives on seven of nine benchmarks each \(Norm0\.730\.73vs\.0\.640\.64\)\. Cross\-family with linear projectors, last\-layer supervision reaches Norm0\.260\.26, versus0\.250\.25for three layers\. Benchmark wins split three to three; one layer retains the aggregate gain with fewer representation targets\. These comparisons test the layer choices motivated by native CKA \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)a,b\); Appendices[D](https://arxiv.org/html/2610.02381#A4)–[E](https://arxiv.org/html/2610.02381#A5)give further comparisons and masking diagnostics\.
#### Training dynamics across teacher panels\.
Figure[5](https://arxiv.org/html/2610.02381#S4.F5)a tracks each regime’s default configuration on500500MATH problems \(88samples each; Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\. From the base score of84\.184\.1, the same\-family run reaches90\.390\.3at step2020and92\.092\.0at step6262; cross\-family reaches86\.286\.2and87\.787\.7, respectively\. Both remain above the base from step2020through the later token\-only phase\.
#### Domain\-pure updates stabilize representation\-only training\.
Routing assigns a teacher to each prompt, but does not determine whether an optimizer step contains one domain or several\. We examine this choice in a same\-family all\-layer representation\-only control, holding teachers, loss, coefficient and budget fixed \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c\)\. Domain\-pure blocks of3232complete all6262steps, reaching91\.791\.7on in\-loop MATH\-500 validation\. Interleaving the same rows collapses MATH\-500 to22\.822\.8at step2020and9\.89\.8at step3030; it remains far below the base at step6262\(43\.243\.2vs\.84\.184\.1\)\. These results show that domain\-pure updates stabilize all\-layer representation\-only training under the tested configuration \(Appendix[F](https://arxiv.org/html/2610.02381#A6)\)\.
#### General\-benchmark performance stays close to the base\.
On three suites absent from the training pool, the same\-family student’s scores remain within1\.11\.1points of the base \(Figure[5](https://arxiv.org/html/2610.02381#S4.F5)b\): ARC\-Challenge and WinoGrande improve, while HellaSwag changes from44\.744\.7to44\.244\.2\. These tests score answer\-option likelihoods without sampling \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\.
#### Learning representation targets from three specialists\.
Figure[5](https://arxiv.org/html/2610.02381#S4.F5)c tracks the same\-family student’s raw representation loss on its training rollouts, normalized within each domain\. At update1010per domain, the math, code and logic losses fall to0\.300\.30,0\.380\.38and0\.390\.39of their first logged values\. All three remain low after representation supervision ends, reaching0\.130\.13,0\.190\.19and0\.150\.15at their final recorded updates\. These losses excludeβ\\beta, so their decline is not simply the prescribed weight decay\. The student reduces its mismatch with all three specialists without maintaining the latent objective throughout training \(measurement details in Appendix[A](https://arxiv.org/html/2610.02381#A1.SS0.SSS0.Px9)\)\.
Figure 6:Domain\-conditioned teacher matching after same\-family distillation\.\(a,b\)Mean next\-token probability cosine on teacher\-disagreement positions: input domains are rows and teachers are columns, on one shared color scale; outlines mark row maxima\.\(c\)Change in matching\-teacher margin \(Latent\-MOPD minus MOPD\-style\) on selected and all positions, with paired95%95\\%instance\-bootstrap intervals\. Appendix[G](https://arxiv.org/html/2610.02381#A7)defines the selection rule, margin and measurement protocol\.
#### Matching the routed specialist after training\.
Figure[6](https://arxiv.org/html/2610.02381#S5.F6)compares the final same\-family students with all three teachers on identical prefixes\. On teacher\-disagreement positions, the math row’s maximum shifts from the code teacher under MOPD\-style to the math teacher under Latent\-MOPD; code and logic retain their domain\-matched maxima\. Matching\-teacher margin measures mean similarity to the domain specialist minus the closest alternative\. The math margin increases by0\.06480\.0648on selected positions and0\.03270\.0327over all positions, with positive paired intervals in both cases \(protocol in Appendix[G](https://arxiv.org/html/2610.02381#A7)\)\.
#### Case studies\.
Appendix[H](https://arxiv.org/html/2610.02381#A8)compares MOPD\-style and Latent\-MOPD solutions with the routed specialist’s reference\. In these examples, Latent\-MOPD preserves the recurrence boundary, rewrite invariant, and counting constraints that MOPD\-style misses\. These differences motivate supervising hidden states, which can encode intermediate reasoning information\([Gurnee et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib12)\)\.
## 6Conclusion
Latent\-MOPD integrates existing specialists through routed late\-layer alignment, shared width projection, domain\-pure updates and per\-teacher crossfades\. The same\-family1\.51\.5B student outperforms token\-only and representation\-only baselines on all nine benchmarks across math, code and logic, and surpasses the per\-benchmark best teacher on five\. With cross\-family teachers, it outperforms both single\-channel baselines on all six benchmarks\. From a parameter\-merged initialization, it improves five of six benchmarks over the initial merge\. Teachers remain frozen; deployment uses only the student, with no projection map\. Appendix[J](https://arxiv.org/html/2610.02381#A10)discusses scope and extensions\.
## References
- Agarwal et al\. \(2024\)Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem\.On\-policy distillation of language models: Learning from self\-generated mistakes\.In*International Conference on Learning Representations*, volume 2024, pp\. 21246–21263, 2024\.
- Ahmad et al\. \(2025\)Wasi Uddin Ahmad, Sean Narenthiran, Somshubra Majumdar, Aleksander Ficek, Siddhartha Jain, Jocelyn Huang, Vahid Noroozi, and Boris Ginsburg\.OpenCodeReasoning: Advancing data distillation for competitive coding\.*arXiv preprint arXiv:2504\.01943*, 2025\.
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al\.Program synthesis with large language models\.*arXiv preprint arXiv:2108\.07732*, 2021\.
- Clark et al\. \(2018\)Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord\.Think you have solved question answering? Try ARC, the AI2 reasoning challenge\.*arXiv preprint arXiv:1803\.05457*, 2018\.
- Dasgupta & Cohn \(2025\)Sayantan Dasgupta and Trevor Cohn\.Improving language model distillation through hidden state matching\.In*International Conference on Learning Representations*, volume 2025, pp\. 19035–19049, 2025\.
- Formont et al\. \(2026\)Philippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger, Jackie CK Cheung, Ismail Ayed, Mohammadhadi Shateri, and Pablo Piantanida\.Learning task\-agnostic representations through multi\-teacher distillation\.*Advances in Neural Information Processing Systems*, 38:109702–109748, 2026\.
- Fu et al\. \(2026a\)Siming Fu, Haojun Xu, Ruizhe He, Zheming Fu, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, et al\.Poly\-OPD: Heterogeneous multi\-teacher on\-policy distillation for capability\-selectable flow models\.*arXiv preprint arXiv:2608\.04349*, 2026a\.
- Fu et al\. \(2026b\)Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan\-ang Gao, Yudong Wang, et al\.Rethinking on\-policy distillation of large language models II: One training example\.*arXiv preprint arXiv:2609\.04172*, 2026b\.
- Gao et al\. \(2026\)Huan\-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei\-Ying Ma, Ya\-Qin Zhang, and Hao Zhou\.Open\-MOPD: Diagnosing and fixing capability imbalance in multi\-teacher on\-policy distillation\.*arXiv preprint arXiv:2608\.19098*, 2026\.
- Gu et al\. \(2024\)Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang\.MiniLLM: Knowledge distillation of large language models\.In*International Conference on Learning Representations*, volume 2024, pp\. 32694–32717, 2024\.
- Guo et al\. \(2025\)Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al\.DeepSeek\-R1: Incentivizing reasoning capability in LLMs via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*, 2025\.
- Gurnee et al\. \(2026\)Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, et al\.Verbalizable representations form a global workspace in language models\.*arXiv preprint arXiv:2607\.15495*, 2026\.
- He et al\. \(2026\)Xixiang He, Xingming Li, Baiqi Wu, Qiyao Sun, Xuanyu Ji, Ao Cheng, and Qingyong Hu\.Learn from whoever is right: Answer\-verified multi\-teacher distillation for multi\-domain LLMs\.*arXiv preprint arXiv:2609\.02548*, 2026\.
- Hendrycks et al\. \(2021\)Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the MATH dataset\.*arXiv preprint arXiv:2103\.03874*, 2021\.
- Ilharco et al\. \(2022\)Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi\.Editing models with task arithmetic\.*arXiv preprint arXiv:2212\.04089*, 2022\.
- Jain et al\. \(2025\)Naman Jain, King Han, Alex Gu, Wen\-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar\-Lezama, Koushik Sen, and Ion Stoica\.LiveCodeBench: Holistic and contamination free evaluation of large language models for code\.In*International Conference on Learning Representations*, volume 2025, pp\. 58791–58831, 2025\.
- Jiao et al\. \(2020\)Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu\.TinyBERT: Distilling BERT for natural language understanding\.In*Findings of the association for computational linguistics: EMNLP 2020*, pp\. 4163–4174, 2020\.
- Kornblith et al\. \(2019\)Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton\.Similarity of neural network representations revisited\.In*International conference on machine learning*, pp\. 3519–3529\. PMLR, 2019\.
- Lewkowycz et al\. \(2022\)Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman\-Solo, et al\.Solving quantitative reasoning problems with language models\.*Advances in neural information processing systems*, 35:3843–3857, 2022\.
- Li et al\. \(2026a\)Menghao Li, Linjie Mu, Yin Wang, Haotian Hu, Yannian Gu, Lujiayi Xue, and Fanyi Wang\.CA\-OPD: Confidence\-aware on\-policy distillation for structured visual prediction\.*arXiv preprint arXiv:2609\.02401*, 2026a\.
- Li et al\. \(2026b\)Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu, Zhengkun Xiao, Ruining Deng, Zhe Jiang, and Shigang Chen\.PR\-OPD: Privileged representation on\-policy self\-distillation for agentic reinforcement learning, 2026b\.URL[https://arxiv\.org/abs/2609\.36642](https://arxiv.org/abs/2609.36642)\.
- Li et al\. \(2026c\)Yuhan Li, Mingxu Zhang, Dazhong Shen, and Ying Sun\.PHF: Privileged hidden flow for on\-policy self\-distillation\.*arXiv preprint arXiv:2606\.29340*, 2026c\.
- Lian et al\. \(2026\)Niu Lian, Alan Chen, Zhehao Yu, Chengzhen Duan, Fazhan Liu, Hui Liu, Pei Fu, Jian Luan, Yaowei Wang, Shu\-Tao Xia, et al\.UI\-MOPD: Multi\-platform on\-policy distillation for continual GUI agent learning\.*arXiv preprint arXiv:2607\.04425*, 2026\.
- Lightman et al\. \(2024\)Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*International Conference on Learning Representations*, volume 2024, pp\. 39578–39601, 2024\.
- Lindsey et al\. \(2025\)Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, et al\.On the biology of a large language model\.*Transformer Circuits Thread*, 2025\.URL[https://transformer\-circuits\.pub/2025/attribution\-graphs/biology\.html](https://transformer-circuits.pub/2025/attribution-graphs/biology.html)\.
- Liu et al\. \(2023\)Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang\.Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation\.*Advances in neural information processing systems*, 36:21558–21572, 2023\.
- Liu et al\. \(2020\)Yuang Liu, Wei Zhang, and Jun Wang\.Adaptive multi\-teacher multi\-level knowledge distillation\.*Neurocomputing*, 415:106–113, 2020\.
- Liu et al\. \(2026\)Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, and Cheng Luo\.Preserving general capabilities during domain specialization with uncertainty\-calibrated MOPD\.*arXiv preprint arXiv:2608\.26735*, 2026\.
- Ma et al\. \(2026\)Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, et al\.MOPD: Multi\-teacher on\-policy distillation for capability integration in LLM post\-training\.*arXiv preprint arXiv:2606\.30406*, 2026\.
- Niu et al\. \(2026\)Yifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, and Jia Li\.Breaking the tokenizer barrier: On\-policy distillation across model families\.*arXiv preprint arXiv:2606\.09456*, 2026\.
- Ouyang et al\. \(2022\)Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al\.Training language models to follow instructions with human feedback\.*Advances in neural information processing systems*, 35:27730–27744, 2022\.
- Romero et al\. \(2015\)Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio\.FitNets: Hints for thin deep nets, 2015\.URL[https://arxiv\.org/abs/1412\.6550](https://arxiv.org/abs/1412.6550)\.
- Sakaguchi et al\. \(2021\)Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi\.WinoGrande: An adversarial Winograd schema challenge at scale\.*Communications of the ACM*, 64\(9\):99–106, 2021\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al\.DeepSeekMath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shen et al\. \(2026\)Ao Shen, Yongheng Zhang, Yinghui Li, Manning Wang, Di Yin, and Xing Sun\.Deep thought alignment: Trajectory\-level latent distillation for video reasoning\.*arXiv preprint arXiv:2608\.16316*, 2026\.
- Sprague et al\. \(2024\)Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett\.MuSR: Testing the limits of chain\-of\-thought with multistep soft reasoning\.In*International Conference on Learning Representations*, volume 2024, pp\. 14670–14728, 2024\.
- Stojanovski et al\. \(2026\)Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kaddour, and Andreas Köpf\.Reasoning Gym: Reasoning environments for reinforcement learning with verifiable rewards\.*Advances in Neural Information Processing Systems*, 38, 2026\.
- Sun et al\. \(2024\)Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu\.Massive activations in large language models\.*arXiv preprint arXiv:2402\.17762*, 2024\.
- Sun et al\. \(2019\)Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu\.Patient knowledge distillation for BERT model compression\.In*Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(EMNLP\-IJCNLP\)*, pp\. 4323–4332, 2019\.
- Sun et al\. \(2026\)Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, and Min Zhang\.D3\-MOPD: Adaptive dynamic domain scheduling for efficient multi\-teacher distillation\.*arXiv preprint arXiv:2608\.24987*, 2026\.
- Suzgun et al\. \(2023\)Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed H Chi, Denny Zhou, et al\.Challenging BIG\-Bench tasks and whether chain\-of\-thought can solve them\.In*Findings of the Association for Computational Linguistics: ACL 2023*, pp\. 13003–13051, 2023\.
- Wan et al\. \(2024\)Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, and Shuming Shi\.Knowledge fusion of large language models\.In*International Conference on Learning Representations*, volume 2024, pp\. 18303–18322, 2024\.
- Wang et al\. \(2020\)Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou\.MiniLM: Deep self\-attention distillation for task\-agnostic compression of pre\-trained transformers\.*Advances in neural information processing systems*, 33:5776–5788, 2020\.
- Wei et al\. \(2026\)Qingyan Wei, Guangzhao Li, Xiaobing Tu, Yinggui Wang, Xiantao Zhang, Jinkui Ren, Xiaohong Liu, and Linfeng Zhang\.STEP\-OPD: Rethinking output targets and internal dynamics in on\-policy distillation for diffusion models\.*arXiv preprint arXiv:2608\.04887*, 2026\.
- Wortsman et al\. \(2022\)Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo\-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al\.Model soups: averaging weights of multiple fine\-tuned models improves accuracy without increasing inference time\.In*International conference on machine learning*, pp\. 23965–23998\. PMLR, 2022\.
- Wu et al\. \(2026\)Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng\-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, et al\.Consolidating RLVR capabilities across domains: A deep dive into fusion paradigms\.*arXiv preprint arXiv:2608\.27409*, 2026\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Yang et al\. \(2026a\)Jie Yang, Zhengyu Fang, Zelin Xu, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Qinghua Liu, Yiwei Cai, and Yan Zheng\.LastOPD: Taming collapse in latent on\-policy distillation, 2026a\.URL[https://arxiv\.org/abs/2609\.28845](https://arxiv.org/abs/2609.28845)\.
- Yang et al\. \(2026b\)Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Junbo Zhao, et al\.OPRD: On\-policy representation distillation\.*arXiv preprint arXiv:2606\.06021*, 2026b\.
- You et al\. \(2017\)Shan You, Chang Xu, Chao Xu, and Dacheng Tao\.Learning from multiple teacher networks\.In*Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining*, pp\. 1285–1294, 2017\.
- Yu et al\. \(2026\)Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al\.DAPO: An open\-source LLM reinforcement learning system at scale\.*Advances in Neural Information Processing Systems*, 38:113222–113244, 2026\.
- Yuan et al\. \(2021\)Fei Yuan, Linjun Shou, Jian Pei, Wutao Lin, Ming Gong, Yan Fu, and Daxin Jiang\.Reinforced multi\-teacher selection for knowledge distillation\.In*Proceedings of the AAAI conference on artificial intelligence*, volume 35, pp\. 14284–14291, 2021\.
- Zellers et al\. \(2019\)Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi\.HellaSwag: Can a machine really finish your sentence?In*Proceedings of the 57th annual meeting of the association for computational linguistics*, pp\. 4791–4800, 2019\.
## Appendix AExperimental setup and reproduction
#### Models, teachers, and routing\.
The student is DeepSeek\-R1\-Distill\-Qwen\-1\.51\.5B\([Guo et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib11)\)\. The same\-family panel is JustRL\-DeepSeek\-1\.51\.5B \(math\), Archer2\-Code\-1\.51\.5B \(code\), and Nemotron\-ProRL\-1\.51\.5B \(logic\), all RL\-tuned from the student’s lineage\.111Public same\-family checkpoints \(ProRL at revisionv2\):[https://huggingface\.co/hbx/JustRL\-DeepSeek\-1\.5B](https://huggingface.co/hbx/JustRL-DeepSeek-1.5B);[https://huggingface\.co/Fate\-Zero/Archer2\.0\-Code\-1\.5B\-Preview](https://huggingface.co/Fate-Zero/Archer2.0-Code-1.5B-Preview);[https://huggingface\.co/nvidia/Nemotron\-Research\-Reasoning\-Qwen\-1\.5B](https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B/tree/v2)\.The cross\-family panel is Skywork\-OR1\-Math\-77B \(math\), AceReason\-Nemotron\-77B \(code\), and DeepSeek\-R1\-Distill\-Qwen\-77B \(logic\)\.222Public cross\-family checkpoints \(Skywork\-OR1\-Math\-7B and AceReason\-Nemotron\-7B are RL\-tuned from DeepSeek\-R1\-Distill\-Qwen\-7B\):[https://huggingface\.co/Skywork/Skywork\-OR1\-Math\-7B](https://huggingface.co/Skywork/Skywork-OR1-Math-7B);[https://huggingface\.co/nvidia/AceReason\-Nemotron\-7B](https://huggingface.co/nvidia/AceReason-Nemotron-7B);[https://huggingface\.co/deepseek\-ai/DeepSeek\-R1\-Distill\-Qwen\-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B)\.All teachers are existing public checkpoints, reused without additional teacher training for this study\. Each training prompt carries a domain label in\{\\\{math, code, logic\}\\\}\. Both channels are routed by that label to the*same*per\-domain teacher: in the same\-family panel, math→\\toJustRL, code→\\toArcher2, and logic→\\toProRL\. The domain\-routed token channel \(MOPD\-style, Equation[1](https://arxiv.org/html/2610.02381#S3.E1)\) scores the tokens of each student response against the teacher for its prompt’s domain, in domain\-pure batches; the domain\-routed representation channel \(Equation[2](https://arxiv.org/html/2610.02381#S3.E2)\) uses that teacher’s selected hidden states as its target\. The parameter merge of Section[3](https://arxiv.org/html/2610.02381#S3)is the uniform weight average of the three same\-family teachers\.
#### Optimization and rollout protocol\.
Every optimizer step samples44responses at temperature1\.01\.0for each of3232prompts and applies a constant learning rate of10−510^\{\-5\}; the6262steps generate7,9367\{,\}936responses, matching the on\-policy budget of LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)\. We use AdamW with betas\(0\.9,0\.999\)\(0\.9,0\.999\), weight decay0\.010\.01, andϵ=10−8\\epsilon=10^\{\-8\}\. The routed token\-only, representation\-only and combined methods share prompts, batches and optimizer, with the teacher panel, initialization, objective and schedule specified for each arm\. Uniform averaging follows the same protocol\. Rollouts and evaluation apply the model’s chat template withenable\_thinking=False; during distillation, teachers score the resulting student responses\. Training responses are capped at16,38416\{,\}384tokens, with repetition penalty1\.01\.0\. The token update uses the student’s top\-1616candidates and the fixed\-advantage surrogate in Equation[1](https://arxiv.org/html/2610.02381#S3.E1)\.
#### Prompt pool and update counts\.
The pool contains17,85617\{,\}856prompts:5,9525\{,\}952each for math, code and logic\. We disable data shuffling and use domain\-pure blocks of3232in math–code–logic order\. The6262\-step budget therefore contains2121,2121and2020updates from the three domains, respectively:672672,672672and640640prompt presentations, for1,9841\{,\}984in total\. Four responses per prompt give128128generated responses per update and7,9367\{,\}936over the run\. These are counts of prompt presentations and generated responses; the pool size does not denote the number of prompts processed within the6262\-step budget\.
#### Token\-loss reduction and accumulation\.
The token channel uses the single\-epoch policy\-gradient surrogate in Equation[1](https://arxiv.org/html/2610.02381#S3.E1)\. For each microbatch, it sums over the1616candidate tokens at each position, then takes a masked mean over valid response positions across its responses; the denominator is the valid position count plus10−810^\{\-8\}\. Padding contributes neither loss nor count\. Responses therefore contribute in proportion to their valid lengths within this calculation\. Domain\-wise advantage scales are computed on the rollout batch before microbatch packing \(Appendix[C](https://arxiv.org/html/2610.02381#A3)\)\. On the eight\-worker training setup, each worker has a local optimizer minibatch of1616responses\. A dynamically packed microbatch withncn\_\{c\}responses contributesnc/16n\_\{c\}/16times its loss to gradient accumulation\. This factor multiplies the complete microbatch loss after applying the configured channel weights\. The actor makes one optimizer update after accumulating the local minibatch, with one optimization epoch per rollout batch\.
#### Position and layer normalization\.
For responsebb, letMbM\_\{b\}be its valid response\-token positions andSb⊆MbS\_\{b\}\\subseteq M\_\{b\}the positions selected for representation supervision\. In Latent\-MOPD,Sb=MbS\_\{b\}=M\_\{b\}same\-family; cross\-family,SbS\_\{b\}contains the lastmin\(2000,\|Mb\|\)\\min\(2000,\|M\_\{b\}\|\)valid response positions\. Padding and unselected positions do not enter the denominator\. Let𝒞\\mathcal\{C\}denote the responses participating in one masked\-mean loss calculation and𝒜\\mathcal\{A\}the selected layer pairs\. Writingℓrep,b,t\(ℓ\)\\ell^\{\(\\ell\)\}\_\{\\mathrm\{rep\},b,t\}for the routed penalty in Equation[2](https://arxiv.org/html/2610.02381#S3.E2)at layer pairℓ\\ell, the position and layer reduction is
R𝒞=1\|𝒜\|∑ℓ∈𝒜∑b∈𝒞∑t∈Sbℓrep,b,t\(ℓ\)∑b∈𝒞\|Sb\|\.R\_\{\\mathcal\{C\}\}=\\frac\{1\}\{\|\\mathcal\{A\}\|\}\\sum\_\{\\ell\\in\\mathcal\{A\}\}\\frac\{\\displaystyle\\sum\_\{b\\in\\mathcal\{C\}\}\\sum\_\{t\\in S\_\{b\}\}\\ell^\{\(\\ell\)\}\_\{\\mathrm\{rep\},b,t\}\}\{\\displaystyle\\sum\_\{b\\in\\mathcal\{C\}\}\|S\_\{b\}\|\}\.\(4\)The selected layer pairs are averaged, so using three layers does not multiply the loss coefficient by three\. Within this calculation, each selected token has equal weight; responses contribute in proportion to their selected\-token counts\. This formula describes the reduction inside the representation\-loss calculation, separately from accumulation across optimizer microbatches and devices\. The token objective uses all valid response positions in both regimes\.
#### Representation targets and crossfade\.
The routed representation penalty is normalized by feature norm and hidden width \(Equation[2](https://arxiv.org/html/2610.02381#S3.E2)\)\. For the crossfade variants, we use fixed coefficientsλ=1\\lambda=1same\-family andλ=2000\\lambda=2000cross\-family, with no running loss\-share normalization\. The crossfade of Equation[3](https://arxiv.org/html/2610.02381#S3.E3)multiplies the representation term byβ\(u\)\\beta\(u\)and the token term byα\(u\)\\alpha\(u\), withTw=10T\_\{w\}=10*per\-teacher*steps for the same\-family panel andTw=7T\_\{w\}=7for the cross\-family one \(the three domains rotate per block, so the fades span3030and2121global optimizer steps\)\. The default same\-family configuration applies the representation term at each of the last three layers and all response positions\. The cross\-family default uses the last layer and the final2,0002\{,\}000response positions, or all positions when the response is shorter\. Each regime also includes a layer\-selection variant with the same coefficient, position scope and crossfade window: last\-layer supervision same\-family, and last\-three\-layer supervision cross\-family\. Layer losses are averaged as above\. The respective per\-worker packing budgets are24,57624\{,\}576and18,43218\{,\}432tokens\.
#### Teacher\-local clock\.
The counteruuin Equation[3](https://arxiv.org/html/2610.02381#S3.E3)is the number of updates already completed for the active teacher, so its first update usesu=0u=0\. Each optimizer update uses a fixed pair\(α\(u\),β\(u\)\)\(\\alpha\(u\),\\beta\(u\)\)across all its responses and token positions; only the active teacher’s counter advances after that update\. Ifs=1,…,62s=1,\\ldots,62is the global step, the math–code–logic rotation givesu=⌊\(s−1\)/3⌋u=\\lfloor\(s\-1\)/3\\rfloorfor the active teacher\. Thus the first three global steps useα=0\\alpha=0andβ=1\\beta=1\. Same\-family, the representation term remains active through step3030and is zero from step3131onward; cross\-family, it remains active through step2121and is zero from step2222onward\. All three teachers receive the same sequence of representation weights during their respective crossfade windows\.
#### Constant\-weight schedule control\.
The*Latent\-MOPD \(constant\)*row in Table[1](https://arxiv.org/html/2610.02381#S3.T1)matches the default same\-family last\-three\-layer configuration except thatα\(u\)=β\(u\)=1\\alpha\(u\)=\\beta\(u\)=1for all6262updates\. It usesλ=1\\lambda=1, all valid response positions, the identity map and a24,57624\{,\}576\-token per\-worker packing budget, with the same teachers, prompt pool, domain\-pure batches and optimizer\. This comparison tests the paired channel schedule against sustained joint supervision\.
#### Representation\-alignment dynamics\.
Figure[5](https://arxiv.org/html/2610.02381#S4.F5)c uses the default same\-family last\-three\-layer run\. At each domain\-pure optimizer step, the logged raw representation loss belongs to the active domain’s teacher and excludes both the coefficientλ\\lambdaand the schedule weightβ\\beta; hereλ=1\\lambda=1\. Letrd,jr\_\{d,j\}denote this logged loss on thejjth update for domaindd\. We plotrd,j/rd,1r\_\{d,j\}/r\_\{d,1\}, with separate denominators for math, code and logic:3\.90603×10−53\.90603\\times 10^\{\-5\},3\.465517×10−63\.465517\\times 10^\{\-6\}and1\.770245×10−51\.770245\\times 10^\{\-5\}, respectively\. These first observations occur at global steps11,22and33; no pre\-update \(step\-00\) value is prepended\. The6262global steps supply2121,2121and2020updates for the three domains\. All recorded points are shown without smoothing\. The first1010updates of each domain use crossfade, andβ=0\\beta=0from its1111th update onward; teachers remain frozen\. We retain the recordedactor/rep\_lossscalar and its logging reduction across microbatch and worker records, rather than recomputing a global token\-pooled loss\. These are measurements on evolving on\-policy rollouts, rather than a fixed validation input set\. Their normalized values describe within\-domain training trajectories and do not compare absolute representation\-error scales between teachers\.
#### Merge\-initialized training\.
All three distilled students in Table[3](https://arxiv.org/html/2610.02381#S4.T3)start from the uniform average of the same\-family teachers and train for6262domain\-pure updates\. Latent\-MOPD usesλ=1\\lambda=1, the last three layers, all valid response positions and a1010\-step per\-teacher crossfade on the three\-domain cycle\. Its per\-worker packing budget is18,43218\{,\}432tokens for policy updates and log\-probability scoring; each generated response is capped at16,38416\{,\}384tokens\. Representation\-only supervises the final three layers withλ=2000\\lambda=2000throughout training\.
#### Projector construction\.
Same\-family runs use the identity map\. The cross\-family default uses a bias\-free linear mapgψ:ℝ1536→ℝ3584g\_\{\\psi\}:\\mathbb\{R\}^\{1536\}\\\!\\to\\\!\\mathbb\{R\}^\{3584\}\. Each worker maintains one map shared across its routed teacher targets and selected layer pairs\. The map is initialized from a closed\-form ridge fit:4242student rollouts provide134,843134\{,\}843domain\-routed pairs of student and teacher last\-layer states at response positions\. A random90%90\\%split suppliesSSandTTforW=\(S⊤S\+λrI\)−1S⊤TW=\(S^\{\\top\}S\+\\lambda\_\{r\}I\)^\{\-1\}S^\{\\top\}T, withλr=10−3dS\\lambda\_\{r\}=10^\{\-3\}d\_\{S\}\. The map is optimized jointly with the student and stored per worker\. The default cross\-family run reloads the ridge initialization at process restarts\. The map is discarded after training; inference uses only the student\.
The last\-layer MLP variant in Table[2](https://arxiv.org/html/2610.02381#S4.T2)replaces the linear map with two bias\-free linear layers,→→35841536\\\!\\to\\\!6144\\\!\\to\\\!3584, with GELU between them\. It uses the same per\-worker map sharing across routed targets and is randomly initialized, whereas the default linear map uses the ridge initialization\. The MLP has31,457,28031\{,\}457\{,\}280parameters versus5,505,0245\{,\}505\{,\}024for the linear map \(5\.71×5\.71\\times\)\. Other settings are unchanged; the comparison changes both projector architecture and initialization\.
#### Storage of selected representations\.
For one routed active teacher, storing one copy each of the selected student and teacher states requiresM=LN\(dSbS\+dTbT\)M=LN\(d\_\{S\}b\_\{S\}\+d\_\{T\}b\_\{T\}\)bytes, whereLLis the number of selected layer pairs,NNis the number of concurrently represented response positions, andbS,bTb\_\{S\},b\_\{T\}are bytes per element\. Holding the other quantities fixed, selecting three rather than2828layers reduces this storage by1−3/28≈89\.3%1\-3/28\\approx 89\.3\\%\. For example, withN=16,384N=16\{,\}384,dS=dT=1,536d\_\{S\}=d\_\{T\}=1\{,\}536and BF16 storage \(bS=bT=2b\_\{S\}=b\_\{T\}=2\), the count is about2\.632\.63GiB for all2828layers and0\.280\.28GiB for the last three; FP32 storage doubles both values\. These are analytic feature\-storage counts, not measured peak GPU memory\. They exclude model weights, optimizer state, attention activations, normalization and projection buffers, and other autograd storage\. The student still backpropagates through the full backbone\. Materializing all hidden states before selecting layers can also retain the full transient allocation, so the storage ratio does not imply the same reduction in training peak memory\.
#### Evaluation and checkpoint selection\.
Main benchmark scores are measured at the final checkpoint, by default at temperature0\.70\.7and top\-pp0\.950\.95; math and reasoning suites use sampled means \(AIME24/25 with1616samples per problem; Minerva with44; BBH, MuSR and Reasoning Gym with11, relying on problem count; LiveCodeBench with44at temperature0\.60\.6\), and the general suites use answer\-option accuracy, using length\-normalized log likelihoods for ARC\-Challenge and HellaSwag and unnormalized log likelihoods for WinoGrande\. All AIME25 evaluations use temperature0\.60\.6,1616samples per problem, top\-pp0\.950\.95and a31,74431\{,\}744\-token generation cap\. AIME24 and Minerva contain3030and272272problems, respectively, and both use a31,74431\{,\}744\-token generation cap\. We never select the best intermediate checkpoint\.
#### In\-loop validation\.
MATH\-500 and GYM\-200 validation use500500and200200fixed problems, respectively, with eight sampled responses per problem, temperature0\.70\.7, top\-pp0\.950\.95and a16,38416\{,\}384\-token generation cap\. Scores average correctness over responses and problems\.
#### Offline evaluation harness\.
For MBPP and MBPP\+, we generate22programs for each of378378tasks, with a16,38416\{,\}384\-token cap, and score the same programs using EvalPlus 0\.3\.1\.333[https://github\.com/evalplus/evalplus](https://github.com/evalplus/evalplus)We report pass@1, averaging correctness over samples and tasks\. MBPP requires passing the base tests; MBPP\+ requires passing both the base and additional tests\. Failed extraction counts as incorrect\. For LiveCodeBench, generations are produced for all165165easy problems dated on or after 2024\-01\-01, and pass@1 is scored on the6767easy problems dated on or after 2024\-08\-01 with the official test suites \(44samples per problem, temperature0\.60\.6, top\-pp0\.950\.95; LiveCodeBench uses a16,38416\{,\}384\-token budget because these long\-chain models truncate heavily on its problems at8,1928\{,\}192tokens\)\. BBH uses all2727task configurations with the first4040problems of each \(1,0801\{,\}080problems, identical for every model\), one sample per problem at temperature0\.70\.7with an8,1928\{,\}192\-token budget, answers extracted from`\\boxed\{\}`with normalized matching\. MuSR uses750750problems spanning its three subtasks \(murder mysteries, object placements and team allocation\), with one sample per problem and an8,1928\{,\}192\-token budget\. GYM uses our frozen Reasoning Gym evaluation set \(3838tasks×25\\times\\,25problems, generator seed42424242,reasoning\-gym0\.1\.19\) under a16,38416\{,\}384\-token budget, scored by each task’s own verifier\. A byte\-level check against the frozen questions excludes one task, leaving925925scored problems \(3737tasks\)\. Generators are frozen to files because seeds do not reproduce across registry versions\. Environment pins:torch2\.8\.0,vllm0\.11\.0,transformers4\.57\.3\.
#### Baseline implementation\.
Our MOPD\-style baseline retains MOPD’s domain routing and uses the same token estimator, prompts, batches, optimizer and budget as Latent\-MOPD\. It is not a replication of the published system\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\), which trains same\-origin teachers at a much larger scale on a different domain mix and uses a policy\-gradient reverse\-KL estimator by default\. The published system’s top\-kkvariant uses teacher\-selected support and a correction for truncation\. Our token channel uses student\-selected support, full\-vocabulary log\-probability ratios and student\-normalized candidate weights in a fixed\-advantage policy\-gradient update \(Appendix[C](https://arxiv.org/html/2610.02381#A3)\)\. Our Norm also uses the headroom\-weighted definition below, not MOPD’s uniform average of per\-domain normalized scores\.
Uniform averaging replaces the routed specialist’s token distribution with the equally weighted mean of all three teachers’ probabilities on each student\-generated response\. It otherwise uses the same training and evaluation protocol as MOPD\-style: teacher checkpoints, student initialization, prompt pool, domain\-pure batch order, optimizer, update budget, token estimator, response limits, numerical precision and microbatch packing\. Both token\-only baselines retain a constant token\-loss weight and are evaluated at the final checkpoint\. Their comparison isolates how the teacher target is formed\.
The same\-family*OPRD\-style \(all layers\)*baseline adapts representation distillation\([Yang et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib49)\)to domain\-routed teachers, supervising all2828layers;*Rep\-only \(last 3\)*and*Rep\-only \(last 1\)*use the final three and one, respectively\. All three use the finalmin\(2000,\|Mb\|\)\\min\(2000,\|M\_\{b\}\|\)valid positions per response, fixedλ=2000\\lambda=2000for6262steps, and an18,43218\{,\}432\-token per\-worker microbatch packing budget\. Cross\-family,*Rep\-only \(last 1\)*matches default Latent\-MOPD in its ridge\-initialized, bias\-free linear projector, per\-worker map sharing, joint student–projector optimization, last\-layer supervision, finalmin\(2000,\|Mb\|\)\\min\(2000,\|M\_\{b\}\|\)valid response positions,λ=2000\\lambda=2000,18,43218\{,\}432\-token per\-worker packing budget and remaining training settings\. All representation\-only baselines retain the representation objective throughout all6262updates, without token loss or crossfade\. The same\-family*Latent\-MOPD \(OPRD\-style\)*variant combines our routed token channel with all2828layers, the last2,0002\{,\}000valid positions,λ=2000\\lambda=2000and an18,43218\{,\}432\-token packing budget\. Both channel weights remain constant for all6262steps\. Its coefficient, layers, positions, schedule and packing budget thus differ from the default\.
#### Headroom\-weighted normalization\.
Within each of Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[3](https://arxiv.org/html/2610.02381#S4.T3), letbjb\_\{j\}be the base\-student score anduju\_\{j\}the best individual teacher’s score on benchmarkjj\. We report
Norm\(x\)=∑j\(xj−bj\)∑j\(uj−bj\)\.\\mathrm\{Norm\}\(x\)=\\frac\{\\sum\_\{j\}\(x\_\{j\}\-b\_\{j\}\)\}\{\\sum\_\{j\}\(u\_\{j\}\-b\_\{j\}\)\}\.\(5\)The sums run over the benchmarks reported in that table: the base student scores00and the per\-benchmark best\-teacher envelope scores11\. This weights each benchmark’s fraction of closed headroom by its available headroom, avoiding disproportionate influence from benchmarks where the teacher is only slightly above the base\.
## Appendix BDatasets and evaluation coverage
The datasets serve three purposes: supplying training prompts, tracking optimization during training, and measuring final\-checkpoint performance\. We describe their task content below; Appendix[A](https://arxiv.org/html/2610.02381#A1)specifies sampling, scoring and harness settings\.
#### Training prompts\.
DAPO\-Math\-17k supplies competition\-style mathematical problems with integer answers\([Yu et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib51)\)\. OpenCodeReasoning provides competitive\-programming questions drawn from multiple programming platforms, together with synthetic reasoning and Python solutions\([Ahmad et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib2)\)\. Reasoning Gym is a library of procedural problem generators and task\-specific answer verifiers, covering such skills as arithmetic, symbolic manipulation, logic and games\([Stojanovski et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib37)\)\. We use these sources for their*prompts*: the student generates its own training trajectories, and the routed teacher provides supervision on those trajectories\. Dataset solutions are not imitation targets, and answer verifiers do not supply a training reward\. The pool contains17,85617\{,\}856prompts,5,9525\{,\}952per domain\. Training and evaluation contain no shared problem instances, as checked before training\. The6262\-step training budget uses1,9841\{,\}984prompt presentations and generates7,9367\{,\}936responses; these counts describe the training budget, not the full pool\.
#### Mathematics\.
MATH\-500 is a500500\-problem subset of the MATH benchmark\([Hendrycks et al\., 2021](https://arxiv.org/html/2610.02381#bib.bib14);[Lightman et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib24)\), whose competition problems span algebra, geometry, number theory and other mathematical topics\. Here it serves as an in\-training validation set for training dynamics and the analysis of batching stability \(Figures[5](https://arxiv.org/html/2610.02381#S4.F5)a and[3](https://arxiv.org/html/2610.02381#S3.F3)c\)\. AIME24 and AIME25 contain problems from the 2024 and 2025 American Invitational Mathematics Examination\. Both examinations use integer\-valued answers\.444Official competition descriptions:[https://maa\.org/maa\-invitational\-competitions/](https://maa.org/maa-invitational-competitions/)\.The Minerva suite tests quantitative problem solving\([Lewkowycz et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib19)\)\.
#### Code generation\.
LiveCodeBench collects recent programming\-contest problems from LeetCode, AtCoder and Codeforces and evaluates executable solutions with tests\([Jain et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib16)\)\. LCB\-e denotes its easy difficulty subset\. Appendix[A](https://arxiv.org/html/2610.02381#A1)specifies the problem selection, date window and scoring procedure used in the main tables\. MBPP consists of short, crowdsourced Python programming problems aimed at entry\-level programming skills\([Austin et al\., 2021](https://arxiv.org/html/2610.02381#bib.bib3)\)\.[MBPP\+](https://github.com/evalplus/evalplus/releases/tag/v0.2.0)uses EvalPlus’s expanded test suites, with dataset corrections and filtering, to examine correctness under more extensive testing\([Liu et al\., 2023](https://arxiv.org/html/2610.02381#bib.bib26)\)\.
#### Logic and structured reasoning\.
Reasoning Gym evaluates generated answers with each task’s own verifier\. Our frozen evaluation set yields925925scored problems after the reproduction checks in Appendix[A](https://arxiv.org/html/2610.02381#A1)\. We evaluate this set with a1616k\-token generation budget and report its scores as GYM\. The batching control \(Appendix[F](https://arxiv.org/html/2610.02381#A6)\) and Figure[8](https://arxiv.org/html/2610.02381#A3.F8)c also use GYM\-200, a fixed200200\-problem Reasoning Gym validation set separate from GYM\. BIG\-Bench Hard \(BBH\) collects challenging language and reasoning tasks, including logical deduction, tracking objects and interpreting structured information\([Suzgun et al\., 2023](https://arxiv.org/html/2610.02381#bib.bib41)\)\. Our harness evaluates2727task configurations with4040problems each, for1,0801\{,\}080problems\. MuSR tests multistep reasoning over natural\-language narratives, such as drawing conclusions from the evidence in a murder mystery\([Sprague et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib36)\)\. It complements procedural puzzles by requiring the model to connect dispersed statements in a longer narrative\.
#### General ability\.
ARC\-Challenge contains challenging multiple\-choice grade\-school science questions\([Clark et al\., 2018](https://arxiv.org/html/2610.02381#bib.bib4)\)\. HellaSwag asks the model to select a plausible continuation of an everyday situation from competing endings, including adversarially selected distractors\([Zellers et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib53)\)\. WinoGrande tests commonsense reference resolution through sentences with a blank and two candidate completions\([Sakaguchi et al\., 2021](https://arxiv.org/html/2610.02381#bib.bib33)\)\. None of these three suites supplies prompts to the distillation pool\. They probe whether training on math, code and logic preserves capabilities beyond the trained domains\. We score them by answer\-option log likelihood, using length\-normalized accuracy for ARC\-Challenge and HellaSwag and accuracy for WinoGrande, not sampled solutions\.
## Appendix CObjectives and supervision designs
#### On\-policy distillation\.
The studentπθ\\pi\_\{\\theta\}and a frozen teacherπT\\pi\_\{T\}share one tokenizer with vocabularyVV\. Given a promptxx, the student generatesy^∼πθ\(⋅∣x\)\\hat\{y\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\), and both models then score every prefixst=\(x,y^<t\)s\_\{t\}=\(x,\\hat\{y\}\_\{<t\}\)of this response\([Agarwal et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib1);[Gu et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib10)\)\. Our token channel uses a policy\-gradient surrogate with rewards computed from a no\-gradient scoring pass\. Letp¯t\\bar\{p\}\_\{t\}denote the student’s distribution from that pass,pθ,tp\_\{\\theta,t\}its differentiable distribution during the update, andpT,tp\_\{T,t\}the teacher distribution\. All three are normalized over the full vocabulary\. We selectVk,tV\_\{k,t\}, the top\-kktokens ofp¯t\\bar\{p\}\_\{t\}, withk=16k=16, and compute, forv∈Vk,tv\\in V\_\{k,t\},
wt,v\\displaystyle w\_\{t,v\}=p¯t\(v\)∑v′∈Vk,tp¯t\(v′\),\\displaystyle=\\frac\{\\bar\{p\}\_\{t\}\(v\)\}\{\\sum\_\{v^\{\\prime\}\\in V\_\{k,t\}\}\\bar\{p\}\_\{t\}\(v^\{\\prime\}\)\},\(6\)rt,v\\displaystyle r\_\{t,v\}=−wt,v\[logp¯t\(v\)−logpT,t\(v\)\]\.\\displaystyle=\-w\_\{t,v\}\\bigl\[\\log\\bar\{p\}\_\{t\}\(v\)\-\\log p\_\{T,t\}\(v\)\\bigr\]\.Thus the candidate weights are normalized within the selected support, while both log\-probabilities in the reward retain their full\-vocabulary normalization\.
We scale the rewards within each prompt\-domain group by their sample standard deviation, pooling all top\-kkcandidate rewards at valid response positions across that group’s responses in the training batch\. Withddthe prompt’s fixed domain, writeσd=max\(stdd\(r\),10−6\)\\sigma\_\{d\}=\\max\(\\operatorname\{std\}\_\{d\}\(r\),10^\{\-6\}\)andAt,v=rt,v/σdA\_\{t,v\}=r\_\{t,v\}/\\sigma\_\{d\}\. This scaling does not subtract the mean; groups with fewer than two entries are left unscaled\. LetMMindex the valid response positions in one loss calculation, across its participating responses\. Writingsg\\mathrm\{sg\}for stop\-gradient, the token objective is
ℒOPD=−1\|M\|∑t∈M∑v∈Vk,tsg\[At,v\]logpθ,t\(v\)\.\\mathcal\{L\}\_\{\\mathrm\{OPD\}\}\\;=\\;\-\\frac\{1\}\{\|M\|\}\\sum\_\{t\\in M\}\\sum\_\{v\\in V\_\{k,t\}\}\\mathrm\{sg\}\[A\_\{t,v\}\]\\log p\_\{\\theta,t\}\(v\)\.\(7\)The candidate dimension is summed and valid response positions are averaged; gradients flow only throughpθ,tp\_\{\\theta,t\}\. Appendix[A](https://arxiv.org/html/2610.02381#A1)specifies microbatch accumulation and the schedule weights\. Our token\-only routed baseline uses this channel without representation supervision\.
#### Latent supervision\.
We adapt the notation of LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)\. At positiontt, letztSz^\{S\}\_\{t\}andztTz^\{T\}\_\{t\}be student and teacher hidden states taken before the LM head,gψg\_\{\\psi\}a projector from the student to the teacher width,ν\(z\)=z/max\(∥z∥2,ϵ\)\\nu\(z\)=z/\\max\(\\lVert z\\rVert\_\{2\},\\epsilon\)the normalization, andsg\\mathrm\{sg\}a teacher\-side stop\-gradient:
ℒrep=1\|M\|dT∑t∈M∥ν\(gψ\(ztS\)\)−sg\[ν\(ztT\)\]∥22\.\\mathcal\{L\}\_\{\\mathrm\{rep\}\}\\;=\\;\\frac\{1\}\{\|M\|\\,d\_\{T\}\}\\sum\_\{t\\in M\}\\big\\lVert\\nu\\\!\\left\(g\_\{\\psi\}\(z^\{S\}\_\{t\}\)\\right\)\-\\mathrm\{sg\}\\\!\\left\[\\nu\(z^\{T\}\_\{t\}\)\\right\]\\big\\rVert\_\{2\}^\{2\}\.\(8\)Normalization acts along the hidden\-feature dimension\. The implementation callsF\.normalizewithout an explicit epsilon, using its defaultϵ=10−12\\epsilon=10^\{\-12\}\. This clamps the norm from below; the10−810^\{\-8\}constant in the masked\-mean denominator \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\) serves a separate purpose\. OPRD\([Yang et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib49)\)matches depth\-paired representations directly for compatible models and through frozen low\-rank projector pairs in OPRD\-Bridge\. LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)applies a normalized hidden\-state loss only at the last\-layer state that feeds the head, with a small trainable projector, and fades out the latent\-loss weight over a short crossfade window\. Both use one teacher\.
#### Multi\-teacher setup\.
We are given a panel ofNNfrozen teachers\{πTi\}i=1N\\\{\\pi\_\{T\_\{i\}\}\\\}\_\{i=1\}^\{N\}, each an expert of the student’s own capacity or larger, together with a training prompt pool in which every prompt carries a domain labeld∈\{1,…,N\}d\\in\\\{1,\\dots,N\\\}indexing the teacher that owns that domain \(math, code, or logic here\)\. Output\-level multi\-teacher distillation can form a target by routing each prompt toπTd\\pi\_\{T\_\{d\}\}\(as in MOPD,[Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\) or by mixing the teachers’ distributions\. Our routed token channel usespT,t=pTd,tp\_\{T,t\}=p\_\{T\_\{d\},t\}in Equations[6](https://arxiv.org/html/2610.02381#A3.E6)–[7](https://arxiv.org/html/2610.02381#A3.E7)\. A probability mixture∑iωipTi,t\\sum\_\{i\}\\omega\_\{i\}\\,p\_\{T\_\{i\},t\}, withωi≥0\\omega\_\{i\}\\geq 0and∑iωi=1\\sum\_\{i\}\\omega\_\{i\}=1, averages the probability each teacher assigns to a token\. When the specialist assigns a token more probability than the other teachers do, averaging lowers that probability\. Routing preserves the selected specialist’s distribution at each position, and is the assignment we adopt in both channels\. The token objective supervises the student’s output distribution; the representation objective also matches selected hidden states before the head\. All teachers contribute across the run through updates to the same student backbone\.
Figure[7](https://arxiv.org/html/2610.02381#A3.F7)illustrates the three supervision designs on the same math prompt\.
Figure 7:Three\-teacher supervision: an illustrative schematic\.All panels use the same math prompt\.\(a\)The math, code and logic teachers contribute equally to the averaged next\-token distribution, with weight1/31/3each\.\(b\)MOPD\-style selects the math teacher’s distribution; the code and logic teachers remain inactive\.\(c\)Latent\-MOPD retains this token target and adds a hidden\-state target from the same math teacher\. Hidden\-feature tiles are separate from the token distributions; their positions do not denote token or layer correspondence\. Values and features are schematic\.
#### Motivation for a transient latent term\.
LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)studies sustained latent supervision in single\-teacher distillation and motivates supervision at the final prediction interface during a short crossfade; on its same\-lineage pair, which matches our student and same\-family math teacher, keeping the latent term active scores higher on MATH\-500 than the crossfade\. We nevertheless start multi\-teacher training from the crossfade, giving each specialist its own update clock; same\-family, it outperforms constant weights \(Norm1\.051\.05vs\.0\.940\.94; Table[1](https://arxiv.org/html/2610.02381#S3.T1)\)\. The evidence for our setting comes from the three multi\-teacher settings in Section[4](https://arxiv.org/html/2610.02381#S4)and the controls in Section[5](https://arxiv.org/html/2610.02381#S5)\. The default recipes use the last three layers same\-family and the last layer cross\-family\. Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)compare one versus three layers within each combined crossfade recipe\. Table[1](https://arxiv.org/html/2610.02381#S3.T1)also compares all\-layer, last\-three\-layer and last\-layer supervision under the constant representation\-only objective, and includes Latent\-MOPD \(OPRD\-style\), an all\-layer variant with constant token and representation weights\.
Figure 8:Layer selection, crossfade and logic\-domain training dynamics\.\(a\)Last\-layer and last\-three\-layer comparisons: representation\-only and Latent\-MOPD in the same\-family \(SF\) setting, and Latent\-MOPD with linear projection in the cross\-family \(CF\) setting\. Norm is computed separately within Tables[1](https://arxiv.org/html/2610.02381#S3.T1)and[2](https://arxiv.org/html/2610.02381#S4.T2); the two regimes have separate axes\.\(b\)Same\-family score gains from crossfade over constant weights, in percentage points, computed from the reported Table[1](https://arxiv.org/html/2610.02381#S3.T1)scores\. Both configurations use last\-three\-layer alignment; the comparison changes the coupled token and representation schedule\.\(c\)Latent\-MOPD on GYM\-200, with88samples per problem: same\-family last\-three\-layer and cross\-family last\-layer linear configurations; the dashed line marks the base\. The validation set is separate from the925925\-problem GYM evaluation in the main tables\.
## Appendix DAdditional results and channel comparisons
#### Same\-family: a broad gain that exceeds individual teachers\.
Default Latent\-MOPD improves over MOPD\-style and all three representation\-only baselines on all nine benchmarks, and exceeds the best teacher on BBH, MuSR, MBPP, MBPP\+ and Minerva\. Norm reaches1\.051\.05, compared with0\.900\.90for MOPD\-style,0\.640\.64for both OPRD\-style \(all layers\) and Rep\-only \(last 1\), and0\.730\.73for Rep\-only \(last 3\)\. Under constant all\-layer representation supervision, adding the token channel improves eight of nine benchmarks, raising Norm from0\.640\.64to0\.940\.94in Latent\-MOPD \(OPRD\-style\)\. Both configurations useλ=2000\\lambda=2000and the last2,0002\{,\}000positions\. The default selective crossfade recipe then improves on this joint variant in eight of nine benchmarks, reaching Norm1\.051\.05\. This comparison evaluates two complete joint configurations, whose differences are listed in Appendix[A](https://arxiv.org/html/2610.02381#A1)\. A separate last\-three\-layer control keeps other settings fixed and sets both channel weights to one throughout\. Crossfade improves eight of nine benchmarks over this constant\-weight control \(Figure[8](https://arxiv.org/html/2610.02381#A3.F8)b\), raising Norm from0\.940\.94to1\.051\.05; the control scores higher on AIME25 \(37\.537\.5vs\.36\.936\.9\)\. This comparison tests the coupled schedule: both the token\-weight ramp and representation\-weight decay\.
Within the same\-family crossfade recipe, last\-three\-layer supervision improves on last\-layer supervision in eight of nine benchmarks \(Norm1\.051\.05vs\.0\.990\.99\)\. Last\-layer supervision scores51\.051\.0on AIME24 versus50\.850\.8for the default\. Both use all valid response positions,λ=1\\lambda=1, the same crossfade schedule and a24,57624\{,\}576\-token packing budget\. The representation\-only layer controls provide a second comparison: last\-three\-layer supervision exceeds both the last\-layer and all\-layer alternatives on seven of nine benchmarks each\. These runs share the same constant coefficient, selected positions and packing budget \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\. Thus, the same\-family preference for three late layers appears under both representation\-only and joint supervision \(Figure[8](https://arxiv.org/html/2610.02381#A3.F8)a\)\.
#### Cross\-family: the gain survives a wider teacher panel\.
The default last\-layer Latent\-MOPD with a linear map improves over MOPD\-style on all six reported benchmarks:\+2\.9\+2\.9and\+1\.5\+1\.5on AIME24/25,\+2\.5\+2\.5and\+1\.7\+1\.7on MBPP\(\+\), and\+4\.9\+4\.9and\+3\.7\+3\.7on GYM and BBH\. Logic has the largest margins, while math and code also improve\. On BBH, this variant exceeds representation\-only training by14\.214\.2points and token\-only MOPD\-style by3\.73\.7points\. The two standalone\-channel comparisons support jointly using token and representation supervision under the tested recipes\. The last\-three\-layer variant also exceeds both single\-channel baselines on all six benchmarks, with Norm0\.250\.25versus0\.260\.26for the default last\-layer variant\. It scores higher on GYM and MBPP\(\+\), while the default leads on BBH and AIME24/25\. Increasing the layer count thus gives no aggregate improvement in this comparison\. Both use the same configured crossfade window and are evaluated at step6262\.
#### Parameter\-merge initialization: gains after distillation\.
Budget\-matched distillation for6262steps improves the parameter merge on five of six reported benchmarks, with gains across math, code and logic\. Norm rises from0\.870\.87to1\.031\.03\. The student exceeds the best individual teacher on BBH, MBPP and MBPP\+, three of the six benchmarks\. This result shows that Latent\-MOPD can improve a student initialized by parameter merging as well as one initialized from the base model\. All comparisons use the final checkpoints and the evaluation protocol in Appendix[A](https://arxiv.org/html/2610.02381#A1)\.
#### Output mixing and routed supervision\.
Latent\-MOPD improves over uniform averaging on every benchmark reported in Table[1](https://arxiv.org/html/2610.02381#S3.T1), including GYM \(from40\.740\.7to52\.452\.4\), BBH \(from60\.060\.0to66\.366\.3\), MBPP \(from61\.161\.1to65\.765\.7\), and MBPP\+ \(from50\.750\.7to55\.755\.7\)\. MOPD\-style records higher logic scores than uniform averaging, reaching51\.851\.8on GYM and65\.465\.4on BBH\. Both baselines use the same training and evaluation protocol \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\)\. In Table[1](https://arxiv.org/html/2610.02381#S3.T1), the default crossfade recipe also improves every math, code and logic score over routed token\-only MOPD\-style\.
#### Logic\-domain training dynamics\.
Figure[8](https://arxiv.org/html/2610.02381#A3.F8)c extends the MATH\-500 analysis in Figure[5](https://arxiv.org/html/2610.02381#S4.F5)a to GYM\-200, the fixed200200\-problem validation set\. Each checkpoint is evaluated with88samples per problem, temperature0\.70\.7, top\-pp0\.950\.95and a16,38416\{,\}384\-token generation cap\. From the initial student’s37\.9437\.94, the same\-family configuration reaches62\.5662\.56at step6262; the cross\-family configuration reaches41\.1941\.19\. Both stay above the base from step2020onward\. These trajectories complement the final\-checkpoint results in Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)by showing how logic performance develops\.
## Appendix ELayer\-wise representation similarity
Figure 9:Depth\-matched CKA after masking high\-activation dimensions\.\(a\)Same\-family, with a magnified vertical scale;\(b\)cross\-family, on the full00–11scale\. Each teacher is compared with the base student on the same192192\-prompt diagnostic pool as Figure[3](https://arxiv.org/html/2610.02381#S3.F3)a,b\. Before centering, the top1%1\\%of dimensions by mean absolute activation are zeroed separately in each model and block\. Curves show linear CKA at matched decoder\-block outputs, before final normalization; shading is one standard deviation over200200prompt\-bootstrap replicates\. The masked cross\-family curves retain a middle\-layer trough, while the same\-family minima occur before the final block\. Hatching marks the selected training depths: the final three blocks in \(a\) and the final block in \(b\)\.#### Models and paired inputs\.
We compare the base DeepSeek\-R1\-Distill\-Qwen\-1\.51\.5B student, before distillation, with all six teachers \(Appendix[A](https://arxiv.org/html/2610.02381#A1)\) on the same fixed base\-student responses\. The diagnostic pool contains192192prompts,6464per domain\. We sample up to128128response\-token positions uniformly per response, retaining all positions in shorter responses\. Inputs are capped at2,5602\{,\}560tokens and positions beyond that cap are excluded, yielding24,56824\{,\}568paired positions:8,1928\{,\}192math,8,1898\{,\}189code and8,1878\{,\}187logic\. Per\-domain comparisons use the corresponding subsets\.
#### Extraction and CKA\.
We collect decoder\-block outputs, numbered11–2828in every model\. The final block is measured*before*final normalization, whereas distillation targets the state after that normalization \(Section[3](https://arxiv.org/html/2610.02381#S3)\)\. Forward passes use bfloat16; extracted features and CKA use float32\. We compare full native feature widths without a learned projector or dimensionality reduction\.
For paired activation matricesXℓ∈ℝn×dSX\_\{\\ell\}\\in\\mathbb\{R\}^\{n\\times d\_\{S\}\}andYℓ∈ℝn×dTY\_\{\\ell\}\\in\\mathbb\{R\}^\{n\\times d\_\{T\}\}at matched blockℓ\\ell, letX~ℓ\\widetilde\{X\}\_\{\\ell\}andY~ℓ\\widetilde\{Y\}\_\{\\ell\}denote their column\-centered versions\. We compute linear CKA\([Kornblith et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib18)\)as
CKA\(Xℓ,Yℓ\)=∥X~ℓ⊤Y~ℓ∥F2∥X~ℓ⊤X~ℓ∥F∥Y~ℓ⊤Y~ℓ∥F\.\\operatorname\{CKA\}\(X\_\{\\ell\},Y\_\{\\ell\}\)=\\frac\{\\lVert\\widetilde\{X\}\_\{\\ell\}^\{\\top\}\\widetilde\{Y\}\_\{\\ell\}\\rVert\_\{F\}^\{2\}\}\{\\lVert\\widetilde\{X\}\_\{\\ell\}^\{\\top\}\\widetilde\{X\}\_\{\\ell\}\\rVert\_\{F\}\\lVert\\widetilde\{Y\}\_\{\\ell\}^\{\\top\}\\widetilde\{Y\}\_\{\\ell\}\\rVert\_\{F\}\}\.\(9\)Each point pools the selected token rows across domains\. Shading gives one empirical standard deviation over200200bootstrap replicates: prompts are sampled with replacement, retaining their selected token rows\. Model checkpoints are fixed; the variation is over diagnostic prompts\.
#### Native profiles\.
Same\-family, mean CKA over blocks11–2525ranges from0\.9870\.987to0\.9920\.992across the three teachers, compared with0\.9450\.945–0\.9620\.962over the final three blocks\. The largest decrease is at the final block, whose CKA ranges from0\.8860\.886to0\.9300\.930\. Cross\-family, all three native curves reach their minimum at block1313\(0\.1160\.116–0\.1240\.124\), return to high similarity by block2121and remain high through block2727, and fall again at the final block \(0\.5360\.536–0\.6350\.635\)\. These native profiles motivate the layer choices: the final three layers within the shared lineage, and the final prediction interface across scales and training lineages, where middle layers differ substantially\. Downstream comparisons evaluate these choices; CKA alone does not determine which layers are sufficient or optimal for transfer\.
#### Sensitivity to high\-activation dimensions\.
We repeat the comparison \(Figure[9](https://arxiv.org/html/2610.02381#A5.F9)\) after zeroing each model’s top1%1\\%of feature dimensions by mean absolute activation, separately for every block and before centering\. This removes1515dimensions in the1\.51\.5B models and3636in the77B models\. The mask is determined on the full diagnostic pool and held fixed for domain subsets and bootstrap replicates\. The cross\-family trough remains at block1313, with minima of0\.4300\.430–0\.4510\.451, while the same\-family final\-block CKA rises to0\.9480\.948–0\.9760\.976\. The masked same\-family minima occur at blocks1818,1919and2222for JustRL, Archer2 and ProRL, respectively\. Thus, masking attenuates the native differences and changes where the same\-family minima lie; the cross\-family middle\-layer trough remains visible under both measurements\.
#### Connection to layer selection\.
The geometry suggests candidate targets; Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)test their value for distillation\. Same\-family, the last three layers outperform the last layer on eight of nine benchmarks under crossfade and seven of nine under representation\-only supervision\. Cross\-family, one and three layers split the benchmark wins, with similar aggregate scores \(Norm0\.260\.26and0\.250\.25\)\. The default therefore uses the smaller last\-layer target\. Together, the diagnostics and downstream controls support panel\-specific layer budgets, rather than layers chosen solely by CKA minima\.
## Appendix FBatch organization and training stability
Uniform averaging and interleaved routed targets are distinct designs: the former averages teacher distributions at each token, while the latter preserves one teacher per prompt and mixes domains within an optimizer step\. In the same\-family setting we compared two runs \(in\-loop validation on MATH\-500 and GYM\-200,88samples per problem; base student84\.1/37\.984\.1/37\.9\): the domain\-pure reference uses the all\-layer representation\-only configuration in Table[1](https://arxiv.org/html/2610.02381#S3.T1), withλ=2000\\lambda=2000, the final2,0002\{,\}000valid response positions and the identity map\. Both runs use representation\-only supervision with*identical*teachers, loss, coefficient and budget, differing*only*in the batch organization \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c\)\. With domain\-pure blocks of3232, training remains stable for6262steps \(91\.7/56\.791\.7/56\.7\)\. With the same1:1:11\{:\}1\{:\}1rows interleaved, so that one optimizer step averages updates aimed at three different teachers, both suites fall below the untrained student by step1010\(79\.8/32\.679\.8/32\.6\) and performance collapses by step2020\(22\.8/4\.122\.8/4\.1\), reaching9\.8/0\.99\.8/0\.9at step3030\. At step6262, performance remains well below the base on both suites \(43\.2/12\.843\.2/12\.8\), showing a lasting deficit within the training budget\. Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c shows the interleaved trajectory over these6262steps\. An earlier interleaved run showed the same early collapse \(21\.4/4\.321\.4/4\.3at step2020\)\. In that run, the representation loss*rises*from7\.9×10−57\.9\\times 10^\{\-5\}to a peak of2\.8×10−42\.8\\times 10^\{\-4\}, without NaNs or gradient overflow in the retained log\. Two additional interleaved runs reproduce the step\-1010signature \(82\.5/35\.882\.5/35\.8;80\.7/34\.180\.7/34\.1\), the second also the step\-2020collapse \(24\.3/4\.324\.3/4\.3\)\.
Per\-prompt routing specifies the teacher for a prompt and allows either domain\-pure or interleaved batches\. Accordingly, MOPD’s teacher replacement experiment\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)and Open\-MOPD’s analysis of capability imbalance\([Gao et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib9)\)address different comparisons from our batching control\. Our result shows that, under the tested protocol, domain\-pure updates keep multi\-teacher training stable and reach higher performance within the matched budget\.
#### The cost of extending interleaved training\.
Figure[10](https://arxiv.org/html/2610.02381#A6.F10)extends the same interleaved trajectory to120120updates\. Even with1\.94×1\.94\\timesthe original budget, MATH\-500 ends at84\.5084\.50\(peak84\.9084\.90at step110110\), close to the base score of84\.1284\.12, while GYM\-200 ends at41\.5641\.56\(peak42\.7542\.75\), a modest gain over37\.9437\.94\. Both remain below the domain\-pure reference already attained at step6262\(91\.7/56\.791\.7/56\.7\)\. The curves largely level off over the final measured updates: from step100100to120120, their endpoint gains are only0\.670\.67and0\.370\.37percentage points, respectively\. Most extra updates offset the collapse, leaving limited gains over the base and a substantial gap to the shorter domain\-pure reference\.
Figure 10:Interleaved representation\-only training under an extended budget\.The same\-family all\-layer control in Figure[3](https://arxiv.org/html/2610.02381#S3.F3)c, continued to120120steps\.\(a\)MATH\-500;\(b\)GYM\-200\. Scores average88samples per problem\. Solid lines connect all measured checkpoints\. Horizontal dashed lines mark the base; vertical dotted lines mark the original6262\-step budget, and diamonds the step\-6262values\. Vertical scales differ\.
## Appendix GTeacher\-output alignment after distillation
We detail Figure[6](https://arxiv.org/html/2610.02381#S5.F6)and extend the comparison across readout depth\. On identical diagnostic prefixes, we compare the final same\-family MOPD\-style and default last\-three\-layer Latent\-MOPD checkpoints from Table[1](https://arxiv.org/html/2610.02381#S3.T1)with all three domain teachers\.
#### Inputs and measurement\.
The diagnostic contains335335positions:9595arithmetic boundaries from4545generated chains, plus three prefix lengths \(50%50\\%,70%70\\%,90%90\\%\) for each of4040code and4040logic prompts from the three\-domain pool\. Arithmetic prefixes provide preceding correct values and withhold the current value\. All models receive identical prefixes without chat templates\. We compute student–teacher cosine similarity between temperature\-one final\-head probability vectors over the shared vocabulary at the last prefix position\. In each domain, we also select positions whose mean pairwise teacher–teacher output cosine is at or below that domain’s median\. This teacher\-only mask retains4848,6060and6060math, code and logic positions, respectively, and is fixed across students\.
#### Matching\-teacher margin\.
For studentaa, domainddand teacherss, letAa,d,sA\_\{a,d,s\}be the mean output\-distribution cosine over the chosen positions in domaindd\. We measure the corresponding teacher’s advantage over the closest alternative teacher as
Ma,d=Aa,d,d−maxs≠dAa,d,s,ΔMd=MLatent\-MOPD,d−MMOPD\-style,d\.M\_\{a,d\}=A\_\{a,d,d\}\-\\max\_\{s\\neq d\}A\_\{a,d,s\},\\qquad\\Delta M\_\{d\}=M\_\{\\text\{Latent\-MOPD\{\}\},d\}\-M\_\{\\text\{MOPD\-style\},d\}\.We bootstrap whole chains or prompts within each domain for3,0003\{,\}000replicates, pairing all their positions and both students\. Each replicate recomputes the margin from position\-weighted means, taking the maximum after averaging\. The selection mask stays fixed; intervals describe variation across diagnostic instances\.
#### Domain\-conditioned alignment\.
Both students match their corresponding code and logic teachers most closely \(Figure[6](https://arxiv.org/html/2610.02381#S5.F6)a,b\)\. The math row’s maximum shifts from the code teacher under MOPD\-style to the math teacher under Latent\-MOPD, raising its margin from−0\.0282\-0\.0282to0\.03660\.0366\. The paired increase is0\.06480\.0648\(95%95\\%interval\[0\.0184,0\.1324\]\[0\.0184,0\.1324\]\); over all positions it is0\.03270\.0327\(\[0\.0077,0\.0659\]\[0\.0077,0\.0659\]\)\. Code and logic increase by0\.00940\.0094and0\.00130\.0013on selected positions, with intervals spanning zero\. These diagnostics complement benchmark results by measuring each recipe’s agreement with the domain specialists\.
### G\.1Teacher matching across readout depth
Each student’s fitted Jacobian lens\([Gurnee et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib12)\)maps intermediate residuals to its final residual basis before final normalization and vocabulary readout\. On identical prefixes, we compare temperature\-one readouts with all three teachers’ final distributions\. Figure[11](https://arxiv.org/html/2610.02381#A7.F11)shows2727fitted readout indices and a separately measured native head\.
Figure 11:Teacher matching across readout depth in the same\-family setting\.\(a\)Latent\-MOPD minus MOPD\-style matching\-teacher margin, by input domain\. Shading and endpoint bars show paired95%95\\%pointwise instance\-bootstrap intervals\.\(b\)Per\-teacher cosine changes on selected math positions at the last four fitted indices and native head; positive values indicate greater similarity under Latent\-MOPD\. Both panels use the same teacher\-disagreement selection rule\. Native\-head measurements are separate from the fitted readouts\.Each student’s lens is calibrated independently on the same150150prompts spanning math, code and logic\. We retain fitted readout indices00–2626and apply each readout at the last position of the diagnostic prefix, evaluating the native head separately\.
For this depth diagnostic, we recompute the teacher\-only selection from this diagnostic’s own teacher forward passes and hold that mask fixed across students and readout depths\. The selected math, code and logic positions cover3535,3535and3434instances, respectively\. The margin is computed as above;3,0003\{,\}000paired instance resamples are shared across depths and students\. At readout index2626, the math margin increases by0\.03370\.0337\(95%95\\%interval\[0\.0117,0\.0612\]\[0\.0117,0\.0612\]\); at the native head, it increases by0\.06550\.0655\(\[0\.0196,0\.1257\]\[0\.0196,0\.1257\]\)\. Code and logic have smaller native\-head changes, with intervals spanning zero\.
Panel \(b\) identifies the teacher\-specific changes behind the math margin\. At index2626, similarity to the math teacher increases by0\.0260\.026, while similarity to the code and logic teachers decreases by0\.0080\.008and0\.0130\.013\. The native head shows the same pattern\. The margin is a relative measure: at index2525, similarity to all three teachers decreases, but the decrease is smaller for the math teacher\. Thus, the late readouts help locate how the two training recipes differ in their domain\-conditioned agreement with the specialists\. The measurement concerns the distributions exposed by the fitted readouts; it does not establish that earlier representations lack useful information\.
## Appendix HCase studies of generated solutions
Routed tokens teach*what*to predict; latent alignment additionally supervises the hidden states used to compute predictions \(Figure[2](https://arxiv.org/html/2610.02381#S1.F2)\)\. We compare MOPD\-style and Latent\-MOPD on three examples, using the routed specialists’ solutions as references\. Excerpts are verbatim with emphasis added; surrounding text is omitted\.
Frozen specialist targets on the student’s prefixWhat \+ howRouted token distributionpTd,t⟶pθ,tp\_\{T\_\{d\},t\}\\;\\longrightarrow\\;p\_\{\\theta,t\}Added: teacher hidden\-state targetsztTd⟶gψ\(ztS\)z^\{T\_\{d\}\}\_\{t\}\\;\\longrightarrow\\;g\_\{\\psi\}\(z^\{S\}\_\{t\}\)Latent\-MOPD combines both targets through crossfade \(Equation[3](https://arxiv.org/html/2610.02381#S3.E3)\)\.
#### \(a\) Code: initializing the computation correctly\.
Same\-family, MBPP\.Compute the Eulerian numbera\(n,m\)a\(n,m\); the supplied test iseulerian\_num\(3, 1\) == 4\. Inner\-loop excerpts:
MOPD\-style: token onlyfor j in range\(i\): if j == 0 or j == i: euler\[i\]\[j\] = 0
Latent\-MOPD: token \+ latentfor j in range\(i\): if j == 0: dp\[i\]\[j\] = 1
Computational structure: the recurrence boundary\.Latent\-MOPD setsa\(i,0\)=1a\(i,0\)=1and gets44, matching Archer2’s closed\-form solution\. MOPD\-style’s zero boundary propagates through the recurrence and returns00\.
#### \(b\) Logic: preserving a state invariant\.
Same\-family, GYM\.Rewrite until no rule applies:
A\# \#A→∅,B\# \#B→∅,A\# \#B→\#B A\#,B\# \#A→\#A B\#\.\\mbox\{\{A\\\# \\\#A\}\}\\to\\varnothing,\\quad\\mbox\{\{B\\\# \\\#B\}\}\\to\\varnothing,\\quad\\mbox\{\{A\\\# \\\#B\}\}\\to\\mbox\{\{\\\#B A\\\#\}\},\\quad\\mbox\{\{B\\\# \\\#A\}\}\\to\\mbox\{\{\\\#A B\\\#\}\}\.The initial state is\#B A\# A\# \#A \#B A\# B\# \#A \#A A\#\. Extracted final answers:
MOPD\-style: token only\#B\\ \#B\\ A\#
Latent\-MOPD: token \+ latent\#B \#BB\#A\#
Computational structure: valid state transitions\.The input has three B\-tokens, and the rules delete B\-tokens only in pairs, so every reachable state keeps an odd number of them\. ProRL and Latent\-MOPD match the unique terminal state; MOPD\-style’s two\-B answer is unreachable even without the backslashes\.
#### \(c\) Math: forming the right counting problem\.
Cross\-family, AIME24\.Count length\-1616paths across an8×88\\times 8grid from lower left to upper right with four direction changes\. Final\-solution excerpts:
MOPD\-style: token only“requires 7 right \(R\) moves and 7 up \(U\) moves, totaling 14 moves\.”Final answer:180\\boxed\{180\}\.
Latent\-MOPD: token \+ latent“Each path consists of8right moves \(R\) and8up moves \(U\), totaling 16 moves\.”Final answer:294\\boxed\{294\}\.
Computational structure: constraints before counting\.Skywork and Latent\-MOPD use88right and88up moves, giving2\(72\)\(71\)=2942\\binom\{7\}\{2\}\\binom\{7\}\{1\}=294\. MOPD\-style’s7\+77\+7formulation gives180180for a length\-1414path\.
With token and representation supervision, Latent\-MOPD gets both the answer and the task’s key constraints right in these examples\. MOPD\-style errs in the recurrence boundary, rewrite invariant or counting setup\. The advantage is visible in the structure of the solutions, illustrating the motivation for adding latent alignment\.
## Appendix IExtended related work
Table[4](https://arxiv.org/html/2610.02381#A9.T4)compares representative methods by whether they use multiple teachers \(Multi\-t\.\), hidden\-state supervision \(Pre\-head\) and student\-generated training samples \(On\-policy\)\.
Table 4:Where multi\-teacher capability integration happens\.Among the methods compared here, Latent\-MOPD combines multiple teachers, hidden\-state supervision and student\-generated training samples\. A dash denotes not applicable\.#### On\-policy distillation\.
In on\-policy distillation, the teacher scores every token of responses that the student samples itself, reducing the train–inference mismatch of imitating fixed teacher\-written text\([Agarwal et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib1);[Gu et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib10)\)\. It is also used in language\-model post\-training alongside supervised fine\-tuning and reinforcement learning\([Yang et al\., 2025](https://arxiv.org/html/2610.02381#bib.bib47)\)\. When teacher and student use different vocabularies, token\-level transfer requires handling that mismatch explicitly\([Niu et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib30)\)\. Our teachers share the student’s tokenizer, allowing us to study representation supervision under the same token\-level interface\. Token\-level OPD variants differ in which tokens receive a teacher signal and how that signal is formed, but their targets all come from the teacher’s output distribution; the hidden states that produce those targets receive no direct supervision\. Latent\-MOPD adds this second source of supervision and routes it across several domain specialists\.
#### Representation\-level distillation\.
Representation matching dates to FitNets\([Romero et al\., 2015](https://arxiv.org/html/2610.02381#bib.bib32)\), whose hint layers guide a narrower student with a teacher’s intermediate outputs; later methods align paired hidden states or attention statistics\([Sun et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib39);[Jiao et al\., 2020](https://arxiv.org/html/2610.02381#bib.bib17);[Wang et al\., 2020](https://arxiv.org/html/2610.02381#bib.bib43);[Dasgupta & Cohn, 2025](https://arxiv.org/html/2610.02381#bib.bib5)\)\. OPRD\([Yang et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib49)\)brings hidden\-state matching to on\-policy distillation: along the student’s own rollouts, it matches layers by relative depth and maps unequal widths through frozen low\-rank projectors\. OPRD also evaluates joint token and representation supervision with one teacher and discusses multi\-model consolidation as a potential application\. LastOPD\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\)reports collapse under sustained latent supervision in cross\-size experiments, and restricts the latent term to the final layer, crossfading to token\-only OPD over a short window\. Latent\-OPD\([Shen et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib35)\)applies a related trajectory\-level latent signal\. PHF\([Li et al\., 2026c](https://arxiv.org/html/2610.02381#bib.bib22)\)and PR\-OPD\([Li et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib21)\)align hidden\-state*flows*and per\-layer states on\-policy, against a privileged copy of the student itself\.*All of these use a single teacher \(or the student itself\) in their experiments\.*Latent\-MOPD directly studies capability integration with*multiple*teachers, selecting routed late\-layer targets and scheduling their supervision with per\-teacher crossfades\. The routed expert provides both hidden\-state and token targets; the default cross\-family configuration uses one shared linear map to match the teacher width\.
#### Multi\-teacher distillation\.
Multi\-teacher distillation trains one student from several teachers\.*Off\-policy*methods combine teachers by averaging or adaptively weighting their predictions\([You et al\., 2017](https://arxiv.org/html/2610.02381#bib.bib50);[Yuan et al\., 2021](https://arxiv.org/html/2610.02381#bib.bib52)\), while feature\-level methods align the student to*several*teachers’ intermediate features or output embeddings on a fixed corpus\([You et al\., 2017](https://arxiv.org/html/2610.02381#bib.bib50);[Liu et al\., 2020](https://arxiv.org/html/2610.02381#bib.bib27);[Formont et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib6)\); for language models, FuseLLM\([Wan et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib42)\)fuses source models through their output distributions\. All of these train on fixed data rather than student\-generated trajectories\. Recent multi\-teacher OPD studies examine domain routing, scheduling and general\-capability retention\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29);[Sun et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib40);[Liu et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib28)\)\. Data\-efficient OPD also extends to multiple teachers\([Fu et al\., 2026b](https://arxiv.org/html/2610.02381#bib.bib8)\)\. UI\-MOPD combines routed token\-level distillation with rule\-based RL rewards\([Lian et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib23)\), while CA\-OPD uses confidence\-aware token targets and teacher\-guided rollouts for structured visual prediction\([Li et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib20)\)\. A comparison of fusion paradigms studies parameter merging, mixed\-domain RL and MOPD\([Wu et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib46)\)\. MT\-SDPO\([He et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib13)\)further studies teacher selection, verifying candidates against the final answer\. These approaches concern which output target the student receives\. Latent\-MOPD adds selected hidden states from the same routed teacher, extending the supervision available on each student\-generated response\.
For image generators, Poly\-OPD\([Fu et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib7)\)combines heterogeneous flow\-model teachers through a pixel bridge and targets in a frozen DINOv2 feature space, while STEP\-OPD\([Wei et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib44)\)distills task\-specialized diffusion teachers on\-policy and aligns the direction and magnitude of student and teacher representation changes across blocks\. Our representation channel matches routed LLM hidden states on student\-generated token prefixes\.
MOPD\([Ma et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib29)\)reports three stages: general SFT, domain\-specific RL, and integration\. Its distillation stage freezes the teachers\. We retain per\-prompt routing and on\-policy scoring, reusing existing specialist checkpoints without additional teacher training for this study\. Our token update uses student top\-kkcandidates and fixed advantages derived from teacher–student log\-probability ratios \(Section[3](https://arxiv.org/html/2610.02381#S3)\); MOPD studies a different top\-kkconstruction alongside its policy\-gradient estimator\. It emphasizes same\-origin teachers and reports degraded transfer after replacing one with a stronger external model\. Our cross\-family experiment integrates separately developed77B specialists into a1\.51\.5B student using both token and latent supervision\. These models differ from the student in scale and post\-training while sharing a tokenizer and Qwen backbone\. Open\-MOPD\([Gao et al\., 2026](https://arxiv.org/html/2610.02381#bib.bib9)\)examines capability imbalance under domain routing\. Our batching analysis addresses a separate implementation choice: whether individually routed prompts from several domains contribute to the same optimizer step\. Per\-prompt routing is compatible with either batch organization, so the effect of batch organization must be tested directly\.
Weight\-space merging provides a complementary route to capability integration\. Parameter averaging\([Wortsman et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib45)\)and task arithmetic\([Ilharco et al\., 2022](https://arxiv.org/html/2610.02381#bib.bib15)\)combine compatible models without distillation\. In Section[4\.4](https://arxiv.org/html/2610.02381#S4.SS4), we initialize the student from a parameter merge for6262steps of Latent\-MOPD to test whether distillation can improve on the merged weights\.
#### Representational similarity and diagnostics\.
Centered kernel alignment \(CKA\)\([Kornblith et al\., 2019](https://arxiv.org/html/2610.02381#bib.bib18)\)compares representations across networks with different widths\. We compare each specialist with the base student at matched block depths on shared student\-generated responses \(Figure[3](https://arxiv.org/html/2610.02381#S3.F3)a,b\)\. The same\-family profiles remain close through most blocks, with a larger native\-CKA drop at the final block; the cross\-family profiles have a pronounced middle\-layer trough and recover in later blocks\. High\-activation dimensions\([Sun et al\., 2024](https://arxiv.org/html/2610.02381#bib.bib38)\)can affect representational similarity\([Yang et al\., 2026a](https://arxiv.org/html/2610.02381#bib.bib48)\); Appendix[E](https://arxiv.org/html/2610.02381#A5)specifies the measurement and examines sensitivity to masking those dimensions\. These diagnostics describe representation geometry before distillation\. CKA alone does not establish which layers suffice for transfer or identify the computations responsible for downstream gains\.
## Appendix JScope and extensions
#### Evaluation setting\.
The experiments compare ways to integrate a fixed set of specialists: output mixing, routed token supervision, representation\-only supervision, our combined objective, and parameter merging\. Appendix[A](https://arxiv.org/html/2610.02381#A1)documents the protocol and the implementation of our MOPD\-style baseline\. The main tables evaluate the final checkpoints\. We use fixed domain labels and set the representation\-loss scale for each setting\. All teachers share the student’s tokenizer\. More diverse teachers, learned routing and support for different tokenizers would extend the setting studied here\.
#### Channel and design comparisons\.
Tables[1](https://arxiv.org/html/2610.02381#S3.T1)–[2](https://arxiv.org/html/2610.02381#S4.T2)test token supervision alone, representation supervision alone, and their combination in Latent\-MOPD\. These comparisons establish the advantage of the joint method over either standalone channel under the reported configurations\. MOPD\-style uses a constant token\-loss weight, representation\-only training retains the latent term throughout, and the default joint method uses crossfade\. The same\-family constant\-weight last\-three\-layer joint control holds the other settings fixed, testing the crossfade schedule as a whole\. Within each regime, the last\-layer and last\-three\-layer Latent\-MOPD variants hold other settings fixed to assess layer selection\. The batching control \(Section[5](https://arxiv.org/html/2610.02381#S5.SS0.SSS0.Px3)\) holds per\-prompt routing fixed and tests how grouping domains within an optimizer step affects stability\.
#### Representation evidence and extensions\.
The*what/how*framing identifies the two supervision targets: token distributions and the hidden states used to compute them\. The case studies \(Appendix[H](https://arxiv.org/html/2610.02381#A8)\) examine computational structure in generated outputs, while depth\-matched CKA describes how each teacher’s representations resemble the base student’s before training \(Appendix[E](https://arxiv.org/html/2610.02381#A5)\)\. The native and masked profiles show how these comparisons depend on layer and activation dimensions\. Measurements of learned projectors and update directions at the states used as training targets \(after final normalization, unlike the CKA readouts\) would help connect this geometry to the improvements from representation supervision\.Similar Articles
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
MOPD proposes a multi-teacher on-policy distillation paradigm for LLM post-training, enabling efficient integration of multiple domain capabilities by distilling specialized RL teachers into a student model using its own rollouts. It outperforms existing methods like Mix-RL and Cascade RL, and has been deployed in industrial-scale models.
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
This paper proposes Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD) to address unbalanced feedback when merging specialist language models, showing performance improvements on benchmarks like mathematics and instruction-following.
Lexicographic Multi-Objective On-Policy Distillation
Cohere researchers introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method that integrates reward-specialized policies under explicit priority orders using gated routing and projected log-policy corrections. Evaluated on 30B-A3B MoE models across three math benchmarks, LMOPD retains far more top-priority accuracy and reasoning gains than baselines while still capturing conciseness improvements.
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
The paper introduces Open-MOPD, a framework that diagnoses and fixes capability imbalance in multi-teacher on-policy distillation by balancing token-level budgets, improving headroom recovery from 35.6% to 83.4% through dynamic allocation and reward refresh.
DOPD: Dual On-policy Distillation
DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.