Scaling Properties of Same-Family On-Policy Distillation

Hugging Face Daily Papers Papers

Summary

This paper studies the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher-student LLM setups, finding a universal 'useful-transfer' regime where held-out accuracy rises linearly with KL divergence, and fitting power laws showing smaller teachers can transfer better at matched gold scores.

*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, G) rises approximately linearly in d=mathrm{KL(π_θVert π_{ref})}, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how G_{peak} and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:17 PM

Paper page - Scaling Properties of Same-Family On-Policy Distillation

Source: https://huggingface.co/papers/2609.32722

Abstract

*Reinforcementlearning(RL)*caninducesubstantialreasoningcapabilitiesinlargelanguagemodels(LLMs),buthowmuchofthiscapabilitytransfersacrossmodelscales,andhowquickly,remainsunclear.Westudythescalingpropertiesof*on-policydistillation(OPD)*across*weak-to-strong*,*same-base*,and*strong-to-weak*teacher--studentsetups.WefindthatearlyOPDtrainingdynamicsuniformlyexhibitaregular*useful-transfer*regime,inwhichheld-outaccuracy(the*goldscore*,G)risesapproximatelylinearlyind=mathrm{KL(π_θVertπ_{ref})},thesquarerootoftoken-levelreverseKLdivergencefromthestudentinitialization.Ineveryobservedweak-to-strongpair,thestudent’speakgoldscoreexceedsitsteacher’sown,soacompactRLexpertcantransfercapabilitytoamuchlargerstudentviaOPD.ToestimateOPDoutcomes,wefit*powerlaws*forhowG_{peak}andtheslopeoftheuseful-transferregimescalewithstudentandteacherparametercountsandwithteachergoldscore.Theselawsshowthatpeakgoldscoreimproveswithteacherscaleonlyuptoroughlythestudent’sscale,andthatatamatchedgoldscoresmallerteacherstransferbetter,soateacher’sscorealonedoesnotdefineitssupervisionvalue.WealsostudythescalingeffectsoftwoOPDvariants,bootstrappingweak-to-strongOPD,andthedegreeofon-policysupervision.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.32722

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### colored-dye/OPD-scaling-checkpoints Reinforcement Learning• Updatedabout 2 hours ago • 1

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32722 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32722 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

On-Policy Distillation (5 minute read)

TLDR AI

This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.

DOPD: Dual On-policy Distillation

Hugging Face Daily Papers

DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.

Weak-to-Strong On-Policy Distillation

arXiv cs.LG

Introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), a framework that improves a strong language model by distilling from multiple weaker models using contrast pairs in logit space, consistently outperforming standard on-policy distillation on math and code benchmarks.