Scaling Properties of Same-Family On-Policy Distillation
Summary
This paper studies the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher-student LLM setups, finding a universal 'useful-transfer' regime where held-out accuracy rises linearly with KL divergence, and fitting power laws showing smaller teachers can transfer better at matched gold scores.
View Cached Full Text
Cached at: 09/30/26, 04:17 PM
Paper page - Scaling Properties of Same-Family On-Policy Distillation
Source: https://huggingface.co/papers/2609.32722
Abstract
*Reinforcementlearning(RL)*caninducesubstantialreasoningcapabilitiesinlargelanguagemodels(LLMs),buthowmuchofthiscapabilitytransfersacrossmodelscales,andhowquickly,remainsunclear.Westudythescalingpropertiesof*on-policydistillation(OPD)*across*weak-to-strong*,*same-base*,and*strong-to-weak*teacher--studentsetups.WefindthatearlyOPDtrainingdynamicsuniformlyexhibitaregular*useful-transfer*regime,inwhichheld-outaccuracy(the*goldscore*,G)risesapproximatelylinearlyind=mathrm{KL(π_θVertπ_{ref})},thesquarerootoftoken-levelreverseKLdivergencefromthestudentinitialization.Ineveryobservedweak-to-strongpair,thestudent’speakgoldscoreexceedsitsteacher’sown,soacompactRLexpertcantransfercapabilitytoamuchlargerstudentviaOPD.ToestimateOPDoutcomes,wefit*powerlaws*forhowG_{peak}andtheslopeoftheuseful-transferregimescalewithstudentandteacherparametercountsandwithteachergoldscore.Theselawsshowthatpeakgoldscoreimproveswithteacherscaleonlyuptoroughlythestudent’sscale,andthatatamatchedgoldscoresmallerteacherstransferbetter,soateacher’sscorealonedoesnotdefineitssupervisionvalue.WealsostudythescalingeffectsoftwoOPDvariants,bootstrappingweak-to-strongOPD,andthedegreeofon-policysupervision.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.32722
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### colored-dye/OPD-scaling-checkpoints Reinforcement Learning• Updatedabout 2 hours ago • 1
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32722 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32722 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
On-Policy Distillation (5 minute read)
This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.
Weak-to-Strong Generalization via Direct On-Policy Distillation
Direct-OPD distills the policy shift from a small model's pre- and post-RL checkpoints to improve a larger student model via on-policy distillation, achieving significant gains without expensive RL on the student.
Calibrating Teacher--Student Discrepancy for On-Policy Distillation
The paper introduces Calibrated On-Policy Distillation (Cal-OPD), a method that estimates the teacher's self-deviation to calibrate teacher-student discrepancies, improving on-policy distillation for mathematical reasoning tasks.
DOPD: Dual On-policy Distillation
DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.
Weak-to-Strong On-Policy Distillation
Introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), a framework that improves a strong language model by distilling from multiple weaker models using contrast pairs in logit space, consistently outperforming standard on-policy distillation on math and code benchmarks.