Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Summary
This paper investigates on-policy distillation in large language models using sparse crosscoders, revealing that it reweights existing features rather than creating new ones, with SFT warm-up playing a role in pre-reweighting features.
View Cached Full Text
Cached at: 09/30/26, 08:20 AM
Paper page - Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Source: https://huggingface.co/papers/2609.35210 Published on Sep 28
·
Submitted byhttps://huggingface.co/zichao1
zichaoon Sep 30
Abstract
On-policydistillation(OPD)isawidelyadoptedpost-trainingtechniqueforLLMreasoning.Itiscommonlybelievedtotransferknowledgefromastrongerteacher,yetwhatOPDactuallydistillsintothestudent’sinternalrepresentationsremainsunclear.Westudythisquestionwithsparsecrosscoders,whichlearnonefeaturedictionarysharedbythestudentbeforeandafterOPDandtheteacher.Standardcrosscoderanalyses,however,identifymodel-specificfeaturesbutcannottellhowamodel’suseofitsfeatureschanges,sinceallmodelsareencodedintoonesetoffeatureactivations.Wethereforeproposetheswapreadout,whichreadseachstudentcheckpoint’sfeatureactivationsonitsown,measuringhowtrainingchangesthestudent’suseofeachfeature,evenforcheckpointsunseenbythecrosscoder.AcrossthreeOPDsettings,wefindthatOPDneithercreatesfeaturesnorpassesontheteacher’sown,andleavesthefiringratesofover98%ofthestudent’sfrequentlyusedfeatureswithin20%.WefurtherexaminetheSFTwarm-upontheteacher’srolloutsthatcommonlyprecedesOPDandmakesitmoreeffective.Ratherthanaddingfeatures,thewarm-upreweightsthesharedonesintwoways.First,italreadyraisesandlowersmanyofthefeaturesthatOPDlaterraisesandlowers,doingpartofOPD’sworkinadvance.Second,itchangesfeaturesthatOPDalonewouldnot,notablythoseforconversationformat,reasoningstyle,andmathematicalnotation,andthesechangespersistthroughOPD.Imposingthisreweightingonadirectlydistilledstudent’sfeatures,withoutchangingitsweights,bringsitsaccuracyclosetothatofthewarmed-upstudent,whereasthesamechangeonshuffledfeaturesdoesnot.Together,thesefindingssuggestthatOPDreweightsexistingfeaturesratherthanacquiringnewones:thestudentlearnsfromtheteacherhowtousethefeaturestheyalreadyshare.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.35210
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35210 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35210 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35210 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
This paper presents a comprehensive empirical study on on-policy distillation for large language models, identifying failure mechanisms like distribution mismatch and optimization instability, and proposing fixes such as stop-gradient objectives and RLVR-adapted teachers.
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.
SFT, RL, and On-Policy Distillation Through a Distributional Lens (19 minute read)
This article analyzes post-training methods for language models through a distributional perspective, comparing how SFT, RL, and on-policy distillation reshape model distributions and impact phenomena like catastrophic forgetting.
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
This paper investigates the parameter-level mechanisms behind the efficiency of On-Policy Distillation (OPD) for large language models, attributing it to early 'foresight' in module allocation and update direction. It proposes EffOPD, a plug-and-play method that accelerates OPD training by 3x without compromising final performance.
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
This paper proposes an empirical 'sparse-to-dense' reward principle for language model post-training, arguing that scarce labeled data should be used with sparse rewards for teacher model discovery and dense rewards for student compression via distillation. The authors demonstrate that this staged approach, bridging sparse RL and on-policy distillation, outperforms direct GRPO on deployment-sized models in math benchmarks.