Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

Hugging Face Daily Papers Papers

Summary

This paper investigates on-policy distillation in large language models using sparse crosscoders, revealing that it reweights existing features rather than creating new ones, with SFT warm-up playing a role in pre-reweighting features.

On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:20 AM

Paper page - Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

Source: https://huggingface.co/papers/2609.35210 Published on Sep 28

·

Submitted byhttps://huggingface.co/zichao1

zichaoon Sep 30

Abstract

On-policydistillation(OPD)isawidelyadoptedpost-trainingtechniqueforLLMreasoning.Itiscommonlybelievedtotransferknowledgefromastrongerteacher,yetwhatOPDactuallydistillsintothestudent’sinternalrepresentationsremainsunclear.Westudythisquestionwithsparsecrosscoders,whichlearnonefeaturedictionarysharedbythestudentbeforeandafterOPDandtheteacher.Standardcrosscoderanalyses,however,identifymodel-specificfeaturesbutcannottellhowamodel’suseofitsfeatureschanges,sinceallmodelsareencodedintoonesetoffeatureactivations.Wethereforeproposetheswapreadout,whichreadseachstudentcheckpoint’sfeatureactivationsonitsown,measuringhowtrainingchangesthestudent’suseofeachfeature,evenforcheckpointsunseenbythecrosscoder.AcrossthreeOPDsettings,wefindthatOPDneithercreatesfeaturesnorpassesontheteacher’sown,andleavesthefiringratesofover98%ofthestudent’sfrequentlyusedfeatureswithin20%.WefurtherexaminetheSFTwarm-upontheteacher’srolloutsthatcommonlyprecedesOPDandmakesitmoreeffective.Ratherthanaddingfeatures,thewarm-upreweightsthesharedonesintwoways.First,italreadyraisesandlowersmanyofthefeaturesthatOPDlaterraisesandlowers,doingpartofOPD’sworkinadvance.Second,itchangesfeaturesthatOPDalonewouldnot,notablythoseforconversationformat,reasoningstyle,andmathematicalnotation,andthesechangespersistthroughOPD.Imposingthisreweightingonadirectlydistilledstudent’sfeatures,withoutchangingitsweights,bringsitsaccuracyclosetothatofthewarmed-upstudent,whereasthesamechangeonshuffledfeaturesdoesnot.Together,thesefindingssuggestthatOPDreweightsexistingfeaturesratherthanacquiringnewones:thestudentlearnsfromtheteacherhowtousethefeaturestheyalreadyshare.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.35210

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35210 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.35210 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35210 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

Hugging Face Daily Papers

This paper presents a comprehensive empirical study on on-policy distillation for large language models, identifying failure mechanisms like distribution mismatch and optimization instability, and proposing fixes such as stop-gradient objectives and RLVR-adapted teachers.

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation

arXiv cs.CL

This paper investigates the parameter-level mechanisms behind the efficiency of On-Policy Distillation (OPD) for large language models, attributing it to early 'foresight' in module allocation and update direction. It proposes EffOPD, a plug-and-play method that accelerates OPD training by 3x without compromising final performance.

Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training

Hugging Face Daily Papers

This paper proposes an empirical 'sparse-to-dense' reward principle for language model post-training, arguing that scarce labeled data should be used with sparse rewards for teacher model discovery and dense rewards for student compression via distillation. The authors demonstrate that this staged approach, bridging sparse RL and on-policy distillation, outperforms direct GRPO on deployment-sized models in math benchmarks.