Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

Hugging Face Daily Papers Papers

Summary

The paper introduces Neighborhood On-Policy Self-Distillation (N-OPSD), which exploits local parameter perturbations of a privileged teacher to pool complementary reference-aligned corrections and route them to student-visited states, improving math reasoning benchmarks (AIME 2024/2025, HMMT Feb 2025) across Qwen3 1.7B/4B/8B models.

On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
Original Article
View Cached Full Text

Cached at: 10/02/26, 08:27 AM

Paper page - Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

Source: https://huggingface.co/papers/2609.39687 Published on Sep 30

·

Submitted byhttps://huggingface.co/dingyii

dingyion Oct 2

Abstract

On-policyself-distillation(OPSD)trainsmathematicalreasoningmodelsusingaprivilegedteacherthatseesareferencesolutionandsupervisesstudent-sampledprefixes.StandardOPSDusesonefixedparametersettingateverystate,butnearbysettingsmayofferadditionalsupervision.Wefindthatlocalparameterperturbationsrevealcomplementaryreference-alignedcorrectionsunderthesamereferencecontext.Differentexpertssupplythesecorrectionsatdifferentreferencepositions.Theirpoolcoversmoresuchpositionsthantheunperturbedprivilegedteacher.WeintroduceNeighborhoodOPSD(N-OPSD)toturnthesecorrectionsintosupervisionatstudent-visitedstates.Offline,greedyselectionbuildsacompactpooloffrozenexpertsbyrewardingfilteredreference-tokengainsbeyondthepool’scurrentbestateachposition.Thehighest-peakexpertneednotprovidethebesttrainingtarget.Onlineroutingthereforeseparatestheanchordirectionfromitslevelofsupport.MaxPeakselectstheanchortoken,andquantileselectionchoosesamongexpertswhosetoptokenmatchesit.Thestudentlearnsfromthechosenexpert’sfullnext-tokendistributionthroughtheclippedforward-KLobjectiveinheritedfromOPSD.WeevaluateonAIME2024,AIME2025,andHMMTFebruary2025.Acrossthreeindependentrunspermethod,NeighborhoodOPSDimprovesthethree-benchmarkAverage@12overOPSDby2.75,1.67,and1.94pointsonQwen3-1.7B,4B,and8B,respectively.Student-prefixcontinuationssupportusingthepoolbeyondthereferencetrajectoriesusedforselection.Matchedablationssupportfilteredreference-tokengainsasaselectioncriterion.Accountingforoverlapwithinthepoolandroutingbystatefurtherimprovestudentaccuracy.Inferenceusesonlythedistilledstudent.

View arXiv pageView PDFProject pageAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.39687 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.39687 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.39687 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

On-Policy Self-Distillation without Any Supervision

Hugging Face Daily Papers

Introduces U-OPSD, an unsupervised on-policy self-distillation method that uses internal consistency and majority-vote pseudo-solutions to improve LLMs without external supervision, matching or exceeding supervised methods on math benchmarks.

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

Hugging Face Daily Papers

This paper analyzes on-policy self-distillation (OPSD) for LLM reasoning and shows that unverified student scaffolds create an imitation gap that worsens with model scale, motivating OASIS, a method that supervises mostly verified by-label on-policy trajectories using model-generated attempts as teacher context, improving Qwen3 1.7B/4B/8B by 3.2-3.8 points on AIME/HMMT benchmarks.

Latent On-Policy Self-Distillation

Hugging Face Daily Papers

This paper introduces Latent On-Policy Self-Distillation (LOPD), a method that makes the teacher's privileged context learnable end-to-end from experience, providing dense token-level supervision to enhance agent performance and efficiency in agentic tool use and code generation.