Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
Summary
The paper introduces Neighborhood On-Policy Self-Distillation (N-OPSD), which exploits local parameter perturbations of a privileged teacher to pool complementary reference-aligned corrections and route them to student-visited states, improving math reasoning benchmarks (AIME 2024/2025, HMMT Feb 2025) across Qwen3 1.7B/4B/8B models.
View Cached Full Text
Cached at: 10/02/26, 08:27 AM
Paper page - Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
Source: https://huggingface.co/papers/2609.39687 Published on Sep 30
·
Submitted byhttps://huggingface.co/dingyii
dingyion Oct 2
Abstract
On-policyself-distillation(OPSD)trainsmathematicalreasoningmodelsusingaprivilegedteacherthatseesareferencesolutionandsupervisesstudent-sampledprefixes.StandardOPSDusesonefixedparametersettingateverystate,butnearbysettingsmayofferadditionalsupervision.Wefindthatlocalparameterperturbationsrevealcomplementaryreference-alignedcorrectionsunderthesamereferencecontext.Differentexpertssupplythesecorrectionsatdifferentreferencepositions.Theirpoolcoversmoresuchpositionsthantheunperturbedprivilegedteacher.WeintroduceNeighborhoodOPSD(N-OPSD)toturnthesecorrectionsintosupervisionatstudent-visitedstates.Offline,greedyselectionbuildsacompactpooloffrozenexpertsbyrewardingfilteredreference-tokengainsbeyondthepool’scurrentbestateachposition.Thehighest-peakexpertneednotprovidethebesttrainingtarget.Onlineroutingthereforeseparatestheanchordirectionfromitslevelofsupport.MaxPeakselectstheanchortoken,andquantileselectionchoosesamongexpertswhosetoptokenmatchesit.Thestudentlearnsfromthechosenexpert’sfullnext-tokendistributionthroughtheclippedforward-KLobjectiveinheritedfromOPSD.WeevaluateonAIME2024,AIME2025,andHMMTFebruary2025.Acrossthreeindependentrunspermethod,NeighborhoodOPSDimprovesthethree-benchmarkAverage@12overOPSDby2.75,1.67,and1.94pointsonQwen3-1.7B,4B,and8B,respectively.Student-prefixcontinuationssupportusingthepoolbeyondthereferencetrajectoriesusedforselection.Matchedablationssupportfilteredreference-tokengainsasaselectioncriterion.Accountingforoverlapwithinthepoolandroutingbystatefurtherimprovestudentaccuracy.Inferenceusesonlythedistilledstudent.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.39687 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.39687 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.39687 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
On-Policy Self-Distillation without Any Supervision
Introduces U-OPSD, an unsupervised on-policy self-distillation method that uses internal consistency and majority-vote pseudo-solutions to improve LLMs without external supervision, matching or exceeding supervised methods on math benchmarks.
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.
Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
This paper analyzes on-policy self-distillation (OPSD) for LLM reasoning and shows that unverified student scaffolds create an imitation gap that worsens with model scale, motivating OASIS, a method that supervises mostly verified by-label on-policy trajectories using model-generated attempts as teacher context, improving Qwen3 1.7B/4B/8B by 3.2-3.8 points on AIME/HMMT benchmarks.
SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
Sign-Gated On-Policy Distillation (SG-OPD) enhances standard on-policy distillation by using a binary verifier as a trust signal for teacher supervision, improving performance on competition-level math reasoning benchmarks.
Latent On-Policy Self-Distillation
This paper introduces Latent On-Policy Self-Distillation (LOPD), a method that makes the teacher's privileged context learnable end-to-end from experience, providing dense token-level supervision to enhance agent performance and efficiency in agentic tool use and code generation.