Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
Summary
This paper investigates whether KL divergence is necessary for on-policy distillation of large language models, showing that preserving update direction suffices and introducing Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to enhance multi-teacher learning.
View Cached Full Text
Cached at: 09/30/26, 12:13 AM
Paper page - Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
Source: https://huggingface.co/papers/2609.33791 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Sincetheadventofknowledgedistillation,KLdivergencehasbeenthestandardlossindistillation.Recently,on-policydistillation(OPD)hasemergedasanefficientpost-trainingparadigmforLLMs.Asadistillationmethod,OPDnaturallyinheritsKLdivergenceasitsstandardloss.However,inthiswork,wefindthatKLdivergencemaynotbenecessaryforOPD.WeshowthatsimplypreservingtheupdatedirectionissufficientforeffectiveOPD.Aslongastheupdatedirectionistowardtheteacher,OPDworks.Moreprecisely,itisnotthedirectionofeverytoken,butthedirectionofasmallsubsetoftokenswheretheteacherandstudentdisagreestrongly.Wefirstshowthatsimplyassigningarewardof(+1)totokenswheretheteacherprobabilityishigherthanthestudentprobabilityand(-1)whereitislower,whichmerelyencouragesupdatestowardtheteacher,reproducesalmostthesametrainingmodeasOPDwithreverseKL.Wefurthershowthatonlythedirectionofasmallsubsetoftokenswithlargeteacher-studentdisagreementiscritical,andtrainingworksaslongastheirupdatedirectionistowardtheteacher,evenifothertokensarepulledawayfromtheteacher.Andasanapplicationofthesefindings,weintroduceConsensusMulti-TeacherOn-PolicyDistillation(C-MOPD)toimproveMulti-TeacherOn-PolicyDistillation(MOPD).UnlikeMOPD,whichrouteseachsampletoasingleteacherandmaycausecapabilityconflictsacrossdomains,C-MOPDletseverysamplebesupervisedbyallteachers.ExperimentsshowthatC-MOPDconsistentlyoutperformsMOPDonbothmathandcodebenchmarks.Ourcodeisavailableathttps://github.com/LeapLabTHU/KL-Free-OPD.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.33791
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33791 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33791 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33791 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
This paper investigates generalization in on-policy distillation for large language models, showing that it transfers reasoning behaviors and that teacher-student origin alignment is crucial, with multi-teacher combinations causing capability trade-offs.
On-Policy Distillation (5 minute read)
This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
This paper presents a comprehensive empirical study on on-policy distillation for large language models, identifying failure mechanisms like distribution mismatch and optimization instability, and proposing fixes such as stop-gradient objectives and RLVR-adapted teachers.
Rethinking the Role of Temperature in Large Language Model Distillation
This paper reexamines the role of temperature in large language model distillation, revealing that temperature asymmetrically benefits forward KL divergence over reverse KL, allowing simple KL methods to match state-of-the-art distillation approaches at higher temperatures.