Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Hugging Face Daily Papers Papers

Summary

This paper investigates whether KL divergence is necessary for on-policy distillation of large language models, showing that preserving update direction suffices and introducing Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to enhance multi-teacher learning.

Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.
Original Article
View Cached Full Text

Cached at: 09/30/26, 12:13 AM

Paper page - Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Source: https://huggingface.co/papers/2609.33791 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Sincetheadventofknowledgedistillation,KLdivergencehasbeenthestandardlossindistillation.Recently,on-policydistillation(OPD)hasemergedasanefficientpost-trainingparadigmforLLMs.Asadistillationmethod,OPDnaturallyinheritsKLdivergenceasitsstandardloss.However,inthiswork,wefindthatKLdivergencemaynotbenecessaryforOPD.WeshowthatsimplypreservingtheupdatedirectionissufficientforeffectiveOPD.Aslongastheupdatedirectionistowardtheteacher,OPDworks.Moreprecisely,itisnotthedirectionofeverytoken,butthedirectionofasmallsubsetoftokenswheretheteacherandstudentdisagreestrongly.Wefirstshowthatsimplyassigningarewardof(+1)totokenswheretheteacherprobabilityishigherthanthestudentprobabilityand(-1)whereitislower,whichmerelyencouragesupdatestowardtheteacher,reproducesalmostthesametrainingmodeasOPDwithreverseKL.Wefurthershowthatonlythedirectionofasmallsubsetoftokenswithlargeteacher-studentdisagreementiscritical,andtrainingworksaslongastheirupdatedirectionistowardtheteacher,evenifothertokensarepulledawayfromtheteacher.Andasanapplicationofthesefindings,weintroduceConsensusMulti-TeacherOn-PolicyDistillation(C-MOPD)toimproveMulti-TeacherOn-PolicyDistillation(MOPD).UnlikeMOPD,whichrouteseachsampletoasingleteacherandmaycausecapabilityconflictsacrossdomains,C-MOPDletseverysamplebesupervisedbyallteachers.ExperimentsshowthatC-MOPDconsistentlyoutperformsMOPDonbothmathandcodebenchmarks.Ourcodeisavailableathttps://github.com/LeapLabTHU/KL-Free-OPD.

View arXiv pageView PDFGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.33791

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33791 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33791 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33791 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

On-Policy Distillation (5 minute read)

TLDR AI

This paper introduces on-policy distillation, which trains a student model on its own trajectories with teacher token-level KL supervision to fix train-inference mismatch, unifying forward-KL, reverse-KL, and JSD losses, with reverse-KL favored for smaller students.

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

Hugging Face Daily Papers

This paper presents a comprehensive empirical study on on-policy distillation for large language models, identifying failure mechanisms like distribution mismatch and optimization instability, and proposing fixes such as stop-gradient objectives and RLVR-adapted teachers.

Rethinking the Role of Temperature in Large Language Model Distillation

arXiv cs.LG

This paper reexamines the role of temperature in large language model distillation, revealing that temperature asymmetrically benefits forward KL divergence over reverse KL, allowing simple KL methods to match state-of-the-art distillation approaches at higher temperatures.