Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Hugging Face Daily Papers Papers

Summary

This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.

Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First, we show that offline KD (caching the teacher's top-K logits once and training the student against the cache) matches online distillation at near-identical training loss while removing the teacher from memory, running about 29\% faster per iteration, and reaching up to 41\% higher throughput on a single H200 GPU. Second, we introduce a fused, chunked KL loss that never materialises the full vocabulary-sized logit tensor, making peak memory linear in the sequence length. This removes the memory spike that otherwise caps context length and lets us train at four times the context (32{,}768 tokens) on a single GPU. A separate output-head-only toy benchmark isolates the loss kernel and confirms its memory and iteration-rate scaling from 4K to 256K tokens. Together these make large-scale healing and hundreds of ablations affordable. We also report supporting ablations on loss design and sequence packing. We release our chunked-loss implementation: https://github.com/CompactifAI/Full-Chunked-KL-Loss.
Original Article
View Cached Full Text

Cached at: 08/10/26, 02:15 PM

Paper page - Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Source: https://huggingface.co/papers/2608.03796

Abstract

Smalllanguagemodelsareoftentheonlyoptionfordeploymentundertightlatency,cost,andon-premisesconstraints,buttheyarerarelytrainedfromscratch:acompressedmodelisusuallyrecoveredthroughknowledgedistillation(KD).Thisrecoverysteplargelydecidesthefinalquality,yetitisexpensive.Wepresentapractitioner’sstudyofhowtomakedistillationtrainingefficient,organisedaroundtwosystemscontributions.First,weshowthatofflineKD(cachingtheteacher’stop-Klogitsonceandtrainingthestudentagainstthecache)matchesonlinedistillationatnear-identicaltraininglosswhileremovingtheteacherfrommemory,runningabout29\%fasterperiteration,andreachingupto41\%higherthroughputonasingleH200GPU.Second,weintroduceafused,chunkedKLlossthatnevermaterialisesthefullvocabulary-sizedlogittensor,makingpeakmemorylinearinthesequencelength.Thisremovesthememoryspikethatotherwisecapscontextlengthandletsustrainatfourtimesthecontext(32{,}768tokens)onasingleGPU.Aseparateoutput-head-onlytoybenchmarkisolatesthelosskernelandconfirmsitsmemoryanditeration-ratescalingfrom4Kto256Ktokens.Togetherthesemakelarge-scalehealingandhundredsofablationsaffordable.Wealsoreportsupportingablationsonlossdesignandsequencepacking.Wereleaseourchunked-lossimplementation:https://github.com/CompactifAI/Full-Chunked-KL-Loss.

View arXiv pageView PDFProject pageGitHub3Add to collection

Get this paper in your agent:

hf papers read 2608\.03796

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper2

#### MultiverseComputingCAI/Hypernova-60B-2605 Text Generation• 59B• Updated3 days ago • 1.45k • 15 #### MultiverseComputingCAI/LittleLamb Text Generation• 0.3B• Updated3 days ago • 1.6k • 9

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.03796 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.03796 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles