Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Summary
This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.
View Cached Full Text
Cached at: 08/10/26, 02:15 PM
Paper page - Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Source: https://huggingface.co/papers/2608.03796
Abstract
Smalllanguagemodelsareoftentheonlyoptionfordeploymentundertightlatency,cost,andon-premisesconstraints,buttheyarerarelytrainedfromscratch:acompressedmodelisusuallyrecoveredthroughknowledgedistillation(KD).Thisrecoverysteplargelydecidesthefinalquality,yetitisexpensive.Wepresentapractitioner’sstudyofhowtomakedistillationtrainingefficient,organisedaroundtwosystemscontributions.First,weshowthatofflineKD(cachingtheteacher’stop-Klogitsonceandtrainingthestudentagainstthecache)matchesonlinedistillationatnear-identicaltraininglosswhileremovingtheteacherfrommemory,runningabout29\%fasterperiteration,andreachingupto41\%higherthroughputonasingleH200GPU.Second,weintroduceafused,chunkedKLlossthatnevermaterialisesthefullvocabulary-sizedlogittensor,makingpeakmemorylinearinthesequencelength.Thisremovesthememoryspikethatotherwisecapscontextlengthandletsustrainatfourtimesthecontext(32{,}768tokens)onasingleGPU.Aseparateoutput-head-onlytoybenchmarkisolatesthelosskernelandconfirmsitsmemoryanditeration-ratescalingfrom4Kto256Ktokens.Togetherthesemakelarge-scalehealingandhundredsofablationsaffordable.Wealsoreportsupportingablationsonlossdesignandsequencepacking.Wereleaseourchunked-lossimplementation:https://github.com/CompactifAI/Full-Chunked-KL-Loss.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2608\.03796
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper2
#### MultiverseComputingCAI/Hypernova-60B-2605 Text Generation• 59B• Updated3 days ago • 1.45k • 15
#### MultiverseComputingCAI/LittleLamb Text Generation• 0.3B• Updated3 days ago • 1.6k • 9
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.03796 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.03796 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Making Knowledge Distillation Cheap Enough to Run at Scale
Multiverse Computing announces a paper on making LLM knowledge distillation cheaper via offline top-K logits and a fused chunked KL loss, cutting VRAM usage for distillation at scale.
FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
The paper proposes FLoKD, an adaptive knowledge-distillation framework for federated LoRA fine-tuning of LLMs over wireless networks, reducing communication overhead by 50-65% while maintaining competitive performance.
Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking
This paper presents a method to fine-tune LLaMA 3 8B as an efficient reranker for Retrieval-Augmented Generation using knowledge distillation and 4-bit quantization, achieving 14-21% gains in retrieval metrics over cross-encoder baselines with reduced inference cost.
Layer-wise Curriculum Learning for Efficient LLM Compression
The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions
This paper proposes a unified approach to interpret knowledge distillation in LLMs using game-theoretic interactions, discovering that distillation sparsifies interactions, and introduces a loss function CIP to improve performance.