kl-loss

Tag

Cards List
#kl-loss

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Hugging Face Daily Papers ↗ · 2026-08-04 Cached

This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.

0 favorites 0 likes
← Back to home

Submit Feedback