Tag
This paper presents a practical study on making knowledge distillation training for LLMs more efficient, introducing offline top-K logits caching and a fused chunked KL loss that reduces memory spikes and enables longer contexts on a single GPU.