masked-distillation

Tag

Cards List
#masked-distillation

Masked Distillation: Internalizing the Chain-of-Thought in Language Models

arXiv cs.AI ↗ · 2026-07-28 Cached

Masked distillation is a knowledge-distillation framework that trains a student LLM to predict only solution tokens while a reasoning teacher provides feedback, aiming to internalize chain-of-thought computation into model parameters. The method shows task-dependent success, working on GSM8K but requiring small scaffolds for harder tasks like Countdown.

0 favorites 0 likes
← Back to home

Submit Feedback