Tag
This paper investigates on-policy distillation in large language models using sparse crosscoders, revealing that it reweights existing features rather than creating new ones, with SFT warm-up playing a role in pre-reweighting features.