Tag
This paper investigates generalization in on-policy distillation for large language models, showing that it transfers reasoning behaviors and that teacher-student origin alignment is crucial, with multi-teacher combinations causing capability trade-offs.