Tag
SAKI introduces a supervision allocation method for on-policy distillation that uses KL-constrained teacher-guided rollouts and maximal coupling to route token-level supervision, improving performance on mathematical reasoning benchmarks for small models.