Tag
This paper investigates how on-policy distillation can cause length inflation due to EOS token mismatches between student and teacher models, and proposes a correction method by aggregating EOS probabilities to reduce response length.