Tag
The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.
Proposes FedEAS, a budget-aware policy for synthetic data augmentation in federated learning that assigns each client an entropy-adaptive per-class generation budget, recovering most accuracy gains of full class balancing while reducing generation cost by 94.1%.
This paper introduces Adaptive-Horizon and Selective-Advantage variants of GRPO that use entropy-based token-level discounting to stabilize training and improve performance on math reasoning tasks, achieving stronger results with lower variance.