entropy-adaptive

Tag

Cards List
#entropy-adaptive

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Hugging Face Daily Papers · 2026-08-31 Cached

The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.

0 favorites 0 likes
#entropy-adaptive

WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning

arXiv cs.LG · 2026-07-09 Cached

Proposes FedEAS, a budget-aware policy for synthetic data augmentation in federated learning that assigns each client an entropy-adaptive per-class generation budget, recovering most accuracy gains of full class balancing while reducing generation cost by 94.1%.

0 favorites 0 likes
#entropy-adaptive

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

arXiv cs.LG · 2026-06-05 Cached

This paper introduces Adaptive-Horizon and Selective-Advantage variants of GRPO that use entropy-based token-level discounting to stabilize training and improve performance on math reasoning tasks, achieving stronger results with lower variance.

0 favorites 0 likes
← Back to home

Submit Feedback