advantage-reweighting

Tag

Cards List
#advantage-reweighting

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Hugging Face Daily Papers · 2026-06-17 Cached

STARE addresses policy entropy collapse in GRPO-based reinforcement learning for large language models by introducing surprisal-guided token-level advantage reweighting and target-entropy regulation, achieving 4%-8% accuracy gains on AIME benchmarks.

0 favorites 0 likes
#advantage-reweighting

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards

Hugging Face Daily Papers · 2026-06-03 Cached

GRAIL introduces gradient-reweighted advantages to improve token-level credit assignment in reinforcement learning for LLM reasoning, outperforming GRPO across multiple models.

0 favorites 0 likes
← Back to home

Submit Feedback