group-in-group-policy-optimization

Tag

Cards List
#group-in-group-policy-optimization

When Does Muon Help Agentic Reinforcement Learning?

Hugging Face Daily Papers · 2026-07-17 Cached

This paper investigates the use of the Muon optimizer in reinforcement learning post-training, finding that applying Muon to hidden weight matrices significantly improves success rates on ALFRED tasks compared to AdamW, with results dependent on the advantage estimator and learning rate.

0 favorites 0 likes
← Back to home

Submit Feedback