Tag
Proposes Progress-conditioned Group Policy Optimization (ProGPO) to overcome credit assignment issues in long-horizon agentic tasks by using first-visit observation coverage when all group samples fail, improving performance on ALFWorld and WebShop.
ProGPO is a learned-critic-free method for step-level advantage estimation in group-based RL for LLM agents, using exact-prefix action comparisons and rollout-based state potentials to improve credit assignment on long-horizon tasks. Experiments on ALFWorld and WebShop with Qwen2.5 models show it outperforms existing agentic RL baselines.