step-level-policy-optimization

Tag

Cards List
#step-level-policy-optimization

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

arXiv cs.AI · 2026-07-14 Cached

EvoCUA-1.5 introduces an online reinforcement learning framework for multi-turn computer-use agents, achieving a 63.2% success rate on OSWorld-Verified and outperforming comparable open-weight models up to 35B parameters through step-level policy optimization and dynamic curriculum learning.

0 favorites 0 likes
← Back to home

Submit Feedback