Tag
Workflow-GYM is a benchmark for evaluating AI agents on long-horizon GUI tasks in professional domains. Experiments show that even top models achieve only ~30% success, revealing significant challenges.
This paper argues that agent skills should incorporate visual information, not just text, and proposes a multimodal skill paradigm combining textual logic with visual support. Experiments show visual skills outperform text-only approaches in visual-centric tasks.