real-world-workflows

Tag

Cards List
#real-world-workflows

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

arXiv cs.AI ↗ · 2026-08-28 Cached

DuMateBench introduces a benchmark derived from anonymized user sessions to evaluate autonomous agents in complex real-world workflows, testing performance under environmental complexities such as insufficiency, instability, and noise.

0 favorites 0 likes
#real-world-workflows

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Hugging Face Daily Papers ↗ · 2026-06-28 Cached

OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.

0 favorites 0 likes
← Back to home

Submit Feedback