Tag
DuMateBench introduces a benchmark derived from anonymized user sessions to evaluate autonomous agents in complex real-world workflows, testing performance under environmental complexities such as insufficiency, instability, and noise.
OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.