Tag
SimTrace is an open-source framework that generates faithful, fine-grained synthetic multimodal user clickstreams using a computer-use agent grounded in anonymized real trajectories and simulated web environments. It outperforms baselines on 7 of 8 fidelity metrics and improves next-action prediction by 11.0% when augmenting real data.
SCOUT introduces a two-stage agentic safety verifier that combines reasoning-intensive rubric generation with tool-intensive evidence gathering to detect unsafe behaviors in computer-use agents, improving unsafe detection on AutoElicit-Bench and reducing unsafe execution rates from 30.2% to 17.2% via test-time reflection.
OSWorld-Science introduces a benchmark and evaluation environment for computer-using VLM agents on scientific software, covering 146 expert-designed tasks across domains like molecular drawing, pathology imaging, statistics, and physical simulation with artifact-based evaluators and a special agent harness.
The paper proposes HybridCUA, a method that combines GUI and CLI interactions for computer-use agents, using a two-stage training framework with supervised fine-tuning and reinforcement learning to improve task accuracy and cross-platform generalizability.
The paper introduces KNOWS, a benchmark for evaluating web agents on complex, long-horizon tasks that involve synthesizing and organizing knowledge into artifacts, revealing that current agents struggle with visual steps and long-horizon reasoning.
CUA-SWE introduces a benchmark, environment, and evaluation pipeline that unifies computer-use agents and coding agents, testing whether frontier agents can complete software engineering tasks requiring visual feedback from running applications, across four SWE domains with deterministic tests.
Cognitive Sharding separates cognitive functions across specialist models to enable computer-use agents to run locally on 16 GB of memory, prioritizing accuracy and error containment over speed.
This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.
This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.
RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.
ERPBench introduces a benchmark for evaluating screenshot-only computer-use agents on a live ERP system, showing that strong general GUI performance does not transfer to enterprise reliability and includes a production harness with human approval for safe deployment.
HazardAuditor introduces an execution-grounded framework for supervising safety in computer-use agents, with Guard Policy Optimization improving safety outcomes by up to 16.5% accuracy over prior methods.
The paper introduces GUI-Primitives, a benchmark of 994 contrastive instruction pairs to diagnose spatial reasoning failures in vision-language models for GUI grounding, revealing that most failures stem from candidate localization rather than relation understanding.
Task Model Induction (TMI) is a method that induces structured task models from raw computer-use traces, improving agent performance by 30% and enabling auditable, reusable skills for real-world workflows.
ComponentBench introduces a benchmark and diagnostic pipeline for evaluating computer-use agents on component-level interactions in modern web UIs, addressing gaps in current evaluation methods by focusing on realistic, short interactions to diagnose failures across models.
This paper presents a deterministic, zero-model pipeline that compiles passively captured screen activity into structured 'activity frames' for agent memory, reducing a day of raw capture to a prompt-ready context block 86× smaller and achieving 98.4% QA accuracy. It also introduces measurements of routine overhead ratio and recurrence to model agent costs.
This paper investigates when hybrid computer-use agents actually choose to use MCP tools versus screenshots, finding that tool availability alone does not guarantee adoption: a reasoning model improves while a non-reasoning model degrades. It also explores training and context compression strategies to close the adoption gap and reduce token costs.
Microsoft Research introduces Echoverse, a set of deep, evolving environments for training computer-use agents. A 9B model trained on these environments nearly doubles its baseline score, coming within 14 points of GPT-5.4, demonstrating that high-fidelity simulation and co-evolution of model, world, and verifier significantly improve agent performance on multi-step workflows.
OSReward introduces a standardized benchmark for evaluating VLM judges on computer-use agent trajectories, revealing that even state-of-the-art models have systematic leniency bias. The authors release OS-Shepherd-9B and 35B reward models that provide reliable, low-cost rewards for CUA tasks.
Echoverse presents a method for generating deep, evolving synthetic environments to train computer-use agents, demonstrating substantial accuracy gains and releasing a benchmark with grounded graders.