computer-use-agents

Tag

Cards List
#computer-use-agents

SimTrace: Grounded Multimodal User Trajectories Generation for Online User Modeling

arXiv cs.AI ↗ · 4d ago Cached

SimTrace is an open-source framework that generates faithful, fine-grained synthetic multimodal user clickstreams using a computer-use agent grounded in anonymized real trajectories and simulated web environments. It outperforms baselines on 7 of 8 fidelity metrics and improves next-action prediction by 11.0% when augmenting real data.

0 favorites 0 likes
#computer-use-agents

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

arXiv cs.CL ↗ · 5d ago Cached

SCOUT introduces a two-stage agentic safety verifier that combines reasoning-intensive rubric generation with tool-intensive evidence gathering to detect unsafe behaviors in computer-use agents, improving unsafe detection on AutoElicit-Bench and reducing unsafe execution rates from 30.2% to 17.2% via test-time reflection.

0 favorites 0 likes
#computer-use-agents

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Hugging Face Daily Papers ↗ · 5d ago Cached

OSWorld-Science introduces a benchmark and evaluation environment for computer-using VLM agents on scientific software, covering 146 expert-designed tasks across domains like molecular drawing, pathology imaging, statistics, and physical simulation with artifact-based evaluators and a special agent harness.

0 favorites 0 likes
#computer-use-agents

HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Hugging Face Daily Papers ↗ · 6d ago Cached

The paper proposes HybridCUA, a method that combines GUI and CLI interactions for computer-use agents, using a two-stage training framework with supervised fine-tuning and reinforcement learning to improve task accuracy and cross-platform generalizability.

0 favorites 0 likes
#computer-use-agents

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

arXiv cs.CL ↗ · 2026-09-28 Cached

The paper introduces KNOWS, a benchmark for evaluating web agents on complex, long-horizon tasks that involve synthesizing and organizing knowledge into artifacts, revealing that current agents struggle with visual steps and long-horizon reasoning.

0 favorites 0 likes
#computer-use-agents

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Hugging Face Daily Papers ↗ · 2026-09-26 Cached

CUA-SWE introduces a benchmark, environment, and evaluation pipeline that unifies computer-use agents and coding agents, testing whether frontier agents can complete software engineering tasks requiring visual feedback from running applications, across four SWE domains with deterministic tests.

0 favorites 0 likes
#computer-use-agents

Cognitive Sharding: Computer-Use Agents Can Run Locally on 16 GB

Reddit r/ArtificialInteligence ↗ · 2026-09-25

Cognitive Sharding separates cognitive functions across specialist models to enable computer-use agents to run locally on 16 GB of memory, prioritizing accuracy and error containment over speed.

0 favorites 0 likes
#computer-use-agents

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

arXiv cs.AI ↗ · 2026-09-24 Cached

This paper introduces CAVEAT, a benchmark for evaluating computer-use agents in incentive-misaligned environments, revealing that agents often fail to preserve user objectives and proposes interventions to improve robustness.

0 favorites 0 likes
#computer-use-agents

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.

0 favorites 0 likes
#computer-use-agents

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Hugging Face Daily Papers ↗ · 2026-09-18 Cached

RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.

0 favorites 0 likes
#computer-use-agents

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

arXiv cs.AI ↗ · 2026-09-17 Cached

ERPBench introduces a benchmark for evaluating screenshot-only computer-use agents on a live ERP system, showing that strong general GUI performance does not transfer to enterprise reliability and includes a production harness with human approval for safe deployment.

0 favorites 0 likes
#computer-use-agents

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

HazardAuditor introduces an execution-grounded framework for supervising safety in computer-use agents, with Guard Policy Optimization improving safety outcomes by up to 16.5% accuracy over prior methods.

0 favorites 0 likes
#computer-use-agents

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

arXiv cs.CL ↗ · 2026-08-25 Cached

The paper introduces GUI-Primitives, a benchmark of 994 contrastive instruction pairs to diagnose spatial reasoning failures in vision-language models for GUI grounding, revealing that most failures stem from candidate localization rather than relation understanding.

0 favorites 0 likes
#computer-use-agents

@dair_ai: This is one of the most effective ways to improve your agentic workflows. If you are building computer-use agents, this…

X AI KOLs Timeline ↗ · 2026-08-22 Cached

Task Model Induction (TMI) is a method that induces structured task models from raw computer-use traces, improving agent performance by 30% and enabling auditable, reusable skills for real-world workflows.

0 favorites 0 likes
#computer-use-agents

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

arXiv cs.AI ↗ · 2026-08-20 Cached

ComponentBench introduces a benchmark and diagnostic pipeline for evaluating computer-use agents on component-level interactions in modern web UIs, addressing gaps in current evaluation methods by focusing on realistic, short interactions to diagnose failures across models.

0 favorites 0 likes
#computer-use-agents

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

arXiv cs.AI ↗ · 2026-08-07 Cached

This paper presents a deterministic, zero-model pipeline that compiles passively captured screen activity into structured 'activity frames' for agent memory, reducing a day of raw capture to a prompt-ready context block 86× smaller and achieving 98.4% QA accuracy. It also introduces measurements of routine overhead ratio and recurrence to model agent costs.

0 favorites 0 likes
#computer-use-agents

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

arXiv cs.AI ↗ · 2026-08-05 Cached

This paper investigates when hybrid computer-use agents actually choose to use MCP tools versus screenshots, finding that tool availability alone does not guarantee adoption: a reasoning model improves while a non-reasoning model degrades. It also explores training and context compression strategies to close the adoption gap and reduce token costs.

0 favorites 0 likes
#computer-use-agents

@MSFTResearch: Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in r…

X AI KOLs Timeline ↗ · 2026-07-30 Cached

Microsoft Research introduces Echoverse, a set of deep, evolving environments for training computer-use agents. A 9B model trained on these environments nearly doubles its baseline score, coming within 14 points of GPT-5.4, demonstrating that high-fidelity simulation and co-evolution of model, world, and verifier significantly improve agent performance on multi-step workflows.

0 favorites 0 likes
#computer-use-agents

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

OSReward introduces a standardized benchmark for evaluating VLM judges on computer-use agent trajectories, revealing that even state-of-the-art models have systematic leniency bias. The authors release OS-Shepherd-9B and 35B reward models that provide reliable, low-cost rewards for CUA tasks.

0 favorites 0 likes
#computer-use-agents

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

Echoverse presents a method for generating deep, evolving synthetic environments to train computer-use agents, demonstrating substantial accuracy gains and releasing a benchmark with grounded graders.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback