MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Summary
MyPCBench evaluates computer-use agents as personal assistants in a simulated Linux desktop environment with real-world web applications, revealing that Claude Opus 4.6 achieves the highest task completion rate of 55.4% while struggling with multi-application tasks and long trajectories.
View Cached Full Text
Cached at: 06/18/26, 03:58 PM
Paper page - MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Source: https://huggingface.co/papers/2606.16748
Abstract
MyPCBench evaluates computer-use agents as personal assistants in a simulated Linux desktop environment with real-world web applications, revealing that Claude Opus 4.6 achieves the highest task completion rate of 55.4% while struggles with multi-application tasks and long trajectories.
Current benchmarks forcomputer-use agentsevaluate models in impersonal environments. This leaves a gap between evaluation and deployment wherepersonal assistantsare expected to work across a user’s wholedigital life, including their context, historical data, and logged-in accounts. This gap is widest onweb tasks, wherelive web evaluationscannot exercise sites that require logging in or personal information, the kind of site a real personal assistant has to drive. We introduce MyPCBench, which testscomputer-use agentsaspersonal assistantson a Linux desktop populated with 17 simulated real-worldweb applicationsand a full desktop stack, all seeded for one canonical persona, Michael Scott from The Office. We define 184 tasks in this environment, each inspired by a real request drawn from the OpenClaw community, and benchmark six closed and open-weight models with a uniform computer+bash tool surface. We find that the best model, Claude Opus 4.6, fully solves 55.4\% of the tasks, the only model above 50\%. Model failures cluster on tasks that span many applications and on long trajectories, where personalization stresses an assistant the most. We release the environment, task set, andagent harnessat https://mypcbench.com.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2606\.16748
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.16748 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.16748 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.16748 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.
WorkBench Revisited: Workplace Agents Two Years On
This paper revisits the WorkBench benchmark for workplace agents two years after its initial release, showing that the best agent (Claude Opus 4.8) now completes 89% of tasks with only 2.5% harmful side effects, compared to GPT-4's 43% completion and 26% harm rate in 2024. It finds that capability and safety improve together, open-weight models have drastically lowered costs, and some basic mistakes persist.
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
ComponentBench introduces a benchmark and diagnostic pipeline for evaluating computer-use agents on component-level interactions in modern web UIs, addressing gaps in current evaluation methods by focusing on realistic, short interactions to diagnose failures across models.
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
WeaveBench is a new benchmark for evaluating computer-use agents across multiple interfaces (GUI, CLI, code) in long-horizon real-world tasks. It reveals that current models achieve only 41.2% PassRate and that outcome-only grading overestimates performance, highlighting significant gaps in evaluation.
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Introduces PPT-Eval, a benchmark of 120 PowerPoint tasks for evaluating computer-use agents, with a rubric-based scoring system that awards partial credit. Strong frontier agents like Claude-4.5-Opus achieve only 45% success rate, highlighting the difficulty of such tasks.