RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Summary
RecreationWorld introduces a scalable and verifiable framework for hybrid computer-use agents, providing environments across five platforms and a benchmark for evaluation.
View Cached Full Text
Cached at: 09/21/26, 03:19 AM
Paper page - RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Source: https://huggingface.co/papers/2609.22000 Published on Sep 18
#1 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Computer-useagents(CUAs)haveadvancedalongtwoseparatelines:graphicalinteractionandsoftwaredevelopmentthroughcodeandthecommandline.Realdigitalworkrequiresboth,interleavedratherthanstackedendtoend.WestudyhybridCUAsthatautonomouslydecidewhentoexploreaninterface,implementsoftware,andrunandvisuallyverifytheirartifacts.WeintroduceRecreationWorld,afive-platformframeworkbuiltaroundrecreation:givenarunningreference,anagentmustdiscoveritsbehaviorandbuildafaithfulimplementationwithnoprescribedworkflow.RecreationWorldprovidesreproducibleenvironmentsonUbuntu,macOS,Windows,Android,andWeb,plusaunifiedharnesswithnativeGUIcontrolandcodingtools.Therunningreferenceservesasanoracleforhiddenbehavioraltests,providingexecution-groundedrewards.Wescaletrajectorygenerationwithhigh-qualityopen-sourceapplications.Modelstrainedonthesetrajectoriesimproveacrossfiveout-of-distributioncodingandhybridcomputer-usebenchmarksandmorefrequentlyverifytheirrenderedoutputs,providingevidenceoftransferbeyondrecreation.Forheld-outevaluation,weintroduceRecreationBench,comprising250diversetasksacrossdomainsandplatforms.Reference-groundedprogrammaticandvisualassertionscoveraction-conditionedoutcomesatmultipleinteractiondepths;eachisvalidatedonthereferenceandbyhumanreviewersbeforethesuiteisfrozenforautomaticscoring.GPT-6Astraleadsat58.1%overall,butpassesallprogrammatictestsonjust2.8%oftasks.Agentsreproducestaticinterfacestructuremorereliablythaninteractionsandcomputedoutputs,whilegeneratedapplicationsremainsmallerandmoremonolithicthantheirreferences.Wereleasethebenchmark,environments,andtestsuites.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.22000
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.22000 in a model README.md to link it from this page.
Datasets citing this paper1
#### Qwen/RecreationBench Viewer• Updated35 minutes ago • 500 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.22000 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Qwen's RecreationWorld Trains Agents to Rebuild Apps (GitHub Repo)
RecreationWorld is a scalable framework for training hybrid AI agents that combine GUI interaction, coding, and visual verification to rebuild applications, with a benchmark suite called RecreationBench.
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
OpenComputer presents a framework for creating verifiable software environments for computer-use agents, integrating state verifiers, self-improving verification layers, task synthesis, and evaluation systems across 33 desktop applications. Experiments show its verifiers align better with human judgment than LLM-as-judge, and frontier agents struggle with end-to-end completion.
Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning
This paper introduces a framework for constructing verified synthetic web environments to improve the training of web agents, demonstrating enhanced performance and transferability across benchmarks.
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
CUA-Gym introduces a scalable pipeline for generating verifiable training environments and tasks for computer-use agents, addressing data scarcity. The resulting dataset and models achieve strong performance on benchmarks like OSWorld-Verified and WebArena.
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
AgentMercury introduces a scalable framework for synthesizing verifiable environments from business scenarios, enabling reinforcement learning agents to improve performance on enterprise and out-of-domain benchmarks.