Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Summary
Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.
View Cached Full Text
Cached at: 07/31/26, 05:52 AM
Paper page - Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Source: https://huggingface.co/papers/2607.28227 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
GUIagentshavethepotentialtobecomeageneralpurposeexecutoroverexistingdigitaldevices.Toadvancethemtowardreal-worlduse,weenvisionagentsthatoperatereliablyonrealdevices,executeworkflowsacrossplatforms,combineGUIinteractionwithCLIexecution,completelong-horizontasks,proactivelyinitiateusefulservices,andautonomouslyimprovetheircapabilitieswithminimalhumaneffort.Guidedbythisvision,wepresentQwen-UI-Agent,areal-worldcentricfoundationGUIagentspanningmobile,computer-use,web,andDeepSearchenvironments.Qwen-UI-Agentcombinesdiversesandboxenvironmentswithalarge-scalereal-devicemobileruntime.ItsunifiedactionspaceinterleavesGUIoperationswithCLIexecutionandgeneratesbatchedactionsinasinglemodelturn.AnAutoResearch-styledataflywheelusesagentstoconstructtasksandenvironments,diagnosefailures,andplansubsequentiterations.OnlineRLsupportstrainingontrajectoriesexceeding100turns,withover10,000concurrentenvironmentsacceleratingrollout.Alightweightharnesslayersupportsproactiveserviceinitiationandstatefulworkflowsacrossmobileandcomputer.Acrossabroadsuiteofevaluations,Qwen-UI-Agentsetsstate-of-the-artperformanceonmobile-usebenchmarkswhiledeliveringcompetitiveperformanceoncomputer-andbrowser-usetasksagainstfrontiermodels,includingOpus4.8,Gemini3.1Pro,andGPT-5.6Sol.Onmobileuse,itachieves82.1%onMobileWorld,92.2%onMobileWorld-Real,and97.5%onAndroidDaily.Oncomputeruse,itachieves79.5%onOSWorld-Verifiedanda40.0%partial-progressscoreonOSWorld-v2.OnbrowseruseandGUIgrounding,itachieves73.6%onWebArenaand81.5%onScreenSpot-Pro,respectively.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.28227
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.28227 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.28227 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.28227 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Qwen3.7: The Agent Frontier (15 minute read)
Alibaba's Qwen team has released Qwen3.7-Max, a proprietary agent-foundation model achieving top scores on multiple benchmarks including Terminal-Bench 2.0, SWE-Pro, and GPQA Diamond, with consistent performance across various code environments.
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
The MAI-UI technical report presents a family of foundation GUI agents in multiple sizes, addressing real-world deployment challenges with a self-evolving data pipeline, device-cloud collaboration, and online RL, achieving state-of-the-art results on GUI grounding and mobile navigation benchmarks.
Qwen3.7-Plus: Multimodal Agent Intelligence (36 minute read)
Qwen3.7-Plus is a multimodal agent model that unifies vision and language for seamless GUI and CLI interactions, now available via Alibaba Cloud Model Studio.
Qwen/Qwen-AgentWorld-35B-A3B
Qwen releases Qwen-AgentWorld-35B-A3B, a native language world model that simulates agentic environments across seven domains via long chain-of-thought reasoning. The model is trained with a three-stage pipeline and supports MCP, Search, Terminal, SWE, Android, Web, and OS interactions.
@NielsRogge: Qwen released a SOTA computer use agent with a technical report It's not on @arxiv, but it is on Papers with Code. You …
Qwen released Qwen-CUA, a native computer-use agent with a 397B-A17B mixture-of-experts backbone, achieving state-of-the-art results on OSWorld-Verified and ranking #2 on WebArena. A technical report is available on Papers with Code.