MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Summary
The MAI-UI technical report presents a family of foundation GUI agents in multiple sizes, addressing real-world deployment challenges with a self-evolving data pipeline, device-cloud collaboration, and online RL, achieving state-of-the-art results on GUI grounding and mobile navigation benchmarks.
View Cached Full Text
Cached at: 08/20/26, 09:54 PM
Paper page - MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
Source: https://huggingface.co/papers/2512.22047
Abstract
ThedevelopmentofGUIagentscouldrevolutionizethenextgenerationofhuman-computerinteraction.Motivatedbythisvision,wepresentMAI-UI,afamilyoffoundationGUIagentsspanningthefullspectrumofsizes,including2B,8B,32B,and235B-A22Bvariants.Weidentifyfourkeychallengestorealisticdeployment:thelackofnativeagent-userinteraction,thelimitsofUI-onlyoperation,theabsenceofapracticaldeploymentarchitecture,andbrittlenessindynamicenvironments.MAI-UIaddressestheseissueswithaunifiedmethodology:aself-evolvingdatapipelinethatexpandsthenavigationdatatoincludeuserinteractionandMCPtoolcalls,anativedevice-cloudcollaborationsystemroutesexecutionbytaskstate,andanonlineRLframeworkwithadvancedoptimizationstoscaleparallelenvironmentsandcontextlength.MAI-UIestablishesnewstate-of-the-artacrossGUIgroundingandmobilenavigation.Ongroundingbenchmarks,itreaches73.5%onScreenSpot-Pro,91.3%onMMBenchGUIL2,70.9%onOSWorld-G,and49.2%onUI-Vision,surpassingGemini-3-ProandSeed1.8onScreenSpot-Pro.OnmobileGUInavigation,itsetsanewSOTAof76.7%onAndroidWorld,surpassingUI-Tars-2,Gemini-2.5-ProandSeed1.8.OnMobileWorld,MAI-UIobtains41.7%successrate,significantlyoutperformingend-to-endGUImodelsandcompetitivewithGemini-3-Probasedagenticframeworks.OuronlineRLexperimentsshowsignificantgainsfromscalingparallelenvironmentsfrom32to512(+5.2points)andincreasingenvironmentstepbudgetfrom15to50(+4.3points).Finally,thenativedevice-cloudcollaborationsystemimproveson-deviceperformanceby33%,reducescloudmodelcallsbyover40%,andpreservesuserprivacy.
View arXiv pageView PDFGitHub2.02kautoAdd to collection
Get this paper in your agent:
hf papers read 2512\.22047
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper6
#### Tongyi-MAI/MAI-UI-8B Image-Text-to-Text• 9B• UpdatedJan 9 • 2.89k • 199
#### mlx-community/MAI-UI-8B-bf16 Image-Text-to-Text• 9B• UpdatedMar 20 • 37
#### mlx-community/MAI-UI-8B-6bit Image-Text-to-Text• 2B• UpdatedMar 20 • 22
#### mlx-community/MAI-UI-8B-4bit Image-Text-to-Text• 2B• UpdatedMar 20 • 32
Browse 6 models citing this paper## Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2512.22047 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2512.22047 in a Space README.md to link it from this page.
Collections including this paper6
Similar Articles
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstrations to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks.
Xiaomi-GUI-0 Technical Report
This technical report presents Xiaomi-GUI-0, a native multimodal GUI agent trained and evaluated in real-device environments using a hybrid infrastructure and a progressive training pipeline, achieving high success rates and improved stability on real-world mobile tasks.
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
UI-TARS-2 is a native GUI-centered agent model that addresses data scalability, multi-turn RL, and environment stability challenges, achieving state-of-the-art results on GUI benchmarks (88.2 on Online-Mind2Web, 47.5 on OSWorld, 50.6 on WindowsAgentArena,73.3 on AndroidWorld) and outperforming Claude and OpenAI agents.
Macaron-A2UI: A Model for Generative UI in Personal Agents
Presents Macaron-A2UI, a model for generative UI in personal agents that synthesizes dynamic interfaces with lightweight executable actions, moving beyond text-only chat. The paper introduces a large-scale corpus, the A2UI-Bench benchmark, and trains models up to 754B parameters using LoRA fine-tuning and reinforcement learning, achieving strong results.