HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
Summary
The paper proposes HybridCUA, a method that combines GUI and CLI interactions for computer-use agents, using a two-stage training framework with supervised fine-tuning and reinforcement learning to improve task accuracy and cross-platform generalizability.
View Cached Full Text
Cached at: 09/30/26, 04:15 AM
Paper page - HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
Source: https://huggingface.co/papers/2609.38008 Authors:
,
,
,
,
,
,
,
,
,
Abstract
Computeruseagents(CUAs)havedemonstratedstrongcapabilitiesincompletingdigitaltasks.However,existingCUAseitherrelysolelyongraphicaluserinterface(GUI)interactions,whichareofteninefficientanderrorprone,oraugmentGUIinteractionswithapplicationspecificAPIsortools,whichrequiresubstantialengineeringeffortandaredifficulttoscaleacrossapplications.WearguethatthenextgenerationofCUAsshouldcombineGUIinteractionswiththecommandlineinterface(CLI),leveragingthegeneralityoftheGUIandtheefficiencyofshellcommands.Acriticalchallenge,however,isthatcurrentmodelsdonotknowwhenorhowtousetheCLIduringtaskexecution.Toaddressthischallenge,wedevelopadataconstructionpipelinethatproducesthreetypesoftrajectories:GUIonly,CLIonly,andinterleavedGUIandCLItrajectories.ThispipelineresultsinHybridCUA-8K,containing5Khybridtrajectoriesand3KverifiedRLVRtasks.Buildingonthesedata,weproposeatrainingframeworkwithtwostages:supervisedfinetuningontheconstructedtrajectories,followedbyreinforcementlearningwithourCLIawarerewardsthatencouragesagentstousetheCLIselectivelyandreliably.ExperimentsshowthatHybridCUA-9Bachieves53.6%accuracyonOSWorld,improvingoverthebasemodelby14.8percentagepoints,andimprovesperformanceonWindowsAgentArenaby4.0percentagepoints.TheseresultsdemonstratetheeffectivenessandcrossplatformgeneralizabilityofthehybridGUIandCLIparadigmforcomputeruseagents.
View arXiv pageView PDFProject pageGitHub3Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38008 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38008 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38008 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
ToolCUA is a new agent framework that optimizes GUI-tool path selection for computer use agents through staged training and reinforcement learning. It achieves state-of-the-art performance on OSWorld-MCP by effectively interleaving GUI actions and high-level tool calls.
Computer-Using Agent
OpenAI introduced the Computer-Using Agent (CUA), a model combining GPT-4o's vision with reinforcement learning to interact with GUIs like a human, powering the new Operator agent. CUA sets new state-of-the-art benchmarks including 38.1% on OSWorld and 58.1% on WebArena, and is available as a research preview for ChatGPT Pro users in the US.
PRO-CUA: Process-Reward Optimization for Computer Use Agents
This paper introduces PRO-CUA, a process-reward optimization framework for training Computer Use Agents (CUAs) using iterative step-level reinforcement learning. The method decouples on-policy environment interaction from policy optimization, enabling dense credit assignment without relying on expert trajectories, and demonstrates effectiveness on live web benchmarks.
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
ClawGUI is an open-source framework for training, evaluating, and deploying GUI agents using reinforcement learning, featuring standardized benchmarks and cross-platform deployment to Android, iOS, and HarmonyOS.
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
GUICrafter introduces a weakly-supervised GUI agent that leverages massive unannotated screenshots and a two-stage curriculum learning framework to reduce reliance on expensive human annotations, achieving competitive performance with advanced systems like UI-TARS using only 0.1% of its data.