HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Hugging Face Daily Papers Papers

Summary

The paper proposes HybridCUA, a method that combines GUI and CLI interactions for computer-use agents, using a two-stage training framework with supervised fine-tuning and reinforcement learning to improve task accuracy and cross-platform generalizability.

Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:15 AM

Paper page - HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Source: https://huggingface.co/papers/2609.38008 Authors:

,

,

,

,

,

,

,

,

,

Abstract

Computeruseagents(CUAs)havedemonstratedstrongcapabilitiesincompletingdigitaltasks.However,existingCUAseitherrelysolelyongraphicaluserinterface(GUI)interactions,whichareofteninefficientanderrorprone,oraugmentGUIinteractionswithapplicationspecificAPIsortools,whichrequiresubstantialengineeringeffortandaredifficulttoscaleacrossapplications.WearguethatthenextgenerationofCUAsshouldcombineGUIinteractionswiththecommandlineinterface(CLI),leveragingthegeneralityoftheGUIandtheefficiencyofshellcommands.Acriticalchallenge,however,isthatcurrentmodelsdonotknowwhenorhowtousetheCLIduringtaskexecution.Toaddressthischallenge,wedevelopadataconstructionpipelinethatproducesthreetypesoftrajectories:GUIonly,CLIonly,andinterleavedGUIandCLItrajectories.ThispipelineresultsinHybridCUA-8K,containing5Khybridtrajectoriesand3KverifiedRLVRtasks.Buildingonthesedata,weproposeatrainingframeworkwithtwostages:supervisedfinetuningontheconstructedtrajectories,followedbyreinforcementlearningwithourCLIawarerewardsthatencouragesagentstousetheCLIselectivelyandreliably.ExperimentsshowthatHybridCUA-9Bachieves53.6%accuracyonOSWorld,improvingoverthebasemodelby14.8percentagepoints,andimprovesperformanceonWindowsAgentArenaby4.0percentagepoints.TheseresultsdemonstratetheeffectivenessandcrossplatformgeneralizabilityofthehybridGUIandCLIparadigmforcomputeruseagents.

View arXiv pageView PDFProject pageGitHub3Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38008 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38008 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38008 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Computer-Using Agent

OpenAI Blog

OpenAI introduced the Computer-Using Agent (CUA), a model combining GPT-4o's vision with reinforcement learning to interact with GUIs like a human, powering the new Operator agent. CUA sets new state-of-the-art benchmarks including 38.1% on OSWorld and 58.1% on WebArena, and is available as a research preview for ChatGPT Pro users in the US.

PRO-CUA: Process-Reward Optimization for Computer Use Agents

arXiv cs.AI

This paper introduces PRO-CUA, a process-reward optimization framework for training Computer Use Agents (CUAs) using iterative step-level reinforcement learning. The method decouples on-policy environment interaction from policy optimization, enabling dense credit assignment without relying on expert trajectories, and demonstrates effectiveness on live web benchmarks.