HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing computer agents that predominantly rely on inefficient graphical user interfaces (GUIs) or costly application programming interfaces (APIs), lacking synergistic capabilities with command-line interfaces (CLIs). To bridge this gap, this work proposes a hybrid GUI-CLI interaction paradigm. By constructing a mixed-trajectory data pipeline and integrating supervised fine-tuning with reinforcement learning, the proposed approach trains agents to dynamically coordinate dual-modal operations, further optimized through a novel CLI-aware reward mechanism. Experimental results demonstrate that this method achieves an accuracy of 53.6% on the OSWorld benchmark, outperforming baselines by 14.8 percentage points, while also yielding significant improvements in WindowsAgentArena. These findings confirm the efficacy of the proposed framework in enabling efficient, cross-platform automated computer operations.
📝 Abstract
Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often inefficient and error prone, or augment GUI interactions with application specific APIs or tools, which require substantial engineering effort and are difficult to scale across applications. We argue that the next generation of CUAs should combine GUI interactions with the command line interface (CLI), leveraging the generality of the GUI and the efficiency of shell commands. A critical challenge, however, is that current models do not know when or how to use the CLI during task execution. To address this challenge, we develop a data construction pipeline that produces three types of trajectories: GUI only, CLI only, and interleaved GUI and CLI trajectories. This pipeline results in HybridCUA-8K, containing 5K hybrid trajectories and 3K verified RLVR tasks. Building on these data, we propose a training framework with two stages: supervised fine tuning on the constructed trajectories, followed by reinforcement learning with our CLI aware rewards that encourages agents to use the CLI selectively and reliably. Experiments show that HybridCUA-9B achieves 53.6% accuracy on OSWorld, improving over the base model by 14.8 percentage points, and improves performance on WindowsAgentArena by 4.0 percentage points. These results demonstrate the effectiveness and cross platform generalizability of the hybrid GUI and CLI paradigm for computer use agents.
Problem

Research questions and friction points this paper is trying to address.

Computer-Use Agents
Graphical User Interface
Command Line Interface
Hybrid Interaction
Task Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computer-Use Agents
Hybrid GUI-CLI
Reinforcement Learning
Data Construction Pipeline
Cross-platform Generalizability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.