IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents

📅 2026-02-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of existing agents deviating from user intent during long-horizon tasks, often leading to redundant handling of routine subproblems, error accumulation, and inefficiency. To mitigate these issues, the authors propose a multi-agent collaborative framework that leverages multi-perspective intent representation learning and skill abstraction to construct an intent-prototype-based skill retrieval mechanism. The framework further incorporates shared plan memory to enable intent-aligned planning and execution. Built upon a Planner-Optimizer-Critic multi-agent architecture, the approach achieves a task success rate of 74.83% and a step efficiency ratio of 0.91 in end-to-end evaluation, significantly outperforming baseline methods based on reinforcement learning and trajectory retrieval, thereby enhancing both stability and efficiency in long-horizon task execution.

Technology Category

Multiagent Systems: Multiagent PlanningIntelligent Robots: Multi-Robot SystemsHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Search and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Computer-use agents operate over long horizons under noisy perception, multi-window contexts, evolving environment states. Existing approaches, from RL-based planners to trajectory retrieval, often drift from user intent and repeatedly solve routine subproblems, leading to error accumulation and inefficiency. We present IntentCUA, a multi-agent computer-use framework designed to stabilize long-horizon execution through intent-aligned plan memory. A Planner, Plan-Optimizer, and Critic coordinate over shared memory that abstracts raw interaction traces into multi-view intent representations and reusable skills. At runtime, intent prototypes retrieve subgroup-aligned skills and inject them into partial plans, reducing redundant re-planning and mitigating error propagation across desktop applications. In end-to-end evaluations, IntentCUA achieved a 74.83% task success rate with a Step Efficiency Ratio of 0.91, outperforming RL-based and trajectory-centric baselines. Ablations show that multi-view intent abstraction and shared plan memory jointly improve execution stability, with the cooperative multi-agent loop providing the largest gains on long-horizon tasks. These results highlight that system-level intent abstraction and memory-grounded coordination are key to reliable and efficient desktop automation in large, dynamic environments.
Problem

Research questions and friction points this paper is trying to address.

intent alignment
long-horizon planning
multi-agent coordination
skill abstraction
error propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

intent-level representation
multi-agent planning
plan memory
skill abstraction
computer-use agents
🔎 Similar Papers
S
Seoyoung Lee
Sookmyung Women’s University, Seoul, Republic of Korea
S
Seobin Yoon
Sookmyung Women’s University, Seoul, Republic of Korea
S
Seongbeen Lee
Sookmyung Women’s University, Seoul, Republic of Korea
Y
Yoojung Chun
Sookmyung Women’s University, Seoul, Republic of Korea
D
Dayoung Park
Sookmyung Women’s University, Seoul, Republic of Korea
D
Doyeon Kim
Sookmyung Women’s University, Seoul, Republic of Korea
Joo Yong Sim
Joo Yong Sim
Electronics and Telecommunications Research Institute
BiophysicsMEMSBioengineeringCell BiologyBiophotonics