Beyond Task Completion: Training Capable and Safe Computer-Use Agents

📅 2026-08-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决计算机使用代理的安全性问题,提出SCOPE方法联合训练任务执行能力和安全决策,并通过SCOPE-Gen生成对齐训练数据。
📝 Abstract
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete ordinary benign tasks, avoid environmental hazards and continue when a safe completion path remains, and refuse when the goal is harmful or no safe path exists. To learn this conditional policy, we develop Safety and Capability Optimization for Policy Execution (SCOPE), which jointly post-trains a CUA for task-execution capability and safety-aware decision making. To provide aligned training data for this joint objective, we further introduce SCOPE-Gen, an automated pipeline that synthesizes verifiable capability tasks and converts them into paired environment-risk variants while preserving their original goals. Using the resulting tasks, we construct SATraj-OS, a trajectory dataset comprising capability demonstrations, safe continuations, and explicit refusals. SCOPE first learns from all three trajectory types through supervised fine-tuning and then further improves task completion through online reinforcement learning. Starting from Qwen3.5-9B, SCOPE-RL achieves a 54.17% task success rate on OSWorld and a 64.30% attack-avoidance rate on OS-BLIND, yielding the best aggregate capability--safety score of 58.80% among the evaluated agents. Ablations reveal asymmetric but complementary roles for the two forms of safety supervision: refusal trajectories account for most of the attack-avoidance gain, whereas risk-handling trajectories preserve greater task utility at comparable attack-avoidance levels.
Problem

Research questions and friction points this paper is trying to address.

Computer-use agents
Safety behavior
Task completion
Risk conditioning
Safe execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

SCOPE
safety-aware decision making
SCOPE-Gen
SATraj-OS
reinforcement learning
🔎 Similar Papers
No similar papers found.
Z
Zeyu Kang
Shanghai Artificial Intelligence Laboratory
Z
Zhenyun Yin
Shanghai Artificial Intelligence Laboratory
Y
Yang Zhang
Shanghai Artificial Intelligence Laboratory
S
Shan He
Shanghai Artificial Intelligence Laboratory
S
Shanzhe Lei
Shanghai Artificial Intelligence Laboratory
Y
Yanjiu Zhong
Hefei University of Technology
X
Xinquan Chen
Shanghai Artificial Intelligence Laboratory
Xuhong Wang
Xuhong Wang
Shanghai Artificial Intelligence Laboratory
LLMKnowledge SystemAI Simulation