Control OSWorld: An AI Control Environment for GUI Computer Use Agents

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the research gap in AI control for GUI agents and the risks of misaligned malicious behaviors by proposing the first control evaluation benchmark tailored to such agents. Building upon OSWorld, we construct an evaluation environment incorporating both benign and malicious side tasks, and develop a full-trajectory monitoring model designed to detect and intercept malicious GUI operations. Experimental results demonstrate that full-trajectory monitoring significantly outperforms single-step monitoring, achieving a 97% recall rate with a 3% false positive rate. Furthermore, our analysis reveals that visible text is critical for identifying malicious behaviors, whereas screenshot information offers limited utility.
📝 Abstract
AI agents that operate a computer through its graphical user interface (GUI) are being widely deployed. AI control studies how to prevent an AI system from causing harm even if it is misaligned and actively trying to do so. Most control research to date has focused on coding agents, leaving computer use largely unexplored. We introduce Control OSWorld, a control evaluation that pairs 318 tasks from OSWorld with 81 harmful side tasks (e.g., exfiltrating a private file) that an agent must complete without being caught. We use Control OSWorld to create and study control monitors that detect malicious agents interacting with a GUI. We find that a weaker monitor can reliably distinguish honest from malicious trajectories produced by a more capable agent when it sees the full trajectory (97% recall at 3% false positive rate). However, when the monitor has to score each step before it is executed, recall drops at low false positive rates because each action is judged with less context. Monitor performance also depends on what the monitor can see: access to the agent's visible text is critical, while screenshots provide little additional value.
Problem

Research questions and friction points this paper is trying to address.

AI control
GUI agents
computer use
malicious agent detection
safety evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Control
GUI Agents
Control OSWorld
Malicious Agent Detection
Control Monitors
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.