Securing Computer-Use Agents Against Branch Steering Attacks

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of Computer-Use Agents (CUAs) to branch manipulation attacks during dynamic planning. To mitigate this, we propose COBRA, a novel security architecture that preemptively enforces capability constraints prior to branching decisions. By integrating an isolated planner, a quarantined large language model (LLM), and predefined capability boundaries, COBRA strictly restricts execution paths, effectively defending against indirect prompt injection without requiring explicit payload injection and thereby overcoming the limitations of conventional dual-LLM architectures. Furthermore, we introduce STEER-Bench as a dedicated evaluation benchmark. Experimental results demonstrate that our approach reduces the attack success rate to 0% while preserving 97% task utility, achieving a robust alignment between security and usability.
📝 Abstract
Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security guarantees - using an isolated Planner LLM (P-LLM) to fix execution paths before processing untrusted inputs via a Quarantined LLM (Q-LLM) - these guarantees break down in graphical environments. Because CUA interaction is inherently dynamic, plans cannot remain data-independent; they must branch based on anticipated runtime web content - covering all possible cases the agent may encounter. This exposes agents to branch steering attacks, where an adversary crafts untrusted data to coerce a CUA down a hazardous, pre-approved branch without injecting explicit instructions. We systematically study branch steering attacks and introduce STEER-Bench (101 tasks across 9 domains), showing high attack success against both standard (94.4%) and vanilla Dual-LLM (89.5%) CUAs. We then propose COBRA, an architecture that pairs trusted branching plans with ahead-of-time capability constraints, strictly bounding the parameters and destinations each branch may execute. On STEER-Bench, COBRA reduces attack success to 0% while retaining 97% benign utility.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computer-Use Agents
Branch Steering Attacks
Indirect Prompt Injection
Dual-LLM Architecture
COBRA
🔎 Similar Papers
No similar papers found.