Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical security vulnerability in AI coding agents wherein inconsistencies between user-approved and actually executed actions can be maliciously exploited. To systematically expose this threat, the work introduces a novel taxonomy termed "approval laundering," identifying six distinct failure modes that bypass user approval. Focusing on credential-binding integrity, the authors propose a defense mechanism based on key-capability tokens. The effectiveness of this approach is validated through PreToolUse hooks and statistical significance testing. Results demonstrate that the proposed token mechanism successfully mitigates delegation-based and temporal laundering attacks, establishing a new paradigm for AI agent security. Nevertheless, limitations persist regarding parameter-level and scope-level protections, indicating avenues for future research.
📝 Abstract
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval authorizes. We show this assumption fails systematically and reproducibly. We introduce Approval Laundering, a taxonomy of six failure modes by which a harness's enforcement mechanism silently substitutes A' for A after approval: Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering. Unlike prior work that evaluates risk classifiers against static corpora or infers implicit authorization boundaries, we study credential-binding integrity: given an already-approved action, does the harness dispatch exactly that action? Instrumenting Claude Code's pre-execution mediation point (PreToolUse), we conduct a controlled, headless, repeated-measures study of all six classes (N=19-20 runs each), reporting a Bound-Gap Rate (BGR) with Wilson confidence intervals and inter-rater agreement (kappa=1.0). We prototype Approval Token, a keyed capability Hk(principal, agent_id, session_id, tool, arguments, scope, expiry) issued by a mediator that never returns the key to the agent, evaluated via paired before/after replay of 118 runs (McNemar's exact test). The token fully eliminates Delegation laundering and, for our seeded session-identity-mismatch construction, Temporal laundering (p<10^-5), but by design leaves Scope laundering unaffected and shows no significant reduction in Argument laundering (p=1): an honest negative result, since these two classes leave every recorded dispatch field unchanged, diverging one process level below what a field-only verifier can observe. We discuss implications for defenses that bind only at the tool-invocation boundary.
Problem

Research questions and friction points this paper is trying to address.

Approval Laundering
AI Coding Agents
Credential-binding Integrity
Execution Binding Failure
Security Boundary
Innovation

Methods, ideas, or system contributions that make the work stand out.

Approval Laundering
Credential-Binding Integrity
Approval Token
AI Coding-Agent Harnesses
Bound-Gap Rate
🔎 Similar Papers
No similar papers found.