🤖 AI Summary
This work addresses the lack of effective evaluation against adaptive, defense-aware attacks in existing defenses against indirect prompt injection for large language model agents. It is the first to unify out-of-band defense mechanisms within classical security frameworks—specifically Biba integrity, reference monitors, and least privilege—and systematically analyzes their capabilities alongside information flow label design. The study further introduces the first adaptive attack evaluation protocol tailored to such defenses. Using Qwen2.5-7B deployed on a single H200 GPU, the authors reproduce and extend the Progent attack, conducting black-box evaluations on the AgentDojo benchmark. Results demonstrate that Progent’s defense reduces average attack success rates from 25.8% to 4.2%, with handcrafted adaptive attacks achieving only 2.6%, thereby confirming the substantial robustness of out-of-band defenses against dynamic adversarial strategies.
📝 Abstract
Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made in-band defenses look strong until adaptive, defense-aware attacks broke twelve of them at over 90% success; we specify the threat model and protocol an adaptive evaluation requires. We then run that protocol as an independent reproduction and extension of Progent's own adaptive-attack analysis, on AgentDojo, with an open-weight agent (Qwen2.5-7B) self-hosted on a single H200, a setting its authors did not test. Averaged over three runs, the defense held: Progent cut mean attack success roughly sixfold (25.8% to 4.2%), and a hand-crafted adaptive attack did not raise it (2.6%). This is one small-scale data point on a weak model with a single black-box attack template; a stronger optimized (white-box GCG) attack remains open. The result is consistent with, but does not establish, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.