🤖 AI Summary
This study addresses the lag in indirect prompt injection defenses for LLM agents caused by rapid tool proliferation. We propose a proactive defense framework grounded in execution boundaries, shifting the defensive focus from attack pattern recognition to stable execution constraints. Specifically, the method employs precompiled authorization contracts to validate output legitimacy and endorse information, while enforcing runtime restrictions through evidence gating and deception exposure mechanisms. This approach uniformly protects diverse capability units without requiring task-specific policies or taint tracking. Experimental evaluations demonstrate that the proposed framework achieves zero attack success rates on five out of six benchmarks. Furthermore, it maintains robustness against adaptive attacks and exhibits strong cross-model generalization capabilities.
📝 Abstract
Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emph{execution boundary}, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emph{evidence-gated prevention} admits an effect only when the contract justifies it, while \emph{deception-based exposure} makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0\% attack success on five of six benchmarks and 0.56\% on the sixth, holds 0\% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at https://github.com/ZhengXR930/APEX_official/tree/official.