APEX: Active Protection at Execution Boundaries for LLM Agents

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lag in indirect prompt injection defenses for LLM agents caused by rapid tool proliferation. We propose a proactive defense framework grounded in execution boundaries, shifting the defensive focus from attack pattern recognition to stable execution constraints. Specifically, the method employs precompiled authorization contracts to validate output legitimacy and endorse information, while enforcing runtime restrictions through evidence gating and deception exposure mechanisms. This approach uniformly protects diverse capability units without requiring task-specific policies or taint tracking. Experimental evaluations demonstrate that the proposed framework achieves zero attack success rates on five out of six benchmarks. Furthermore, it maintains robustness against adaptive attacks and exhibits strong cross-model generalization capabilities.
📝 Abstract
Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emph{execution boundary}, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emph{evidence-gated prevention} admits an effect only when the contract justifies it, while \emph{deception-based exposure} makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0\% attack success on five of six benchmarks and 0.56\% on the sixth, holds 0\% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at https://github.com/ZhengXR930/APEX_official/tree/official.
Problem

Research questions and friction points this paper is trying to address.

Indirect prompt injection
LLM agents
execution boundary
adversarial instructions
agent security
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Execution Boundary
Authorization Contract
Evidence-gated Prevention
Deception-based Exposure
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xinran Zheng
Tsinghua University
X
Xin Fan Guo
Imperial College London
Z
Zhiqiang Hao
Nanjing University
F
Fan Yang
The Chinese University of Hong Kong
X
Xingzhi Qian
University College London
Jiawei Du
Jiawei Du
National Taiwan University; ex-Intern @ Samsung Research
Speech processingNeural codingGenerative AIAI security
J
Jinfeng Xu
University of British Columbia
Zheng Xing
Zheng Xing
Master of Science, Imperial College London
LLM
Shuo Yang
Shuo Yang
The University of Hong Kong
Xingjun Wang
Xingjun Wang
Professor@Tsinghua University