🤖 AI Summary
This study addresses the challenge in large language model (LLM) agents of distinguishing instructions from data and localizing intervention points under indirect prompt injection. By employing counterfactual role probes, component-level activation patching, and trajectory-independent interventions via AgentDojo, this work leverages causal tracing to reveal a fundamental divergence between the readability of role signals and the intervenability of agent behavior. It is the first to explicitly distinguish readable role signals from effective behavioral interventions, demonstrating that cross-channel transfer of intervention directions is inherently difficult. Furthermore, the study systematically analyzes how network depth and positional factors influence attack success. Results indicate that single-point editing fails in deeper layers, whereas wide-span and repeated editing strategies significantly reduce attack success rates, offering actionable insights for securing LLM agents against indirect prompt injection.
📝 Abstract
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.