Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates where large language models decide to violate system instructions under prompt injection attacks. By integrating causal activation patching, linear probing, and low-rank subspace decomposition, the authors localize decision-making mechanisms layer by layer. The findings reveal that although attack information is decodable in early layers, the ultimate decision authority concentrates within bottleneck layers in the model's latter stages. Furthermore, compliance mechanisms occupy a compact, cross-architecturally stable linear subspace. Patching these bottleneck layers reverses 77%–92% of non-compliant behaviors. Detection based on this locus outperforms early-layer classifiers while exhibiting robustness against obfuscation, confirming that the causal peak constitutes the optimal detection site for prompt injection defenses.
📝 Abstract
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the rules? Using layer-by-layer causal activation patching across five models (4B to 32B parameters), we find a clear dissociation: attack information is linearly decodable from the first layer, yet causal leverage over the model's behavior is negligible until a late-layer bottleneck in the final third of the network. Patching this bottleneck reverses compliance in 77--92\% of cases. We show that the compliance mechanism occupies a compact linear subspace (rank-8 in 4B and 14B models, scaling to rank-64 at 32B) and is architecturally stable across varying model families. Finally, we validate our mechanistic account by showing that this causal peak layer is also the representationally optimal site for detecting attacks, outperforming early-layer classifiers that degrade under surface-level obfuscation such as leetspeak substitution. This alignment between causal leverage and detection performance provides converging evidence that the late-layer bottleneck captures decision-relevant computation rather than merely reflecting an artifact of the intervention.
Problem

Research questions and friction points this paper is trying to address.

Prompt Injection
Large Language Models
Mechanistic Interpretability
Rule Compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt injection
mechanistic interpretability
causal activation patching
late-layer bottleneck
linear subspace
🔎 Similar Papers
No similar papers found.