CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language model agents to indirect prompt injection attacks by proposing a fine-tuning-free, inference-time defense mechanism. The core innovation lies in introducing an "always-on" residual stream directional editing technique: a harmful behavior direction vector is fitted using paired data and, after validation through causal and capability gating, directly subtracted from the residual stream during the prefill phase to suppress malicious instruction execution. Experiments demonstrate that this method reduces attack success rates below 0.17 across five open-source models while preserving over 93% benign utility, and it effectively withstands adaptive attacks.
📝 Abstract
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.
Problem

Research questions and friction points this paper is trying to address.

Indirect prompt injection
LLM agent security
Instructional takeover
Parameter manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Activation Steering
Residual Stream
Inference-time Defense
LLM Agent Security
🔎 Similar Papers