Where Steering Signals Come From: Activation Source Selection in Activation Steering

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic analysis regarding the upstream sources of steering signals in activation intervention research, which has limited intervention efficacy. By fixing downstream intervention conditions and systematically manipulating source context and activation reading strategies, the work identifies the "execution boundary state" as a critical source of effective steering signals. To enhance signal purity and stability, the authors propose a tail-truncation method that disentangles prompt and continuation semantics. Experiments across three instruction-tuned models and four steering tasks demonstrate that judicious selection of source activations substantially improves intervention performance, with execution boundary states consistently outperforming contexts containing only target behaviors.
📝 Abstract
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.
Problem

Research questions and friction points this paper is trying to address.

activation steering
source selection
steering signals
execution-boundary states
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation steering
activation source selection
execution-boundary states
tail subtraction
representation alignment