🤖 AI Summary
Traditional activation steering requires continuous intervention throughout generation, incurring substantial computational overhead and often degrading the model's general capabilities. This work proposes prefix steering, which applies short-horizon intervention exclusively during the initial generation phase to control model behavior. Theoretically, under a fixed-state attention assumption, we establish that a single-token intervention suffices to alter the entire generation trajectory, challenging the prevailing token-by-token steering paradigm. Empirically, we validate this approach by applying existing steering directions and operators over short spans. Experimental results across multiple models and tasks demonstrate that such brief interventions preserve most of the intended control while better maintaining general capabilities, outperforming several full-sequence steering strategies.
📝 Abstract
Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under fixed-state attention assumptions, we establish sufficient conditions for single- and multi-token steering to match prompt-induced attention-head outputs, and characterize how changes in input representations affect this match and its approximation error. This attention-level connection leads us to ask whether, at the behavioral level, steering can also guide subsequent generation through a brief initial intervention. We study Prefix Steering, which applies existing steering directions and operators over a short span starting at the final prompt token, with no further direct intervention afterward. We examine how intervention duration and strength jointly shape the control-capability trade-off. Across four models and five tasks, intervention over a short span, even a single token, often retains much of full steering's behavioral control while better preserving general capabilities, offering a trade-off competitive with, and in some settings better than, prompting and alternative steering-strength policies. Prefix Steering also remains effective on final-answer formatting tasks after reasoning, suggesting that a brief initial intervention can influence behavior expressed well after steering ends. These findings challenge the common practice of steering every generated token and motivate a more dynamical view of activation steering, in which a brief intervention can alter the trajectory of subsequent generation without continued intervention.