🤖 AI Summary
This study addresses the unclear internal representational mechanisms underlying refusal behavior in activation steering of large language models. By decomposing refusal steering into component-level interventions through sparse subset analysis and residual stream dimension decoupling, this work identifies sparse attention and MLP subsets capable of reproducing the behavioral effect. It provides the first evidence that refusal is assembled by structured, identifiable sparse mechanisms rather than diffusely encoded representations, thereby establishing the privileged basis hypothesis. Experiments demonstrate that retaining only 28–48% of components or 50% of residual stream dimensions preserves 85–101% of the original steering efficacy, revealing a dual sparsity of the signal across both components and dimensions. The implementation code has been made publicly available.
📝 Abstract
Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in https://github.com/wang-research-lab/Refusal_Mechanisms.