🤖 AI Summary
This study addresses the limitations of conventional activation steering methods, which neglect local geometric structures and lack input adaptivity, by proposing a kernelized activation steering framework. This framework maps intervention operations into a reproducing kernel Hilbert space, implicitly constructing activation-dependent nonlinear steering fields via kernel evaluation, thereby overcoming the linear vector constraints inherent in existing approaches. As a unified paradigm, it enables position-aware adaptive model control without requiring additional training. Evaluated on large language model jailbreak defense and image style transfer tasks, the proposed method matches or surpasses current baselines, significantly enhancing both model expressiveness and control flexibility.
📝 Abstract
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local geometry of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space. KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, while richer kernels enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS outperforms or is on par with the existing methods.