Steering by Influence: Curvature Aware Data Weighting for Activation Steering

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of existing activation steering methods to noise interference and their limitation in encoding token-level rather than topic-level semantics. We propose a precise topic-level activation steering framework that leverages influence functions to identify the most topic-representative samples. Furthermore, we incorporate loss landscape curvature information for perceptual weighting to transcend superficial similarity constraints, and integrate optimal transport techniques to achieve precise guidance within the activation space. Experimental results demonstrate that our approach significantly outperforms baselines across toxicity suppression, concept induction, and truthfulness tasks. The proposed method effectively enhances steering performance while maintaining stable model perplexity and MMLU accuracy.
📝 Abstract
Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than thematic concepts. In this work, we steer towards examples that most express a concept thematically, rather than towards an expectation over all. We identify these examples using influence functions, which estimate how much each data point contributes to a model's representation of a concept. Unlike simple model activation similarity, they incorporate the curvature of the model's loss landscape, allowing them to capture concept-relevant relationships beyond superficial token-level similarity. We then propose influence-weighted activation transport, which uses optimal transport to steer activations of non-concept text towards those of concept text, weighting concept examples by their influence scores. We evaluate on toxicity suppression (Jigsaw), object-based concept induction (OneSec) and truthfulness induction (TruthfulQA), outperforming existing activation-transport baselines. We track capability after steering using perplexity and MMLU accuracy, finding that our method improves steering while largely preserving model quality. We further show that influence functions capture concept-relevant information that activation-based methods miss with the two approaches ranking data points significantly differently. Together, these results demonstrate the value of curvature-aware influence information for activation steering.
Problem

Research questions and friction points this paper is trying to address.

activation steering
concept representation
influence functions
language models
inference-time control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Influence Functions
Activation Steering
Curvature Aware
Optimal Transport
Data Weighting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
James A. E. Dixon
Machine Learning Research Group, Department of Engineering Science, University of Oxford
S
Stephen J. Roberts
Machine Learning Research Group, Department of Engineering Science, University of Oxford
Francesco Quinzan
Francesco Quinzan
University of Oxford