Minimally Invasive Steering of Language Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses output distribution shift and generation quality degradation caused by unregularized optimization in pre-logit steering, proposing MISVO for test-time reward adaptation with frozen models. Theoretically, it employs Fisher information to quantify distributional sensitivity, derives an exact decomposition of the sequence-level KL divergence gradient, and proves that suffix terms are second-order infinitesimals. Methodologically, it designs a local KL geometric penalty, computes analytical gradients via matrix-vector products over the frozen language model head, and implements position-specific interventions. Experiments demonstrate that across preference and code tasks on 1B- to 14B-parameter models, MISVO achieves the highest average reward in six of seven settings while effectively preserving generation diversity and coherence.
📝 Abstract
Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.
Problem

Research questions and friction points this paper is trying to address.

pre-logit steering
language models
reward optimization
output distribution
generation quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Minimally Invasive Steering
Fisher Information Geometry
Pre-logit Steering
KL Divergence Decomposition
Test-time Adaptation
🔎 Similar Papers
2024-02-15International Conference on Machine LearningCitations: 14