🤖 AI Summary
This study addresses the limitations of existing large model intervention methods, which rely on linear and context-independent assumptions that induce information bottlenecks and hinder the disentanglement of complex behaviors. To overcome these constraints, this work discards linear geometric assumptions and proposes an adaptation mechanism based on attention projection matrices. By integrating low-rank adapters (LoRA) with a preference optimization framework, the approach learns nonlinear, context-dependent intervention strategies that achieve structured control while preserving trajectory dynamics. Notably, the proposed method supports zero-shot transfer to unseen concepts and out-of-distribution scenarios. Extensive evaluations demonstrate that it matches or surpasses strong baselines across text generation and agent benchmarks, thereby validating the intrinsic relationship between generative performance and agent behavior, as well as ensuring enhanced safety.
📝 Abstract
Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns nonlinear interventions with context-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept-geometry assumption. Framed as preference-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero-shot to unseen concepts and out-of-distribution contexts. We find that, despite using low-rank adapters, MetaSteer induces structured, context-dependent changes in hidden-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature. We evaluate MetaSteer on three controlled text-generation benchmarks and three agentic settings across multiple model families and scales. MetaSteer matches or outperforms strong task-specific steering baselines on most aggregate comparisons in the zero-shot regime. Across the evaluated settings, stronger text-generation steering is associated with stronger agentic steering performance. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering.