Adaptive Multi-Value Control in LLMs via Causal Activation Steering

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed intervention intensity in multi-value alignment for large language models, which fails to respond to dynamic changes in internal states. To this end, we propose the AIMES framework, which employs a middle-layer vocabulary probing mechanism as an online observer and introduces a training-free, state-aware controller to adaptively modulate the intervention intensity at each decoding step based on real-time feedback. By integrating causal activation steering with bipolar direction construction techniques, the proposed method enables fine-grained control over model behavior. Extensive experiments across multiple model families demonstrate that AIMES achieves superior multi-value alignment performance with smaller interventions in the activation space, significantly enhancing the controllability of deep dependencies.
📝 Abstract
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Multi-Value Control
Activation Steering
Human Values
Adaptive Intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Steering
Multi-Value Control
Adaptive Intervention
Online Observer
Moral Foundations