Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in large language model preference alignment where base models already possess target capabilities but fail to express them reliably. To this end, this work proposes a direct hidden state alignment framework that parses preference representations via Residual Competition Maps (RCMs) and employs Causal Activation State Transition (CAST) to intervene in local activations during inference while keeping the base model frozen. By shifting the adaptation space from model weights to hidden states, the approach enables lightweight control without retraining. Remarkably, the proposed method achieves performance comparable to Direct Preference Optimization (DPO) using only 256 to 16,384 parameters. Furthermore, it supports pluggable, multi-domain preference control and functions complementarily with DPO, offering a highly parameter-efficient alternative for aligning large language models.
📝 Abstract
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
Problem

Research questions and friction points this paper is trying to address.

preference alignment
hidden-state representation
behavioral expression
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Competition Maps
Direct Hidden-State Alignment
Causal Activation State Transition
Preference Alignment
Hidden-State Intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fansheng Zhang
Chengdu University
S
Shengran Guo
North Carolina State University
Z
Zexiao Wang
Fudan University
L
Liang Yuan
Australian Catholic University
Jiyuan Chen
Jiyuan Chen
The Hong Kong Polytechnic University
Spatial-Temporal ModelingGraph Neural NetworkComputer Vision
R
Ruikun Luo
University of Macau