🤖 AI Summary
This study addresses the challenge in large language model preference alignment where base models already possess target capabilities but fail to express them reliably. To this end, this work proposes a direct hidden state alignment framework that parses preference representations via Residual Competition Maps (RCMs) and employs Causal Activation State Transition (CAST) to intervene in local activations during inference while keeping the base model frozen. By shifting the adaptation space from model weights to hidden states, the approach enables lightweight control without retraining. Remarkably, the proposed method achieves performance comparable to Direct Preference Optimization (DPO) using only 256 to 16,384 parameters. Furthermore, it supports pluggable, multi-domain preference control and functions complementarily with DPO, offering a highly parameter-efficient alternative for aligning large language models.
📝 Abstract
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.