HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent behavioral errors in aligned language models, such as refusing benign requests and erroneously invoking tools, noting that existing methods overlook behavior-related information already encoded within model internals. To this end, it proposes HeadEdit, a framework demonstrating for the first time that target behaviors can be linearly decoded from hidden states. By extracting behavior-relevant low-rank subspaces and leveraging a frozen unembedding matrix to generate vocabulary-level corrections, the method achieves implicit adaptive guidance without parameter updates or manual specification of target tokens. Extensive evaluations across nine configurations spanning three task categories and multiple model families show comprehensive behavioral improvements with minimal inference overhead and no degradation in general capabilities. Furthermore, the learned subspaces are transferable after fine-tuning, yielding additional performance gains.
📝 Abstract
Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.
Problem

Research questions and friction points this paper is trying to address.

behavioral errors
language model alignment
unembedding matrix
hidden state representation
behavior calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient-free calibration
unembedding matrix
low-rank behavioral subspace
adaptive steering
representation decoding
🔎 Similar Papers
No similar papers found.