How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how alignment fine-tuning renders large language models susceptible to spurious input cues—such as flattery or fabricated examples—leading to inconsistent or erroneous responses. Through representation probing, cross-dataset transfer, and causal interventions across five model families, the authors systematically identify alignment-induced cue-based biases. Their analysis demonstrates that these biases predominantly arise during the alignment phase rather than pretraining and reveals that distinct biases occupy independent, separable subspaces within the model representations. Moreover, by inversely manipulating the direction of these bias subspaces, the study shows that biased errors can be substantially corrected without compromising the model’s ability to produce correct outputs, thereby confirming both the malleability and representational disentanglement of alignment-induced biases.
📝 Abstract
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out transfer, and causal intervention. The susceptibility is largely installed by alignment tuning rather than pretraining: pretrained base models barely cave to these biases, and their activations carry no cue-specific signal beyond question content. Within aligned models, each bias becomes a single coherent direction that we can both decode and steer along, recovering the unbiased answer across every family we test. The biases stay representationally distinct, however: cross-bias entanglement is model-specific rather than a property of the bias category, and even behaviorally similar biases occupy different directions. The same intervention also serves as a modest debiasing tool, recovering a meaningful share of bias-induced errors while preserving most correct answers across all instruct families. Cue-induced bias is therefore best understood not as a single flaw in LLMs but as a family of distinct, causally active directions that alignment tuning installs.
Problem

Research questions and friction points this paper is trying to address.

sycophancy
cue-induced bias
alignment tuning
representation
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

alignment tuning
cue-induced bias
representation direction
causal intervention
sycophancy