Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
This study addresses the tendency of language models to erroneously defer to authoritative sources, a phenomenon fundamentally distinct from conventional user sycophancy. By employing causal intervention and activation direction fitting, this work disentangles the model’s response mechanisms toward verified sources and user inputs. It provides the first demonstration that source deference and user agreement are behaviorally non-interchangeable, proposing an independent evaluation framework accordingly. A key contribution is the identification of an “authority direction” representation that transfers across datasets such as Trivia and PIQA. Ablating this direction reduces erroneous compliance rates by 65–80 percentage points without compromising performance on benchmarks like MMLU-Pro.