Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of LLM-as-a-judge systems to superficial feature interference, which compromises judgment accuracy and obscures whether internal representations encode genuine preferences. To investigate this, we employ linear probing on internal model activations to recover authentic preference signals masked by final verdicts. By integrating causal intervention with residualization techniques, we localize bias-inducing pathways and introduce a diagnostic metric based on the predictive power of superficial features to delineate the conditions under which signal recovery is applicable. Evaluated across multiple dataset benchmarks, our probing approach significantly outperforms direct judgments in accuracy and effectively enhances the quality of labels used for preference learning. This work establishes a novel paradigm for improving the reliability of LLM-based evaluation frameworks.
📝 Abstract
LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but not reflected in the output? We study this across 64 open-weight evaluators and 14 datasets, including causal interventions on 41 judges (editing activations mid-run to see whether the verdict changes). On LLMBar, built so the superficially better answer is the worse one, the verdicts of 50 judges agree with human labels only 0.456 of the time, even after averaging both answer orders. Yet a small probe on the same judges'activations, with no weight updates, reaches 0.846, and 0.686 once surface features such as length and position are residualized out (0.507 with shuffled labels). The gap holds across eight benchmarks and model families, but is not universal: a score of how well surface features alone predict the human label, computed before any probe is trained, predicts the size of the gain (Spearman rho = 0.90). On rubric tasks that score one answer at a time, leaving no surface cue to exploit, reading the internals gives no advantage. The interventions also show that editing activations mid-network already changes the verdict, before it can be read off directly, and locate the pathways carrying position and length bias. At the same label budget, the recovered signal lets a judge flag cases where it is likely wrong and yields better labels for preference learning. A wrong verdict, then, does not mean the judge lacks the information, and a simple diagnostic shows when it is worth recovering.
Problem

Research questions and friction points this paper is trying to address.

LLM judges
preference signals
internal representations
surface feature bias
evaluation reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM judges
activation probing
causal intervention
preference signals
surface bias
🔎 Similar Papers
No similar papers found.