🤖 AI Summary
This study addresses the vulnerability of LLM-as-a-Judge paradigms to judgment reversals caused by candidate ordering changes, where conventional re-evaluation strategies incur prohibitive computational overhead. To mitigate this, we propose a method that predicts positional bias directly from internal residual stream activations. By employing regularized linear probes with nested grouped cross-validation, our approach anticipates order-sensitive reversals prior to initial judgments without requiring re-evaluation. This work contributes a paradigm shift in detecting sequential bias from internal model states, surpassing baselines reliant on external features such as confidence scores. Empirical results demonstrate AUROC scores ranging from 0.621 to 0.850, alongside robust cross-dataset generalization capabilities.
📝 Abstract
The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whether residual stream activations recorded immediately before the initial verdict can predict a flip. We use nested grouped cross-validation to evaluate regularized linear probes on 534 JudgeBench pairs for three Qwen3 judges and Llama-3.1-8B. The linear probes achieve AUROCs of .621-.850 and outperform a combined baseline that uses verbalized confidence, verdict-label logits, response lengths, and the judge's initial choice by .062-.113 AUROC. Linear probes trained on JudgeBench and then frozen achieve AUROCs of .685-.853 on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration. These results show that pre-verdict activations support prediction of susceptibility to candidate order and outperform the non-activation predictors evaluated here.