VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the miscalibrated reliability estimation and catastrophic forgetting that arise when LLM judges adapt to new feedback due to their reliance on output-level confidence. To this end, we propose VERA, a framework that innovatively constructs a judgment-conditioned reliability axis by leveraging hidden-layer activations to distinguish correct from incorrect responses within groups, thereby enabling precise reliability estimation to guide periodic adaptive updates. Furthermore, VERA introduces reliability-based error correction and residual replay mechanisms to effectively mitigate forgetting. Experimental results demonstrate that 8B and 14B models equipped with VERA outperform the strongest baselines across four public benchmarks, achieving relative improvements of up to 23.01% and a 16.1% gain in focal-class recall.
📝 Abstract
Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose VERA, a VErdict-conditioned Reliability Axis that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using VERA as a control signal, we develop a VERA-guided periodic adaptation framework that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to 23.01%. The framework also improves focal-class recall by up to 16.1% relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.
Problem

Research questions and friction points this paper is trying to address.

LLM judges
judgment reliability
confidence estimation
adaptive learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verdict-Conditioned Reliability
Hidden Activations
Adaptive LLM Judges
Periodic Adaptation Framework
Reliability-Residual Replay
Q
Qiushui Xu
Penn State University
S
Syamil Mohd Razak
Amazon
Tao Yuan
Tao Yuan
University of California, Los Angeles
Computer VisionArtificial Intelligence
P
Piotr Habas
Amazon