Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the version-dependence bias inherent in fixed LLM judges evaluating agent iterations, which risks misjudgment and violates error invariance under varying task conditions. Leveraging the SWE-bench and tau-bench benchmarks alongside statistical multiple-testing corrections and bootstrap confidence intervals, we reveal that stronger model capabilities paradoxically increase false positive rates and cause cross-domain calibration failures. We demonstrate that paired auditing outperforms standalone judging or legacy-version calibration. Empirically, 32 of 60 evaluation units exhibit detectable discrepancies, while legacy-version calibration inflates errors to 19.5 points, underscoring the necessity of human review for reliable agent assessment.
📝 Abstract
Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned agent-task cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare improvements that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge confidently reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and reference outcome, while a task-solvability prediction from AgentRewardBench reverses sign in SWE-bench. Transporting old-version calibration raises mean absolute comparison error on SWE-bench from 3.8 to 19.5 percentage points; 24.6% of ratio-bootstrap draws are undefined near the correction boundary. A tuned paired audit narrows a classical interval by only about 5% at 80 labeled tasks. A randomized three-arm test does not support the predicted increase in false acceptance from showing the agent's final report (all three Holm-adjusted p-values = 1.0). These results favor explicit reference standards and paired audits of current outputs over judge-only release decisions or transported old-version calibration.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
Agent Evaluation
Version-Dependent Error
Evaluation Reliability
SWE-bench
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-Judge Error
Version-Dependent Evaluation
Agent Evaluation
False Acceptance
Paired Audits