🤖 AI Summary
This study addresses the version-dependence bias inherent in fixed LLM judges evaluating agent iterations, which risks misjudgment and violates error invariance under varying task conditions. Leveraging the SWE-bench and tau-bench benchmarks alongside statistical multiple-testing corrections and bootstrap confidence intervals, we reveal that stronger model capabilities paradoxically increase false positive rates and cause cross-domain calibration failures. We demonstrate that paired auditing outperforms standalone judging or legacy-version calibration. Empirically, 32 of 60 evaluation units exhibit detectable discrepancies, while legacy-version calibration inflates errors to 19.5 points, underscoring the necessity of human review for reliable agent assessment.
📝 Abstract
Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned agent-task cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare improvements that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge confidently reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and reference outcome, while a task-solvability prediction from AgentRewardBench reverses sign in SWE-bench. Transporting old-version calibration raises mean absolute comparison error on SWE-bench from 3.8 to 19.5 percentage points; 24.6% of ratio-bootstrap draws are undefined near the correction boundary. A tuned paired audit narrows a classical interval by only about 5% at 80 labeled tasks. A randomized three-arm test does not support the predicted increase in false acceptance from showing the agent's final report (all three Holm-adjusted p-values = 1.0). These results favor explicit reference standards and paired audits of current outputs over judge-only release decisions or transported old-version calibration.