Certified Selective Automation of LLM Agent Evaluation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of error-rate guarantees in automated judges for LLM agent evaluation and the failure of conventional statistical certificates caused by trajectory correlations. To overcome these limitations, we propose task-level bootstrap certificates, introducing the first inference framework capable of handling correlated clustered data and breaking through the safety bottleneck of i.i.d. assumptions under high coverage. Furthermore, we train a 4B-parameter log-probability judge using supervised fine-tuning combined with rejection-weighted GRPO, enabling safe automated evaluation and label-free self-training under strictly controlled error-rate budgets. Experiments demonstrate that this framework certifies 30%–59% automation rates on tool-calling and web corpora while maintaining negligible pseudo-label contamination. Notably, it also exhibits strong zero-shot cross-domain transferability, offering a statistically rigorous approach to scalable agent evaluation.
📝 Abstract
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.
Problem

Research questions and friction points this paper is trying to address.

LLM agent evaluation
selective automation
certified error rate
correlated clusters
automatic judge
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-level bootstrap certificate
selective automation
reject-weighted GRPO
self-training filter
LLM agent evaluation
C
Chengguang Gan
Independent Researcher
Y
Yunhao Liang
University of Chinese Academy of Sciences
Qinghao Zhang
Qinghao Zhang
Department of Electrical Engineering, Tsinghua University
ReliabilityPower electronicsjunction temperaturecondition monitoringlifetime prediction
S
Shiwen Ni
Shenzhen University of Advanced Technology