Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of repeated evaluation under hard budgets, where benchmark scores are accurate yet uncertainty certification remains difficult. It characterizes the optimal uncertainty width under fixed grids and hard budgets, proposing a task-coverage design to balance precision against cost. By establishing tight bounds for adaptive strategies and introducing joint mean-divergence intervals, the approach transforms task coverage into an explicit design variable. Furthermore, it integrates randomized subset designs, divergence certificates, and finite-budget inference techniques to achieve efficient estimation. Evaluated on LiveCodeBench, the proposed method reduces mean squared error by 87% and narrows confidence interval widths by 30.6%.
📝 Abstract
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < α\le 1/12$, the optimal expected width on the worst pure cohort is $Θ_{α,L}([M(t+1)]^{-1/2})$ when every task is observed and $Θ_{α,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.
Problem

Research questions and friction points this paper is trying to address.

repeated evaluation
hard budget
uncertainty quantification
benchmark evaluation
task coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hard-budget repeated evaluation
Disagreement certificates
Task coverage
Joint mean/disagreement interval
Sharp asymptotic bounds
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yezhou Cheng
Independent
R
Runjia Du
Independent
Z
Zeming Liu
Independent
Q
Qibai Chen
Independent
H
Hang Lyu
Independent
Y
Yilan Wei
Northwestern University
Yankai Zeng
Yankai Zeng
University of Texas at Dallas
B
Bojun Lin
Pinterest, Inc.