Holdout Best-of-N: Unbiased Evaluation and Its Cost

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reward overestimation bias arising from reusing selection scores in Best-of-N evaluation by investigating unbiased estimation methods based on a fixed scoring matrix. Theoretically, it establishes the necessary and sufficient conditions for unbiasedness when the number of fresh scores J is strictly less than the total number of scores K. Methodologically, this work proposes a Holdout strategy, derives its optimal risk bound and asymptotic constant, and integrates statistical estimation, minimax risk analysis, and cyclic averaging algorithms. Computationally, the proposed approach achieves O(MK log M) complexity and attains a risk order of σ²/√K under specific conditions, yielding a convergence rate that significantly outperforms existing biased methods.
📝 Abstract
Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$. At $J=K-1$, the selector deepens as $K$ grows. For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $σ^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $σ^2/K$. For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant $1/(π\sqrt2)$, which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in $O(MK\log M)$ operations. At fixed selector depth, cyclic evaluation of bounded scores has $O(K^{-1})$ risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all-$K$ policy.
Problem

Research questions and friction points this paper is trying to address.

Best-of-N
unbiased evaluation
holdout estimation
selection bias
minimax risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

Holdout Best-of-N
Unbiased Evaluation
Minimax Risk
Cyclic Average
Selector Depth
🔎 Similar Papers
2024-07-08arXiv.orgCitations: 1