🤖 AI Summary
This study addresses the challenge of evaluating Best-of-N (BoN) strategies under sample-only access, where intractable density ratios pose a fundamental bottleneck. To overcome this, we propose a sample-specific framework that leverages order statistics to reformulate density ratios as estimable score-ranking probabilities. Building upon this formulation, we construct a doubly robust BoN-DR estimator and establish valid asymptotic inference alongside a no-harm guarantee for selection rules. By integrating order statistics theory with off-policy evaluation, our method achieves accurate policy value estimation and optimal sampling budget selection on both synthetic data and the GSM8K benchmark. Ultimately, this work provides rigorous theoretical foundations for the safe evaluation of BoN strategies.
📝 Abstract
Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.