Speculative Evaluation of Stochastic LLMs

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high evaluation costs of large language models (LLMs) and the limitation that uniform sampling overlooks inter-task variance discrepancies. To this end, it proposes a stratified Bayesian Neyman allocation strategy designed to minimize the variance of benchmark mean estimation under a fixed budget. Methodologically, the approach jointly optimizes pilot sample sizes and allocation weights while introducing an asynchronous speculative execution mechanism to alleviate synchronization bottlenecks. Experimental results demonstrate that the proposed strategy reduces estimation variance by 12.8%–33.6% compared to uniform sampling, significantly outperforming multiple baseline methods. Ultimately, this work establishes a new paradigm for efficient LLM evaluation.
📝 Abstract
Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation with a Hierarchical Bayesian Neyman (HBN) policy with pilot size and stage weight jointly chosen ex ante. It runs a short uniform pilot, pools per-task success counts with a hierarchical Bayesian model, and uses posterior expectations of task-level sampling variances for exact positive-integer Neyman allocation. To mitigate the pilot synchronization barrier, HBN-async speculatively executes continuations from partial pilot feedback and retains those selected by the final allocation. Across six checkpoints and 18 benchmark groups, we evaluate 107 nondegenerate benchmark-checkpoint profiles. For rollout budgets of 8-64 per task, Speculative Evaluation reduces variance relative to Uniform by 12.8%-33.6% on average across profiles, outperforming hindsight-tuned empirical and independent Bayesian baselines. Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits.
Problem

Research questions and friction points this paper is trying to address.

Stochastic LLMs
Evaluation variance
Rollout budget
Benchmark evaluation
Neyman allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Evaluation
Hierarchical Bayesian Neyman Allocation
Stochastic LLMs
Variance Reduction
Asynchronous Pilot
Q
Qianli Shen
Alibaba Group
X
Xiang Li
National University of Singapore
R
Ruomeng Ding
University of North Carolina at Chapel Hill
Y
Yanxi Chen
Alibaba Group
Daoyuan Chen
Daoyuan Chen
Alibaba Group
Efficient Machine LearningHuman-Centric MLLarge Language ModelsMultimodality
Yaliang Li
Yaliang Li
Alibaba Group
Machine Learning