π€ AI Summary
This study addresses the limitation that offline evaluation cannot determine whether modifications to large language models are suitable for direct deployment, while also failing to distinguish between trial-level and unit-level surrogacy. To this end, this work formally establishes the logical independence between trial-level and Prentice surrogacy for the first time. It derives closed-form decision boundaries that correct for noise under a bivariate normal model, incorporates variance-subtraction covariance recovery techniques, and analyzes risks associated with judge bias and distributional drift. Empirical validation on the Upworthy dataset demonstrates a weak correlation between offline evaluations and online outcomes, confirming that offline metrics alone are insufficient to replace A/B testing in practice.
π Abstract
Offline evals increasingly guide changes to LLM-powered products, but showing that an eval agrees with human judgment does not tell a team which changes it can ship without an A/B test. That decision depends on whether eval treatment effects predict online treatment effects across a class of interventions. We call this requirement trial-level surrogacy and show that it is logically independent of unit-level (Prentice) surrogacy. In historical launch logs, sampling noise in both effect estimates attenuates the observed relationship. Under a bivariate normal working model, subtracting the reported within-intervention variances recovers the between-intervention covariance, and the corrected moments give a closed-form ship/test/kill rule that bounds the probability of shipping a change with a non-positive online effect. LLM judges add a separate distortion. When the judge is equally accurate on both arms, eval effects are only rescaled, but arm-dependent accuracy can reverse their sign. Once deployed, the eval-to-outcome relationship can no longer be observed in the ship and kill regions, so drift there becomes undetectable without randomized audit experiments. On 183 interventions from the Upworthy Research Archive, the pooled eval-to-outcome relationship is weak despite precise measurement on both sides, and the illustrative gate sends every intervention to an A/B test.