When Can You Ship on Evals Alone? Trial-Level Surrogacy for Offline Evaluation of LLM Systems

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation that offline evaluation cannot determine whether modifications to large language models are suitable for direct deployment, while also failing to distinguish between trial-level and unit-level surrogacy. To this end, this work formally establishes the logical independence between trial-level and Prentice surrogacy for the first time. It derives closed-form decision boundaries that correct for noise under a bivariate normal model, incorporates variance-subtraction covariance recovery techniques, and analyzes risks associated with judge bias and distributional drift. Empirical validation on the Upworthy dataset demonstrates a weak correlation between offline evaluations and online outcomes, confirming that offline metrics alone are insufficient to replace A/B testing in practice.
πŸ“ Abstract
Offline evals increasingly guide changes to LLM-powered products, but showing that an eval agrees with human judgment does not tell a team which changes it can ship without an A/B test. That decision depends on whether eval treatment effects predict online treatment effects across a class of interventions. We call this requirement trial-level surrogacy and show that it is logically independent of unit-level (Prentice) surrogacy. In historical launch logs, sampling noise in both effect estimates attenuates the observed relationship. Under a bivariate normal working model, subtracting the reported within-intervention variances recovers the between-intervention covariance, and the corrected moments give a closed-form ship/test/kill rule that bounds the probability of shipping a change with a non-positive online effect. LLM judges add a separate distortion. When the judge is equally accurate on both arms, eval effects are only rescaled, but arm-dependent accuracy can reverse their sign. Once deployed, the eval-to-outcome relationship can no longer be observed in the ship and kill regions, so drift there becomes undetectable without randomized audit experiments. On 183 interventions from the Upworthy Research Archive, the pooled eval-to-outcome relationship is weak despite precise measurement on both sides, and the illustrative gate sends every intervention to an A/B test.
Problem

Research questions and friction points this paper is trying to address.

offline evaluation
trial-level surrogacy
LLM systems
A/B testing
treatment effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

trial-level surrogacy
offline evaluation
LLM systems
closed-form decision rule
evaluator bias
M
MΓ₯rten Schultzberg
Spotify, Stockholm, Sweden