The Noise Is the Signal: Correlated Sampling Error Is Rank-Informative for Proxy Metric Selection

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of surrogate metric selection in A/B testing, where high noise and limited sample sizes compromise the reliability of north star metrics. Departing from conventional approaches that treat shared sampling error between surrogate and north star metrics as contamination requiring correction, this work proposes a paradigm-shifting "noise-as-signal" framework. We demonstrate that such shared errors encode critical ranking information. Through disjoint subset estimation, Spearman rank correlation analysis, and validation across 262 real-world experiments, we find that shared sampling error aligns strongly with true metric rankings (correlation coefficient of 0.65), and removing it significantly degrades surrogate ranking accuracy. This research establishes the positive value of error correlation for model selection, offering online platforms a cost-effective evaluation strategy for surrogate metric assessment.
📝 Abstract
North-star metrics such as customer lifetime value are often too slow and noisy to decide a short A/B test. Teams therefore rely on a proxy metric, commonly chosen by how closely its effects tracked the north star's across past experiments. Validating that choice, or any method for making it, is hard: the only benchmark is the noisy north star, and the number of available past experiments is limited. In addition, proxy and north-star effects are estimated on the same customers, so their sampling errors are correlated. Recent work at major experimentation platforms removes this shared error as contamination, improving estimates of the true-effect covariance. Choosing a proxy, however, is a ranking problem, and a better estimate need not give a better ranking. We measure agreement free of shared error by estimating the two effects on disjoint random halves of each experiment's customers. In an archive of 262 experiments and 69 candidate proxies, the shared error ranks the candidates in a similar order to this agreement (Spearman correlation 0.65): it carries information about proxy quality. The more of it a correction removes, the worse the ranking because removal discards part of the signal but leaves the main sources of ranking noise, the noisy north star and the limited number of experiments, untouched. Archive-calibrated simulations, in which the correct ranking is known, confirm this even when every correction receives the true sampling covariance. Held-out real experiments, evaluated on disjoint customer halves so that shared error cannot bias the comparison, closely reproduce the predicted ordering (Spearman correlation 0.93). Correction can still pay off with more experiments, but the number needed rises steeply with the north star's noise. We map this crossover and give platform teams three inexpensive checks for deciding from their own archive whether and how strongly to correct.
Problem

Research questions and friction points this paper is trying to address.

proxy metric selection
north-star metrics
A/B testing
correlated sampling error
ranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proxy Metric Selection
Correlated Sampling Error
A/B Testing
North-star Metrics
Ranking Problem
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sandro Provenzano
Zalando SE