🤖 AI Summary
In online A/B testing, the relationship between proxy metrics and long-term objectives often breaks down due to user heterogeneity, and relying solely on global correlations can lead to erroneous decisions. This work proposes PROXIMA, a novel framework that introduces a decision-consistency-oriented diagnostic approach for evaluating proxy metrics through three dimensions: normalized effect correlation, directional accuracy, and subgroup vulnerability rate. Integrating causal inference, subgroup analysis, and sensitivity testing, PROXIMA is validated across 80 simulated experiments on the Criteo and KuaiRec datasets. Results show an average decision accuracy of 98.4%; while the subgroup vulnerability rate is markedly higher in recommendation scenarios (68%) than in advertising (13%), directional accuracy exceeds 96% in both, effectively identifying subpopulations where proxy metrics fail.
📝 Abstract
Online A/B testing at scale relies on proxy metrics -- short-term, easily-measured signals used in place of slow-moving long-term outcomes. When the proxy-outcome relationship is heterogeneous across user segments, aggregate correlation can mask directional failures akin to Simpson's Paradox, leading to costly ship/no-ship errors. We introduce PROXIMA (Proxy Metric Validation Framework for Online Experiments), a lightweight diagnostic framework that scores proxy reliability through a composite of three complementary dimensions: normalised effect correlation, directional accuracy, and segment-level fragility rate. Unlike surrogate-index approaches that predict long-term treatment effects, PROXIMA directly audits whether a candidate proxy leads to correct launch decisions and flags the user segments where it fails. We validate PROXIMA on two public datasets -- the Criteo Uplift corpus (14M observations, advertising) and KuaiRec (7K users, video recommendation) -- using 80 simulated A/B tests. Early engagement metrics achieve a composite reliability of 0.80 on Criteo and 0.62 on KuaiRec, yielding 98.4% average decision agreement with an oracle policy. Fragility analysis reveals that recommendation domains exhibit substantially higher segment-level heterogeneity (68% fragility) than advertising (13%), yet directional accuracy remains above 96% in both cases. A sensitivity analysis over the weight space confirms that no single component suffices and that the composite provides substantially better discrimination between reliable and unreliable proxies than correlation alone. Code and reproduction scripts are available at: https://github.com/Avinash-Amudala/PROXIMA