PROXIMA: A Reliability Scoring Framework for Proxy Metrics in Online Controlled Experiments

📅 2026-04-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In online A/B testing, the relationship between proxy metrics and long-term objectives often breaks down due to user heterogeneity, and relying solely on global correlations can lead to erroneous decisions. This work proposes PROXIMA, a novel framework that introduces a decision-consistency-oriented diagnostic approach for evaluating proxy metrics through three dimensions: normalized effect correlation, directional accuracy, and subgroup vulnerability rate. Integrating causal inference, subgroup analysis, and sensitivity testing, PROXIMA is validated across 80 simulated experiments on the Criteo and KuaiRec datasets. Results show an average decision accuracy of 98.4%; while the subgroup vulnerability rate is markedly higher in recommendation scenarios (68%) than in advertising (13%), directional accuracy exceeds 96% in both, effectively identifying subpopulations where proxy metrics fail.

Technology Category

Reasoning under Uncertainty: CausalityMachine Learning: Online Learning & BanditsSearch and Optimization: Algorithm Configuration

Application Category

User Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successResponsible Web: Human-perceived consequences of algorithmic deployment on the webSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Online A/B testing at scale relies on proxy metrics -- short-term, easily-measured signals used in place of slow-moving long-term outcomes. When the proxy-outcome relationship is heterogeneous across user segments, aggregate correlation can mask directional failures akin to Simpson's Paradox, leading to costly ship/no-ship errors. We introduce PROXIMA (Proxy Metric Validation Framework for Online Experiments), a lightweight diagnostic framework that scores proxy reliability through a composite of three complementary dimensions: normalised effect correlation, directional accuracy, and segment-level fragility rate. Unlike surrogate-index approaches that predict long-term treatment effects, PROXIMA directly audits whether a candidate proxy leads to correct launch decisions and flags the user segments where it fails. We validate PROXIMA on two public datasets -- the Criteo Uplift corpus (14M observations, advertising) and KuaiRec (7K users, video recommendation) -- using 80 simulated A/B tests. Early engagement metrics achieve a composite reliability of 0.80 on Criteo and 0.62 on KuaiRec, yielding 98.4% average decision agreement with an oracle policy. Fragility analysis reveals that recommendation domains exhibit substantially higher segment-level heterogeneity (68% fragility) than advertising (13%), yet directional accuracy remains above 96% in both cases. A sensitivity analysis over the weight space confirms that no single component suffices and that the composite provides substantially better discrimination between reliable and unreliable proxies than correlation alone. Code and reproduction scripts are available at: https://github.com/Avinash-Amudala/PROXIMA
Problem

Research questions and friction points this paper is trying to address.

proxy metrics
online controlled experiments
Simpson's Paradox
heterogeneous treatment effects
decision reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

proxy metrics
online controlled experiments
reliability scoring
heterogeneous treatment effects
Simpson's Paradox
🔎 Similar Papers
No similar papers found.
A
Avinash Amudala
Rochester Institute of Technology