construct proxy variables

Designs and constructs measurable substitute variables that represent unobserved, hard-to-measure, or missing constructs by selecting candidate indicators, engineering or aggregating features, and specifying transformation or combination rules. Evaluates and documents proxy validity, reliability, bias, and sensitivity, and analyzes the effects of proxy choices on downstream models, inference, and decisions.

constructproxyvariables

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the bias and invalid inference that arise when data-driven proxies—such as embeddings from fine-tuned machine learning models—fail to capture the true dimensions of product differentiation in high-dimensional unstructured data like text or images, leading to inaccurate counterfactual demand predictions. To resolve this, the paper proposes a bias-correction method that integrates econometric counterfactual frameworks with machine learning embeddings, accommodating data-dependent proxy variables while enabling analytical standard error computation. The approach is applicable to both unstructured data and traditional product attributes. A lightweight statistical correction algorithm, accompanied by diagnostic tools, facilitates unbiased and efficient counterfactual inference at either the market or individual level. Simulations and empirical applications demonstrate that the method substantially improves the accuracy of substitution effect estimates and ensures valid statistical inference.

counterfactualsdemand estimationdifferentiated products

Active multiple testing with proxy p-values and e-values

Feb 08, 2025
ZX
Ziyu Xu
🏛️ Carnegie Mellon University

In resource-constrained multiple testing, exhaustive evaluation of all hypotheses and computation of exact test statistics (e.g., via experiments or precise calculations) is infeasible. Method: This paper proposes a surrogate-driven active testing framework that leverages auxiliary information—such as expert judgment, ML predictions, or historical data—to construct surrogate test statistics. It dynamically decides whether to invoke costly exact tests; otherwise, it substitutes the surrogate values directly. Contribution/Results: The framework is the first to enable compatible p-value and e-value constructions under arbitrary dependence structures—without requiring independence between surrogates and true statistics—while provably controlling the false discovery rate (FDR). By unifying active learning, multiple testing theory, and e-value theory, it achieves both theoretical rigor and practical utility. Empirical evaluation on scCRISPR causal effect analysis demonstrates a 32% increase in discoveries and a 68% reduction in computational cost compared to exhaustive testing, under identical FDR constraints.

Develops active multiple testing using proxy p-values and e-valuesEnsures false discovery rate control while maintaining high powerSelects hypotheses to query based on proxy statistics to save resources

Missing Value Knockoffs

Feb 26, 2022
DK
Deniz Koyuncu
🏛️ Rensselaer Polytechnic Institute

Existing variable selection methods struggle to control the false discovery rate (FDR) under missing data, while model-X knockoffs—though theoretically guaranteed to control FDR—cannot directly accommodate missing values. This work establishes, for the first time, the theoretical FDR controllability of knockoffs in the presence of missing data. We propose three novel paradigms: (i) posterior sampling-based imputation and knockoff reuse, (ii) knockoff generation restricted to observed variables only, and (iii) joint latent-variable imputation and knockoff construction. Our approaches integrate Bayesian posterior sampling, univariate imputation, and latent-variable modeling, and we rigorously prove that they satisfy FDR ≤ α under standard assumptions. Extensive experiments demonstrate precise FDR control across diverse missingness mechanisms (MCAR, MAR, MNAR), variable correlation structures, and sample sizes, while achieving high statistical power and substantially reduced computational complexity compared to existing alternatives.

Extending model-x knockoffs framework to handle missing dataPreserving false selection guarantees with imputation methodsReducing computational complexity for latent variable models

Potential weights and implicit causal designs in linear regression

Jul 30, 2024
JC
Jiafeng Chen
🏛️ Stanford University

When linear regression is used to estimate treatment effects in quasi-experiments, its causal interpretation rests on implicit assumptions—specifically, under what conditions does the regression coefficient represent a comparable contrast of individual potential outcomes? Method: We formally introduce the concept of “latent weights” to characterize regression’s implicit weighting of unobserved counterfactuals; derive necessary linear constraints on treatment assignment for causal interpretability; and define the “implicit causal design set,” unifying and extending existing theoretical frameworks. Our approach integrates design-based inference, counterfactual modeling, and linear constraint analysis. Contribution: We establish a necessary conditions framework for causal interpretation of regression, provide operational transparency diagnostics, and deliver novel theoretical justification for widely used—but previously under-justified—regression specifications, including covariate-adjusted regression.

Characterize implications of causal linear regression interpretationIdentify implicit designs for true causal treatment assignmentUnify theoretical results across diverse regression settings

Synthetic Controls for Experimental Design

Aug 04, 2021
AA
Alberto Abadie
🏛️ MIT | Boston University

In large-scale aggregate-unit experiments (e.g., markets), conventional randomized treatment assignment often yields severe baseline imbalance due to extremely few treated units, leading to biased causal estimates. To address this, we systematically integrate the synthetic control method into experimental design, proposing a non-randomized treatment allocation mechanism: dynamically constructing a weighted synthetic control group based on pre-treatment covariates. We further develop配套 components—including counterfactual prediction, distance-driven unit matching, robust variance estimation, and a novel confidence interval construction procedure. Theoretically, our estimator is proven consistent and asymptotically normal. Empirically, it reduces estimation bias by 40–65% relative to standard randomization and substantially improves statistical power. Our core contribution is a new causal inference paradigm for small-N aggregate experiments—rigorous in inference, unbiased under mild assumptions, and highly interpretable.

Addresses experimental design for large aggregate unitsProposes synthetic control designs for accurate estimationReduces bias in treated and control group selection

Latest Papers

What's happening recently
View more

Although proxy metrics are commonly employed to substitute for hard-to-observe primary outcomes, their systematic biases often undermine inference validity and distort confidence intervals. This work proposes the *proxymate* framework—a systematic, modular four-layer diagnostic and correction system encompassing representativeness, unit, estimation, and domain levels—that maps specific failure modes to targeted remediation strategies. Implemented through hierarchical diagnostics, bias-correction algorithms, and an open-source Python toolkit, the framework supports diverse applications including experimentation, monitoring, and prevalence estimation. Validated across thousands of Meta experiments and multiple product lines on millions of proxy–primary outcome pairs, *proxymate* significantly enhances inference reliability and accelerates decision-making.

confidence interval calibrationproxy outcomessurrogate endpoints

In online A/B testing, the relationship between proxy metrics and long-term objectives often breaks down due to user heterogeneity, and relying solely on global correlations can lead to erroneous decisions. This work proposes PROXIMA, a novel framework that introduces a decision-consistency-oriented diagnostic approach for evaluating proxy metrics through three dimensions: normalized effect correlation, directional accuracy, and subgroup vulnerability rate. Integrating causal inference, subgroup analysis, and sensitivity testing, PROXIMA is validated across 80 simulated experiments on the Criteo and KuaiRec datasets. Results show an average decision accuracy of 98.4%; while the subgroup vulnerability rate is markedly higher in recommendation scenarios (68%) than in advertising (13%), directional accuracy exceeds 96% in both, effectively identifying subpopulations where proxy metrics fail.

decision reliabilityheterogeneous treatment effectsonline controlled experiments

This study addresses the frequent conflict between surrogate metrics and core North Star metrics in online experimentation, where a principled framework for their integrated decision-making has been lacking. The authors propose a dynamic optimal fusion method that establishes, for the first time, a formal trade-off framework unifying metric quality and experimental power within a single decision system: it prioritizes the North Star metric when experimental power increases, while upweighting the surrogate metric when its validity is high. The approach estimates optimal fusion weights from historical experiment data by jointly leveraging statistical power analysis and surrogate metric validity assessment. Deployed on Netflix’s experimentation platform, this method significantly enhances decision efficiency and the precision of resource allocation in online experiments.

A/B testingexperiment designnorth star metric

This study addresses the challenge of biased causal effect estimation under unmeasured confounding, where existing proxy methods suffer from ill-posed inverse problems or overly strong assumptions. To overcome these limitations, this work proposes the Proximal Balancing method (PROBE algorithm), which extends classical covariate balancing to confounders observed solely through proxies. By learning low-dimensional summaries of covariates and proxies, PROBE achieves treatment group comparability without specifying proxy roles or solving inverse problems, thereby supporting complex data modalities such as high-dimensional images. The proposed approach effectively corrects estimation bias while maintaining broad applicability. Extensive evaluations on multi-dimensional scenarios and real-world datasets validate its effectiveness. Furthermore, this project establishes rigorous statistical identification theory alongside finite-sample guarantees, significantly enhancing the robustness of causal inference in the presence of unmeasured confounding.

causal effect estimationobservational dataproximal causal inference

This study addresses the pervasive issue of estimation bias and invalid inference that arises when machine learning–generated proxy variables are directly employed in downstream econometric models. The authors propose a novel identification framework that leverages two data sources: a downstream sample containing covariates and the proxy, and an auxiliary validation sample comprising the proxy alongside its true target variable. Using the proxy as a bridge, they construct a locally identified model based on unconditional optimal transport. Crucially, this approach does not require the upstream machine learning estimator to be consistent or to satisfy specific convergence rates, nor does it rely on a fully observed validation sample. Valid asymptotic inference with correct size is achieved through analytically derived critical values, eliminating the need for resampling. Monte Carlo simulations demonstrate that the method maintains accurate size control and yields informative confidence sets across a range of proxy prediction accuracies.

data combinationeconometric inferencemachine learning proxies

Hot Scholars

ZY

Zihan Yu

MSc Student at Imperial College London
Computer VisionMedical AI
SK

Samuel Kaski

Director, ELLIS Institute Finland; Professor, Aalto University and University of Manchester
Probabilistic machine learningAI4ScienceCollaborative AI
SZ

Sichen Zhao

PhD student, RMIT University
Computer Science
SJ

Sabina J. Sloman

Department of Computer Science, University of Manchester
experimental designactive learningmodel misspecificationBayesian inference
EN

Eyal Neuman

Imperial College London
Stochastic processesMathematical finance