Planning for Gold: Sample Splitting for Valid Powerful Design of Observational Studies

📅 2024-06-02
📈 Citations: 3
Influential: 0
📄 PDF

career value

179K/year
🤖 AI Summary
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.

Technology Category

Application Category

📝 Abstract
Observational studies are valuable tools for inferring causal effects in the absence of controlled experiments. However, these studies may be biased due to the presence of some relevant, unmeasured set of covariates. The design of an observational study has a prominent effect on its sensitivity to hidden biases, and the best design may not be apparent without examining the data. One approach to facilitate a data-inspired design is to split the sample into a planning sample for choosing the design and an analysis sample for making inferences. We devise a powerful and flexible method for selecting outcomes in the planning sample when an unknown number of outcomes are affected by the treatment. We investigate the theoretical properties of our method and conduct extensive simulations that demonstrate pronounced benefits, especially at higher levels of allowance for unmeasured confounding. Finally, we demonstrate our method in an observational study of the multi-dimensional impacts of a devastating flood in Bangladesh.
Problem

Research questions and friction points this paper is trying to address.

Addresses bias from unmeasured covariates in observational studies
Develops hypothesis screening method using split sample design
Enables valid causal inference despite unmeasured confounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Split samples for hypothesis screening in observational studies
Flexible method for selecting hypotheses with unmeasured confounding
Powerful inference under concerns of hidden biases
🔎 Similar Papers
No similar papers found.
W
William Bekerman
Department of Statistics and Data Science, University of Pennsylvania, Philadelphia, PA, USA
A
Abhinandan Dalal
Department of Statistics and Data Science, University of Pennsylvania, Philadelphia, PA, USA
C
C. Ninno
Retired, Formerly of the World Bank, District of Columbia, USA
D
Dylan S. Small
Department of Statistics and Data Science, University of Pennsylvania, Philadelphia, PA, USA