Score
Estimating units' treatment-assignment probabilities and applying matching or weighting schemes to control for confounders, enabling causal interpretation of observed associations and design of experimental or placebo-matching strategies to test cohort-specific effects.
When the positivity assumption fails, conventional causal estimands—such as the average treatment effect (ATE), average treatment effect on the treated (ATT), and average treatment effect on the controls (ATC)—are not identifiable. To address this, we propose weighted causal estimands (WATE, WATT, WATC) as a robust alternative that preserves internal validity. We systematically review the theoretical foundations and recent advances in propensity score weighting, clarify principles for selecting target estimands, and provide practical guidance—including estimation implementation, post-weighting balance diagnostics, and hands-on tutorials using the R package *ChiPS*. Simulation studies demonstrate the method’s validity and robustness. We further illustrate its feasibility and utility through two real-world applications: (1) estimating the effect of smoking on blood lead levels using NHANES data, and (2) assessing the impact of sex work history on HIV incidence among transgender women in South Africa. This work advances reliable causal inference under non-ideal data conditions.
While covariate-adjusted estimators (e.g., linear regression, matching) are widely used in causal inference, it remains unclear whether rerandomization—despite such adjustments—still delivers meaningful benefits, particularly in finite samples where existing asymptotic theory fails to capture non-precision advantages (e.g., estimator consistency) or practical performance. Method: We conduct large-scale simulation studies to systematically evaluate rerandomization’s impact on estimation precision, statistical power, confidence interval coverage, and consistency across multiple estimators, complemented by theoretical analysis. Contribution/Results: Rerandomization substantially improves finite-sample estimation accuracy, robustness, and consistency of causal effect estimates; enhances statistical power; and reduces false-positive rates—even when covariate adjustment is already employed. These gains extend beyond asymptotic guarantees, offering practitioners a principled, efficiency-enhancing, and reliability-improving design strategy for randomized experiments.
To address confounding bias arising from covariate distribution imbalance in causal inference from observational data, this paper proposes a two-stage interpretable matching framework. In the first stage, exact matching is performed on all covariates to ensure baseline comparability. In the second stage, the least significant confounders are iteratively removed based on feature importance, and an interpretable distance metric learning approach is introduced to quantify proximity with respect to the removed variables. The method simultaneously ensures multivariate overlap and unbiased estimation of conditional average treatment effects (CATE), while substantially enhancing matching transparency and robustness. Experiments on synthetic datasets and real-world CDC healthcare data demonstrate that the proposed approach significantly reduces CATE estimation bias, improves high-dimensional overlap between treatment and control groups, and exhibits strong computational scalability.
This paper addresses the longstanding challenge of estimating post-matching or post-reweighting treatment-effect variance—particularly for average treatment effects on the treated (ATT)—under settings with small treated samples and control-group reuse. We propose a computationally efficient, theoretically robust unifying framework that matches treated units only to control units, avoiding symmetric matching or full reweighting. Our method enables valid inference for population-level causal parameters while preserving finite-sample reliability. Crucially, we develop the first variance estimator that is both asymptotically efficient and computationally feasible, compatible with widely used methods including radius matching, *k*-nearest-neighbor matching, propensity score matching, and stable balancing weights. Under novel regularity conditions, our asymptotic theory guarantees statistical validity. Simulations demonstrate that our 95% confidence intervals achieve nominal coverage consistently, markedly outperforming bootstrap-based alternatives (which drop as low as 61%). The methodology is implemented in the R package `scmatch2`.
Multifactor causal inference in observational studies faces two key challenges: sparse or missing factor combinations hinder interaction effect identification, while covariate and factor imbalance induces estimation bias. This paper proposes Doubly Balanced Weighting (DBW), the first method to elevate factor balance to theoretical parity with covariate balance in observational factorial studies. DBW jointly optimizes an inverse-probability-weighting objective to simultaneously balance both covariate distributions and factor combination distributions. It enables robust estimation of main and interaction effects—even under missing treatment combinations—and provides consistent asymptotic variance estimation. Simulation and empirical analyses demonstrate that DBW substantially improves estimation accuracy and achieves nominal 95% confidence interval coverage. The framework offers a theoretically grounded, generalizable solution for causal inference in multifactor observational studies.
This study addresses two fundamental challenges in evaluating individualized treatment benefit predictors (TBPs) from observational data: nonidentifiability and confounding bias. Methodologically, we first establish that confounding bias propagates nonlinearly and unpredictably in TBP evaluation; we then develop a novel identifiability framework grounded solely in observable data, leveraging latent-variable reconstruction—including the benefit concentration index and moderate calibration curve—to derive causal identifiability expressions for discrimination and calibration metrics. We theoretically prove identifiability under partial confounder control and quantify the systematic failure of conventional causal intuition in this setting. Our contributions provide a new paradigm and practical toolkit for robust TBP evaluation in clinical decision support.
This study addresses the problem of extrapolating causal effects from multi-site randomized controlled trials (RCTs) to a new target site with baseline survey data only. To handle site-level population heterogeneity and unobserved confounding, we propose modeling baseline covariates as functional data—thereby capturing site-specific confounding structures—for the first time. We then develop a design-oriented, nonparametric method to construct an optimal finite-dimensional feature space, ensuring optimal convergence rates for conditional average treatment effect (CATE) estimation. Our approach integrates functional data analysis, nonparametric regression, and causal transfer learning theory. Evaluated across five integrated multi-site RCTs on cash transfer programs, the method significantly improves prediction accuracy of treatment effects at target sites and quantifies the estimation gain attributable to adaptive transfer.
This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.
In observational causal inference, accurate identification of confounding variables is critical for reliable causal estimates. This work proposes ConfoundingSHAP, a novel method that uniquely adapts Shapley values to quantify the confounding strength of covariates. By formulating a Shapley game specifically tailored to confounding effects and designing an appropriate value function, the approach precisely measures each variable’s contribution to bias in causal estimation. Furthermore, it integrates TabPFN to enable efficient and scalable evaluation of adjustment sets without repeated model retraining. Empirical results across multiple datasets demonstrate that ConfoundingSHAP accurately identifies key confounders and provides interpretable, trustworthy insights into the sources of confounding.
Existing matching-based causal inference methods are limited to single treatment types (e.g., binary, ordinal, or continuous) or binary multi-factor treatments, rendering them inadequate for real-world policy evaluation involving multi-factor treatments with continuous or ordinal components. To address this gap, we propose the first general matching framework supporting non-binary, multi-factor treatments. Our approach employs a two-stage non-bipartite matching procedure to construct comparable unit sets, enabling unbiased estimation of both main and interaction effects. We introduce the generalized factorial Neyman estimator—unifying factorial structure modeling with arbitrary treatment types—and develop a Fisher- and Neyman-type randomization inference framework, augmented by a covariate-driven variance tuning method. Evaluated on nationwide U.S. county-level COVID-19 data, the framework successfully identifies causal effects of work/non-work mobility reduction on disease transmission and drug-related outcomes, demonstrating its validity and robustness.
Clinical studies often exhibit systematic discrepancies between the sample and the target population, inducing extrapolation bias. This paper investigates the robustness of inverse probability sampling weighting (IPSW) under misspecified target populations: even with correctly specified models, IPSW yields systematic bias if the selected target population fails to represent the actual inferential population. Through simulation experiments across diverse real-world covariate distributions and selection mechanisms, we quantify how deviation of the target population from representativeness affects the estimation of the population average treatment effect (PATE). Results demonstrate that bias increases monotonically with the degree of target-population mismatch—and in severe cases, IPSW performs worse than unweighted estimation. To our knowledge, this is the first systematic study revealing that target-population selection constitutes a foundational design decision in causal extrapolation, whose impact can surpass that of model misspecification—providing a critical methodological warning for causal inference beyond the study sample.
This study addresses the multiple testing problem in matched observational studies with a single intervention and multiple endpoints. We propose a robust method that jointly controls the false discovery rate (FDR) and quantifies unmeasured confounding bias. Our key innovation is the first integration of FDR control with formal sensitivity analysis, achieved via integer programming and a hierarchical screening strategy to efficiently compute sensitivity sets—i.e., subsets of hypotheses remaining significant under varying magnitudes of unmeasured confounding—enabling conservative estimation of the true positive rate (TPR). The method supports simultaneous inference across the entire hypothesis space, balancing statistical power and robustness. Simulation studies and an empirical application investigating long-term effects of childhood abuse demonstrate that our approach reliably identifies high-confidence endpoint subsets even under substantial hidden bias, substantially improving the reproducibility and interpretability of exploratory analyses.