Score
Designs and implements methods and preprocessing pipelines to identify, measure, and control confounding when estimating associations or causal effects, using covariate selection, adjustment, matching, weighting, stratification, and time‑varying confounding techniques. Builds and applies balance diagnostics and sensitivity analyses—standardized mean differences, covariate distribution plots, post‑weighting/matching tests, multi‑baseline controls, and residual confounding estimation—and iterates preprocessing or model choices to isolate the target effect from confounds.
In observational causal inference, weighting methods mitigate covariate imbalance but often inflate variance estimates and yield overly conservative standard errors. This paper proposes augmenting weighted regression with main effects of covariates and their interactions with the treatment variable, integrated with residualization and parametric model augmentation to form a unified inferential framework. We establish, for the first time under design-based, model-based, and finite-sample-corrected superpopulation sampling assumptions, that this approach yields asymptotically valid and more precise standard errors. Theory, simulations, and multiple empirical applications demonstrate substantially narrower confidence intervals—on average 15–30% shorter—with improved inferential accuracy and robustness to both exact and approximately balanced weights. The key innovation lies in achieving simultaneous gains in statistical efficiency and asymptotic validity at minimal variance cost.
In causal inference, conventional covariate balance tests suffer from inflated false-positive rates when irrelevant covariates are imbalanced and exhibit low sensitivity to imbalance in potential outcomes. To address these limitations, we propose a conditional balance test grounded in prognostic covariate importance—explicitly incorporating each covariate’s predictive strength for potential outcomes into the test weighting scheme. This enables joint optimization of statistical power and false-positive control. Our method employs a standardized regression-weighted mean difference test, supported by theory-driven weight construction and a Monte Carlo simulation validation framework. We provide theoretical guarantees of improved statistical power. Simulation studies demonstrate that our approach achieves substantially higher detection power than global balance tests under potential outcome imbalance, while reducing the false rejection rate due to irrelevant covariate imbalance by over 40%.
This paper addresses the challenge of confounder selection in observational studies by proposing an interactive, iterative method that requires neither a pre-specified causal graph nor a complete set of candidate variables. Grounded in latent projection theory, the method dynamically expands a causal graph through successive user-provided local adjustment sets and automatically identifies a minimal “principal adjustment set,” thereby determining whether confounding is controllable. Its key contributions are threefold: (1) it is the first approach to achieve sound and complete confounding control assessment without prior structural assumptions on the causal graph; (2) it makes no assumptions about causal relationships among potential confounders; and (3) it bridges theoretical rigor with practical feasibility. Both theoretical analysis and empirical evaluation demonstrate that, under correct user feedback, the algorithm accurately identifies admissible adjustment sets and correctly determines confounding controllability.
This paper addresses the sensitivity of weighted least squares (WLS) estimators in causal inference to unobserved confounding. We propose a general, assumption-lean sensitivity analysis framework grounded in omitted-variable bias theory. Our key innovation is the introduction of the *weighted partial R²*, a measure that quantifies both the direction and magnitude of bias induced by unobserved confounders on weighted causal estimates. From this, we derive a distribution-free sensitivity statistic and an interpretable bound on the strength of unmeasured confounding—expressed as a multiplier relative to observed covariates. The method accommodates arbitrary weighting schemes—including inverse probability weighting (IPW), augmented IPW (AIPW), and stabilized weights—without requiring parametric model specifications or distributional assumptions. An accompanying R package, *weightsense*, enables automated sensitivity reporting and adjusted inference. By providing a transparent, reproducible, and broadly applicable tool for assessing robustness, our framework significantly advances the reliability evaluation of weighted causal estimators.
To address the challenge of causal effect estimation under high-dimensional confounding, this study systematically compares the Augmented Inverse Probability Weighting (AIPW) and Targeted Maximum Likelihood Estimation (TMLE) estimators under cross-fitting and data-adaptive modeling (including LASSO, random forests, and gradient boosting machines). Our key contributions are threefold: First, we demonstrate—novelly in a real-world, complex epidemiological setting—that cross-fitting is critical for accurate standard error calibration and nominal confidence interval coverage. Second, full-library Super Learner ensemble learning substantially reduces both bias (−37%) and variance (−29%) relative to single-model approaches. Third, while TMLE achieves point estimation accuracy comparable to AIPW, it exhibits superior stability across specifications. Collectively, these results establish the joint use of cross-fitting, full-library Super Learner, and TMLE as a robust and statistically reliable paradigm for causal inference under high-dimensional confounding.
This study addresses a key challenge in randomized controlled trials: how to effectively leverage covariate adjustment to improve the precision of average treatment effect estimation while satisfying regulatory requirements and ensuring statistical validity. The authors propose a prespecified, transparent, and reproducible covariate adjustment framework that, for the first time, integrates data-adaptive methods and machine learning into a regulatory-compliant analytical pipeline. By combining model-misspecification-robust estimation with semiparametric efficiency theory, the approach consistently outperforms unadjusted analyses without compromising causal interpretability or statistical validity. It substantially enhances estimation precision, increases statistical power, and yields narrower confidence intervals.
This study addresses the challenge of efficiently and accurately adjusting for covariates in randomized clinical trials with time-to-event endpoints while simultaneously preserving the validity of the log-rank test and improving marginal hazard ratio estimation. The authors establish, for the first time, a first-order asymptotic equivalence between balancing weighting methods—such as stabilized balancing weights and entropy balancing—and augmentation approaches in time-to-event analysis, demonstrating that calibrated weighting achieves estimation efficiency comparable to augmentation without requiring outcome modeling. By integrating propensity score weighting with augmented log-rank scores, they propose a variance estimator that controls finite-sample type I error inflation. Theoretical and simulation results show substantial efficiency gains when covariates are strongly prognostic, and the method’s practical utility is confirmed through application to the REWIND cardiovascular trial.