doubly robust estimation

Constructs and analyzes estimators (e.g., augmented inverse probability weighting / AIPW or other double-robust estimators) that combine a model for the outcome with a model for treatment/selection/censoring so the estimator is consistent when either model is correctly specified. Uses these estimators to estimate causal effects in settings such as mediation, repeated events, and longitudinal studies with time-varying confounding or adherence by incorporating appropriate time-varying treatment and censoring models and augmentation terms.

doublyrobustestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses robust estimation of causal treatment effects—particularly the average treatment effect (ATE)—in survival analysis. We systematically review and empirically compare six RMST-based methods: two-stage G-formula, augmented inverse probability of censoring weighting–augmented inverse probability weighting (AIPCW-AIPTW), Buckley-James, causal survival forests, direct RMST modeling, and model misspecification analysis. For the first time, we conduct a unified finite-sample evaluation of their robustness, stability, and sensitivity to model misspecification. Results indicate that the two-stage G-formula, AIPCW-AIPTW, Buckley-James, and causal survival forests achieve the strongest overall performance. We provide a practical, evidence-based estimator selection guideline tailored to real-world applications and publicly release all implementation code. This work fills a critical gap in the empirical comparison and applied recommendation of causal survival estimators.

Comparing performance of various estimators through experimentsEstimating treatment effects in causal survival analysisProviding practical guidelines for selecting appropriate methods

Adjusted Nelson--Aalen estimators by inverse treatment probability weighting with an estimated propensity score

Oct 01, 2024
YD
Yuhao Deng
🏛️ University of Michigan | University of Washington

This paper addresses unbiased estimation of the marginal counterfactual cumulative incidence function (CIF) in observational studies under competing risks. We propose a modified Nelson–Aalen estimator based on inverse probability weighting (IPW), incorporating estimated propensity scores. Our key contribution is the first rigorous asymptotic theory for such an adjusted estimator within the competing risks framework: we explicitly characterize the additional variability induced by propensity score estimation and prove its asymptotically negligible impact on final inference. Leveraging counting process theory and influence function analysis, we derive influence functions for both the counterfactual cumulative hazard and the CIF, establishing consistency and asymptotic normality of the estimator. Simulation studies and real-data analyses demonstrate low bias, robust performance across scenarios, and accurate standard error estimation. The proposed method thus provides a theoretically sound and practically viable tool for causal inference in competing risks settings.

Assessing impact of propensity score uncertainty on estimator variationDeriving influence functions for hazard and incidence with estimated propensity scoresEstimating counterfactual cumulative incidence with IPW in competing risks

Combining Experimental and Observational Data for Identification and Estimation of Long-Term Causal Effects

Jan 26, 2022
AG
AmirEmad Ghassami
🏛️ Boston University | Stanford University | University of California, Irvine | Johns Hopkins University | University of Pennsylvania

This study addresses the challenge of identifying and estimating long-term causal effects when observational data suffer from unmeasured confounding and experimental data provide only short-term outcomes—neither source alone supports valid long-term causal inference. To bridge this gap, we propose three novel data fusion frameworks: (i) the equal-confounding-bias assumption, (ii) the partially observable equal-association assumption, and (iii) a proximal causal inference framework leveraging proxy variables—thereby relaxing restrictive external validity assumptions. Integrating proximal causal reasoning, influence function estimation, and potential outcomes modeling, our methods deliver unbiased, efficient, and doubly robust estimators for both the average treatment effect (ATE) and average treatment effect on the treated (ATT). For each framework, we establish rigorous identifiability conditions, construct explicit estimators, and prove their consistency and asymptotic normality. Moreover, the estimators exhibit robustness to certain forms of model misspecification.

Addressing unobserved confounding in observational study dataDeveloping data fusion methods combining experimental and observational datasetsEstimating long-term causal effects with short-term experimental data

Causal machine learning methods and use of sample splitting in settings with high-dimensional confounding

May 24, 2024
SE
Susan Ellul
🏛️ Murdoch Children’s Research Institute | University of Melbourne | Ghent University

To address the challenge of causal effect estimation under high-dimensional confounding, this study systematically compares the Augmented Inverse Probability Weighting (AIPW) and Targeted Maximum Likelihood Estimation (TMLE) estimators under cross-fitting and data-adaptive modeling (including LASSO, random forests, and gradient boosting machines). Our key contributions are threefold: First, we demonstrate—novelly in a real-world, complex epidemiological setting—that cross-fitting is critical for accurate standard error calibration and nominal confidence interval coverage. Second, full-library Super Learner ensemble learning substantially reduces both bias (−37%) and variance (−29%) relative to single-model approaches. Third, while TMLE achieves point estimation accuracy comparable to AIPW, it exhibits superior stability across specifications. Collectively, these results establish the joint use of cross-fitting, full-library Super Learner, and TMLE as a robust and statistically reliable paradigm for causal inference under high-dimensional confounding.

Assessing Super Learner library impact on bias and variance reductionEstimating causal effects with high-dimensional confounding in observational studiesEvaluating doubly robust methods like AIPW and TMLE with cross-fitting

Propensity weighting plus adjustment in proportional hazards model is not doubly robust.

Oct 24, 2023
EE
Erin E Gabriel
🏛️ University of Copenhagen | Uppsala University | Ghent University | Karolinska Institutet

Combining inverse probability weighting (IPW) with multivariable Cox regression is commonly—but incorrectly—regarded as doubly robust; rigorous analysis shows it achieves consistency only under the null causal effect assumption and lacks double robustness in general, even with regression standardization. Method: We propose the first doubly robust estimation framework for both pointwise survival probability differences and the full survival curve, introducing two novel doubly robust estimators for survival differences and one for the survival curve. All estimators remain consistent when either the outcome model (e.g., Cox, Weibull, or flexible parametric proportional hazards models) or the treatment assignment model is correctly specified. Results: We establish theoretical consistency and asymptotic normality, validate finite-sample performance and failure modes via extensive simulations, and demonstrate superior robustness and efficiency over conventional IPW–Cox approaches. A comprehensive R package is provided, enabling estimation, inference, and practical implementation.

Combining Cox model and propensity weighting lacks double robustnessNo double robustness in causal effect estimation for survival modelsProposing alternative doubly robust estimators for survival analysis

Latest Papers

What's happening recently
View more

This study addresses the bias arising from negative weights in staggered difference-in-differences (DiD) designs, which distorts the average of heterogeneous treatment effects—particularly when time-varying covariates are present. Within a model-agnostic framework, the paper nonparametrically defines group-time, group, period, and dynamic average treatment effects for the first time. To this end, the authors propose an augmented inverse variance-weighted (AIVW) estimator that integrates the augmented inverse probability weighting (AIPW) principle, achieving double robustness: consistent estimation of target parameters is maintained even if either the outcome or propensity score model is misspecified. The asymptotic variance is derived via influence functions, enabling valid causal inference with time-varying covariates. Simulations and an empirical application to China’s college admissions policy demonstrate that the estimator performs reliably in finite samples and effectively identifies the causal effects of parallel志愿 and immediate enrollment policies.

heterogeneous effectsnegative weightsstaggered difference-in-differences

Estimating causal mediation effects remains challenging in complex longitudinal settings involving time-varying mediators, time-dependent confounders, and time-to-event outcomes. This study systematically evaluates, for the first time, the applicability and limitations of difference methods in such contexts and compares them with the parametric mediation g-formula. By integrating Cox, Aalen, and accelerated failure time (AFT) models to handle time-varying covariates, the analysis reveals that the Aalen-based difference method yields unbiased estimates in the absence of treatment-induced time-dependent confounding; the Cox-based difference method is approximately unbiased only under rare outcomes; and the AFT-based difference method is generally biased due to non-collapsibility. In contrast, the g-formula consistently produces unbiased estimates across all scenarios. This work clarifies the key assumptions and sources of bias underlying difference methods, offering methodological guidance for mediation analysis with complex survival data.

causal mediation analysisindirect effecttime-dependent confounder

This study addresses the challenge of estimating the causal effect of longitudinal treatment strategies on survival outcomes using electronic health records, where monitoring frequency of covariates varies across patients, variable types, and time, and may itself carry information about underlying health status—potentially biasing conventional causal inference methods. For the first time, monitoring indicators are formally treated as time-varying confounders. The authors integrate inverse probability weighting, G-computation, and longitudinal targeted maximum likelihood estimation (TMLE) to develop a unified framework for causal effect estimation under informative monitoring. This approach substantially reduces bias arising from ignoring the monitoring mechanism and extends the applicability of both static and dynamic treatment strategies. Simulations confirm its validity, and an application to real-world ICU data demonstrates its ability to accurately assess the impact of different mechanical ventilation initiation strategies on mortality.

causal inferenceelectronic health recordsinformative monitoring

When longitudinal outcomes are truncated by death, causal effects become challenging to define and estimate, and existing methods often lack clear causal assumptions and appropriate estimands. This study develops a unified framework that clarifies the definitional challenges and identification assumptions underlying various causal estimands in the presence of truncation by death. It proposes an integrated characterization combining stratum-specific average causal effects with restricted mean survival time, thereby revealing the intrinsically multifactorial nature of the problem. Building on Bayesian inference, the authors derive corresponding estimation procedures and evaluate their performance through simulations and real data from a randomized controlled trial on amyotrophic lateral sclerosis. The results demonstrate that the proposed framework yields a more comprehensive and accurate assessment of treatment effects.

causal estimanddeath censoringlongitudinal study

Hot Scholars

EH

Edward H. Kennedy

Associate Professor of Statistics & Data Science, Carnegie Mellon University
causal inferencenonparametricsmachine learninghealth & public policy
SF

Stefan Feuerriegel

Professor, LMU Munich
AI in ManagementBusiness AnalyticsComputational Social ScienceAI for Good
VS

Vasilis Syrgkanis

Assistant Professor, Stanford University
Machine LearningCausal InferenceEconometricsGame Theory
HZ

Houssam Zenati

Gatsby Computational Neuroscience Unit, University College of London
causal inferenceoff-policy learningsequential learning
AL

Alex Luedtke

Member of the Faculty, Harvard University
causal inferencemachine learningsemiparametricsautomated estimation