Targeted maximum likelihood estimation for longitudinal two-stage designs with outcome subsampling

📅 2026-07-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of data missingness in longitudinal two-stage studies due to participant dropout and outcome subsampling, where conventional inverse probability weighting methods suffer from low efficiency and fail to leverage covariate information. The authors propose two novel approaches by integrating such designs into the longitudinal targeted maximum likelihood estimation (LTMLE) framework: first, an IPCW-LTMLE that incorporates known sampling weights, and second, an unweighted LTMLE that treats the sampling indicator as an intervention node within sequential regression. Theoretical and simulation results demonstrate that the proposed LTMLE methods reduce variance by up to 73% compared to weighted Kaplan–Meier estimators—typically achieving 30–50% gains—with IPCW-LTMLE further improving efficiency by 20–35%. When combined with cross-fitted variance estimation, nominal confidence interval coverage is restored from as low as 76% to the desired level, substantially enhancing inferential validity.
📝 Abstract
We consider efficient estimation of causal parameters in longitudinal two-stage designs with outcome subsampling, motivated by resampling designs in HIV-related mortality studies. In these studies, many participants become lost to follow-up; resampling designs address this by tracing a subset of lost individuals to ascertain their outcomes. Analyses often use inverse-probability-weighted Kaplan-Meier (wKM) estimators that discard longitudinal covariate information and suffer from efficiency losses. We note that resampling designs are an instance of a broader class: two-stage designs with outcome subsampling, in which a first stage collects some data on all participants and a second stage collects outcome information on a selected subset. This connection motivates two estimators. First, drawing on inverse probability of censoring weighted targeted maximum likelihood estimation (IPCW-TMLE) for two-stage designs, we develop its longitudinal extension, IPCW longitudinal TMLE (IPCW-LTMLE) and show that estimating and targeting the known second-stage sampling weights yields variance reductions of up to 36% over the use of known sampling probabilities. Second, given that inverse weighting sacrifices efficiency, we propose an LTMLE that incorporates the second-stage sampling indicator as an intervention node in the sequential regression framework, returning to plug-in estimation and avoiding inverse weighting entirely. Simulations across sample sizes show that LTMLE achieves up to 73% lower variance than wKM with known sampling weights, with reductions of 30-50% common across settings, while IPCW-LTMLE achieves consistent gains of 20-35%. We further demonstrate that cross-fitted variance estimation is essential for valid inference: standard variance estimators yield confidence interval coverage as low as 76%, while our cross-fitted variants consistently restore coverage to nominal levels.
Problem

Research questions and friction points this paper is trying to address.

longitudinal two-stage designs
outcome subsampling
causal parameter estimation
efficiency loss
lost to follow-up
Innovation

Methods, ideas, or system contributions that make the work stand out.

longitudinal TMLE
two-stage design
outcome subsampling
cross-fitted variance estimation
causal inference
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Kirsten E. Landsiedel
Kirsten E. Landsiedel
University of California, Berkeley
causal inferencetargeted machine learningmissing data
M
Maya L. Petersen
Division of Biostatistics, School of Public Health, University of California, Berkeley, CA 94720, USA
M
Mark J. van der Laan
Division of Biostatistics, School of Public Health, University of California, Berkeley, CA 94720, USA