Hazard-Based Targeted Maximum Likelihood Estimation for Survival in Resampling Designs

πŸ“… 2025-11-19
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
In resource-limited settings, high loss-to-follow-up (LTFU) among HIV patients leads to severe underestimation of mortality by conventional survival estimators (e.g., weighted Kaplan–Meier). This paper introduces targeted maximum likelihood estimation (TMLE) into a resampling-based survival analysis framework, proposing a risk-based longitudinal covariate integration method. It employs dynamic conditional hazard modeling combined with stratified inverse probability of censoring weighting (IPCW-TMLE) to robustly correct for non-random LTFU bias. The estimator leverages known resampling probabilities to ensure consistency and accommodates variable follow-up times. Simulation studies demonstrate that the proposed method reduces estimation variance by 55% relative to benchmark approaches while maintaining nominal confidence interval coverage, thereby substantially improving statistical efficiency and accuracy.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationIntelligent Robots: State EstimationReasoning under Uncertainty: Causality

Application Category

User Modeling, Personalization and Recommendation: User privacy protection in personalized systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSecurity and Privacy: Large-scale security measurements
πŸ“ Abstract
Survival is a key metric for evaluating standards of care for people living with HIV. In resource-limited settings, high rates of loss to follow-up (LTFU) often result in underestimation of mortality when only observed deaths are considered. Resampling, which tracks a subset of LTFU patients to ascertain their outcomes, mitigates bias and improves survival estimates. However, common estimators for survival in resampling designs, such as weighted Kaplan-Meier (KM), fail to leverage covariate information collected during repeated clinic visits, even though this information is highly predictive of survival. We propose a Targeted Maximum Likelihood Estimator (TMLE) for survival in resampling designs, which addresses these limitations by leveraging baseline and longitudinal covariates to achieve greater efficiency. Our TMLE is a plug-in estimator and is robust to misspecification of the initial model for the conditional hazard of death, guaranteeing consistency of our estimator due to known resampling probabilities. We present: (1) a fully efficient TMLE for data from resampling studies with fixed follow-up time for all participants and (2) an inverse probability of censoring weighted (IPCW) TMLE that accounts for varied follow-up times by stratifying on patients with sufficient follow-up to evaluate survival. This IPCW-TMLE can be made highly efficient through nonparametric or targeted estimation of the follow-up censoring mechanism. In simulations, our TMLE reduced variance by up to 55% compared with the commonly used weighted KM estimator while preserving nominal confidence interval coverage. These findings demonstrate the potential of our TMLE to improve survival estimation in resampling designs, offering a robust and resource-efficient framework for HIV research. Keywords: Resampling designs, Survival analysis, Targeted Maximum Likelihood Estimation, Inverse probability weighting
Problem

Research questions and friction points this paper is trying to address.

Estimating HIV survival with bias from loss to follow-up in resampling designs
Improving efficiency by leveraging longitudinal covariate data in survival estimation
Developing robust estimators that handle variable follow-up times in HIV studies
Innovation

Methods, ideas, or system contributions that make the work stand out.

TMLE leverages baseline and longitudinal covariates
Plug-in estimator robust to initial model misspecification
IPCW-TMLE handles varied follow-up times efficiently
πŸ”Ž Similar Papers
Kirsten E. Landsiedel
Kirsten E. Landsiedel
University of California, Berkeley
causal inferencetargeted machine learningmissing data
R
Rachael V. Phillips
Division of Biostatistics, University of California, Berkeley
M
Maya L. Petersen
Division of Biostatistics, University of California, Berkeley
M
Mark J. van der Laan
Division of Biostatistics, University of California, Berkeley