Semiparametric Regression for Misclassified Competing Risks Data

πŸ“… 2026-05-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

165K/year
πŸ€– AI Summary
This study addresses the critical issue of event-type misclassification in competing risks analysis, which can induce substantial estimation bias, particularly when internal validation data are unavailable. The authors propose a novel semiparametric regression approach that, for the first time, incorporates misclassification probabilities from external validation studies in the absence of internal validation samples. By employing B-spline sieve pseudo-likelihood functions, the method jointly models cause-specific hazards while accounting for misclassification. This framework circumvents the need for costly internal validation procedures. Theoretical analysis establishes the consistency of the resulting estimators, and simulation studies demonstrate robust finite-sample performance, outperforming existing methods. The approach is successfully applied to correct for death misclassification in an HIV cohort study in sub-Saharan Africa.
πŸ“ Abstract
The analysis of competing risks data is often complicated by misclassification of the cause of failure. This issue can lead to seriously biased estimates and invalid conclusions. One way to deal with such misclassification is to use a gold-standard cause of failure ascertainment procedure in a subset of the non-right-censored participants (internal validation sample) along with methods for missing data to deal with the missing gold-standard ascertainments. However, this approach can be costly and time-consuming and, therefore, cannot be implemented in many studies. In this work, we propose a semiparametric regression analysis methodology for the case where no internal validation sample exists. Our approach leverages estimates of the misclassification probabilities from an external validation study to adjust for misclassification in the study at hand. These probabilities are incorporated in a B-spline-based sieve pseudo-likelihood function, which is maximized to jointly estimate models for all event types. Using empirical process theory, we show that the proposed estimator is consistent. Extensive simulation experiments demonstrate that the method performs well with realistic sample sizes and provides substantially more efficient estimates compared to previously proposed approaches. The methodology is applied to competing risks data from a large HIV observational study in sub-Saharan Africa, where event type is misclassified due to significant death under-reporting.
Problem

Research questions and friction points this paper is trying to address.

competing risks
misclassification
cause of failure
validation sample
semiparametric regression
Innovation

Methods, ideas, or system contributions that make the work stand out.

semiparametric regression
competing risks
misclassification
external validation
sieve pseudo-likelihood
πŸ”Ž Similar Papers
No similar papers found.
T
Theofanis Balanos
Department of Biostatistics and Health Data Science, Fairbanks School of Public Health and School of Medicine, Indiana University Indianapolis, Indianapolis, IN, USA
C
Constantin T. Yiannoutsos
Department of Epidemiology and Biostatistics, CUNY Graduate School of Public Health and Health Policy, City University of New York, New York, NY, USA
F
Felix M. Pabon-Rodriguez
Department of Biostatistics and Health Data Science, Fairbanks School of Public Health and School of Medicine, Indiana University Indianapolis, Indianapolis, IN, USA
H
Hongmei Nan
Department of Epidemiology, Richard M. Fairbanks School of Public Health, Indiana University Indianapolis, IN, USA
G
Giorgos Bakoyannis
Department of Biostatistics and Health Data Science, Fairbanks School of Public Health and School of Medicine, Indiana University Indianapolis, Indianapolis, IN, USA