survival analysis

Statistical modeling and evaluation of time-to-event data, hazard functions, and competing risks to predict event timing and incidence. It covers formulating stochastic-process representations, cause-specific hazards, and losses appropriate for censored or multi-outcome clinical prediction tasks.

survivalanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of jointly predicting event types and their occurrence times under competing risks by proposing a Bayesian nonparametric approach within a multi-state modeling framework. The method employs a prior defined through a hierarchical completely random measure to model transition probabilities. By deriving the joint marginal and posterior distributions of the observed data and latent random partitions, the work constructs, for the first time in a Bayesian nonparametric setting, theoretically grounded “prediction curves” that enable joint inference on the time-evolving probabilities of competing events and identifies a class of conditionally conjugate priors. Simulation studies and real-world clinical data analyses demonstrate that the proposed approach effectively estimates survival functions, cause-specific hazard rates, and subdistribution functions, confirming its strong predictive performance and practical utility.

Bayesian nonparametricscompeting risksmulti-state modeling

Competing Risks: Impact on Risk Estimation and Algorithmic Fairness

Aug 07, 2025
VJ
Vincent Jeanselme
🏛️ Columbia University | University of Cambridge

This work exposes systemic bias and algorithmic unfairness arising from misclassifying competing risks as censoring in survival analysis: such misclassification systematically overestimates individual hazard, with disparities amplified across subpopulations due to heterogeneous competing event incidence. We formally define this misclassification problem, propose an analytically tractable error estimation framework, and derive theoretical bounds on the direction and magnitude of bias. Using real-world cardiovascular clinical data, we empirically quantify its dual detrimental impact on predictive calibration and fairness metrics—including equalized odds and equal opportunity. Results demonstrate that ignoring competing risks not only distorts population-level survival estimates but also disproportionately harms high-risk subgroups (e.g., elderly, low-income patients) through miscalibrated risk predictions. This study establishes a theoretical foundation and practical guidelines for developing survival prediction models that are both statistically accurate and equitable.

Competing risks amplify disparities in algorithmic risk assessmentIgnoring competing risks disproportionately harms high-risk individualsMisclassifying competing risks as censoring biases survival estimates

In studies with competing risks, conventional incidence measures are susceptible to censoring mechanisms and often fail to accurately reflect the true burden of events. This work proposes the Average Cause-Specific Hazard (ACSH), a novel incidence metric based on survival weighting that does not rely on assumptions about the censoring distribution. The authors develop a nonparametric estimator for ACSH and a corresponding two-sample comparison procedure. ACSH retains the intuitive interpretability of incidence rates while fully eliminating bias induced by censoring, and enables group comparisons without requiring strong modeling assumptions. Simulation studies demonstrate its favorable finite-sample properties, and its application to the CANVAS trial yields robust and interpretable estimates of between-group differences.

cause-specific hazardcensoringcompeting risks

Time-Dependent Pseudo $oldsymbol{R^2}$ for Assessing Predictive Performance in Competing Risks Data

Jul 20, 2025
ZZ
Zian Zhuang
🏛️ University of California at Los Angeles | City University of Hong Kong | University of Southern California

For right-censored time-to-event data with competing risks, conventional predictive performance metrics—such as the C-index, Brier score, and time-dependent AUC—exhibit weak discriminative power and poor stability in model comparison. To address this, we propose a time-varying pseudo-R² metric specifically designed for the cumulative incidence function (CIF), the first of its kind to define an overall, time-restricted pseudo-coefficient of determination. We derive a sample estimator based on pseudo-observations and establish its asymptotic theory—consistency and asymptotic normality. The proposed metric is interpretable, sensitive to temporal dynamics, and highly discriminative among competing models. Extensive simulations and real-data analyses demonstrate its superior performance across diverse censoring mechanisms and competing-risk configurations, significantly outperforming existing metrics. This work provides a robust, theoretically grounded framework for evaluating predictive models under competing risks.

Address limitations of conventional predictive performance metricsDevelop robust metrics for competing risks time-to-event dataIntroduce time-dependent pseudo R² for predictive cumulative incidence function

PyDTS: A Python Package for Discrete-Time Survival (Regularized) Regression with Competing Risks

Apr 12, 2022
TM
T. Meir
🏛️ Technion - Israel Institute of Technology | Tel Aviv University

Continuous-time survival models yield biased estimates when applied to discretized or interval-censored event times—common in clinical and administrative data—and fail to properly account for competing risks. Method: We propose the first Python-based, regularized (LASSO/elastic net) semiparametric discrete-time competing risks regression framework, built upon a discrete-time Cox-type formulation. It integrates synthetic data generation, model fitting, and comprehensive evaluation modules, enabling robust high-dimensional modeling and variable selection. Contribution/Results: Evaluated on MIMIC-IV clinical data, the method accurately predicts length of hospital stay. Monte Carlo simulations confirm unbiased parameter estimation and significantly superior predictive performance over mainstream continuous-time approaches—including the Fine–Gray model. This work fills a critical gap in the Python ecosystem for discrete competing risks modeling, combining theoretical rigor with practical scalability and reproducibility.

Addresses bias from standard continuous-time survival analysis modelsModels discrete-time survival data with competing risks scenariosProvides regularized estimation methods for competing risks survival analysis

Latest Papers

What's happening recently
View more

This study addresses the critical issue of event-type misclassification in competing risks analysis, which can induce substantial estimation bias, particularly when internal validation data are unavailable. The authors propose a novel semiparametric regression approach that, for the first time, incorporates misclassification probabilities from external validation studies in the absence of internal validation samples. By employing B-spline sieve pseudo-likelihood functions, the method jointly models cause-specific hazards while accounting for misclassification. This framework circumvents the need for costly internal validation procedures. Theoretical analysis establishes the consistency of the resulting estimators, and simulation studies demonstrate robust finite-sample performance, outperforming existing methods. The approach is successfully applied to correct for death misclassification in an HIV cohort study in sub-Saharan Africa.

cause of failurecompeting risksmisclassification

This study addresses the lack of accessible modeling tools for time-to-event data observed on dual time scales by introducing an R package that offers the first open-source, user-friendly solution for fitting smoothly varying baseline hazard functions across two temporal dimensions. Built upon two-dimensional P-spline smoothing, the proposed method seamlessly integrates with both proportional hazards models and their competing risks extensions, enabling flexible modeling and efficient parameter estimation. The utility of the approach is demonstrated through an application to postoperative breast cancer follow-up data, where it effectively enhances risk estimation and visualization. This work substantially advances the practical capabilities of multi-time-scale survival analysis, making sophisticated modeling more accessible to applied researchers.

epidemiological studieshazard modelsmultiple time scales

This study addresses the limitation of conventional variable selection methods, which are typically confined to a single hazard structure—such as proportional hazards (PH) or accelerated failure time (AFT)—and struggle to simultaneously identify relevant covariates and an appropriate risk model. Within the generalized hazard (GH) model framework, this work proposes a Bayesian joint selection approach that concurrently infers active variables and the underlying hazard structure. The method innovatively introduces two GH-adapted g-priors and incorporates a hierarchical prior that accounts for both multiplicity adjustment and model complexity penalization, thereby ensuring selection consistency. An extended Add-Delete-Swap MCMC algorithm is developed to efficiently explore the joint space of variables and hazard structures. Simulation studies and real-data analyses demonstrate that the proposed method accurately recovers the true model and outperforms existing approaches in both variable selection and hazard structure identification.

Bayesian variable selectionGeneral Hazard modelhazard structure selection

This study addresses the limitations of traditional joint models in clinical longitudinal studies, where repeatedly measured biomarkers or quality-of-life outcomes are often associated with event times but constrained by the proportional hazards assumption, hindering interpretability on the time scale. The authors propose a class of Bayesian semiparametric accelerated failure time joint models that integrate linear mixed-effects models for the longitudinal process and employ Bernstein polynomials to flexibly model the baseline hazard. A time-warping rescaling strategy is introduced to enhance numerical stability and parameter identifiability. By relaxing the proportional hazards assumption, the proposed approach offers more intuitive time-scale interpretations. Simulation studies demonstrate that, when event risk depends on underlying longitudinal trajectories, the method yields more accurate estimates of treatment effects compared to separate modeling approaches and exhibits excellent finite-sample performance.

accelerated failure timeinformative censoringjoint modelling

This study addresses the critical need for reliable prediction intervals for future event counts in interim analyses of clinical trials, where accurate timing of analyses depends on such forecasts. The authors propose a novel, general-purpose prediction framework tailored to this setting, extending prediction interval methodology from reliability engineering to patient-level survival data. The approach accommodates various parametric survival models, covariate adjustment, staggered enrollment, and modeling of dependence between enrollment and censoring processes. Within a frequentist framework, it estimates the conditional distribution of future events using a bootstrap-based procedure. Extensive simulations demonstrate favorable performance, and the method is successfully applied to a Phase III trial in pediatric acute lymphoblastic leukemia, yielding reliable interval predictions for event counts at prespecified calendar dates—thereby filling a notable gap in existing methodology.

clinical trialsevent predictioninterim analysis

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
RG

Russell Greiner

Professor of Computing Science, University of Alberta; CIFAR AI Chair
Artificial IntelligenceMachine LearningSurvival PredictionMedical Informatics
NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory
JF

Junyi Fan

University of Southern California
machine learning
SC

Shuheng Chen

University of Southern California
Machine LearningData SciencePredictive AnalyticsClinical Prediction