Score
Designs, builds, and evaluates time-to-event (duration) statistical models and analyses that estimate survival functions and hazards from censored data, including nonparametric Kaplan‑Meier estimation, duration regression, and Cox proportional hazards models; this work includes simulating survival data, handling censoring and ties, computing survival metrics, confidence intervals and plots, and comparing groups with survival tests. It also implements and assesses extensions such as time‑varying covariate and time‑varying survival modeling, tests of the proportional hazards assumption, estimation of hazard ratios, survival‑aware clustering, and prospective aesa procedures for prospective evaluation.
This study addresses the limitations of overly restrictive parametric assumptions on baseline functions and the challenges of model selection in survival analysis. We propose a semi-parametric time-to-event regression method based on Bernstein polynomials, implemented in the R package spsurv. By avoiding prespecified baseline distributions, this approach enables smooth estimation of the baseline hazard and supports both Bayesian inference and maximum likelihood estimation via Stan. It unifies the interfaces for proportional hazards, proportional odds, and accelerated failure time models while preserving intuitive effect interpretations. Monte Carlo simulations demonstrate the method's robustness in finite samples, and its practical utility is illustrated through an application to oncology clinical trial data. Overall, this work provides a flexible and efficient tool for survival modeling.
Traditional parametric approaches often fail to accurately model complex time-to-event data in clinical trials due to their reliance on prespecified hazard function forms. This work proposes a joint simulation framework based on empirical copulas that integrates nonparametric reconstruction with parametric tail modeling. The method handles right censoring via conditional Kaplan–Meier imputation, models marginal distributions through log-scale location–scale transformations and power distortion of quantile functions, and preserves multivariate rank correlation structures using a Gaussian copula. Remarkably, the approach can reproduce treatment-group survival curves using only a few target quantiles. Applied to a non-small cell lung cancer trial, it successfully replicated both overall survival and progression-free survival curves, yielding a simulated censored Kendall’s tau of 0.522—closely approximating the observed value of 0.549.
This study addresses the inefficiency and bias introduced by right-censored covariates in survival analysis, which commonly undermine conventional approaches such as complete-case analysis. Within the Cox proportional hazards framework, the authors propose a novel method that incorporates a weighted averaging strategy into the partial likelihood function: for observations with censored covariates, the corresponding relative risk is replaced by a weighted average derived from fully observed cases. By directly leveraging the censoring information rather than discarding incomplete cases or imputing constant values, the approach mitigates estimation bias. Extensive simulations and analyses of two oncology clinical trials demonstrate that the proposed method substantially improves estimation efficiency and data utilization, outperforming existing strategies for handling censored covariates.
In insurance analytics, risk assessment for newly enrolled populations is hindered by two interrelated challenges: reporting delays (inducing right-censoring) and administrative constraints limiting follow-up duration—both rendering true event times unobservable. To address this, we propose a parametric proportional hazards model that jointly models event occurrence time and reporting delay time. We introduce latent variables representing the underlying event status and perform marginalization over these latent states to enable valid inference. A two-stage estimation framework based on the EM algorithm is developed, ensuring asymptotic consistency and numerical stability. Our method substantially improves both timeliness and accuracy of risk prediction for new cohorts. Extensive simulations and empirical analysis demonstrate rapid convergence, robust performance, and practical utility. The approach establishes a novel, interpretable, and implementable paradigm for survival analysis under reporting delay.
Traditional accelerated failure time (AFT) models assume constant covariate effects across all survival quantiles, limiting their ability to capture effect heterogeneity. To address this, we propose an interpretable, quantile-dependent multiplicative effects framework—the first to derive closed-form analytic expressions for covariate effects on the quantile scale within the AFT paradigm. Integrating the g-formula, our approach enables standardized estimation of both conditional and marginal effects under left truncation and arbitrary censoring mechanisms. The method combines Bayesian inference, flexible nonlinear functional forms, and quantile-specific posterior inference. Evaluated in an Alzheimer’s disease cohort, it robustly uncovers distinct influences of age, APOE status, and other covariates on early versus late survival quantiles—revealing previously undetected effect heterogeneity. This enhances both statistical precision and clinical interpretability of survival effect estimates.
This study addresses the limitations of traditional joint models in clinical longitudinal studies, where repeatedly measured biomarkers or quality-of-life outcomes are often associated with event times but constrained by the proportional hazards assumption, hindering interpretability on the time scale. The authors propose a class of Bayesian semiparametric accelerated failure time joint models that integrate linear mixed-effects models for the longitudinal process and employ Bernstein polynomials to flexibly model the baseline hazard. A time-warping rescaling strategy is introduced to enhance numerical stability and parameter identifiability. By relaxing the proportional hazards assumption, the proposed approach offers more intuitive time-scale interpretations. Simulation studies demonstrate that, when event risk depends on underlying longitudinal trajectories, the method yields more accurate estimates of treatment effects compared to separate modeling approaches and exhibits excellent finite-sample performance.
This study addresses bias in estimating treatment duration effects from observational survival data arising from time-varying confounding and immortal time bias. Building on the clone–censor–weight (CCW) framework to emulate a target trial, it clarifies core identifying assumptions and distinguishes artificial from natural censoring. Under scenarios involving baseline and time-varying confounders, the authors systematically compare the performance of inverse probability of censoring weighting (IPCW), G-formula, and doubly robust estimators. The approach is applied to a breast cancer cohort to evaluate the comparative effectiveness of two versus five years of adjuvant tamoxifen therapy. Results reveal substantial estimation uncertainty, underscoring the necessity of sensitivity analyses and offering methodological guidance for causal inference in complex observational studies.
This study addresses the challenges posed by unverifiable dependence structures and overly restrictive proportional intensity assumptions in settings involving recurrent events subject to competing terminal events. The authors propose a weighted likelihood–based semiparametric regression model that directly targets the marginal mean of recurrent events while accommodating external time-varying covariates. Innovatively integrating weighted nonparametric maximum likelihood estimation (NPMLE), sandwich variance estimation, and a novel simulation algorithm tailored to the marginal mean, the method avoids untestable assumptions about inter-event dependence and circumvents traditional proportional intensity constraints. Demonstrating robust performance under independent right censoring, the approach is successfully applied to the STATCOPE clinical trial to enable personalized prediction of acute exacerbation frequency in COPD patients and to reassess the efficacy of simvastatin specifically in GOLD stage 4 patients.
This study addresses the challenge of jointly predicting event types and their occurrence times under competing risks by proposing a Bayesian nonparametric approach within a multi-state modeling framework. The method employs a prior defined through a hierarchical completely random measure to model transition probabilities. By deriving the joint marginal and posterior distributions of the observed data and latent random partitions, the work constructs, for the first time in a Bayesian nonparametric setting, theoretically grounded “prediction curves” that enable joint inference on the time-evolving probabilities of competing events and identifies a class of conditionally conjugate priors. Simulation studies and real-world clinical data analyses demonstrate that the proposed approach effectively estimates survival functions, cause-specific hazard rates, and subdistribution functions, confirming its strong predictive performance and practical utility.
Traditional survival analysis methods are often limited by structural assumptions on the hazard function or discretization of time. This work proposes the first application of denoising diffusion models to continuous-time survival analysis, directly modeling the conditional joint distribution of observed event times and censoring indicators in a transformed target space, thereby enabling flexible, nonparametric modeling without restrictive assumptions. The approach integrates a normalized log-time transformation, a continuous Gaussian mixture representation of the censoring indicator, and Kaplan–Meier-based reconstruction, substantially improving calibration and predictive performance. Evaluated across ten real-world datasets, the method achieves competitive results in terms of concordance index (C-index), time-dependent AUC, and Brier score; furthermore, on synthetic data, it more accurately recovers the true underlying continuous survival distribution.