Score
Designs, builds, and analyzes models for time-to-event and duration data, including construction of person-period/panel event-history datasets, handling censoring and time-varying covariates, and choosing discrete- or continuous-time formulations. Implements and interprets hazard, survival, and transition models (e.g., Cox, parametric survival, logistic person‑period) to estimate entry/exit probabilities, durations, and to compare hazards or odds across groups.
In survival analysis, interval censoring of covariates—such as HIV infection time known only to lie within a follow-up interval—induces bias in causal effect estimation. Method: We propose the first identifiable discrete-time parametric joint model that simultaneously characterizes the interval-censored covariate, the primary event (all-cause mortality), and the censoring mechanism. Our approach integrates discrete-time survival modeling, a joint likelihood framework, the EM algorithm, and Monte Carlo integration to ensure strict parameter identifiability. Contribution/Results: Applied to a South African prospective cohort, the model yields an accurate estimate of the causal effect of HIV status on all-cause mortality (HR = 2.17, 95% CI: 1.83–2.57), substantially outperforming ad hoc alternatives—including ignoring censoring or using proxy variables. This work establishes a theoretically sound and practically viable paradigm for causal inference under interval-censored covariates.
In insurance analytics, risk assessment for newly enrolled populations is hindered by two interrelated challenges: reporting delays (inducing right-censoring) and administrative constraints limiting follow-up duration—both rendering true event times unobservable. To address this, we propose a parametric proportional hazards model that jointly models event occurrence time and reporting delay time. We introduce latent variables representing the underlying event status and perform marginalization over these latent states to enable valid inference. A two-stage estimation framework based on the EM algorithm is developed, ensuring asymptotic consistency and numerical stability. Our method substantially improves both timeliness and accuracy of risk prediction for new cohorts. Extensive simulations and empirical analysis demonstrate rapid convergence, robust performance, and practical utility. The approach establishes a novel, interpretable, and implementable paradigm for survival analysis under reporting delay.
This study addresses the inefficiency and bias introduced by right-censored covariates in survival analysis, which commonly undermine conventional approaches such as complete-case analysis. Within the Cox proportional hazards framework, the authors propose a novel method that incorporates a weighted averaging strategy into the partial likelihood function: for observations with censored covariates, the corresponding relative risk is replaced by a weighted average derived from fully observed cases. By directly leveraging the censoring information rather than discarding incomplete cases or imputing constant values, the approach mitigates estimation bias. Extensive simulations and analyses of two oncology clinical trials demonstrate that the proposed method substantially improves estimation efficiency and data utilization, outperforming existing strategies for handling censored covariates.
This paper addresses the fragmentation and poor interpretability of survival models in elderly fall-risk prediction. We propose a unified survival analysis framework that systematically integrates Poisson regression, exponential regression, and the Cox proportional hazards model. Crucially, we provide the first rigorous proof that, under specific stratification and temporal discretization assumptions, Poisson regression is mathematically equivalent to a discrete-time approximation of the Cox model—thereby bridging a long-standing theoretical gap between these paradigms. The framework enables joint modeling of risk prediction, covariate attribution, and event-time estimation without requiring deep learning–based post-hoc interpretation or multi-task training. Evaluated on real-world clinical data, our approach achieves high predictive accuracy, strong clinical interpretability, and deployment robustness—significantly enhancing the practical utility and trustworthiness of survival analysis in clinical decision support.
Continuous-time survival models yield biased estimates when applied to discretized or interval-censored event times—common in clinical and administrative data—and fail to properly account for competing risks. Method: We propose the first Python-based, regularized (LASSO/elastic net) semiparametric discrete-time competing risks regression framework, built upon a discrete-time Cox-type formulation. It integrates synthetic data generation, model fitting, and comprehensive evaluation modules, enabling robust high-dimensional modeling and variable selection. Contribution/Results: Evaluated on MIMIC-IV clinical data, the method accurately predicts length of hospital stay. Monte Carlo simulations confirm unbiased parameter estimation and significantly superior predictive performance over mainstream continuous-time approaches—including the Fine–Gray model. This work fills a critical gap in the Python ecosystem for discrete competing risks modeling, combining theoretical rigor with practical scalability and reproducibility.
This study addresses the failure of conventional inference methods in parametric survival models under model misspecification, where both the estimand and limiting distribution may deviate from their nominal forms. Adopting a model-agnostic perspective, the work integrates asymptotic theory with a robust inference framework, extending influence functions to censored data settings and developing bootstrap strategies tailored to complex life history models—including regression frameworks and Markov chains. The research rigorously characterizes the true probability limit and asymptotic distribution of estimators under misspecification, proposes novel approaches to enhance inferential robustness, and demonstrates through both theoretical analysis and simulation studies the effectiveness and applicability of the proposed framework in complex survival analysis scenarios.
This study addresses the lack of intuitive and interpretable graphical diagnostic tools for parametric models in survival analysis by proposing a local normalized hazard comparison curve. The method constructs a test process that, under correct model specification, is approximately standard normal at each time point by locally normalizing the difference between nonparametric and parametric hazard rate estimates. It is applicable to a range of models, including exponential, Weibull, Gompertz, Gamma, and parametric Cox regression, substantially enhancing the interpretability and practical utility of goodness-of-fit assessment. Leveraging stochastic process theory, the authors develop an S-Plus implementation that demonstrates strong empirical power in both simulation studies and real-data applications, and the approach naturally extends to discrete-time settings.
This study addresses the limitations of overly restrictive parametric assumptions on baseline functions and the challenges of model selection in survival analysis. We propose a semi-parametric time-to-event regression method based on Bernstein polynomials, implemented in the R package spsurv. By avoiding prespecified baseline distributions, this approach enables smooth estimation of the baseline hazard and supports both Bayesian inference and maximum likelihood estimation via Stan. It unifies the interfaces for proportional hazards, proportional odds, and accelerated failure time models while preserving intuitive effect interpretations. Monte Carlo simulations demonstrate the method's robustness in finite samples, and its practical utility is illustrated through an application to oncology clinical trial data. Overall, this work provides a flexible and efficient tool for survival modeling.
Traditional parametric approaches often fail to accurately model complex time-to-event data in clinical trials due to their reliance on prespecified hazard function forms. This work proposes a joint simulation framework based on empirical copulas that integrates nonparametric reconstruction with parametric tail modeling. The method handles right censoring via conditional Kaplan–Meier imputation, models marginal distributions through log-scale location–scale transformations and power distortion of quantile functions, and preserves multivariate rank correlation structures using a Gaussian copula. Remarkably, the approach can reproduce treatment-group survival curves using only a few target quantiles. Applied to a non-small cell lung cancer trial, it successfully replicated both overall survival and progression-free survival curves, yielding a simulated censored Kendall’s tau of 0.522—closely approximating the observed value of 0.549.
This study addresses the limitations of traditional joint models in clinical longitudinal studies, where repeatedly measured biomarkers or quality-of-life outcomes are often associated with event times but constrained by the proportional hazards assumption, hindering interpretability on the time scale. The authors propose a class of Bayesian semiparametric accelerated failure time joint models that integrate linear mixed-effects models for the longitudinal process and employ Bernstein polynomials to flexibly model the baseline hazard. A time-warping rescaling strategy is introduced to enhance numerical stability and parameter identifiability. By relaxing the proportional hazards assumption, the proposed approach offers more intuitive time-scale interpretations. Simulation studies demonstrate that, when event risk depends on underlying longitudinal trajectories, the method yields more accurate estimates of treatment effects compared to separate modeling approaches and exhibits excellent finite-sample performance.