Score
Constructs and fits accelerated failure time (AFT) regression models for time-to-event (survival) data that model event time (often on the log scale) as a function of covariates to estimate effects such as treatment effects on survival time. Performs estimation, inference, model selection and diagnostics to interpret multiplicative effects on time, assess parameter collapsibility and sensitivity to omitted covariates, and compare AFT-based estimates with hazard-based approaches.
Most existing Accelerated Failure Time (AFT) models assume a linear relationship between covariates and log-survival time, limiting their ability to capture complex nonlinear effects. To address this, we propose DeepR-AFT—the first end-to-end semiparametric AFT framework that couples deep neural networks with a Gehan-type rank-based loss function, enabling direct modeling of high-dimensional, nonlinear covariate effects on log-survival time without prespecifying the error distribution. Our method incorporates a subsampling optimization strategy to enhance training stability and integrates an interpretability module to improve model transparency. Extensive experiments on simulated data and three real-world right-censored datasets demonstrate that DeepR-AFT significantly outperforms classical parametric and semiparametric AFT models, particularly in scenarios involving nonlinear covariate structures and high-dimensional feature spaces, achieving superior predictive accuracy and robustness.
Semiparametric accelerated failure time (AFT) models lack systematic diagnostic tools; existing methods inadequately assess overall model adequacy, link function specification, and functional forms of covariates. Method: We propose the first comprehensive diagnostic framework specifically for semiparametric AFT models, implemented in the R package `afttest`. The framework employs Kolmogorov-type test statistics based on transformed aggregated martingale residual processes, with the multiplier bootstrap used to efficiently approximate the null distribution. It further provides visualization tools to compare observed residual paths against simulated ones. Contribution/Results: The method supports simultaneous hypothesis testing under both rank-based and least-squares estimation, enabling the first joint diagnostic assessment of key AFT model assumptions—including linearity, proportional effects, and link function correctness. Empirical evaluation on the Mayo primary biliary cirrhosis dataset demonstrates high statistical power and interpretability.
Traditional accelerated failure time (AFT) models assume constant covariate effects across all survival quantiles, limiting their ability to capture effect heterogeneity. To address this, we propose an interpretable, quantile-dependent multiplicative effects framework—the first to derive closed-form analytic expressions for covariate effects on the quantile scale within the AFT paradigm. Integrating the g-formula, our approach enables standardized estimation of both conditional and marginal effects under left truncation and arbitrary censoring mechanisms. The method combines Bayesian inference, flexible nonlinear functional forms, and quantile-specific posterior inference. Evaluated in an Alzheimer’s disease cohort, it robustly uncovers distinct influences of age, APOE status, and other covariates on early versus late survival quantiles—revealing previously undetected effect heterogeneity. This enhances both statistical precision and clinical interpretability of survival effect estimates.
This study addresses the challenge of robust inference for a target exposure variable in partially linear accelerated failure time (AFT) models with right-censored data. The authors propose, for the first time, a rank-based debiased machine learning framework that integrates orthogonalized rank-based U-statistics, censoring-corrected influence functions, and a block-pair cross-fitting strategy. This approach overcomes two critical limitations: the lack of Neyman orthogonality in conventional U-statistics and the inapplicability of standard cross-fitting under censoring. The method accommodates flexible covariate adjustment and enables valid inference even in finite samples. Extensive simulations and an application to electronic health records from the All of Us Research Program demonstrate its strong statistical performance and practical utility.
Continuous-time survival models yield biased estimates when applied to discretized or interval-censored event times—common in clinical and administrative data—and fail to properly account for competing risks. Method: We propose the first Python-based, regularized (LASSO/elastic net) semiparametric discrete-time competing risks regression framework, built upon a discrete-time Cox-type formulation. It integrates synthetic data generation, model fitting, and comprehensive evaluation modules, enabling robust high-dimensional modeling and variable selection. Contribution/Results: Evaluated on MIMIC-IV clinical data, the method accurately predicts length of hospital stay. Monte Carlo simulations confirm unbiased parameter estimation and significantly superior predictive performance over mainstream continuous-time approaches—including the Fine–Gray model. This work fills a critical gap in the Python ecosystem for discrete competing risks modeling, combining theoretical rigor with practical scalability and reproducibility.
This study addresses the limitations of overly restrictive parametric assumptions on baseline functions and the challenges of model selection in survival analysis. We propose a semi-parametric time-to-event regression method based on Bernstein polynomials, implemented in the R package spsurv. By avoiding prespecified baseline distributions, this approach enables smooth estimation of the baseline hazard and supports both Bayesian inference and maximum likelihood estimation via Stan. It unifies the interfaces for proportional hazards, proportional odds, and accelerated failure time models while preserving intuitive effect interpretations. Monte Carlo simulations demonstrate the method's robustness in finite samples, and its practical utility is illustrated through an application to oncology clinical trial data. Overall, this work provides a flexible and efficient tool for survival modeling.
Existing tabular foundation models struggle to effectively model censored data in survival analysis, often leading to biased predictions. This work proposes SurvPFN—the first successful extension of Prior-data Fitted Networks (PFNs) to survival analysis—by pretraining on millions of synthetic survival tasks and framing the problem as continuous-time distribution regression that explicitly accounts for censoring. SurvPFN requires no task-specific architecture, feature engineering, or dataset-specific fine-tuning. It leverages Weibull-distributed event times, a non-informative censoring mechanism, and a novel censored negative log-likelihood loss. Evaluated on real-world datasets from SurvSet, SurvPFN matches or surpasses the performance of established classical and deep learning baselines in survival prediction.
This work addresses the challenge of survival analysis with right-censored data by proposing a training-free, end-to-end approach that introduces tabular foundation models (TFMs) into survival analysis for the first time. The method integrates the accelerated failure time (AFT) model with the Buckley–James estimator to perform nonparametric, iterative imputation of censored times within context. Requiring only the estimation of a single scalar parameter and no task-specific training, it effectively handles right-censoring while maintaining computational simplicity. On standard benchmarks, the proposed approach matches or even surpasses the performance of trained Cox regression and parametric AFT models, substantially enhancing the practicality and competitiveness of zero-shot survival regression.
This study addresses the limitations of traditional joint models in clinical longitudinal studies, where repeatedly measured biomarkers or quality-of-life outcomes are often associated with event times but constrained by the proportional hazards assumption, hindering interpretability on the time scale. The authors propose a class of Bayesian semiparametric accelerated failure time joint models that integrate linear mixed-effects models for the longitudinal process and employ Bernstein polynomials to flexibly model the baseline hazard. A time-warping rescaling strategy is introduced to enhance numerical stability and parameter identifiability. By relaxing the proportional hazards assumption, the proposed approach offers more intuitive time-scale interpretations. Simulation studies demonstrate that, when event risk depends on underlying longitudinal trajectories, the method yields more accurate estimates of treatment effects compared to separate modeling approaches and exhibits excellent finite-sample performance.
This study addresses causal inference for time-to-event outcomes in the presence of unmeasured confounding and right censoring. The authors propose a novel identification strategy based on instrumental variable interactions, tailored to structural accelerated failure time models, which circumvents conventional instrumental variable validity assumptions. By integrating augmented inverse probability of censoring weighting, generalized empirical likelihood estimation, and Neyman-orthogonal moment conditions, the method achieves double robustness and orthogonality under multiple weak moment conditions. Theoretical analysis establishes consistency and asymptotic normality of the resulting estimator. Extensive simulations demonstrate superior performance across varying censoring rates and instrumental variable configurations, and the approach is successfully applied to real-world data from the UK Biobank.