Score
Statistical techniques for handling censored and interval-censored observations in longitudinal and time-to-event data, accounting for limits of detection, differing assay scales and sampling schedules, and for defining unbiased comparison windows and appropriate duration/regression models.
In survival analysis, interval censoring of covariates—such as HIV infection time known only to lie within a follow-up interval—induces bias in causal effect estimation. Method: We propose the first identifiable discrete-time parametric joint model that simultaneously characterizes the interval-censored covariate, the primary event (all-cause mortality), and the censoring mechanism. Our approach integrates discrete-time survival modeling, a joint likelihood framework, the EM algorithm, and Monte Carlo integration to ensure strict parameter identifiability. Contribution/Results: Applied to a South African prospective cohort, the model yields an accurate estimate of the causal effect of HIV status on all-cause mortality (HR = 2.17, 95% CI: 1.83–2.57), substantially outperforming ad hoc alternatives—including ignoring censoring or using proxy variables. This work establishes a theoretically sound and practically viable paradigm for causal inference under interval-censored covariates.
In clinical practice, disease onset times are often only known to lie within follow-up intervals (interval censoring), yet existing illness–death three-state models frequently ignore this feature, leading to biased evaluation of discrimination performance (e.g., time-dependent AUC). This study first systematically demonstrates that interval censoring induces substantial underestimation of time-specific AUC. We establish the necessity of jointly accommodating interval-censored structures both in model fitting and in discrimination assessment. Using simulations and real-world soft-tissue sarcoma data, we evaluate four approaches: Weibull parametric, M-spline–smoothed hazards, piecewise-constant hazards (msm), and a naïve time-dependent Cox model ignoring censoring. Results show that ignoring interval censoring underestimates dynamic AUC by over 12% on average; appropriately accounting for it markedly improves estimation accuracy. Our work provides a methodological benchmark for robust discrimination evaluation of high-dimensional longitudinal prediction models under interval censoring.
This study addresses the challenging problem of parameter estimation in generalized linear models with interval-censored covariates—a common yet methodologically underdeveloped issue in biomedical research. The authors propose GELc, a novel likelihood-based semiparametric approach that, for the first time, integrates an augmented Turnbull nonparametric estimator into the generalized linear modeling framework. By combining maximum likelihood estimation with asymptotic theory, the method establishes estimators that are consistent and asymptotically normal, enabling valid standard error computation. Extensive simulations demonstrate favorable finite-sample performance, with confidence intervals achieving nominal coverage rates. The practical utility of GELc is further confirmed through two real-data applications. The proposed methodology is publicly available as the R package ICenCov.
This study addresses regression modeling for recurrent events in the presence of competing risks when both administrative and random right censoring coexist. The authors propose a unified unbiased estimation strategy that integrates risk-set adjustments with inverse probability of censoring weighting (IPCW) to separately account for the two distinct censoring mechanisms, thereby avoiding excessive modeling assumptions about the censoring process. For the first time, this approach systematically combines administrative and random censoring to yield consistent estimators for parameters in both Ghosh–Lin and Fine–Gray models. Under minimal modeling assumptions, the method substantially enhances estimation robustness and applicability, making it well-suited for large-scale registry databases and clinical trial data.
Interval-censored data are prevalent in survival analysis, yet conventional boosting methods struggle to handle them effectively. This work proposes a nonparametric boosting approach tailored for such data, uniquely integrating an unbiased transformation with functional gradient descent. By designing a customized loss function and employing imputation of response variables, the method enables scalable regression and classification modeling. Theoretical analysis elucidates the trade-off between mean squared error and optimality, while empirical experiments demonstrate that the approach is robust under finite-sample settings and substantially improves predictive accuracy. These advantages underscore its practical utility in domains such as medicine and engineering.
This study addresses the low power and poor stability of conventional Wald-type tests for interval-censored data in small-sample settings by proposing the first likelihood ratio test framework based on spline sieves. The method integrates spline sieve modeling with likelihood ratio testing and rigorously derives its asymptotic distribution, ensuring both theoretical validity and practical robustness. Simulation studies demonstrate that the proposed approach substantially improves statistical power while maintaining proper control of Type I error rates. Its applicability and advantages are further confirmed through analysis of real clinical data.
This study addresses the challenge of misclassified disease status in interval-censored data, particularly when diagnostic errors or biomarker measurement bias are compounded by terminal events such as death. The authors propose a novel semiparametric Cox proportional hazards model that explicitly incorporates diagnostic sensitivity and specificity into the modeling framework to account for uncertainty in disease onset times. The approach simultaneously accommodates terminal events and postmortem confirmation of disease status. Parameter estimation is achieved via nonparametric maximum likelihood combined with an efficient EM algorithm, yielding estimators that attain the semiparametric efficiency bound asymptotically. Simulation studies demonstrate favorable finite-sample performance. In an empirical analysis of Alzheimer’s disease, amyloid-beta was found to be significantly associated with AD onset, while tau protein emerged as a predictor of both AD incidence and mortality risk.
This study addresses the limitations of traditional joint models in clinical longitudinal studies, where repeatedly measured biomarkers or quality-of-life outcomes are often associated with event times but constrained by the proportional hazards assumption, hindering interpretability on the time scale. The authors propose a class of Bayesian semiparametric accelerated failure time joint models that integrate linear mixed-effects models for the longitudinal process and employ Bernstein polynomials to flexibly model the baseline hazard. A time-warping rescaling strategy is introduced to enhance numerical stability and parameter identifiability. By relaxing the proportional hazards assumption, the proposed approach offers more intuitive time-scale interpretations. Simulation studies demonstrate that, when event risk depends on underlying longitudinal trajectories, the method yields more accurate estimates of treatment effects compared to separate modeling approaches and exhibits excellent finite-sample performance.
This study addresses the lack of accessible modeling tools for time-to-event data observed on dual time scales by introducing an R package that offers the first open-source, user-friendly solution for fitting smoothly varying baseline hazard functions across two temporal dimensions. Built upon two-dimensional P-spline smoothing, the proposed method seamlessly integrates with both proportional hazards models and their competing risks extensions, enabling flexible modeling and efficient parameter estimation. The utility of the approach is demonstrated through an application to postoperative breast cancer follow-up data, where it effectively enhances risk estimation and visualization. This work substantially advances the practical capabilities of multi-time-scale survival analysis, making sophisticated modeling more accessible to applied researchers.
Modeling left- or right-censored response variables arising from detection limits remains challenging; conventional approaches—including complete-case analysis, single-value imputation, and parametric Tobit regression—suffer from low efficiency or sensitivity to misspecification of the error distribution. Method: We propose a robust semiparametric rank-based accelerated failure time (AFT) regression framework that yields consistent slope estimates without requiring specification of the error distribution. Contribution/Results: We develop the first unified simulation framework to systematically evaluate model robustness under censoring. Under 10%–60% censoring rates and misspecified error distributions (normal, Weibull, log-normal), our estimator achieves bias <0.05 and relative efficiency >85%, substantially outperforming Tobit, Weibull AFT, and Cox models. We establish the rank-based AFT model as the default robust method for detection-limit data, eliminating reliance on prior distributional assumptions and enhancing reliability and generalizability of censored data analysis in biomedical and environmental research.