Qualify-Then-Borrow: A Five-Step Framework for Bayesian Borrowing Beyond Outcome Agreement
本文提出Qualify-Then-Borrow框架,通过五个步骤评估并选择性地利用外部数据进行贝叶斯借用,以解决动态借用方法中的信息兼容性和有效性问题。
本文提出Qualify-Then-Borrow框架,通过五个步骤评估并选择性地利用外部数据进行贝叶斯借用,以解决动态借用方法中的信息兼容性和有效性问题。
研究使用视觉-语言模型从RGB视频估计手动物料搬运任务中的外部手部力,以解决职业物理暴露评估中连续力测量的问题。
This study addresses the longstanding challenge in time-to-event trials of balancing inferential maturity with practical feasibility during interim monitoring, where conventional event-driven or enrollment-driven approaches suffer from inherent limitations. The authors propose the WCR framework, which reframes interim monitoring as an information–time alignment problem. By fixing the cohort size and calibrating follow-up requirements, WCR enables continued enrollment while synchronizing interim analyses with information maturity, reserving later-enrolled patients for the final analysis. The framework explicitly distinguishes follow-up constraints between landmark survival estimators and proportional hazards models, jointly calibrates design parameters and decision thresholds, and integrates constrained optimization, simulation-based calibration, and Bayesian methods, implemented in the open-source R package WCRBayesDesign. In simulations of rare pediatric oncology trials, WCR substantially improves the stability and interpretability of interim analysis timing while rigorously controlling Type I error and maintaining power, outperforming existing strategies.
In survival analysis, traditional models assume all individuals will eventually experience the event of interest. However, advances in therapeutics have led to multiple clinical contexts with potentially curative therapies, and in these contexts, certain individuals may never experience the event. Statisticians have developed cure models as a methodology to address this challenge. Nonetheless, despite significant statistical advances in cure models, we have seen more limited uptake in biomedical applications, and we hypothesize that this is caused by limited guidance in the appropriate application of cure models. Cure models require specific identifiability conditions for valid parameter estimation, and previous reports have demonstrated significant issues with the inappropriate application of cure models. Existing tutorials for cure models focus on model implementation and either assume or provide only limited guidance on whether cure modeling is appropriate for the given dataset. This tutorial addresses this gap by describing a systematic procedure that integrates clinical judgment, visual inspection of Kaplan-Meier curves, and quantitative evaluation. We provide a worked example using data from a randomized clinical trial in acute myeloid leukemia, and we also summarize findings from a series of other datasets of hematopoietic cell transplantation to suggest broad practical guidance for choosing to apply cure models. By systematically evaluating cure model appropriateness before fitting these models, researchers can achieve more reliable survival analysis and improved clinical decision-making.
This study addresses the limitations of conventional two-sample time-to-event tests in the presence of long-term survivors (L-TS) under non-proportional hazards, where ignoring the cured fraction can substantially reduce statistical power and where the impact of follow-up duration remains inadequately characterized. Through extensive Monte Carlo simulations under neutral scenarios, the authors systematically compare the type I error and power of traditional tests (e.g., log-rank), non-proportional hazards adjustments, and correctly specified parametric cure models across varying sample sizes, follow-up durations, and effect magnitudes. They find that when both groups contain L-TS, the power of conventional methods varies non-monotonically with follow-up time, whereas parametric cure models exhibit monotonically increasing power. A numerical tool is proposed to predict this non-monotonic behavior, thereby informing optimal follow-up design. The results demonstrate that parametric cure models offer superior performance under prolonged follow-up, establishing a new paradigm for trial design in settings with L-TS.
本文提出Qualify-Then-Borrow框架,通过五个步骤评估并选择性地利用外部数据进行贝叶斯借用,以解决动态借用方法中的信息兼容性和有效性问题。
研究使用视觉-语言模型从RGB视频估计手动物料搬运任务中的外部手部力,以解决职业物理暴露评估中连续力测量的问题。
This study addresses the longstanding challenge in time-to-event trials of balancing inferential maturity with practical feasibility during interim monitoring, where conventional event-driven or enrollment-driven approaches suffer from inherent limitations. The authors propose the WCR framework, which reframes interim monitoring as an information–time alignment problem. By fixing the cohort size and calibrating follow-up requirements, WCR enables continued enrollment while synchronizing interim analyses with information maturity, reserving later-enrolled patients for the final analysis. The framework explicitly distinguishes follow-up constraints between landmark survival estimators and proportional hazards models, jointly calibrates design parameters and decision thresholds, and integrates constrained optimization, simulation-based calibration, and Bayesian methods, implemented in the open-source R package WCRBayesDesign. In simulations of rare pediatric oncology trials, WCR substantially improves the stability and interpretability of interim analysis timing while rigorously controlling Type I error and maintaining power, outperforming existing strategies.
In survival analysis, traditional models assume all individuals will eventually experience the event of interest. However, advances in therapeutics have led to multiple clinical contexts with potentially curative therapies, and in these contexts, certain individuals may never experience the event. Statisticians have developed cure models as a methodology to address this challenge. Nonetheless, despite significant statistical advances in cure models, we have seen more limited uptake in biomedical applications, and we hypothesize that this is caused by limited guidance in the appropriate application of cure models. Cure models require specific identifiability conditions for valid parameter estimation, and previous reports have demonstrated significant issues with the inappropriate application of cure models. Existing tutorials for cure models focus on model implementation and either assume or provide only limited guidance on whether cure modeling is appropriate for the given dataset. This tutorial addresses this gap by describing a systematic procedure that integrates clinical judgment, visual inspection of Kaplan-Meier curves, and quantitative evaluation. We provide a worked example using data from a randomized clinical trial in acute myeloid leukemia, and we also summarize findings from a series of other datasets of hematopoietic cell transplantation to suggest broad practical guidance for choosing to apply cure models. By systematically evaluating cure model appropriateness before fitting these models, researchers can achieve more reliable survival analysis and improved clinical decision-making.
This study addresses the limitations of conventional two-sample time-to-event tests in the presence of long-term survivors (L-TS) under non-proportional hazards, where ignoring the cured fraction can substantially reduce statistical power and where the impact of follow-up duration remains inadequately characterized. Through extensive Monte Carlo simulations under neutral scenarios, the authors systematically compare the type I error and power of traditional tests (e.g., log-rank), non-proportional hazards adjustments, and correctly specified parametric cure models across varying sample sizes, follow-up durations, and effect magnitudes. They find that when both groups contain L-TS, the power of conventional methods varies non-monotonically with follow-up time, whereas parametric cure models exhibit monotonically increasing power. A numerical tool is proposed to predict this non-monotonic behavior, thereby informing optimal follow-up design. The results demonstrate that parametric cure models offer superior performance under prolonged follow-up, establishing a new paradigm for trial design in settings with L-TS.