๐ค AI Summary
This study addresses the critical yet underexamined practice of โborrowing treatment effectsโ (rather than individual patient data) in Bayesian treatment effect extrapolation. We conduct the first large-scale frequentist simulation study to systematically evaluate the operating characteristics of four methods: conditional power priors, robust mixture priors, post-hoc pooling, and p-value power priors. Results demonstrate that conditional power priors and robust mixture priors achieve the best overall performance in terms of success probability, bias control, and nominal coverage of credible intervals. In contrast, post-hoc pooling and p-value power priors exhibit substantial upward bias and severe undercoverage. This work fills a key methodological gap by providing the first frequentist evaluation framework for Bayesian extrapolation at the treatment-effect level. It delivers empirical evidence and practical guidance for selecting appropriate extrapolation strategies in confirmatory trials.
๐ Abstract
Extrapolating treatment effects from related studies is a promising strategy for designing and analyzing clinical trials in situations where achieving an adequate sample size is challenging. Bayesian methods are well-suited for this purpose, as they enable the synthesis of prior information through the use of prior distributions. While the operating characteristics of Bayesian approaches for borrowing data from control arms have been extensively studied, methods that borrow treatment effects -- quantities derived from the comparison between two arms -- remain less well understood. In this paper, we present the findings of an extensive simulation study designed to address this gap. We evaluate the frequentist operating characteristics of these methods, including the probability of success, mean squared error, bias, precision, and credible interval coverage. Our results provide insights into the strengths and limitations of existing methods in the context of confirmatory trials. In particular, we show that the Conditional Power Prior and the Robust Mixture Prior perform better overall, while the test-then-pool variants and the p-value-based power prior display suboptimal performance.