🤖 AI Summary
This study investigates the impact of repeated interim analyses on operating characteristics in Bayesian clinical trials, challenging the common misconception that Bayesian methods are inherently immune to multiplicity issues. Using Monte Carlo simulations based on normally distributed outcomes, the authors compare the performance of Bayesian and frequentist frameworks under group-sequential designs, evaluating the roles of futility stopping rules and multiplicity adjustments. The findings reveal that, without explicit multiplicity correction, the type I error rate inflates with increasing numbers of interim analyses. Moreover, the informational value of Bayesian conclusions diminishes with more frequent looks and is highly sensitive to prior specification, lacking the strict control afforded by frequentist error-rate guarantees. This work offers a novel perspective on evaluating Bayesian operating characteristics and underscores the necessity of carefully addressing multiplicity in repeated analyses within Bayesian trial designs.
📝 Abstract
While Bayesian methods are increasingly used in clinical research, confusion persists as to whether Bayesian designs are affected by repeated interim analyses, and how such effects should be evaluated. We aimed to clarify this question by evaluating both frequentist and Bayesian operating characteristics of group-sequential trials. We conducted simulation studies with normally distributed outcomes, examining designs with repeated analyses, with and without futility stopping rules and multiplicity adjustments. We show that repeated interim analyses alter both frequentist and Bayesian operating characteristics in group-sequential trials, regardless of the inferential framework adopted. Without proper adjustment for multiplicity, the Type I error rate increases with the number of analyses. Bayesian operating characteristics such as the risk of erroneous conclusions and the informative value of an efficacy conclusion are meaningful alternatives to classical frequentist metrics, but they are sensitive to prior divergence between stakeholders, and this sensitivity grows with the number of analyses. Even under full prior agreement, the informative value of an efficacy conclusion is reduced. While it is possible to control frequentist operating characteristics of Bayesian trials with appropriate methods, it is not possible to guarantee such control for Bayesian operating characteristics because they are prior-dependent and different stakeholders may adopt different priors. Contrary to claims that Bayesian inference is immune to multiplicity, our results show that Bayesian clinical trials are no less affected by repeated analyses than frequentist ones, regardless of the framework used to evaluate them.