🤖 AI Summary
This study addresses the high computational burden in Bayesian group sequential designs when jointly calibrating decision thresholds and design skeletons. Conventional approaches rely on time-consuming simulations or approximate posterior computations, which hinder practical implementation. To overcome this, the authors propose a semi-simulation framework that leverages a finite mixture of conjugate priors to enable closed-form posterior updates and low-dimensional numerical integration. By performing a single Monte Carlo precomputation to cache posterior tail probabilities across all candidate analysis times, the method drastically reduces redundant calculations. It accommodates multi-endpoint settings—particularly binary outcomes—using posterior probability–based decision rules. Applied to the ADRENAL trial redesign, the approach replicates the operating characteristics of BATSS and adaptr with speedups of several hundred to thousand-fold, rendering efficient calibration of confirmatory Bayesian group sequential designs routinely feasible.
📝 Abstract
Bayesian group sequential designs (GSDs) extend frequentist GSDs with interpretable decision-making and external evidence borrowing, but their use is limited by the computational burden of design-stage operating-characteristic evaluation. Conventional methods simulate virtual trials with Markov chain Monte Carlo or approximate analytical posterior updates at each interim look, making joint calibration of decision thresholds and design skeletons impractical on commodity hardware. Here we introduce a semi-simulation framework with two innovations. First, finite conjugate-mixture priors replace the posterior computation for each look (``per-look'') with closed-form conjugate updates and low-dimensional numerical integration for decision-rule tail probabilities. Second, a precomputation strategy caches per-look posterior tail probabilities from a single Monte Carlo pass at the union of all candidate analysis times, and each design in the calibration grid is evaluated against the same cache by a sub-second sweep, with no further simulation cost. The framework supports posterior-probability decision rules with multiple efficacy and futility criteria under either binding or non-binding futility, and derives closed-form per-look updates for binary, continuous, count and time-to-event endpoints, with benchmarking here focused on the binary endpoint. When applied to re-design the ADRENAL trial, with up to nine analyses, the framework reproduces the operating characteristics of BATSS and adaptr within Monte Carlo error while running, per GSD, approximately $7\times$ to $16\times$ faster than adaptr at a matched budget (several hundredfold at the million-trial calibration budget) and $3{,}700\times$ to $6{,}600\times$ faster than BATSS. This brings routine Bayesian GSD calibration within computational reach for confirmatory trials.