Score
Designs and carries out analyses that characterize and select operating points for statistical procedures and decision rules, including estimation and reporting of type I error rates, statistical power, sensitivity and specificity, and calibration metrics. This skill includes comparing frequentist and Bayesian decision metrics, analyzing trade‑offs between sensitivity and specificity, and choosing validation‑derived thresholds (for example targeting high specificity) for deployment or triage use.
The area under the ROC curve (AUC) is commonly interpreted as the probability that a classifier ranks a randomly chosen positive instance higher than a negative one; however, this interpretation relies on specific assumptions whose violation can introduce bias that has not been rigorously characterized. This work systematically reviews the relevant literature and, drawing on probabilistic and statistical methods, provides the first rigorous proof of the conditions under which this probabilistic interpretation holds. Furthermore, when these assumptions are violated, the study derives a computable upper bound on the resulting deviation. By establishing a solid theoretical foundation for the probabilistic interpretation of AUC and offering explicit error bounds, this research significantly enhances the reliability and practical applicability of ROC analysis.
This study addresses the need for more reliable regulatory decision-making in Bayesian clinical trials by systematically calibrating Bayesian success criteria to control decision errors. It establishes the first theoretical correspondence between Bayesian decision error metrics and frequentist operating characteristics—specifically Type I and Type II error rates—and proposes a practical calibration strategy grounded in this relationship. The approach is illustrated through a case study on a revascularization trial in cardiogenic shock. To facilitate adoption under the FDA’s emerging Bayesian framework, the authors also developed an interactive Shiny web application that enables sponsors and regulators to efficiently and reliably formulate decisions while maintaining rigorous error control.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
Conventional Monte Carlo simulation for evaluating operating characteristics (e.g., decision accuracy) and determining sample size in Bayesian clinical trials with clustered data and multiple endpoints is computationally expensive and inefficient. Method: We derive, for the first time, an analytical functional relationship between posterior probability and sample size within a Bayesian hierarchical framework that accommodates clustering and multiple endpoints. This enables full operating characteristic curve extrapolation from only two Monte Carlo simulations. We further quantify how simulation variability affects recommended sample sizes. Contribution/Results: By integrating Bayesian hierarchical modeling, cluster-aware inference, and theoretical derivation, our approach drastically reduces computational burden. It is validated on real-world cluster-randomized, adaptive, multi-endpoint Bayesian trials, demonstrating robustness and enabling rapid, reliable sample size determination without sacrificing statistical rigor.
This study addresses the lack of a systematic framework for identifying critical input variables and conducting sensitivity analysis under uncertainty in complex simulations, particularly in military decision-making contexts. The authors propose a unified sensitivity analysis framework that integrates local and global methods—including variance-based, derivative-based, screening, and uncertainty quantification techniques—and strategically maps these approaches to specific decision objectives such as factor prioritization, fixing, variance reduction, and mapping. Innovatively, the framework introduces a “sensitivity audit” mechanism to enhance traceability of model assumptions and promote responsible model usage. By providing a structured guide for high-dimensional, complex simulation systems, this work significantly improves model interpretability, transparency, and the credibility of decisions derived from such models.
This study addresses the lack of systematic guidance on the influence of prior hyperparameters in Bayesian Go/No-Go decision-making for early-phase clinical trials with dual endpoints. The authors propose a calibrated bivariate prior specification framework that constructs skeptical and optimistic priors by assigning target probabilities to predefined decision regions, requiring only the selection of a central location and a precision parameter κ. Theoretically, for any κ > 0, there exists a unique scaling parameter λ₀ that achieves calibration, and κ governs the prior’s discriminative power, supporting adaptive κ strategies. Under a normal–inverse-Wishart model (with default ν₀ = 2), simulations show that at κ = 10, the Go-rate difference between optimistic and skeptical priors reaches 0.56 while maintaining a false-positive rate below 0.01. In a real lupus dataset, optimistic priors with high κ yield Go rates up to three times those of skeptical priors, even with small sample sizes.
This work proposes a Bayesian group sequential design with dynamic information borrowing that reconciles the efficiency gains of incorporating historical data with the regulatory requirement of strict Type I error control. By establishing an explicit correspondence between posterior probability decision rules and uniformly most powerful frequentist tests, the method employs dual thresholds at each interim analysis: one to guarantee exact Type I error control and another to adaptively borrow strength from historical data when appropriate. This approach uniquely unifies frequentist error calibration with Bayesian information borrowing without sacrificing power. Numerical experiments demonstrate that the design maintains the nominal Type I error rate while substantially improving statistical power, and it has been successfully implemented in the design of a Phase III tuberculosis prevention trial integrating historical data from both adult and pediatric populations.
To address low detection sensitivity and frequent omission of rare classes in heterogeneous populations, this paper proposes a covariate-dependent adaptive thresholding method. The method establishes, for the first time, a theoretical framework for the optimal adaptive threshold function, derives its asymptotic properties via nonparametric estimation, and proves a central limit theorem to support statistical inference under conditional mean and variance assumptions. It achieves precise rare-event identification through standardized statistic monitoring, proportional rule modeling, nonparametric function estimation, and bootstrap-based uncertainty quantification. Simulation and empirical studies demonstrate that, compared with conventional fixed-threshold approaches, the proposed method significantly improves alarm sensitivity while maintaining strong false-alarm control and robust inferential performance. This work provides a novel paradigm for dynamic monitoring and screening in imbalanced data settings.