Score
Designs and evaluates procedures that produce calibrated probability or interval estimates and confidence scores that achieve intended nominal coverage under repeated sampling. This includes building recalibration methods for model outputs, constructing and validating confidence intervals and high‑probability bounds, selecting thresholds to control error rates or coverage, estimating acceptance‑conditioned error rates, and applying held‑out calibration data and frequentist coverage‑control techniques to reduce biased confidence assignments across groups.
This study addresses the challenge of constructing valid confidence intervals in two-stage adaptive enrichment clinical trials, where patient subgroups are selected based on interim data, thereby compromising the nominal coverage of conventional intervals. The authors propose a novel method that constructs confidence intervals conditional on the interim selection decision, leveraging conditional inference and inversion of uniformly most accurate unbiased (UMAU) tests to guarantee exact coverage within the selected subgroup. The approach is broadly applicable to various adaptive enrichment designs and is implemented via an efficient numerical algorithm. Extensive simulation studies demonstrate that the proposed intervals consistently achieve the desired coverage probability across diverse design configurations, substantially outperforming existing methods in both validity and precision.
In multi-group comparisons, constructing confidence intervals for a data-selected target group (e.g., the group with the largest sample mean) while enforcing strict conditional coverage leads to infinite expected interval width under the normal means model, rendering inference meaningless. To address this, we propose a novel empirical Bayes framework that employs selection-adjusted priors and data-driven shrinkage estimation to approximate the oracle procedure achieving exact selective coverage. Our method guarantees finite expected width over reasonable parameter regimes and delivers high-accuracy approximate selective coverage. Numerical experiments demonstrate substantially improved coverage compared to existing approaches, with only a modest increase in interval width. Our key contributions are: (i) the first formal identification of the width pathology inherent in strict conditional coverage; and (ii) the construction of the first computationally tractable empirical Bayes inference scheme that simultaneously ensures finite expected width, computational feasibility, and rigorous approximate selective coverage guarantees.
This paper addresses the lack of robust design foundations for sensitivity analysis in finite-population causal inference. Methodologically, it introduces a novel sensitivity analysis framework grounded in the experimental design distribution—first integrating design-based distributions with partial identification theory to construct model-free, non-asymptotic confidence intervals for the average treatment effect (ATE). It further reinterprets the role of randomization in sensitivity analysis and provides a new design-driven rationale for covariate balance checks. Key contributions include: (1) model-free, finite-population inference under heterogeneous treatment effects; (2) robust ATE confidence intervals with clear identification-theoretic interpretation; and (3) empirical validation across three real-world applications, demonstrating reliability and practicality in small-sample and highly heterogeneous settings.
Approximate Bayesian inference often underestimates true uncertainty due to posterior credible intervals that are excessively narrow. This work proposes two simulation-based calibration (SBC)-driven methods for recalibrating approximate posteriors, systematically leveraging the SBC framework to adjust the width of posterior uncertainty intervals and achieve marginal calibration. The approach is applicable to complex model structures, including hierarchical models, and demonstrates consistent efficacy across diverse experimental settings by meaningfully widening posterior intervals. As a result, the proposed recalibration substantially enhances the calibration accuracy and reliability of approximate Bayesian inference.
Classical algorithms for strongly convex stochastic optimization achieve fast convergence (O(1/√n)) but suffer from asymptotically non-negligible bias, violating the conditions required for a valid central limit theorem (CLT) and thus impeding asymptotically efficient statistical inference. Method: We propose the first dual-objective algorithm that simultaneously guarantees fast convergence and a provable CLT. Our approach integrates stochastic approximation, asymptotic statistical inference, and adaptive experimental design into a unified framework that ensures asymptotic normality of the estimator. Contribution/Results: We establish theoretical guarantees that the algorithm retains the O(1/√n) convergence rate while satisfying the CLT. Numerical experiments demonstrate substantial improvements over existing methods in estimation accuracy, confidence interval coverage, and identification of optimal treatment parameters. The method provides a new paradigm for continuous, parameterized A/B testing in online platforms—balancing optimization efficiency with statistical reliability.
This work proposes a Bayesian group sequential design with dynamic information borrowing that reconciles the efficiency gains of incorporating historical data with the regulatory requirement of strict Type I error control. By establishing an explicit correspondence between posterior probability decision rules and uniformly most powerful frequentist tests, the method employs dual thresholds at each interim analysis: one to guarantee exact Type I error control and another to adaptively borrow strength from historical data when appropriate. This approach uniquely unifies frequentist error calibration with Bayesian information borrowing without sacrificing power. Numerical experiments demonstrate that the design maintains the nominal Type I error rate while substantially improving statistical power, and it has been successfully implemented in the design of a Phase III tuberculosis prevention trial integrating historical data from both adult and pediatric populations.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.