calibrate confidence estimates

Designs and evaluates procedures that produce calibrated probability or interval estimates and confidence scores that achieve intended nominal coverage under repeated sampling. This includes building recalibration methods for model outputs, constructing and validating confidence intervals and high‑probability bounds, selecting thresholds to control error rates or coverage, estimating acceptance‑conditioned error rates, and applying held‑out calibration data and frequentist coverage‑control techniques to reduce biased confidence assignments across groups.

calibrateconfidenceestimates

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$179K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of constructing valid confidence intervals in two-stage adaptive enrichment clinical trials, where patient subgroups are selected based on interim data, thereby compromising the nominal coverage of conventional intervals. The authors propose a novel method that constructs confidence intervals conditional on the interim selection decision, leveraging conditional inference and inversion of uniformly most accurate unbiased (UMAU) tests to guarantee exact coverage within the selected subgroup. The approach is broadly applicable to various adaptive enrichment designs and is implemented via an efficient numerical algorithm. Extensive simulation studies demonstrate that the proposed intervals consistently achieve the desired coverage probability across diverse design configurations, substantially outperforming existing methods in both validity and precision.

adaptive enrichment designsconfidence intervalscoverage probability

Selective and marginal selective inference for exceptional groups

Sep 16, 2025
PH
Peter Hoff
🏛️ Duke University

In multi-group comparisons, constructing confidence intervals for a data-selected target group (e.g., the group with the largest sample mean) while enforcing strict conditional coverage leads to infinite expected interval width under the normal means model, rendering inference meaningless. To address this, we propose a novel empirical Bayes framework that employs selection-adjusted priors and data-driven shrinkage estimation to approximate the oracle procedure achieving exact selective coverage. Our method guarantees finite expected width over reasonable parameter regimes and delivers high-accuracy approximate selective coverage. Numerical experiments demonstrate substantially improved coverage compared to existing approaches, with only a modest increase in interval width. Our key contributions are: (i) the first formal identification of the width pathology inherent in strict conditional coverage; and (ii) the construction of the first computationally tractable empirical Bayes inference scheme that simultaneously ensures finite expected width, computational feasibility, and rigorous approximate selective coverage guarantees.

Addresses infinite expected width in selective confidence intervalsCompares selective versus marginal coverage control trade-offsDevelops empirical Bayes methods for approximate selective coverage

Finite Population Identification and Design-Based Sensitivity Analysis

Apr 19, 2025
BK
Brendan Kline
🏛️ University of Texas at Austin | Duke University

This paper addresses the lack of robust design foundations for sensitivity analysis in finite-population causal inference. Methodologically, it introduces a novel sensitivity analysis framework grounded in the experimental design distribution—first integrating design-based distributions with partial identification theory to construct model-free, non-asymptotic confidence intervals for the average treatment effect (ATE). It further reinterprets the role of randomization in sensitivity analysis and provides a new design-driven rationale for covariate balance checks. Key contributions include: (1) model-free, finite-population inference under heterogeneous treatment effects; (2) robust ATE confidence intervals with clear identification-theoretic interpretation; and (3) empirical validation across three real-world applications, demonstrating reliability and practicality in small-sample and highly heterogeneous settings.

Analyzes randomization role and motivates covariate balance examinationConstructs design-based confidence intervals for heterogeneous treatment effectsDevelops sensitivity analysis using design distributions for finite populations

Approximate Bayesian inference often underestimates true uncertainty due to posterior credible intervals that are excessively narrow. This work proposes two simulation-based calibration (SBC)-driven methods for recalibrating approximate posteriors, systematically leveraging the SBC framework to adjust the width of posterior uncertainty intervals and achieve marginal calibration. The approach is applicable to complex model structures, including hierarchical models, and demonstrates consistent efficacy across diverse experimental settings by meaningfully widening posterior intervals. As a result, the proposed recalibration substantially enhances the calibration accuracy and reliability of approximate Bayesian inference.

approximate posteriorBayesian inferenceposterior recalibration

Adaptive A/B Tests and Simultaneous Treatment Parameter Optimization

Oct 13, 2022
YW
Yuhang Wu
🏛️ University of California, Berkeley | Amazon.com Inc

Classical algorithms for strongly convex stochastic optimization achieve fast convergence (O(1/√n)) but suffer from asymptotically non-negligible bias, violating the conditions required for a valid central limit theorem (CLT) and thus impeding asymptotically efficient statistical inference. Method: We propose the first dual-objective algorithm that simultaneously guarantees fast convergence and a provable CLT. Our approach integrates stochastic approximation, asymptotic statistical inference, and adaptive experimental design into a unified framework that ensures asymptotic normality of the estimator. Contribution/Results: We establish theoretical guarantees that the algorithm retains the O(1/√n) convergence rate while satisfying the CLT. Numerical experiments demonstrate substantial improvements over existing methods in estimation accuracy, confidence interval coverage, and identification of optimal treatment parameters. The method provides a new paradigm for continuous, parameterized A/B testing in online platforms—balancing optimization efficiency with statistical reliability.

Addresses non-vanishing bias in stochastic optimization statistical inferenceDevelops algorithm maintaining fast convergence with valid central limit theoremEnables reliable confidence intervals for optimal objective value estimation

Latest Papers

What's happening recently
View more

This work proposes a Bayesian group sequential design with dynamic information borrowing that reconciles the efficiency gains of incorporating historical data with the regulatory requirement of strict Type I error control. By establishing an explicit correspondence between posterior probability decision rules and uniformly most powerful frequentist tests, the method employs dual thresholds at each interim analysis: one to guarantee exact Type I error control and another to adaptively borrow strength from historical data when appropriate. This approach uniquely unifies frequentist error calibration with Bayesian information borrowing without sacrificing power. Numerical experiments demonstrate that the design maintains the nominal Type I error rate while substantially improving statistical power, and it has been successfully implemented in the design of a Phase III tuberculosis prevention trial integrating historical data from both adult and pediatric populations.

Bayesian analysisclinical trialsdynamic borrowing

Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.

Bayesian decision proceduresexperimental designgeneralized posteriors

This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.

distributional realignmenthigh-frequency monitoringintervention effect

Hot Scholars

KB

Keping Bi

Institute of Computing Technology, Chinese Academy of Sciences
Information Retrieval
JL

Jian Liang

Kuaishou Inc.
transfer learninggraph learning
OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
XC

Xueqi Cheng

Ph.D. student, Florida State University
Data miningLLMGNNComputational social science
PC

Pengzhou Cheng

Shanghai Jiao Tong University
Cyber SecurityArtificial Intelligence SecurityLLM ReasoningAgent