🤖 AI Summary
This work addresses the interpretational difficulty in classical frequentist inference, where assigning a post-hoc coverage probability to a specific confidence interval is traditionally prohibited. The authors model the coverage event of a confidence interval as a Bernoulli random variable and treat the nominal confidence level \(1-\alpha\) as a probabilistic prediction for this event, evaluated via strictly proper scoring rules. Within a frequentist framework, they combine decision theory, pivotal quantities, and parameter-free statistics to prove that \(1-\alpha\) is the unique optimal constant prediction. Furthermore, in unbounded translation-invariant models, they construct improved predictive forms—conditioned on ancillary statistics such as relative interval width—that yield non-constant but superior conditional coverage probabilities. This approach preserves frequentist validity while offering a coherent post-hoc probabilistic interpretation of confidence intervals, thereby resolving the classic interpretational paradox.
📝 Abstract
What, if anything, should a frequentist say about a single realized confidence interval (CI) and its chance of having covered the parameter? Jerzy Neyman's original answer was to refuse any nondegenerate probability for coverage ex post and, instead, to "state that the interval covers". In this paper I argue that the usual frequentist machinery already supports a different reading. I treat the coverage event as a Bernoulli random variable, with the nominal level 1-alpha as its design-based success probability, and view "confidence" as a probability forecast for that Bernoulli outcome. Using strictly proper scoring rules, I show that 1-alpha is the unique optimal constant forecast for coverage, both before and after observing the data, and that it remains optimal post-trial in common unbounded, translation-invariant models with pivot-based CIs. When the design yields a theta-free statistic--such as the relative width of the interval in a finite-window uniform model--the conditional coverage given that statistic provides a nonconstant, design-based refinement of 1-alpha that strictly improves predictive performance. Two thought experiments, a Monty Hall-style shell game and the "lost submarine" example of Morey et al. (2016), illustrate how this perspective resolves familiar interpretational puzzles about CIs without appealing to priors or single-case subjective degrees of belief. I conclude with simple "what to do when you see an interval" guidance for applied work and some implications for teaching confidence intervals as tools for forecasting long-run coverage.
Keywords: Confidence intervals, coverage probability, proper scoring rules, probabilistic forecasting, frequentist inference
Disclaimer: The findings and conclusions in this report are those of the author and do not necessarily represent the official position of the Centers for Disease Control and Prevention