🤖 AI Summary
Standard likelihood-based inference can yield overconfident, vanishingly narrow confidence intervals under model misspecification, leading to unreliable uncertainty quantification. This work proposes a "well-calibrated likelihood" approach that dynamically rescales inference by dividing the likelihood ratio statistic by a goodness-of-fit statistic evaluated at the best-fitting point. The method recovers classical results when the model is correctly specified and, under misspecification, produces confidence intervals that honestly reflect the precision limit imposed by model bias. It is the first framework to integrate goodness-of-fit directly into likelihood inference and supports both binned and unbinned implementations—the latter leveraging classifier two-sample tests. Empirical validation on Gaussian models and strong coupling constant estimation from simulated electron-positron collision data demonstrates its robustness, achieving adaptive and reliable uncertainty quantification in the presence of model misspecification.
📝 Abstract
Likelihood-based inference in particle physics, and in the physical sciences more broadly, relies on the assumption that the model accurately describes the data. When the model is misspecified, though, standard confidence intervals shrink to zero width with increasing data, producing overconfident and potentially misleading constraints. We propose the well-tempered likelihood, which divides the likelihood-ratio test statistic by a goodness-of-fit (GOF) statistic evaluated at the best-fit point. Under correct specification, the GOF is $\mathcal{O}(1)$ and standard inference is recovered. Under misspecification, both the likelihood and the GOF scale as $\mathcal{O}(N)$ for $N$ data points, so their ratio remains $\mathcal{O}(1)$ and the resulting well-tempered confidence interval self-limits at a floor determined by the model's inadequacy, essentially reducing the effective sample size. In other words, \textit{all models are correct, as long as your dataset is small enough}. We present binned and unbinned formulations of the well-tempered likelihood---the latter based on a classifier two-sample test---and demonstrate our method on a Gaussian example and on a measurement of the strong coupling constant using synthetic electron-positron collisions.