🤖 AI Summary
This study addresses the computational bottleneck in evaluating the calibration of nested uncertainty sets within expensive simulation models. To overcome this limitation, the work proposes an efficient calibration assessment method grounded in a Bayesian framework and the Dirichlet-Multinomial model. By exploiting the nested structure, the approach directly processes interval outputs without requiring access to the full predictive distribution. Furthermore, it incorporates Bayes factor testing for statistical inference, substantially reducing the number of independent simulations needed. The proposed method successfully detects model miscalibration in data assimilation tasks under limited simulation budgets. Overall, this work significantly lowers computational costs while demonstrating both the effectiveness and practical utility of the proposed approach for calibrating complex simulation systems.
📝 Abstract
We present a Bayesian framework for assessing the calibration of nested uncertainty sets from simulation studies. Calibration is typically evaluated by repeatedly simulating data from a known model and comparing empirical coverage probabilities with their nominal values, an approach that often requires hundreds or thousands of independent model evaluations to obtain reliable estimates. Such computational demands can be prohibitive when the forward models are expensive to evaluate, for example in data assimilation using a large Earth-system model, or may not be applicable, for exmaple when the output of the simulation is a single or multiple uncertainty sets rather than a full distribution. We present a method that leverages the nested structure of uncertainty sets across nominal coverage levels to assess the calibration of these models. The method, which is based on a Dirichlet-Multinomial model, allows for testing both exact and region-based notions of model calibration in an efficient manner using Bayes factors. Unlike methods that require a full predictive distribution, this method works for procedures that directly produce prediction or confidence intervals. We demonstrate the method on a data assimilation problem containing model form error and show that it is capable of detecting miscalibration even when there are only a limited number of simulation studies available.