Can we trust our models? Epistemic calibration in second-order classification

📅 2026-06-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in existing model calibration methods, which assess only the reliability of predicted probabilities and fail to evaluate whether the estimated epistemic uncertainty itself is trustworthy—particularly in second-order classification tasks. To bridge this gap, the paper introduces the notion of *cognitive calibration*, a stronger criterion than classical calibration, which measures whether a model’s reported epistemic uncertainty faithfully reflects the dispersion of its predictive distribution around the true label. The authors formalize an evaluation framework and propose the Expected Epistemic Calibration Error (EECE) as a consistent estimator of the true cognitive calibration error (TECE). This reveals failure modes invisible to conventional metrics and leads to an impossibility theorem. Empirical results demonstrate that cognitive calibration provides a coherent and meaningful evaluation standard, under which different uncertainty quantification methods exhibit markedly distinct behaviors despite similar predictive performance.
📝 Abstract
Uncertainty estimation is critical for deploying machine learning models in high-stakes settings. However, classical calibration only assesses the reliability of predicted probabilities and does not evaluate whether epistemic uncertainty estimates are themselves trustworthy. This limitation is particularly relevant for second-order classification models. We introduce epistemic calibration, a principled criterion that measures whether reported epistemic uncertainty faithfully reflects the dispersion of model predictions around the ground truth. We show that epistemic calibration is a strictly stronger notion than classical calibration and captures failure modes invisible to standard metrics. We relate this work to the existing literature through an impossibility theorem that holds under the epistemic calibration hypothesis. To operationalize this concept, we propose the Expected Epistemic Calibration Error (EECE), which we prove to be a consistent estimator of a True Epistemic Calibration Error (TECE). Experiments across a broad range of uncertainty quantification methods show that epistemic calibration is a coherent and meaningful criterion and reveal substantial differences across methods, despite similar predictive performance.
Problem

Research questions and friction points this paper is trying to address.

epistemic calibration
uncertainty estimation
second-order classification
model trustworthiness
calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

epistemic calibration
uncertainty quantification
second-order classification
Expected Epistemic Calibration Error
calibration