Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical misalignment between Bayesian optimal strategies and commonly used uncertainty scores in tasks such as out-of-distribution detection and active learning. The authors propose a framework based on epistemic reject options, formulating selective prediction as a constrained optimization problem balancing coverage, expected risk, and regret. They introduce the attainable risk–regret–coverage surface as a joint diagnostic metric to disentangle uncertainty sources and evaluate their practical utility. Through decision-theoretic analysis, uncertainty disentanglement, optimization of selective prediction, and benchmarking on densely human-annotated datasets, the study reveals that rankings derived from decision theory often differ substantially—and sometimes inversely—from those obtained via conventional proxy tasks, demonstrating that traditional correlation-based measures fail to accurately reflect real-world effectiveness.
📝 Abstract
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
Problem

Research questions and friction points this paper is trying to address.

epistemic uncertainty
out-of-distribution detection
active learning
uncertainty disentanglement
selective prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

epistemic uncertainty
selective prediction
regret minimization
uncertainty disentanglement
decision-theoretic evaluation