🤖 AI Summary
Existing methods struggle to consistently define and accurately evaluate the separation between aleatoric and epistemic uncertainty, often relying on imperfect proxy tasks due to the absence of ground-truth uncertainty targets. This work proposes a unified definition of uncertainty as the pointwise posterior risk—the expected loss of a predictor with respect to the true function distribution given observed data—thereby integrating Bayesian functional uncertainty with estimation bias. Building on this formulation, we introduce the first semi-synthetic benchmark that provides direct access to ground-truth uncertainty targets, eliminating dependence on proxy tasks. Experiments reveal that predictive accuracy does not necessarily correlate with uncertainty reliability, enabling clear identification of methods aligned with true uncertainty while exposing their sensitivity to data and modeling choices.
📝 Abstract
Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric uncertainty, yet these uncertainty types are not defined consistently across the literature, making it difficult to assess whether a method produces accurate uncertainty estimates. Evaluation is further complicated by the fact that ground-truth epistemic uncertainty is typically unavailable. Existing benchmarks therefore mostly rely on proxy tasks such as out-of-distribution detection, which do not provide complete ground-truth uncertainty targets and offer limited insight into the structure and quality of uncertainty estimates. We propose a unified definition of uncertainty as pointwise posterior risk, the expected loss of a predictor under the distribution of plausible ground-truth functions given the data. This view combines Bayesian uncertainty over functions with estimator-dependent deviations from the posterior mean, capturing effects such as misspecification and optimization error. This formulation constitutes the foundation of a theory-backed benchmark that enables direct computation of oracle epistemic and aleatoric uncertainty using semi-synthetic datasets with real covariates and known generative processes. By avoiding proxy evaluations, the benchmark enables fine-grained analysis of uncertainty estimates. Empirically, we find that accurate prediction does not guarantee reliable uncertainty disentanglement. The benchmark reveals practically useful differences between methods, identifying approaches with meaningful alignment to oracle uncertainty targets while exposing sensitivity to datasets and modeling choices.