🤖 AI Summary
This work addresses the challenge of unreliable predictive uncertainty estimation in deep neural networks, which undermines their trustworthiness in safety-critical applications. The paper presents a systematic survey of uncertainty quantification methods, with a focus on ensemble and approximate Bayesian techniques, and introduces a decoupled “method–metric” framework that unifies the generation of predictive distributions and the aggregation of uncertainties. By integrating diverse approaches—including Bayesian neural networks, Monte Carlo Dropout, deep and efficient ensembles, single-forward methods, evidential networks, conformal prediction, and post-hoc calibration—the study establishes a unified taxonomy and evaluation benchmark. This enables a clear delineation of each method’s theoretical foundations, implementation strategies, empirical performance, and limitations, while also outlining promising directions for uncertainty research in large language models.
📝 Abstract
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.