Score
Designs and trains models that represent predictive uncertainty using evidential outputs (e.g., Dirichlet-based beliefs) and integrates adversarial training so those uncertainty estimates remain meaningful under input perturbations. This includes constructing evidence-based uncertainty losses and adversarial-example training schedules to preserve clean and robust accuracy and to support reliable selective classification when faced with attacks.
Deep learning models frequently produce high-confidence mispredictions in safety-critical applications, while existing uncertainty quantification (UQ) methods lack principled differentiation between data uncertainty (aleatoric) and model uncertainty (epistemic), hindering informed method selection. Method: We propose the first UQ taxonomy grounded in a dual-dimensional uncertainty source framework—explicitly distinguishing data- versus model-related uncertainty—and systematically evaluate major UQ paradigms—including Bayesian approximations, ensembles, calibration techniques, and deep generative models—characterizing their modeling assumptions, computational costs, and applicability domains. Contribution/Results: We establish formal mapping principles linking UQ methods to downstream tasks (e.g., active learning, robust decision-making, reinforcement learning) and introduce a structured evaluation matrix that bridges the “source–method–task” triad—a gap unaddressed in prior surveys. This work delivers an interpretable, deployment-ready UQ guidance framework for high-stakes AI systems.
This work addresses the often-overlooked degradation of predictive uncertainty quality caused by conventional adversarial training, which undermines selective classification performance despite improving model robustness. The study systematically reveals, for the first time, the adverse impact of adversarial training on uncertainty calibration and ranking. To mitigate this issue, the authors propose Evidence-based Adversarial Training (EV-AT), a novel approach grounded in evidential theory that jointly optimizes standard accuracy and uncertainty reliability in the Dirichlet parameter space. EV-AT employs an evidential loss combined with a robust evidential alignment loss to enforce consistency between predictions on clean and adversarial examples. Extensive experiments across multiple datasets and threat models demonstrate that EV-AT significantly outperforms existing methods, simultaneously enhancing both robust accuracy and selective classification performance, thereby advancing the Pareto frontier of the robustness–uncertainty trade-off.
This work addresses the unreliable and unstable estimation of predictive uncertainty from softmax outputs of neural network classifiers, which adversely affects downstream task performance. The authors propose a novel approach that does not rely on evidential learning losses; instead, it explicitly estimates the parameters of a Dirichlet distribution by aggregating multiple softmax outputs and combining the method of moments with optional maximum likelihood optimization. This formulation effectively decouples uncertainty modeling from classification training. Evaluated across multiple datasets, the method significantly improves the quality of uncertainty estimates and achieves superior performance in tasks such as confidence calibration and selective classification, demonstrating both robustness and practical applicability.
In safety-critical applications, adversarial attacks can maliciously manipulate model uncertainty estimates—over- or under-estimating confidence—thereby undermining decision reliability and system usability. This work is the first to theoretically and empirically demonstrate that standard adversarial training methods (e.g., TRADES, PGD) inherently improve the robustness of uncertainty estimation, obviating the need for dedicated uncertainty-specific defenses. We conduct a systematic evaluation across multiple adversarially robust models on CIFAR-10 and ImageNet using the RobustBench benchmark, showing substantial gains in resilience against uncertainty-targeted attacks: uncertainty calibration error decreases by up to 42% compared to standard models. We further provide theoretical analysis proving that this robustness arises from an implicit regularization effect induced by adversarial perturbations on the confidence margin—effectively tightening the bounds of predictive confidence. Our findings bridge adversarial robustness and reliable uncertainty quantification, offering a principled, unified approach to trustworthy AI in high-stakes settings.
Deep learning models often exhibit overconfidence under distributional shift, and existing post-hoc calibration methods fail to fundamentally address this issue. This paper proposes a retraining-free meta-model calibration framework—Post-hoc Evidential Learning (PEL)—which freezes the backbone network, identifies discriminative regions via feature saliency analysis, and constructs a noise-driven curriculum to explicitly guide the model to learn *when* it is uncertain and *how* to quantify uncertainty. PEL introduces no architectural or parametric modifications to the original model; instead, it learns an uncertainty representation mechanism solely from calibration data. Evaluated across multiple benchmarks, PEL improves out-of-distribution detection and adversarial example detection performance by approximately 77% and 80%, respectively, significantly surpassing current state-of-the-art methods. The approach achieves high reliability while maintaining zero intrusiveness—preserving model integrity and deployment compatibility.
Deep learning models often produce overconfident yet incorrect predictions on out-of-distribution (OOD) or adversarial inputs, undermining reliability in safety-critical applications. To address the overconfidence of evidential deep learning (EDL) under input perturbations, we propose C-EDL—a lightweight, post-hoc uncertainty calibration framework. Its core innovation is a representation conflict-aware calibration mechanism: leveraging task-preserving input transformations to extract feature-level conflict signals, and integrating them into Dirichlet-distributed evidence modeling—enabling dynamic confidence recalibration without model retraining. Experiments across multiple datasets and attack types demonstrate that C-EDL significantly improves uncertainty quantification quality: OOD detection coverage decreases by 55%, adversarial coverage drops by 90%, while in-distribution accuracy remains stable. C-EDL consistently outperforms existing EDL variants and mainstream baselines in comprehensive uncertainty-aware evaluation metrics.
This work addresses the computational complexity and implementation challenges associated with computing Dirichlet expectation targets in evidential deep learning (EDL). To overcome these issues, the authors propose a first-order empirical risk minimization approximation based on a plug-in loss evaluated at the Dirichlet mean, which substantially simplifies uncertainty modeling and training procedures. Notably, this approach is the first to formally incorporate standard softmax classifiers into the EDL theoretical framework and introduces a general strategy for plug-in loss approximation. Experiments on the Google Speech Commands dataset demonstrate that the proposed method achieves predictive accuracy and selective prediction performance comparable to classical EDL while significantly reducing implementation complexity. Furthermore, it enables, for the first time in speech recognition tasks, an EDL-driven analysis of the trade-off between coverage and accuracy.
This work addresses critical limitations in traditional evidential deep learning, where the KL penalty suppresses evidence only for negative classes, leading to uncontrolled evidence growth and degraded uncertainty quantification, while the common choice of Dirichlet parameter α = e + 1 lacks theoretical justification. To overcome these issues, the authors reformulate evidential deep learning through variational inference, proposing the first VI-EDL framework. They derive an evidence lower bound (ELBO) that effectively regularizes evidence magnitude and establish a generalization error bound, rigorously proving that α = e + 1 minimizes this bound. The method achieves state-of-the-art performance on standard vision and medical datasets and demonstrates superior uncertainty quantification in out-of-distribution detection, noise identification, and autonomous driving tasks.
This work addresses the lack of a unified theoretical framework in existing evidential deep learning (EDL) approaches, which obscures their intrinsic connections and design principles. By casting EDL within a generalized Bayesian perspective for the first time, this study elucidates the fundamental nature of distributional uncertainty and systematically dissects the interplay among prior specification, posterior updating, and training objectives. Building on this insight, the authors propose a unified and extensible Generalized Evidential Deep Learning (GEDL) framework. Through component-wise decoupling, GEDL not only integrates and generalizes existing methods but also achieves state-of-the-art performance in classification, uncertainty quantification, and out-of-distribution detection, all while enjoying a rigorous theoretical foundation.
This study addresses the overconfidence and uncertainty estimation failures of evidential deep learning under adversarial and out-of-distribution (OOD) inputs. To this end, it proposes CLEAR, a lightweight post-hoc calibration framework that requires no retraining. This method introduces a novel task-agnostic latent consistency mechanism that leverages calibration data to characterize the geometric structure of the latent space, dynamically rectifying evidence strength through perturbed view generation and conflict measurement. Consequently, CLEAR significantly enhances anomalous input detection without altering base predictions. Evaluated on the ImageNet-to-CUB shift, it improves OOD and adversarial AUROC by 8.29% and 5.01%, respectively, while achieving 17.4× faster inference than comparable methods and preserving multi-task performance.
This work addresses the challenge of unreliable predictive uncertainty estimation in deep neural networks, which undermines their trustworthiness in safety-critical applications. The paper presents a systematic survey of uncertainty quantification methods, with a focus on ensemble and approximate Bayesian techniques, and introduces a decoupled “method–metric” framework that unifies the generation of predictive distributions and the aggregation of uncertainties. By integrating diverse approaches—including Bayesian neural networks, Monte Carlo Dropout, deep and efficient ensembles, single-forward methods, evidential networks, conformal prediction, and post-hoc calibration—the study establishes a unified taxonomy and evaluation benchmark. This enables a clear delineation of each method’s theoretical foundations, implementation strategies, empirical performance, and limitations, while also outlining promising directions for uncertainty research in large language models.