evidential adversarial training

Designs and trains models that represent predictive uncertainty using evidential outputs (e.g., Dirichlet-based beliefs) and integrates adversarial training so those uncertainty estimates remain meaningful under input perturbations. This includes constructing evidence-based uncertainty losses and adversarial-example training schedules to preserve clean and robust accuracy and to support reliable selective classification when faced with attacks.

evidentialadversarialtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the often-overlooked degradation of predictive uncertainty quality caused by conventional adversarial training, which undermines selective classification performance despite improving model robustness. The study systematically reveals, for the first time, the adverse impact of adversarial training on uncertainty calibration and ranking. To mitigate this issue, the authors propose Evidence-based Adversarial Training (EV-AT), a novel approach grounded in evidential theory that jointly optimizes standard accuracy and uncertainty reliability in the Dirichlet parameter space. EV-AT employs an evidential loss combined with a robust evidential alignment loss to enforce consistency between predictions on clean and adversarial examples. Extensive experiments across multiple datasets and threat models demonstrate that EV-AT significantly outperforms existing methods, simultaneously enhancing both robust accuracy and selective classification performance, thereby advancing the Pareto frontier of the robustness–uncertainty trade-off.

adversarial robustnessdeep neural networkspredictive uncertainty

This work addresses the unreliable and unstable estimation of predictive uncertainty from softmax outputs of neural network classifiers, which adversely affects downstream task performance. The authors propose a novel approach that does not rely on evidential learning losses; instead, it explicitly estimates the parameters of a Dirichlet distribution by aggregating multiple softmax outputs and combining the method of moments with optional maximum likelihood optimization. This formulation effectively decouples uncertainty modeling from classification training. Evaluated across multiple datasets, the method significantly improves the quality of uncertainty estimates and achieves superior performance in tasks such as confidence calibration and selective classification, demonstrating both robustness and practical applicability.

Dirichlet modelingensemble methodsneural network classifiers

On the Robustness of Adversarial Training Against Uncertainty Attacks

Oct 29, 2024
EL
Emanuele Ledda
🏛️ Sapienza University of Rome | University of Genova | University of Cagliari

In safety-critical applications, adversarial attacks can maliciously manipulate model uncertainty estimates—over- or under-estimating confidence—thereby undermining decision reliability and system usability. This work is the first to theoretically and empirically demonstrate that standard adversarial training methods (e.g., TRADES, PGD) inherently improve the robustness of uncertainty estimation, obviating the need for dedicated uncertainty-specific defenses. We conduct a systematic evaluation across multiple adversarially robust models on CIFAR-10 and ImageNet using the RobustBench benchmark, showing substantial gains in resilience against uncertainty-targeted attacks: uncertainty calibration error decreases by up to 42% compared to standard models. We further provide theoretical analysis proving that this robustness arises from an implicit regularization effect induced by adversarial perturbations on the confidence margin—effectively tightening the bounds of predictive confidence. Our findings bridge adversarial robustness and reliable uncertainty quantification, offering a principled, unified approach to trustworthy AI in high-stakes settings.

Defending adversarial examples for trustworthy uncertainty measuresEnsuring robust uncertainty estimates against adversarial attacksEvaluating adversarial-robust models on security-sensitive applications

Guided Uncertainty Learning Using a Post-Hoc Evidential Meta-Model

Sep 29, 2025
CB
Charmaine Barker
🏛️ University of York

Deep learning models often exhibit overconfidence under distributional shift, and existing post-hoc calibration methods fail to fundamentally address this issue. This paper proposes a retraining-free meta-model calibration framework—Post-hoc Evidential Learning (PEL)—which freezes the backbone network, identifies discriminative regions via feature saliency analysis, and constructs a noise-driven curriculum to explicitly guide the model to learn *when* it is uncertain and *how* to quantify uncertainty. PEL introduces no architectural or parametric modifications to the original model; instead, it learns an uncertainty representation mechanism solely from calibration data. Evaluated across multiple benchmarks, PEL improves out-of-distribution detection and adversarial example detection performance by approximately 77% and 80%, respectively, significantly surpassing current state-of-the-art methods. The approach achieves high reliability while maintaining zero intrusiveness—preserving model integrity and deployment compatibility.

Existing post-hoc approaches fail to teach models when to be uncertainLightweight meta-model learns how and when to express uncertainty without retrainingReliable uncertainty quantification under distributional shift remains challenging

Deep learning models often produce overconfident yet incorrect predictions on out-of-distribution (OOD) or adversarial inputs, undermining reliability in safety-critical applications. To address the overconfidence of evidential deep learning (EDL) under input perturbations, we propose C-EDL—a lightweight, post-hoc uncertainty calibration framework. Its core innovation is a representation conflict-aware calibration mechanism: leveraging task-preserving input transformations to extract feature-level conflict signals, and integrating them into Dirichlet-distributed evidence modeling—enabling dynamic confidence recalibration without model retraining. Experiments across multiple datasets and attack types demonstrate that C-EDL significantly improves uncertainty quantification quality: OOD detection coverage decreases by 55%, adversarial coverage drops by 90%, while in-distribution accuracy remains stable. C-EDL consistently outperforms existing EDL variants and mainstream baselines in comprehensive uncertainty-aware evaluation metrics.

Enhances robustness against out-of-distribution and adversarial inputsImproves adversarial uncertainty quantification in Evidential Deep LearningReduces overconfident errors without retraining the model

Latest Papers

What's happening recently
View more

This work addresses the computational complexity and implementation challenges associated with computing Dirichlet expectation targets in evidential deep learning (EDL). To overcome these issues, the authors propose a first-order empirical risk minimization approximation based on a plug-in loss evaluated at the Dirichlet mean, which substantially simplifies uncertainty modeling and training procedures. Notably, this approach is the first to formally incorporate standard softmax classifiers into the EDL theoretical framework and introduces a general strategy for plug-in loss approximation. Experiments on the Google Speech Commands dataset demonstrate that the proposed method achieves predictive accuracy and selective prediction performance comparable to classical EDL while significantly reducing implementation complexity. Furthermore, it enables, for the first time in speech recognition tasks, an EDL-driven analysis of the trade-off between coverage and accuracy.

Dirichlet DistributionEvidential Deep LearningPlug-in Losses

This work addresses critical limitations in traditional evidential deep learning, where the KL penalty suppresses evidence only for negative classes, leading to uncontrolled evidence growth and degraded uncertainty quantification, while the common choice of Dirichlet parameter α = e + 1 lacks theoretical justification. To overcome these issues, the authors reformulate evidential deep learning through variational inference, proposing the first VI-EDL framework. They derive an evidence lower bound (ELBO) that effectively regularizes evidence magnitude and establish a generalization error bound, rigorously proving that α = e + 1 minimizes this bound. The method achieves state-of-the-art performance on standard vision and medical datasets and demonstrates superior uncertainty quantification in out-of-distribution detection, noise identification, and autonomous driving tasks.

Dirichlet distributionepistemic uncertaintyEvidential Deep Learning

This work addresses the lack of a unified theoretical framework in existing evidential deep learning (EDL) approaches, which obscures their intrinsic connections and design principles. By casting EDL within a generalized Bayesian perspective for the first time, this study elucidates the fundamental nature of distributional uncertainty and systematically dissects the interplay among prior specification, posterior updating, and training objectives. Building on this insight, the authors propose a unified and extensible Generalized Evidential Deep Learning (GEDL) framework. Through component-wise decoupling, GEDL not only integrates and generalizes existing methods but also achieves state-of-the-art performance in classification, uncertainty quantification, and out-of-distribution detection, all while enjoying a rigorous theoretical foundation.

Bayesian frameworkdistributional uncertaintyEvidential Deep Learning

This study addresses the overconfidence and uncertainty estimation failures of evidential deep learning under adversarial and out-of-distribution (OOD) inputs. To this end, it proposes CLEAR, a lightweight post-hoc calibration framework that requires no retraining. This method introduces a novel task-agnostic latent consistency mechanism that leverages calibration data to characterize the geometric structure of the latent space, dynamically rectifying evidence strength through perturbed view generation and conflict measurement. Consequently, CLEAR significantly enhances anomalous input detection without altering base predictions. Evaluated on the ImageNet-to-CUB shift, it improves OOD and adversarial AUROC by 8.29% and 5.01%, respectively, while achieving 17.4× faster inference than comparable methods and preserving multi-task performance.

Adversarial RobustnessEvidential Deep LearningLatent Consistency

This work addresses the challenge of unreliable predictive uncertainty estimation in deep neural networks, which undermines their trustworthiness in safety-critical applications. The paper presents a systematic survey of uncertainty quantification methods, with a focus on ensemble and approximate Bayesian techniques, and introduces a decoupled “method–metric” framework that unifies the generation of predictive distributions and the aggregation of uncertainties. By integrating diverse approaches—including Bayesian neural networks, Monte Carlo Dropout, deep and efficient ensembles, single-forward methods, evidential networks, conformal prediction, and post-hoc calibration—the study establishes a unified taxonomy and evaluation benchmark. This enables a clear delineation of each method’s theoretical foundations, implementation strategies, empirical performance, and limitations, while also outlining promising directions for uncertainty research in large language models.

Deep LearningPredictive ConfidenceSafety-Critical Systems

Hot Scholars

JH

Jin-Hee Cho

Computer Science Department, Virginia Tech
AI-based cybersecuritydecision making under uncertaintynetwork science
MF

Michael Felsberg

Professor of Computer Vision, Linköping University
Computer VisionMachine LearningRobot Vision
ZG

Zaiwang Gu

Institute for Infocomm Research, A*STAR, Singapore
object detectionmedical image analysis,
KM

Kira Maag

Heinrich-Heine-University Düsseldorf
Computer VisionDeep Learning