Score
Designs and implements predictive models that produce point estimates for continuous or discrete targets together with evidence-based probabilistic uncertainty quantified via evidential distributions (e.g., Dirichlet for categorical outputs or conjugate evidential priors for regression), yielding per-query uncertainty scores and calibrated confidence measures. Builds training objectives, inference procedures, and evaluation analyses for evidential learning (EDL), uncertainty calibration and robustness to degraded inputs, and integrates these uncertainty estimates into downstream decisions such as filtering, ranking, or risk-sensitive choices.
This work addresses the challenge in evidential deep learning of disentangling epistemic and aleatoric uncertainty under distributional shift, where standard approaches often exhibit overconfidence on out-of-distribution samples. The authors propose DIP-EDL, a novel method that explicitly decouples class prediction from uncertainty magnitude through a density-aware pseudo-count mechanism, modeling the conditional label distribution and marginal covariate density separately. Built upon a hierarchical Bayesian framework, DIP-EDL integrates amortized variational inference with Dirichlet parameterization to achieve, for the first time in evidential deep learning, asymptotically identifiable separation of the two uncertainty types. Experiments demonstrate that DIP-EDL significantly improves calibration, robustness, and interpretability on out-of-distribution data while preserving strong predictive performance in high-density regions.
Existing Evidential Deep Learning (EDL) relies on the standard Dirichlet distribution to model class-wise uncertainty, but its strong parametric assumptions hinder generalization and robustness under complex, long-tailed, and noisy data conditions. To address this, we propose Flexible Evidential Deep Learning (F-EDL), which replaces the standard Dirichlet with the generalized Dirichlet distribution—thereby relaxing structural constraints and enabling more adaptive, flexible modeling of class probability uncertainty. F-EDL employs a neural network to directly predict the parameters of the generalized Dirichlet, preserving computational efficiency while significantly improving robustness to distributional shifts. Extensive experiments demonstrate that F-EDL achieves state-of-the-art performance in uncertainty estimation across standard, long-tailed, and noisy benchmarks. Notably, it enhances model reliability and trustworthiness in high-stakes applications where calibrated uncertainty is critical.
This work addresses the unreliable and unstable estimation of predictive uncertainty from softmax outputs of neural network classifiers, which adversely affects downstream task performance. The authors propose a novel approach that does not rely on evidential learning losses; instead, it explicitly estimates the parameters of a Dirichlet distribution by aggregating multiple softmax outputs and combining the method of moments with optional maximum likelihood optimization. This formulation effectively decouples uncertainty modeling from classification training. Evaluated across multiple datasets, the method significantly improves the quality of uncertainty estimates and achieves superior performance in tasks such as confidence calibration and selective classification, demonstrating both robustness and practical applicability.
This work addresses the lack of a unified theoretical framework in existing evidential deep learning (EDL) approaches, which obscures their intrinsic connections and design principles. By casting EDL within a generalized Bayesian perspective for the first time, this study elucidates the fundamental nature of distributional uncertainty and systematically dissects the interplay among prior specification, posterior updating, and training objectives. Building on this insight, the authors propose a unified and extensible Generalized Evidential Deep Learning (GEDL) framework. Through component-wise decoupling, GEDL not only integrates and generalizes existing methods but also achieves state-of-the-art performance in classification, uncertainty quantification, and out-of-distribution detection, all while enjoying a rigorous theoretical foundation.
This work addresses the challenge of entangled uncertainty sources and the difficulty of disentangling pointwise statistical risk in predictive modeling. We propose a unified generative framework based on approximate Bayesian inference that, for the first time, establishes an explicit, interpretable decomposition linking pointwise statistical risk to two fundamental uncertainty types: aleatoric uncertainty (arising from inherent data noise) and epistemic uncertainty (stemming from model ignorance). The framework jointly generates multiple uncertainty measures while ensuring semantic consistency across them. Experiments on image benchmarks demonstrate significant improvements in out-of-distribution detection and misclassification identification, achieving higher AUROC scores compared to existing methods. Our approach thus provides robust, quantifiable uncertainty estimates essential for downstream uncertainty-aware tasks such as active learning, safe decision-making, and model debugging.
This work addresses critical limitations in traditional evidential deep learning, where the KL penalty suppresses evidence only for negative classes, leading to uncontrolled evidence growth and degraded uncertainty quantification, while the common choice of Dirichlet parameter α = e + 1 lacks theoretical justification. To overcome these issues, the authors reformulate evidential deep learning through variational inference, proposing the first VI-EDL framework. They derive an evidence lower bound (ELBO) that effectively regularizes evidence magnitude and establish a generalization error bound, rigorously proving that α = e + 1 minimizes this bound. The method achieves state-of-the-art performance on standard vision and medical datasets and demonstrates superior uncertainty quantification in out-of-distribution detection, noise identification, and autonomous driving tasks.
This work addresses the computational complexity and implementation challenges associated with computing Dirichlet expectation targets in evidential deep learning (EDL). To overcome these issues, the authors propose a first-order empirical risk minimization approximation based on a plug-in loss evaluated at the Dirichlet mean, which substantially simplifies uncertainty modeling and training procedures. Notably, this approach is the first to formally incorporate standard softmax classifiers into the EDL theoretical framework and introduces a general strategy for plug-in loss approximation. Experiments on the Google Speech Commands dataset demonstrate that the proposed method achieves predictive accuracy and selective prediction performance comparable to classical EDL while significantly reducing implementation complexity. Furthermore, it enables, for the first time in speech recognition tasks, an EDL-driven analysis of the trade-off between coverage and accuracy.
This work addresses the lack of explicit characterization of the interplay among information, reliability, and uncertainty in existing probabilistic forecast calibration methods. For any proper scoring rule, the authors propose the first general triple-decomposition framework grounded in information algebra and conditional entropy theory, which rigorously decomposes predictive loss into three distinct components: reliability (calibration error), information loss, and irreducible uncertainty. This framework uniquely quantifies the information loss incurred when mapping features to predictive scores and provides a unified interpretation of post-hoc calibration, model ensembling, and boosting strategies. In classification tasks, the approach is successfully applied to calibration evaluation, model aggregation, and staged training, clearly disentangling each component’s contribution to overall predictive uncertainty.
This work addresses a critical limitation in existing model calibration methods, which assess only the reliability of predicted probabilities and fail to evaluate whether the estimated epistemic uncertainty itself is trustworthy—particularly in second-order classification tasks. To bridge this gap, the paper introduces the notion of *cognitive calibration*, a stronger criterion than classical calibration, which measures whether a model’s reported epistemic uncertainty faithfully reflects the dispersion of its predictive distribution around the true label. The authors formalize an evaluation framework and propose the Expected Epistemic Calibration Error (EECE) as a consistent estimator of the true cognitive calibration error (TECE). This reveals failure modes invisible to conventional metrics and leads to an impossibility theorem. Empirical results demonstrate that cognitive calibration provides a coherent and meaningful evaluation standard, under which different uncertainty quantification methods exhibit markedly distinct behaviors despite similar predictive performance.
This work addresses the widespread lack of reliable confidence estimation in pretrained models and the high computational cost and poor compatibility of existing uncertainty quantification methods. To overcome these limitations, the authors propose ETN, a lightweight post-processing module that applies sample-dependent affine transformations in logit space and interprets the transformed outputs as parameters of a Dirichlet distribution. This approach enables, for the first time, the conversion of any pretrained model into an evidential model without retraining. Evaluated on both image classification and large language model question-answering tasks, ETN significantly outperforms existing post-hoc baselines in uncertainty estimation quality under both in-distribution and out-of-distribution settings, while introducing negligible computational overhead—thus achieving an effective balance among accuracy, efficiency, and deployment practicality.