🤖 AI Summary
This work addresses the limitation of conventional post-hoc explainable AI methods, which produce deterministic attribution maps that fail to capture the inherent uncertainty in explanations derived from Bayesian neural networks—thereby hindering trustworthy decision-making in high-stakes scenarios. The paper introduces, for the first time, a formal notion of an “explanation distribution” and establishes a unified framework by pushing forward the Bayesian posterior through a Lipschitz-continuous attribution operator into explanation space. It further proposes a family of Uncertainty-Aware Relevance Attribution Operators (UA-RAO), enabling diverse statistical summaries such as means and quantiles, with both Monte Carlo tractability and theoretical guarantees via Wasserstein approximation. Evaluated on a 15-class power quality disturbance classification task, the integration of deep ensembles with UA-RAO significantly improves attribution localization accuracy, reveals uncertainty patterns invisible to point estimates, and demonstrates strong generalization on real-world signals.
📝 Abstract
Post-hoc explainable AI (XAI) methods typically produce deterministic attribution maps, whereas Bayesian neural networks (BNNs) induce a distribution over explanations. Capturing the variability of this distribution is important for uncertainty-aware decision-making. This paper formalises the \emph{explanation distribution} as the push-forward measure of the BNN posterior through any Lipschitz-continuous attribution operator. It further proposes the uncertainty-aware relevance attribution operator (UA-RAO), a general family of operators that summarises the explanation distribution using the mean, variance, coefficient of variation, quantiles, and set-theoretic aggregation measures. Theoretical support is provided through Monte Carlo accessibility and Wasserstein approximation bounds. The framework is evaluated on a 15-class power quality disturbance (PQD) classification benchmark, comparing three BNN approximations paired with three attribution operators using relevance mass accuracy and intersection-over-union as localisation metrics. Results show that deep ensembles with the mean UA-RAO improve localisation over the deterministic baseline, while other UA-RAO summaries reveal uncertainty patterns absent from point-estimate attributions. Qualitative results on measured signals further suggest that these patterns generalise beyond the synthetic training distribution. The framework is domain-agnostic and can be applied to any BNN paired with a Lipschitz-continuous attribution operator.