Score
Designs, builds, and evaluates explainability methods that tie model attributions and intermediate computations to physically interpretable causes, constraints, or quantum‑inspired representations so that explanations map to real-world physical processes. Produces uncertainty‑aware visualizations of decision maps, explicates mechanisms for mitigating observation/illumination biases, and develops robustness‑aware explanations that account for noise and measurement effects.
Low trustworthiness, difficulty in error diagnosis, and weak human-AI collaboration arise from the “black-box” nature of machine learning in physics. Method: This study systematically constructs the first multidimensional interpretability classification framework tailored to physical sciences, integrating philosophical reflection with technical practice. It unifies diverse interpretability approaches—including surrogate modeling, attention mechanisms, symbolic regression, causal inference, and physics-informed constraint embedding—across condensed matter, high-energy, astrophysical, and statistical physics. Contribution/Results: We identify a fundamental interpretability–performance trade-off law and establish verifiable, reproducible evaluation metrics and application guidelines. The framework elevates interpretable AI from a mere analytical tool to a core paradigm of scientific intelligence, enabling human-understandable, automated scientific discovery.
This study clarifies the conceptual confusion in physics-oriented machine learning between “interpretability”—referring to model transparency—and “explainability,” which denotes the capacity to map onto domain knowledge. It delineates the boundaries of these two notions and examines their trade-offs in terms of expressive power and adaptability. Through conceptual analysis and the construction of a unifying framework, complemented by a systematic review of both intrinsic and post-hoc explanation methods, the work advocates for integrating interpretability and explainability into scientific modeling paradigms. Crucially, it underscores the central role of task formulation and intervention design in model development. By establishing a clear conceptual foundation and methodological guidance, this research advances the principled integration of machine learning models with scientific reasoning in physics.
To address declining model interpretability and consequent trust deficits in complex machine learning systems, this paper proposes a trustworthy XAI framework grounded in uncertainty decoupling. Methodologically, it is the first to disentangle aleatoric from epistemic uncertainty and leverages epistemic uncertainty as a reliability metric for explanations: (i) as a rejection threshold for low-fidelity explanations, and (ii) as a dynamic signal to adapt explanation strategies—e.g., feature attribution or counterfactual generation—to local uncertainty conditions. The framework integrates Bayesian neural networks, rigorous uncertainty quantification, and state-of-the-art XAI techniques. Extensive experiments across diverse models—including traditional machine learning and deep neural networks—demonstrate substantial improvements in explanation stability and robustness. Crucially, it effectively filters unreliable explanations, thereby enhancing user trust in AI-driven decisions.
This study addresses the challenge that existing algorithmic explanations are often poorly understood and misapplied by non-expert users due to semantic ambiguity and insufficient contextual information, leading to a disconnect between explanations and actual decision-making. To bridge this gap, the authors propose an “Explanation Card” framework that augments widely used interpretability methods—such as SHAP and counterfactual explanations—with structured metadata specifying their applicability boundaries, robustness properties, and user-oriented interpretation guidance. By shifting explanatory responsibility from end users to explanation providers, this approach enhances the practical utility and regulatory compliance of model explanations, aligning with the transparency requirements of the EU AI Act. Empirical evaluations demonstrate that Explanation Cards significantly improve users’ comprehension accuracy of complex model explanations and effectively flag scenarios where explanations are unreliable, thereby facilitating trustworthy real-world deployment of algorithmic systems.
This study addresses the challenge of effective attribution in large-scale hybrid cyber-physical Internet-of-Things systems, where traditional causal explanation methods struggle due to their reliance on explicit directed graphs and difficulties handling feedback loops and partial observability. To overcome these limitations, this work proposes a statistical mechanics–inspired undirected energy-based modeling framework that captures dependency structures among variables and analyzes shifts in the energy landscape to enable structured attribution without reconstructing a causal graph. The approach introduces a novel energy-landscape–based dependency-aware mechanism capable of reasoning about perturbation effects in systems with mixed continuous-discrete variables. Experiments on an industrial IoT platform demonstrate that the method significantly outperforms state-of-the-art graph-based approaches in attribution accuracy, robustness, and scalability, making it well-suited for high-dimensional cyber-physical and socio-technical systems.
This work addresses the lack of aleatoric uncertainty attribution in predictive modeling and the limitations of existing explanation methods—namely their reliance on Bayesian modeling or generative auxiliaries, resulting in poor generalizability. We propose the first lightweight, plug-and-play framework for aleatoric uncertainty attribution. Methodologically, we model regression outputs as Gaussian distributions and directly perform interpretability analysis on the predicted variance, leveraging standard XAI tools (e.g., Grad-CAM, SHAP) to localize uncertainty-driving factors. Crucially, our approach requires no architectural modifications or Bayesian inference, greatly enhancing deployment flexibility and reliability. Experiments on synthetic and real-world tabular and image datasets demonstrate that our method consistently outperforms state-of-the-art complex baselines in explanation fidelity. To our knowledge, it is the first to achieve a principled unification of high interpretability and architectural simplicity.
This work addresses the limitation of existing deep learning approaches to uncertainty quantification, which struggle to distinguish between uncertainty arising from missing evidence (vacuity) and that stemming from conflicting evidence (dissonance), while also lacking spatial interpretability. To overcome this, the paper introduces the Uncertainty Activation Map (UAM) framework, which uniquely integrates the concepts of vacuity and dissonance from subjective logic with Full Gradient-based class activation mapping (FullGrad). By leveraging evidential deep learning, UAM generates spatially resolved visualizations of uncertainty that are both theoretically grounded and intuitively interpretable. The method effectively localizes the spatial origins of different uncertainty types across multiple benchmark datasets, offering an explainable visual feedback mechanism for assessing model reliability in complex vision tasks.
This work addresses the limitation of conventional post-hoc explainable AI methods, which produce deterministic attribution maps that fail to capture the inherent uncertainty in explanations derived from Bayesian neural networks—thereby hindering trustworthy decision-making in high-stakes scenarios. The paper introduces, for the first time, a formal notion of an “explanation distribution” and establishes a unified framework by pushing forward the Bayesian posterior through a Lipschitz-continuous attribution operator into explanation space. It further proposes a family of Uncertainty-Aware Relevance Attribution Operators (UA-RAO), enabling diverse statistical summaries such as means and quantiles, with both Monte Carlo tractability and theoretical guarantees via Wasserstein approximation. Evaluated on a 15-class power quality disturbance classification task, the integration of deep ensembles with UA-RAO significantly improves attribution localization accuracy, reveals uncertainty patterns invisible to point estimates, and demonstrates strong generalization on real-world signals.
This work addresses the lack of a unified theoretical foundation in interpretable machine learning, which has led to fragmented methodologies and inconsistent evaluation criteria. By introducing Lagrangian mechanics into this domain for the first time, the paper proposes a general theoretical framework grounded in user-oriented interpretability. Through systematic analysis of symmetries and constraints, the approach derives optimal interpretable models by minimizing a suitably defined Lagrangian. This deductive methodology not only unifies existing techniques under a coherent theoretical umbrella but also reveals novel research directions. It has successfully informed the design of core programming interfaces, mitigated limitations of current methods, and established a rigorous theoretical basis for interpretability education and interdisciplinary integration.
This work addresses the unreliability of existing medical image diagnosis models that often rely on non-causal or clinically irrelevant visual cues. To enhance trustworthiness, the authors propose a systematic framework that integrates explanation-aware loss directly into the end-to-end training objective by incorporating saliency-based interpretability supervision. A custom explanation loss function jointly optimizes diagnostic accuracy and spatial fidelity of model explanations. The study introduces two quantitative metrics—annotation coverage and saliency precision—to evaluate explanation quality and uncover the trade-off between explanation loss strength and model performance. Experiments on a chest X-ray dataset demonstrate that the proposed method achieves diagnostic accuracy comparable to baseline models while significantly improving spatial alignment between model-generated explanations and clinical annotations.
This study addresses the challenge of distinguishing whether observable patterns in latent variable models arise from genuine reasoning mechanisms or superficial correlations. The work conceptualizes “latent thought” as an intrinsic computational process rather than a post hoc interpretive construct and introduces a systematic evaluation framework combining causal interventions, low-rank geometric analysis, control model design, and latent state decoding. Findings reveal that similar patterns emerge even in control models lacking reasoning capabilities, and that the latent variables genuinely influencing behavior exhibit gradient effects concentrated within a low-dimensional subspace. The research underscores the necessity of integrating causal testing with controlled experimentation in interpretability analyses and establishes a new paradigm for rigorously evaluating internal model mechanisms.