Score
Designs, builds, and evaluates methods and tools that generate, visualize, and quantify diagnostic, post-hoc explanations for machine learning models — including local saliency and class-activation maps (e.g., Grad-CAM), embedding-space diagnostics, finite-sample uncertainty estimates, and overlay visualizations — and constructs evaluation metrics and benchmarks to compare explainability methods. Analyzes those explanations to localize model-relevant regions or components, quantify the significance and uncertainty of saliency regions, summarize findings for developers, and recommend targeted repair actions.
This work addresses the lack of a unified framework in existing feature attribution methods, which leads to opaque assumptions, incomparable results, and susceptibility to failure modes. The authors propose the first unified mathematical framework for locally additive attributions, systematically integrating Shapley values, path integrals, gradient-based methods, perturbation approaches, and CAM-style techniques through five core dimensions: value functions, reference points, paths, perturbation distributions, and conservation rules. Through axiomatic analysis, comparative matrices, and formal modeling, the study elucidates how attribution outcomes depend critically on underlying assumptions and establishes causal links between methodological choices and characteristic failure modes. To enhance rigor, the paper concludes with a ten-item reporting checklist designed to substantially improve the transparency, reproducibility, and reliability of attribution research.
Large vision models suffer from poor interpretability, and conventional attribution methods merely localize attended regions without providing semantic explanations. Method: This paper proposes a dual-path framework—“aligning with human reasoning” and “concept-level explanation”—featuring the first formally guaranteed attribution method based on verification-aware perturbation analysis (EVA), and the CRAFT-MACO collaborative framework for automatic discovery, importance quantification, and visualization of internal model concepts. It integrates algorithmic stability metrics, Sobol indices computed via quasi-Monte Carlo sampling, 1-Lipschitz function-space optimization, and concept activation mapping. Results: The approach enables interactive, concept-level explanations across all 1,000 ImageNet classes on ResNet, significantly improving explanation fidelity and human consistency—particularly in complex scenes—while offering rigorous theoretical guarantees and actionable semantic insights.
This work addresses the lack of unified, cross-type evaluation for input feature attribution methods. We propose the first automated, horizontal evaluation framework covering token-level, token-interaction-level, and span-interaction-level explanation methods. Grounded in four diagnostic properties—faithfulness, stability, interpretability, and robustness—the framework integrates diverse techniques including Shapley values, Integrated Gradients, Bivariate Shapley, attention mechanisms, and Louvain-based Span Interactions. It is systematically validated on two benchmark datasets (SST-2 and BoolQ) and two foundational models (BERT and RoBERTa). Results demonstrate that span-interaction explanations significantly outperform conventional approaches across most metrics, revealing their previously underappreciated potential; moreover, the three explanation types exhibit complementary strengths. This study establishes a reproducible, principled benchmark for scientifically selecting and improving attribution methods.
This paper addresses the evaluation of explanation quality in image classification heatmaps, proposing a novel assessment paradigm that jointly considers accuracy and stability. For accuracy, we introduce the Weighting Game metric—a first-of-its-kind quantitative measure evaluating how well class-relevant explanations align with ground-truth semantic segmentation masks. For stability, we design a geometric transformation–based similarity comparison method using scaling and translation to quantify heatmap robustness against input perturbations. Our framework integrates class activation mapping (CAM), segmentation mask matching, and statistical analysis, and is systematically evaluated across mainstream CAM methods and diverse model architectures. Experiments reveal that explanation quality is strongly architecture-dependent, providing empirical guidance for selecting appropriate interpretability methods. This work advances heatmap evaluation from qualitative inspection toward a reproducible, comparable, and quantitative paradigm.
This study evaluates the faithfulness and localization reliability of Grad-CAM for interpreting lung cancer classification in chest CT scans. It presents the first systematic comparison between convolutional architectures (ResNet, DenseNet, EfficientNet) and Vision Transformers (ViT) in terms of explanation consistency, localization accuracy, and robustness to perturbations, thereby elucidating how attention mechanisms influence interpretability. The findings reveal that Grad-CAM yields robust explanations in convolutional models but suffers from distorted visualizations in ViT due to its non-local attention, with significant inter-model discrepancies in localization performance—raising concerns about its clinical generalizability. To address these issues, this work proposes a model-aware interpretability evaluation framework, offering a new perspective toward the trustworthy deployment of medical AI systems.
Existing explanation methods for image classifiers lack rigorous formal definitions of causality and explanation, relying predominantly on heuristic strategies. Method: This paper introduces the Halpern–Pearl theory of actual causality to black-box image classification interpretability—the first systematic application of this causal framework to the domain. We propose REX, a causally grounded explanation generation framework that formally defines “cause” and “explanation,” designs a provably terminating algorithm for approximating minimal explanations, and implements an iterative solving mechanism with controllable computational complexity. Contribution/Results: The implemented tool REX outperforms state-of-the-art black-box explanation methods across explanation compactness, computational efficiency, and standard quality metrics (e.g., fidelity, stability, and comprehensibility). Experiments demonstrate that REX produces the most concise explanations and achieves the fastest convergence. This work establishes a rigorous causal foundation for explainable AI while delivering a practical, scalable technical solution.
This work addresses the unreliability of existing medical image diagnosis models that often rely on non-causal or clinically irrelevant visual cues. To enhance trustworthiness, the authors propose a systematic framework that integrates explanation-aware loss directly into the end-to-end training objective by incorporating saliency-based interpretability supervision. A custom explanation loss function jointly optimizes diagnostic accuracy and spatial fidelity of model explanations. The study introduces two quantitative metrics—annotation coverage and saliency precision—to evaluate explanation quality and uncover the trade-off between explanation loss strength and model performance. Experiments on a chest X-ray dataset demonstrate that the proposed method achieves diagnostic accuracy comparable to baseline models while significantly improving spatial alignment between model-generated explanations and clinical annotations.
Current visual explanations for skin disease classification models lack systematic, clinically relevant evaluation. This work proposes an explainability assessment framework grounded in large language models (LLMs), integrating Grad-CAM heatmaps with state-of-the-art LLMs—including GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6—guided by progressive, structured prompt engineering to evaluate both lesion localization accuracy and explanation credibility. For the first time, a domain-adapted LLM-as-a-Judge paradigm is introduced into dermatological AI explainability assessment. Validated on models such as EfficientNet-B0, MobileNetV3, and ResNet18, the approach significantly enhances alignment between automated evaluations and clinical judgments, enabling quantifiable and standardized measurement of explanation quality.
This work addresses the lack of a systematic synthesis and unified evaluation framework for class activation mapping (CAM) methods across diverse model architectures and eras. Surveying 57 foundational studies since 2016, it proposes a cohesive taxonomy grounded in attribution mechanisms, architectural dependencies, and evaluation objectives, encompassing visualization techniques from CNNs to Transformers and foundation models such as CLIP, DINO, and SAM. The study traces the evolution of CAM from single-layer, low-resolution explanations toward multi-layer, probabilistic, token-aware, and foundation-model-aware paradigms. It systematically summarizes key contributions and persistent challenges across method categories and highlights the current fragmentation in evaluation protocols, offering methodological guidance and a benchmark reference for future research on interpretable vision models.
This study addresses a critical gap in clustering interpretability: existing post-hoc explanation methods primarily focus on feature importance or instance-level explanations and struggle to reliably uncover structured patterns within clusters. To systematically evaluate this limitation, the authors conduct the first controlled assessment of multiple explanation techniques—including random forest permutation importance, LIME, and principal component analysis—in synthetic datasets where ground-truth structured patterns are explicitly embedded. Results demonstrate that while these methods partially recover relevant features, none consistently identifies all types of predefined patterns. This reveals a fundamental shortcoming of current interpretability tools in capturing pattern-level cluster structure and underscores the urgent need for dedicated methods designed specifically for detecting and explaining such intra-cluster patterns.