Score
Designs and implements pipelines to extract, visualise, and analyse attention, activation and saliency maps (including CAM/Grad-CAM, attention rollout, RISE and token- or spatial-level maps). Builds quantitative and qualitative evaluation methods — e.g., IoU/localisation agreement, activation overlap, cross-method and cross-architecture comparisons, and validation of explanations for cascaded or composite outputs — to assess interpretability and fidelity of those maps.
This study systematically investigates the consistency of evaluation metrics for XAI saliency maps—specifically LIME, Grad-CAM, and Guided Backpropagation. Through a large-scale user study (N=166), it conducts the first cross-paradigm, unified comparison of three distinct evaluation frameworks: subjective trust, objective improvement in model understanding, and quantitative mathematical metrics. Results reveal substantial inconsistency across these frameworks: Grad-CAM most effectively enhances users’ objective understanding of model behavior, while Guided Backpropagation achieves the highest scores on mathematical fidelity metrics; notably, several widely used mathematical metrics exhibit significant negative correlations with actual user understanding—challenging the prevailing assumption that high mathematical scores imply superior explainability. The study identifies paradigmatic fragmentation as a core issue in XAI evaluation and provides empirical evidence and methodological insights for developing human-centered, multi-dimensional evaluation frameworks grounded in both cognitive validity and technical rigor.
This study addresses the lack of standardized implementation of Grad-CAM in Vision Transformers (ViTs), which has led to ambiguous formulations, poor reproducibility, and inconsistent interpretations. Through a systematic review of 175 relevant works, this paper proposes the first descriptive taxonomy for Grad-CAM variants tailored to ViTs, clarifying implicit assumptions and inconsistencies across critical components—namely feature token selection, gradient target specification, spatial reconstruction, and aggregation strategies. The analysis reveals that most existing studies inadequately document implementation details, demonstrating that the adaptation of Grad-CAM to ViTs is far from a trivial extension of its CNN-based counterpart. By establishing a clear and rigorous methodological framework, this work aims to enhance transparency, reproducibility, and reliability in explainable artificial intelligence for transformer-based vision models.
Existing Class Activation Mapping (CAM) methods lack explicit modeling of user intent and domain knowledge, making it difficult to flexibly control heatmap properties such as localization accuracy, faithfulness, and robustness. Method: We propose SyCAM, the first metric-driven framework for automatic synthesis of CAM expressions. Under formal grammar constraints, SyCAM jointly optimizes multiple objectives—including Intersection-over-Union (IoU), faithfulness, and robustness—via differentiable symbolic expression search and grammar-guided program synthesis, generating interpretable and customizable activation expressions. Contribution/Results: SyCAM breaks away from rigid, hand-crafted CAM formulas, enabling on-demand customization of heatmap behavior. Evaluated on ResNet50, VGG16, and VGG19, it achieves an average 12.7% improvement across target metrics. The framework significantly enhances flexibility, controllability, and interpretability of heatmap generation, establishing a new paradigm for adaptive, user-aligned visual explanation.
This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.
This study addresses the lack of systematic and empirically validated evaluation methods for attention maps in speaker recognition tasks. To this end, we propose a modified Random Input Sampling for Evaluation (Modified RISE-eval) algorithm that enables more accurate quantification of attention map quality and facilitates a systematic comparison between GradCAM and LayerCAM under varying conditions. Experimental results demonstrate that both methods exhibit distinct strengths, and the proposed evaluation framework effectively uncovers their underlying decision-making rationales. The findings substantiate the practical utility and reliability of our approach in assessing attention mechanisms within speaker recognition systems.
This work addresses the lack of a systematic synthesis and unified evaluation framework for class activation mapping (CAM) methods across diverse model architectures and eras. Surveying 57 foundational studies since 2016, it proposes a cohesive taxonomy grounded in attribution mechanisms, architectural dependencies, and evaluation objectives, encompassing visualization techniques from CNNs to Transformers and foundation models such as CLIP, DINO, and SAM. The study traces the evolution of CAM from single-layer, low-resolution explanations toward multi-layer, probabilistic, token-aware, and foundation-model-aware paradigms. It systematically summarizes key contributions and persistent challenges across method categories and highlights the current fragmentation in evaluation protocols, offering methodological guidance and a benchmark reference for future research on interpretable vision models.
研究通过乳腺MRI案例探讨了基于显著性图的深度学习解释方法可能误导临床医生的问题,评估了多种显著性方法,并指出当前方法存在的挑战及需要更稳健、标准化评估框架的需求。
Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.
The mechanisms underlying the emergence of semantic structure during diffusion model generation remain poorly understood, and existing approaches struggle to simultaneously capture the dynamic evolution of attention across both spatial and temporal dimensions. This work proposes a novel visual analytics framework that, for the first time, integrates timestep-indexed token-level cross-attention maps with data-driven phase identification. By combining time-series clustering, quantitative attention metrics, and interactive visualization, the framework enables structured analysis of attention dynamics in Stable Diffusion–like models. Evaluated on a benchmark of 60 structured prompts, it reveals interpretable patterns of attention evolution, effectively supporting human-in-the-loop understanding and control of the generative process.
This study addresses the significant performance degradation of current vision-language models in emergency scenarios under visually degraded conditions such as smoke, haze, and thermal imaging, where both grounding and visual question answering capabilities deteriorate markedly and exhibit model-dependent responses to linguistic feedback. To systematically evaluate robustness, the authors introduce the RefCOCO-Degraded dataset and assess prominent multimodal large language models—including Gemini, Qwen2-VL, and BLIP-2—across four degradation types. The work reveals two key findings: the previously unreported “thermal imaging paradox” and BLIP-2’s heightened susceptibility to hallucination under image degradation. Experimental results demonstrate that iterative language feedback improves Gemini’s performance in thermal imaging by 47.3%, whereas Qwen2-VL suffers a 5.1% decline, and further uncover that standard cropping strategies can severely fail in thermal settings.