Score
Design methods and tools that compute, aggregate, analyze, and visualize saliency maps or feature-importance scores for inputs and models, using gradient-based approaches (e.g., gradients, integrated gradients), model-based techniques, or dimensionality-reduction methods (PCA, LDA), and that can produce instance-level or universal saliency predictions such as attention or gaze-like maps. Build pipelines that aggregate saliency across scenes or datasets, evaluate and quantify spatially concentrated discriminative evidence, and apply those saliency outputs to guide selection of influential input regions and to inform saliency-guided, saliency-aware, or saliency-explored training and model refinement.
Saliency maps are widely used for visual explanation in deep learning, yet a lack of consensus between their intended explanatory purpose and users’ query types leads to inaccurate evaluation and limited applicability. Method: We propose the Reference Frame × Granularity (RF×G) taxonomy—the first systematic framework disentangling explanation intent (pointwise vs. contrastive) from semantic granularity (class-level vs. group-level)—and uncover cognitive biases inherent in existing faithfulness metrics. Based on this theory, we design four novel faithfulness evaluation metrics and establish a cross-dimensional benchmark across ten saliency methods, four model architectures, and three datasets. Results: Experiments reveal significant deficiencies in mainstream methods for contrastive explanation and group-level semantic modeling. This work provides an interpretable AI framework aligned with human cognition and delivers reproducible, theoretically grounded evaluation tools.
Deep learning excels in image analysis but suffers from poor interpretability, hindering its trustworthy deployment in safety-critical applications. This paper systematically surveys four mainstream explainable AI (xAI) paradigms in computer vision: saliency maps, concept bottleneck models, prototype-driven methods, and hybrid approaches—unifying their underlying mechanisms, applicability boundaries, and inherent limitations. We propose a novel, multi-dimensional evaluation framework integrating fidelity, stability, and human agreement to quantitatively compare explanation quality and computational cost across methods. Our key contribution is a task-aware xAI selection guideline—the first structured framework encompassing theoretical foundations, technical pathways, and empirical validation. This work advances model transparency and provides systematic support for deploying xAI in high-stakes visual domains such as medical imaging and surveillance.
This study systematically investigates the consistency of evaluation metrics for XAI saliency maps—specifically LIME, Grad-CAM, and Guided Backpropagation. Through a large-scale user study (N=166), it conducts the first cross-paradigm, unified comparison of three distinct evaluation frameworks: subjective trust, objective improvement in model understanding, and quantitative mathematical metrics. Results reveal substantial inconsistency across these frameworks: Grad-CAM most effectively enhances users’ objective understanding of model behavior, while Guided Backpropagation achieves the highest scores on mathematical fidelity metrics; notably, several widely used mathematical metrics exhibit significant negative correlations with actual user understanding—challenging the prevailing assumption that high mathematical scores imply superior explainability. The study identifies paradigmatic fragmentation as a core issue in XAI evaluation and provides empirical evidence and methodological insights for developing human-centered, multi-dimensional evaluation frameworks grounded in both cognitive validity and technical rigor.
Existing saliency map methods often produce explanations with high generalizability but low discriminability, failing to precisely localize class-specific decision evidence. To address this, we propose the Cross-Class Attribution Fusion (CCAF) framework—the first to explicitly integrate inter-class attributions by decoupling shared features from discriminative ones. CCAF enhances attribution specificity in a model- and method-agnostic manner via plug-and-play aggregation of gradient- and perturbation-based attributions, guided by counterfactual contrastive masking. The framework comprises three stages: attribution aggregation, class-contrastive masking, and randomized robustness validation. Evaluated on grid-pointing localization and randomized sanity checks, CCAF significantly improves discriminative accuracy across mainstream attribution methods. It reliably identifies both class-discriminative and class-shared visual evidence on multiple benchmarks, advancing the fidelity and interpretability of post-hoc explanations.
This work addresses the limitations of existing saliency-guided training approaches in presentation attack detection (PAD), which rely on costly, domain-specific methods to obtain saliency maps and thus hinder broad applicability. For the first time, the study introduces classical dimensionality reduction techniques—Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA)—to generate saliency maps directly from raw training data, eliminating the need for manual annotations or domain-specific knowledge. By doing so, it overcomes critical bottlenecks in cost, generalizability, and scalability inherent in prior saliency-guided frameworks. The proposed method achieves competitive performance across multiple established and emerging PAD tasks, outperforming baseline approaches and even matching certain state-of-the-art methods, all without requiring additional resources or customized tools.
Large vision models suffer from poor interpretability, and conventional attribution methods merely localize attended regions without providing semantic explanations. Method: This paper proposes a dual-path framework—“aligning with human reasoning” and “concept-level explanation”—featuring the first formally guaranteed attribution method based on verification-aware perturbation analysis (EVA), and the CRAFT-MACO collaborative framework for automatic discovery, importance quantification, and visualization of internal model concepts. It integrates algorithmic stability metrics, Sobol indices computed via quasi-Monte Carlo sampling, 1-Lipschitz function-space optimization, and concept activation mapping. Results: The approach enables interactive, concept-level explanations across all 1,000 ImageNet classes on ResNet, significantly improving explanation fidelity and human consistency—particularly in complex scenes—while offering rigorous theoretical guarantees and actionable semantic insights.
This study addresses the misalignment between conventional chart similarity metrics and human perceptual judgment in information visualization, proposing a deep feature–based similarity assessment method to enhance visualization retrieval and recommendation systems. Methodologically, it systematically compares five ImageNet-pretrained CNN architectures (e.g., ResNet, VGG) against the traditional MS-SSIM metric, integrating multi-scale structural similarity analysis and crowdsourced perceptual experiments to evaluate performance on scatterplot and visual channel similarity tasks. Key contributions include: (1) the first empirical demonstration that ImageNet-pretrained features significantly outperform fine-tuned MS-SSIM; (2) evidence that high-level semantic features better align with human visual perception than low-level statistical features; and (3) establishment of a transferable, perception-aligned similarity metric foundation for visualization analysis tools.
This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.
This work addresses the limitation of existing visual saliency models, which predominantly rely on the free-viewing assumption and thus struggle to capture task-driven attentional patterns. To overcome this, the authors propose a task-driven saliency prediction model that explicitly integrates natural language descriptions of task semantics into the visual attention mechanism for the first time, establishing a task-conditioned architecture for saliency prediction. Experimental results demonstrate that the proposed approach effectively captures the dynamic shifts in human attention across different tasks and significantly improves prediction accuracy in task-oriented viewing scenarios compared to conventional methods.
This study addresses the challenge that existing text-to-image models struggle to control visual attention allocation and lack methods for enhancing target saliency without requiring visual priors. To this end, this work pioneers the task of target saliency enhancement by proposing the GazeME framework. Motivated by insights into relative saliency, the framework introduces a lightweight learnable token insertion mechanism. Through saliency token prompting and a Saliency Prior Marker Activation strategy, it precisely modulates object saliency in a prior-free manner. Furthermore, a dedicated dataset is constructed to validate the proposed approach. Experimental results demonstrate that GazeME effectively enhances the visual prominence of target objects while preserving text-semantic alignment and overall image generation quality.
Traditional fixed-bandwidth isotropic Gaussian kernel density estimation (KDE) struggles to meet the demands of sample-specific evaluation of eye fixation density maps. This work proposes a hybrid fixation density estimation method that integrates adaptive-bandwidth KDE—guided by Abramson’s rule—with center bias, uniform distribution, and state-of-the-art saliency models. Notably, it is the first to incorporate semantic-aware components into adaptive KDE and employs leave-one-subject-out cross-validation to optimize parameters per image. By departing from a decades-old paradigm, the approach significantly improves inter-observer consistency across multiple benchmarks: median log-likelihood gains range from 5% to 15%, AUC increases by up to 2 percentage points, and critical failure cases show improvements exceeding 25%.
Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.