saliency aggregation

Design methods and tools that compute, aggregate, analyze, and visualize saliency maps or feature-importance scores for inputs and models, using gradient-based approaches (e.g., gradients, integrated gradients), model-based techniques, or dimensionality-reduction methods (PCA, LDA), and that can produce instance-level or universal saliency predictions such as attention or gaze-like maps. Build pipelines that aggregate saliency across scenes or datasets, evaluate and quantify spatially concentrated discriminative evidence, and apply those saliency outputs to guide selection of influential input regions and to inform saliency-guided, saliency-aware, or saliency-explored training and model refinement.

saliencyaggregation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

xAI-CV: An Overview of Explainable Artificial Intelligence in Computer Vision

Sep 23, 2025
NV
Nguyen Van Tu
🏛️ University of Science

Deep learning excels in image analysis but suffers from poor interpretability, hindering its trustworthy deployment in safety-critical applications. This paper systematically surveys four mainstream explainable AI (xAI) paradigms in computer vision: saliency maps, concept bottleneck models, prototype-driven methods, and hybrid approaches—unifying their underlying mechanisms, applicability boundaries, and inherent limitations. We propose a novel, multi-dimensional evaluation framework integrating fidelity, stability, and human agreement to quantitatively compare explanation quality and computational cost across methods. Our key contribution is a task-aware xAI selection guideline—the first structured framework encompassing theoretical foundations, technical pathways, and empirical validation. This work advances model transparency and provides systematic support for deploying xAI in high-stakes visual domains such as medical imaging and surveillance.

Deep learning creates black-box models in computer vision tasksLack of interpretability raises reliability concerns in critical applicationsSurveying explainable AI methods to understand AI decision processes

Must-Read Papers

Most classic and influential ideas
View more

What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI)

Apr 23, 2025
FK
Felix Kares
🏛️ Saarland University | University of Bayreuth | University of Freiburg

This study systematically investigates the consistency of evaluation metrics for XAI saliency maps—specifically LIME, Grad-CAM, and Guided Backpropagation. Through a large-scale user study (N=166), it conducts the first cross-paradigm, unified comparison of three distinct evaluation frameworks: subjective trust, objective improvement in model understanding, and quantitative mathematical metrics. Results reveal substantial inconsistency across these frameworks: Grad-CAM most effectively enhances users’ objective understanding of model behavior, while Guided Backpropagation achieves the highest scores on mathematical fidelity metrics; notably, several widely used mathematical metrics exhibit significant negative correlations with actual user understanding—challenging the prevailing assumption that high mathematical scores imply superior explainability. The study identifies paradigmatic fragmentation as a core issue in XAI evaluation and provides empirical evidence and methodological insights for developing human-centered, multi-dimensional evaluation frameworks grounded in both cognitive validity and technical rigor.

Assessing correlation between mathematical and user measuresComparing user trust and model understanding metricsEvaluating effectiveness of saliency maps in XAI

Now you see me! A framework for obtaining class-relevant saliency maps

Mar 10, 2025
NP
Nils Philipp Walter
🏛️ CISPA Helmholtz Center for Information Security | Max Planck Institute for Informatics

Existing saliency map methods often produce explanations with high generalizability but low discriminability, failing to precisely localize class-specific decision evidence. To address this, we propose the Cross-Class Attribution Fusion (CCAF) framework—the first to explicitly integrate inter-class attributions by decoupling shared features from discriminative ones. CCAF enhances attribution specificity in a model- and method-agnostic manner via plug-and-play aggregation of gradient- and perturbation-based attributions, guided by counterfactual contrastive masking. The framework comprises three stages: attribution aggregation, class-contrastive masking, and randomized robustness validation. Evaluated on grid-pointing localization and randomized sanity checks, CCAF significantly improves discriminative accuracy across mainstream attribution methods. It reliably identifies both class-discriminative and class-shared visual evidence on multiple benchmarks, advancing the fidelity and interpretability of post-hoc explanations.

Enhancing transparency in high-stakes decision-making contextsIdentifying class-relevant features in model predictionsImproving specificity of saliency maps for neural networks

This work addresses the limitations of existing saliency-guided training approaches in presentation attack detection (PAD), which rely on costly, domain-specific methods to obtain saliency maps and thus hinder broad applicability. For the first time, the study introduces classical dimensionality reduction techniques—Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA)—to generate saliency maps directly from raw training data, eliminating the need for manual annotations or domain-specific knowledge. By doing so, it overcomes critical bottlenecks in cost, generalizability, and scalability inherent in prior saliency-guided frameworks. The proposed method achieves competitive performance across multiple established and emerging PAD tasks, outperforming baseline approaches and even matching certain state-of-the-art methods, all without requiring additional resources or customized tools.

biometric presentation attack detectiondomain specificitysaliency acquisition

Large vision models suffer from poor interpretability, and conventional attribution methods merely localize attended regions without providing semantic explanations. Method: This paper proposes a dual-path framework—“aligning with human reasoning” and “concept-level explanation”—featuring the first formally guaranteed attribution method based on verification-aware perturbation analysis (EVA), and the CRAFT-MACO collaborative framework for automatic discovery, importance quantification, and visualization of internal model concepts. It integrates algorithmic stability metrics, Sobol indices computed via quasi-Monte Carlo sampling, 1-Lipschitz function-space optimization, and concept activation mapping. Results: The approach enables interactive, concept-level explanations across all 1,000 ImageNet classes on ResNet, significantly improving explanation fidelity and human consistency—particularly in complex scenes—while offering rigorous theoretical guarantees and actionable semantic insights.

Complex ImagesImage RecognitionModel Interpretability

Seeing Eye to AI? Applying Deep-Feature-Based Similarity Metrics to Information Visualization

Feb 28, 2025
SL
Sheng Long
🏛️ Northwestern University | Utrecht University

This study addresses the misalignment between conventional chart similarity metrics and human perceptual judgment in information visualization, proposing a deep feature–based similarity assessment method to enhance visualization retrieval and recommendation systems. Methodologically, it systematically compares five ImageNet-pretrained CNN architectures (e.g., ResNet, VGG) against the traditional MS-SSIM metric, integrating multi-scale structural similarity analysis and crowdsourced perceptual experiments to evaluate performance on scatterplot and visual channel similarity tasks. Key contributions include: (1) the first empirical demonstration that ImageNet-pretrained features significantly outperform fine-tuned MS-SSIM; (2) evidence that high-level semantic features better align with human visual perception than low-level statistical features; and (3) establishment of a transferable, perception-aligned similarity metric foundation for visualization analysis tools.

Explores deep-feature-based similarity metrics for visualization similarity judgments.Extends similarity metrics using ML architectures and pre-trained weight sets.Improves visualization tools by enhancing similarity assessments with deep-feature metrics.

Latest Papers

What's happening recently
View more

This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.

attention mechanismmodel interpretabilitysaliency maps

This work addresses the limitation of existing visual saliency models, which predominantly rely on the free-viewing assumption and thus struggle to capture task-driven attentional patterns. To overcome this, the authors propose a task-driven saliency prediction model that explicitly integrates natural language descriptions of task semantics into the visual attention mechanism for the first time, establishing a task-conditioned architecture for saliency prediction. Experimental results demonstrate that the proposed approach effectively captures the dynamic shifts in human attention across different tasks and significantly improves prediction accuracy in task-oriented viewing scenarios compared to conventional methods.

goal-directed attentionnatural-language taskstask-driven attention

This study addresses the challenge that existing text-to-image models struggle to control visual attention allocation and lack methods for enhancing target saliency without requiring visual priors. To this end, this work pioneers the task of target saliency enhancement by proposing the GazeME framework. Motivated by insights into relative saliency, the framework introduces a lightweight learnable token insertion mechanism. Through saliency token prompting and a Saliency Prior Marker Activation strategy, it precisely modulates object saliency in a prior-free manner. Furthermore, a dedicated dataset is constructed to validate the proposed approach. Experimental results demonstrate that GazeME effectively enhances the visual prominence of target objects while preserving text-semantic alignment and overall image generation quality.

Target Saliency BoostingText-to-Image GenerationVisual Attention

Traditional fixed-bandwidth isotropic Gaussian kernel density estimation (KDE) struggles to meet the demands of sample-specific evaluation of eye fixation density maps. This work proposes a hybrid fixation density estimation method that integrates adaptive-bandwidth KDE—guided by Abramson’s rule—with center bias, uniform distribution, and state-of-the-art saliency models. Notably, it is the first to incorporate semantic-aware components into adaptive KDE and employs leave-one-subject-out cross-validation to optimize parameters per image. By departing from a decades-old paradigm, the approach significantly improves inter-observer consistency across multiple benchmarks: median log-likelihood gains range from 5% to 15%, AUC increases by up to 2 percentage points, and critical failure cases show improvements exceeding 25%.

eye-trackingfixation densityinterobserver consistency

Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.

graphical perceptionhuman visionimage-based modeling

Hot Scholars

AC

Adam Czajka

University of Notre Dame
BiometricsComputer VisionIris RecognitionPresentation Attack Detection
NR

Nidhi Rastogi

Assistant Professor, Rochester Institute of Technology, NY
CybersecurityArtificial IntelligenceAutonomous VehiclesGraph Analytics
DB

Dipkamal Bhusal

Rochester Institute of Technology
InterpretabilityExplainabilityDeep LearningReliable AI
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
ZH

Zi Huang

PhD Candidate
Deep Learning