Score
Designs and implements methods that aggregate per-object or per-pixel relevance scores into spatial heatmaps (including Gaussian-blurred and perception-aware variants) and produces visualizations that convey model attention or importance across an image or scene. Builds pipelines to generate, normalize, smooth, and analyze these heatmaps for interpretability and to support downstream modules or decision-making.
This paper addresses the evaluation of explanation quality in image classification heatmaps, proposing a novel assessment paradigm that jointly considers accuracy and stability. For accuracy, we introduce the Weighting Game metric—a first-of-its-kind quantitative measure evaluating how well class-relevant explanations align with ground-truth semantic segmentation masks. For stability, we design a geometric transformation–based similarity comparison method using scaling and translation to quantify heatmap robustness against input perturbations. Our framework integrates class activation mapping (CAM), segmentation mask matching, and statistical analysis, and is systematically evaluated across mainstream CAM methods and diverse model architectures. Experiments reveal that explanation quality is strongly architecture-dependent, providing empirical guidance for selecting appropriate interpretability methods. This work advances heatmap evaluation from qualitative inspection toward a reproducible, comparable, and quantitative paradigm.
To address perceptual distortion caused by downsampling in large-scale scatterplot visualization, this paper proposes a human visual perception–guided sampling paradigm. Methodologically, it introduces saliency detection into visualization sampling for the first time, constructing a perception-enhanced dataset and a perception-aware similarity metric. We design two algorithms: PAwS (exact) and ApproPAwS (approximate), balancing sampling accuracy and computational efficiency. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in perceptual similarity; ApproPAwS achieves up to 100× speedup with negligible loss in visual fidelity. A user study further confirms its substantial subjective preference advantage. This work establishes a novel theoretical framework and practical toolkit for perception-oriented data visualization sampling.
Existing Class Activation Mapping (CAM) methods lack explicit modeling of user intent and domain knowledge, making it difficult to flexibly control heatmap properties such as localization accuracy, faithfulness, and robustness. Method: We propose SyCAM, the first metric-driven framework for automatic synthesis of CAM expressions. Under formal grammar constraints, SyCAM jointly optimizes multiple objectives—including Intersection-over-Union (IoU), faithfulness, and robustness—via differentiable symbolic expression search and grammar-guided program synthesis, generating interpretable and customizable activation expressions. Contribution/Results: SyCAM breaks away from rigid, hand-crafted CAM formulas, enabling on-demand customization of heatmap behavior. Evaluated on ResNet50, VGG16, and VGG19, it achieves an average 12.7% improvement across target metrics. The framework significantly enhances flexibility, controllability, and interpretability of heatmap generation, establishing a new paradigm for adaptive, user-aligned visual explanation.
This work addresses the lack of systematic evaluation of explanation methods in multiple instance learning (MIL) models widely used in computational pathology, which commonly rely on attention heatmaps for interpretability. The authors propose a general, annotation-free framework to comprehensively benchmark various explanation techniques—including Layer-wise Relevance Propagation (LRP), Integrated Gradients, and Single Perturbation—across classification, regression, and survival analysis tasks under diverse architectures such as Attention, Transformer, and Mamba. Their large-scale evaluation reveals that both model architecture and task type significantly influence explanation quality, with LRP, Integrated Gradients, and Single Perturbation consistently outperforming conventional attention heatmaps. Furthermore, the high-performing heatmaps are correlated with spatial transcriptomics data to validate their biological relevance and uncover divergent decision strategies among models in predicting HPV infection status.
Vision Transformers (ViTs) suffer from high computational overhead due to global self-attention, especially when modeling large receptive fields—compromising both efficiency and representational capacity. To address this, we propose the Heat Conduction Operator (HCO), the first approach to incorporate the physical heat conduction equation into visual representation learning: image patches are treated as thermal sources, and long-range semantic dependencies are modeled via heat diffusion dynamics. HCO is physically interpretable, supports global receptive fields, and achieves only *O*(*N*<sup>1.5</sup>) computational complexity. By leveraging DCT/IDCT for accelerated frequency-domain heat diffusion, HCO enables lightweight, end-to-end differentiable thermodynamic modeling. Experiments across multiple vision tasks demonstrate consistent superiority over ViTs: 37% faster high-resolution inference, 52% fewer FLOPs, and 41% reduced GPU memory consumption.
This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.
Traditional research on graphical perception has predominantly evaluated visualizations from an encoding perspective, often overlooking the fact that the human visual system processes pixel-based images, thereby creating a disconnect between evaluation and actual perception. This work proposes treating visualizations as images and, for the first time, systematically integrates summary statistic vision theory by employing computational vision models that take pixels as input to model the perceptual process from the decoding end. The approach not only successfully reproduces established findings in graphical perception but also sensitively predicts perceptual changes induced by subtle variations in data distributions or design choices, demonstrating the effectiveness and potential of image-based vision models for evaluating visualizations.
Current research in visualization perception lacks a generalizable computational framework capable of predicting user performance across novel combinations of visualizations and tasks. This work proposes modeling visualization interpretation as a sequence of composable and reusable visual decoding operators. By decomposing tasks into chart-agnostic perceptual operations and characterizing the error properties of each operator through a hierarchical Bayesian model, the approach enables accurate prediction of performance on unseen tasks. In preregistered experiments involving PDF/CDF charts, the method successfully predicted both bias and variance in mean estimation from scatterplots using a specific operator composition strategy, significantly outperforming five alternative strategies. These results demonstrate the framework’s predictive validity and theoretical promise for generalizing across visualization types and perceptual tasks.
Current medical vision-language models (VLMs) produce attention heatmaps that lack causal validation, making it difficult to ascertain whether these maps genuinely reflect the critical image regions underlying model predictions. This work proposes the first multidimensional evaluation framework integrating clinical annotations with causal perturbations to systematically assess the faithfulness of VLM attention. The framework evaluates region overlap with radiologist-annotated areas, attribution quality within masked regions, and performance under 16×16 image patch occlusion. Results reveal that none of the evaluated VLMs simultaneously satisfy the dual criteria of effectively leveraging visual information and concentrating attention on clinically relevant regions. In contrast, all specialized chest X-ray (CXR) classifiers pass the assessment, exposing a fundamental deficiency in the explainability of existing medical VLMs.