Score
Designs, builds, and evaluates methods and tools that produce, visualize, and automatically score visual explanations for model predictions — including saliency maps and other local explanation outputs, as well as grounded explanations that link model outputs to input evidence. Analyzes alignment between explanations and reference evidence, measures trustworthiness and consistency across models and methods, and implements automated metrics or human-evaluation protocols to compare and benchmark local explanation techniques.
Visual foundation models (e.g., ViT, SAM, diffusion models) suffer from limited interpretability, hindering trust and deployment. Method: This work establishes the first comprehensive XAI research framework tailored to visual foundation models, leveraging bibliometric analysis and topic modeling to systematically categorize and critically assess attribution methods, concept activation, surrogate models, and visualization techniques—explicitly distinguishing their dual roles as *targets of explanation* and *explanation tools*. It identifies the fundamental tension among scalability, fidelity, and generalizability, and proposes standardized evaluation dimensions. Contribution/Results: The study delivers the first XAI research map for visual foundation models, offering a structured taxonomy, critical insights into methodological trade-offs, and foundational guidelines for advancing both theoretical understanding and practical implementation of explainability in this rapidly evolving domain.
Large vision models suffer from poor interpretability, and conventional attribution methods merely localize attended regions without providing semantic explanations. Method: This paper proposes a dual-path framework—“aligning with human reasoning” and “concept-level explanation”—featuring the first formally guaranteed attribution method based on verification-aware perturbation analysis (EVA), and the CRAFT-MACO collaborative framework for automatic discovery, importance quantification, and visualization of internal model concepts. It integrates algorithmic stability metrics, Sobol indices computed via quasi-Monte Carlo sampling, 1-Lipschitz function-space optimization, and concept activation mapping. Results: The approach enables interactive, concept-level explanations across all 1,000 ImageNet classes on ResNet, significantly improving explanation fidelity and human consistency—particularly in complex scenes—while offering rigorous theoretical guarantees and actionable semantic insights.
Current evaluation of visual explanation methods is hindered by scarce human-annotated ground-truth explanations and the absence of standardized, comprehensive evaluation protocols—limiting simultaneous assessment of alignment (faithfulness to model behavior) and causality (counterfactual robustness). To address this, we introduce the first standardized benchmark for explainability in image classification: it comprises eight cross-domain datasets, each with human-annotated explanations. We propose a unified evaluation framework that systematically integrates six quantitative metrics—including Infidelity, ROAR, and Faithfulness—to enable fair, comparable assessment of both post-hoc and intrinsic explanation methods. Furthermore, we release an open-source, multi-dataset, multi-metric evaluation platform. Empirical evaluation across four datasets and eight state-of-the-art methods demonstrates the benchmark’s discriminative power and robustness. All code and datasets are publicly available.
Existing evaluation methods for model explanations—relying either on ground-truth annotations or strong model-sensitivity assumptions—are fundamentally limited by the absence of reliable, human-validated explanation labels. To address this, we propose AXE, the first ground-truth-agnostic and model-agnostic framework for evaluating local feature importance explanations. AXE’s core contribution lies in formalizing three axiomatic principles—consistency, stability, and separability—and deriving unsupervised, quantitative metrics from them. These metrics enable principled, standalone assessment of explanation quality and facilitate detection of “fairwashing”—i.e., spurious explanations that mask model bias. Extensive experiments across diverse models (e.g., LLMs, tree-based, and neural networks) and benchmark datasets demonstrate that AXE consistently outperforms baseline methods dependent on ground truth or sensitivity analysis. The implementation is publicly available.
This study investigates the discrepancies between large language models (LLMs) and human cognition in high-level chart understanding, with a focus on interpreting designer intent and extracting complex data patterns. Through qualitative user studies, it systematically compares the higher-order interpretation strategies employed by humans and LLMs on line charts, bar charts, and scatter plots, while analyzing LLM outputs and reasoning pathways under three distinct prompting conditions. The work reveals, for the first time, that LLMs consistently adopt a structured enumeration strategy rather than constructing coherent trend-based narratives, and their explanatory patterns remain remarkably stable across different prompts. In contrast, humans demonstrate a superior ability to synthesize holistic, narrative-driven interpretations. These findings highlight fundamental mechanistic limitations of LLMs in visual reasoning and offer critical insights for future model design.
This paper addresses the evaluation of explanation quality in image classification heatmaps, proposing a novel assessment paradigm that jointly considers accuracy and stability. For accuracy, we introduce the Weighting Game metric—a first-of-its-kind quantitative measure evaluating how well class-relevant explanations align with ground-truth semantic segmentation masks. For stability, we design a geometric transformation–based similarity comparison method using scaling and translation to quantify heatmap robustness against input perturbations. Our framework integrates class activation mapping (CAM), segmentation mask matching, and statistical analysis, and is systematically evaluated across mainstream CAM methods and diverse model architectures. Experiments reveal that explanation quality is strongly architecture-dependent, providing empirical guidance for selecting appropriate interpretability methods. This work advances heatmap evaluation from qualitative inspection toward a reproducible, comparable, and quantitative paradigm.
This work addresses a critical limitation of existing explainable AI methods, which predominantly offer passive attribution and thus fail to support practitioners in effectively intervening on model behavior. To bridge this gap, the authors propose an interactive analysis workflow that integrates sparse autoencoder (SAE)-based attribution with activation intervention, introducing activation steering into human-in-the-loop debugging for the first time and enabling a paradigm shift from observation to active intervention. Through semi-structured interviews with eight domain experts, the study reveals that users commonly engage in intervention-based hypothesis testing, primarily employing component suppression strategies and grounding their trust in model responses rather than the plausibility of explanations. The research also uncovers key risks—including ripple effects and limited instance-level generalizability of corrections—thereby charting a new path toward trustworthy AI debugging.
This study addresses the challenge that local interpretability methods often produce seemingly plausible yet unfaithful explanations for complex tabular data. To rigorously evaluate explanation fidelity, robustness, and complexity, the authors construct a comprehensive benchmarking framework encompassing multiple models and datasets, incorporating for the first time a prediction-consistency grouping strategy. They systematically assess prominent methods—including LIME, Kernel SHAP, and feature ablation—across 32 tabular datasets. The findings reveal that explanation quality shows no significant correlation with model accuracy; instead, it is predominantly influenced by data complexity and feature distribution, particularly on samples where models consistently err. This work provides a novel perspective and empirical foundation for trustworthy evaluation in explainable AI.
This work reveals the high sensitivity of fidelity evaluation metrics (e.g., Insertion/Deletion) in eXplainable AI (XAI) to baseline selection: varying baselines induce inconsistent rankings—and even reversal of optimality judgments—among attribution methods. The core issue lies in the inherent trade-off faced by conventional baselines: they struggle to simultaneously achieve *complete removal of target feature information* and *distributional plausibility* (i.e., avoidance of out-of-distribution, OOD, inputs). To address this, we formally define these two ideal baseline properties for the first time and propose the first model-dependent, learnable *feature-visualizing baseline*, generated via activation inversion to produce model-adaptive baseline images. Extensive experiments across multiple models, datasets, and mainstream attribution methods demonstrate that all standard fixed baselines fail to satisfy both properties concurrently. Our learned baseline substantially alleviates the OOD–information-removal trade-off, yielding significantly more consistent and reliable fidelity evaluations.
This study investigates how visual and textual explanation formats in educational recommender systems interact with users’ individual characteristics to influence their perceptions of transparency, trust, sense of control, and satisfaction. Drawing on a within-subjects experiment with 54 participants, the research integrates multidimensional user traits—including the Big Five personality dimensions, need for cognition, decision-making styles, familiarity with visualizations, and technical expertise—and analyzes the data using robust linear mixed-effects models. The findings reveal, for the first time, that the moderating effects of these personal characteristics on explanation preferences are limited. Instead, the study proposes general design principles for visual explanations: those that are concise, interactive, highly selective, and intuitive consistently enhance user perceptions across diverse individuals, with minimal susceptibility to individual differences.
This work addresses the unreliability of existing medical image diagnosis models that often rely on non-causal or clinically irrelevant visual cues. To enhance trustworthiness, the authors propose a systematic framework that integrates explanation-aware loss directly into the end-to-end training objective by incorporating saliency-based interpretability supervision. A custom explanation loss function jointly optimizes diagnostic accuracy and spatial fidelity of model explanations. The study introduces two quantitative metrics—annotation coverage and saliency precision—to evaluate explanation quality and uncover the trade-off between explanation loss strength and model performance. Experiments on a chest X-ray dataset demonstrate that the proposed method achieves diagnostic accuracy comparable to baseline models while significantly improving spatial alignment between model-generated explanations and clinical annotations.