Score
Design and implement analyses and visualizations that attribute a model’s attention to individual input tokens and map those token-level contributions across attention heads and layers; compute normalized per-token attention or activation metrics (e.g., z-scoring), perform cue-to-layer alignment, and produce scores or plots that identify which tokens drive particular model outputs or behaviors.
Existing attention visualization methods often rely on specific model architectures and incur high computational costs, lacking lightweight and general-purpose tools for token importance analysis. This work proposes a model-agnostic attribution method that incurs no additional overhead by perturbing inputs and introducing a three-matrix analytical framework: the Angular Deviation Matrix, Magnitude Deviation Matrix, and Dimensional Importance Matrix. These matrices respectively capture semantic directional shifts, magnitude changes, and dimensional contributions, enabling fine-grained and mathematically rigorous assessment of token importance. The approach demonstrates strong efficiency and interpretability across multiple large language models, and the authors release their code to support reproducible research.
This work addresses the challenge of enhancing interpretability in language and vision models by effectively leveraging information encoded in Transformer attention mechanisms. Existing XAI methods suffer from limitations in both local feature attribution and global concept-level analysis. To overcome these, we propose two novel approaches: (1) integrating attention weights into the Shapley value decomposition framework via an attention-weighted feature function for more precise local attribution; and (2) combining Concept Activation Vectors (CAVs) with directional derivatives to introduce an attention-guided concept sensitivity metric, enabling interpretable global semantic analysis. Our methodology unifies cooperative game theory and gradient-based analysis, ensuring theoretical rigor and practical feasibility. Extensive experiments across multiple standard benchmarks demonstrate significant improvements in explanation accuracy and consistency. The proposed framework provides a unified, scalable, and attention-enhanced XAI paradigm for Transformer-based models.
How do language models—lacking explicit visual pretraining—achieve image understanding? Method: We systematically analyze 16 multimodal large language models (MLLMs) spanning four architectural families and four parameter scales. Introducing the concept of “vision-preferring attention heads,” we identify such heads via attention behavior analysis, statistical modeling of attention weights, and cross-scale ablation experiments, empirically validating their strong, consistent focus on visual tokens. Contribution: We are the first to discover and formally define this generalizable, modular visual-perception substructure within LLMs. Our work reveals the pivotal role of attention mechanisms in cross-modal adaptation, demonstrating how vision-preferring heads mediate text–vision alignment. This provides an interpretable, spatially localizable mechanism underlying joint text–vision representation learning, thereby advancing research toward transparent, controllable, and analyzable multimodal foundation models.
This work addresses the limited interpretability of current vision-language models (VLMs) in understanding data visualizations, which hinders verification of whether their reasoning focuses on semantically relevant image regions. The authors propose a lightweight diagnostic saliency mapping method that, for the first time, aggregates attention weights across all layers and attention heads of a Transformer model with respect to visual tokens and back-projects them onto the image patch grid to establish direct correspondences between generated text and specific image regions. This approach requires no gradient computation and efficiently produces causally faithful explanations. Experimental results demonstrate that the resulting saliency maps accurately highlight the regions attended by the model, and deletion tests confirm their causal fidelity to the model’s behavior.
Current visual language models (VLMs) lack rigorous, human-centered evaluation for data visualization literacy. Method: This study introduces the first standardized, human-oriented assessment framework to systematically evaluate eight VLMs across six chart comprehension tasks, benchmarking their zero-shot performance against human responses via expert annotation, error-pattern analysis, correlation testing, and multi-dimensional behavioral comparison. Contribution/Results: All models significantly underperform humans—even under relaxed scoring criteria—exhibiting only weak behavioral correlation and systematically distinct error distributions unrelated to known human cognitive biases. The work reveals fundamental limitations in VLMs’ visualization understanding and pioneers the integration of human cognitive assessment paradigms into AI evaluation, thereby establishing a novel foundation for cognitively grounded modeling and interpretability research in visual language understanding.
研究探讨了模型素养作为视觉分析性能的额外评价因素,通过控制实验发现模型-任务准确性和视觉分析任务准确性之间存在正相关关系。
This study addresses a critical gap in visualization research by examining how divided attention in real-world multitasking scenarios affects users’ interpretation of visualizations—contrary to the prevailing single-task assumption in existing literature. Through two behavioral experiments integrated with the Linear Ballistic Accumulator (LBA) cognitive model, the work systematically compares user performance under single- and dual-task conditions when interpreting visual designs that either align with or violate viewer expectations (e.g., in color schemes or spatial-semantic mappings). The findings reveal, for the first time, that distraction significantly amplifies the impact of expectation consistency on response times, accuracy, and time-constrained judgment capabilities. These results underscore the crucial importance of expectation-aligned design in authentic multitasking contexts and demonstrate how process-oriented computational modeling can elucidate the underlying cognitive mechanisms.
The mechanisms underlying the emergence of semantic structure during diffusion model generation remain poorly understood, and existing approaches struggle to simultaneously capture the dynamic evolution of attention across both spatial and temporal dimensions. This work proposes a novel visual analytics framework that, for the first time, integrates timestep-indexed token-level cross-attention maps with data-driven phase identification. By combining time-series clustering, quantitative attention metrics, and interactive visualization, the framework enables structured analysis of attention dynamics in Stable Diffusion–like models. Evaluated on a benchmark of 60 structured prompts, it reveals interpretable patterns of attention evolution, effectively supporting human-in-the-loop understanding and control of the generative process.
This study investigates how large language models detect internal activation perturbations and localize the positions of such changes. Through concept vector injection, attention head intervention, QK/OV computation analysis, and comparative experiments across multiple model families, it systematically dissects the underlying mechanisms of attention heads. The work makes two primary contributions: first, it identifies for the first time distinct gating heads responsible for detecting changes and routing heads responsible for selecting positions, revealing their mutually inhibitory interaction; second, it elucidates the relationship between localization precision and attention response magnitude, establishing the micro-level mechanisms that support introspective detection within these models.
This work addresses a critical limitation of existing explainable AI methods, which predominantly offer passive attribution and thus fail to support practitioners in effectively intervening on model behavior. To bridge this gap, the authors propose an interactive analysis workflow that integrates sparse autoencoder (SAE)-based attribution with activation intervention, introducing activation steering into human-in-the-loop debugging for the first time and enabling a paradigm shift from observation to active intervention. Through semi-structured interviews with eight domain experts, the study reveals that users commonly engage in intervention-based hypothesis testing, primarily employing component suppression strategies and grounding their trust in model responses rather than the plausibility of explanations. The research also uncovers key risks—including ripple effects and limited instance-level generalizability of corrections—thereby charting a new path toward trustworthy AI debugging.