Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

📅 2025-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current research on interpretability of multimodal attention models faces two key bottlenecks: (1) insufficient characterization of cross-modal interaction mechanisms, and (2) a lack of systematic, consistent evaluation methodologies that account for modality-specific cognition and contextual factors. To address these issues, we conduct a systematic literature review (2020–early 2024) analyzing attention-based multimodal models—and their associated interpretation techniques—in both vision-language and unimodal language settings. Our analysis identifies critical gaps: inadequate interaction modeling and the absence of standardized evaluation criteria. Building on these findings, we propose a normative framework for interpretable multimodal AI, comprising three pillars: principled evaluation metric design, standardized reporting guidelines, and context-aware explanation principles. This framework enhances transparency, comparability, and credibility in interpretability research, thereby advancing the development of robust, responsible multimodal AI systems.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsComputer Vision: Multi-modal VisionMachine Learning: Multimodal Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the demand for explainable artificial intelligence (XAI) has spurred a growing body of research aimed at interpreting the complex decision-making processes of these models. This systematic literature review analyzes research published between January 2020 and early 2024 that focuses on the explainability of multimodal models. Framed within the broader goals of XAI, we examine the literature across multiple dimensions, including model architecture, modalities involved, explanation algorithms and evaluation methodologies. Our analysis reveals that the majority of studies are concentrated on vision-language and language-only models, with attention-based techniques being the most commonly employed for explanation. However, these methods often fall short in capturing the full spectrum of interactions between modalities, a challenge further compounded by the architectural heterogeneity across domains. Importantly, we find that evaluation methods for XAI in multimodal settings are largely non-systematic, lacking consistency, robustness, and consideration for modality-specific cognitive and contextual factors. Based on these findings, we provide a comprehensive set of recommendations aimed at promoting rigorous, transparent, and standardized evaluation and reporting practices in multimodal XAI research. Our goal is to support future research in more interpretable, accountable, and responsible mulitmodal AI systems, with explainability at their core.
Problem

Research questions and friction points this paper is trying to address.

Analyzing explainability in multimodal attention-based models
Addressing inconsistent evaluation methods in multimodal XAI
Improving understanding of cross-modal interactions in AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic review on multimodal attention-based models
Focus on explainability in vision-language models
Recommendations for standardized XAI evaluation methods
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Md Raisul Kibria
Faculty of Science and Engineering, Information Technology, Åbo Akademi University, Turku, 20500, Finland
S
Sébastien Lafond
Faculty of Science and Engineering, Information Technology, Åbo Akademi University, Turku, 20500, Finland
Janan Arslan
Janan Arslan
Senior Research Engineer, Scientist, Consultant, Teacher, CTO
AImathematics & statisticsexplainable AIophthalmologyoncology