🤖 AI Summary
Current research on interpretability of multimodal attention models faces two key bottlenecks: (1) insufficient characterization of cross-modal interaction mechanisms, and (2) a lack of systematic, consistent evaluation methodologies that account for modality-specific cognition and contextual factors. To address these issues, we conduct a systematic literature review (2020–early 2024) analyzing attention-based multimodal models—and their associated interpretation techniques—in both vision-language and unimodal language settings. Our analysis identifies critical gaps: inadequate interaction modeling and the absence of standardized evaluation criteria. Building on these findings, we propose a normative framework for interpretable multimodal AI, comprising three pillars: principled evaluation metric design, standardized reporting guidelines, and context-aware explanation principles. This framework enhances transparency, comparability, and credibility in interpretability research, thereby advancing the development of robust, responsible multimodal AI systems.
📝 Abstract
Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the demand for explainable artificial intelligence (XAI) has spurred a growing body of research aimed at interpreting the complex decision-making processes of these models. This systematic literature review analyzes research published between January 2020 and early 2024 that focuses on the explainability of multimodal models. Framed within the broader goals of XAI, we examine the literature across multiple dimensions, including model architecture, modalities involved, explanation algorithms and evaluation methodologies. Our analysis reveals that the majority of studies are concentrated on vision-language and language-only models, with attention-based techniques being the most commonly employed for explanation. However, these methods often fall short in capturing the full spectrum of interactions between modalities, a challenge further compounded by the architectural heterogeneity across domains. Importantly, we find that evaluation methods for XAI in multimodal settings are largely non-systematic, lacking consistency, robustness, and consideration for modality-specific cognitive and contextual factors. Based on these findings, we provide a comprehensive set of recommendations aimed at promoting rigorous, transparent, and standardized evaluation and reporting practices in multimodal XAI research. Our goal is to support future research in more interpretable, accountable, and responsible mulitmodal AI systems, with explainability at their core.