T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data
This study addresses the trust and adoption challenges arising from the opacity of temporal multimodal AI models by proposing a domain-agnostic, hierarchical interpretability framework. Methodologically, it employs Transformers to handle irregular sequences and integrates Shapley values, attention mechanisms, and gradient-based attribution to elucidate prediction logic across temporal, modal, and feature dimensions. This approach supports interactive decision-making while enabling explanation generation within seconds. Experimental evaluations on in vitro fertilization (IVF) outcome prediction and wheat yield estimation tasks demonstrate the framework's superior performance, with ablation studies confirming the critical role of temporal modeling. By effectively bridging fragmented gaps in existing literature, this work provides a unified interpretability solution for complex temporal multimodal systems.