π€ AI Summary
This study addresses the trust and adoption challenges arising from the opacity of temporal multimodal AI models by proposing a domain-agnostic, hierarchical interpretability framework. Methodologically, it employs Transformers to handle irregular sequences and integrates Shapley values, attention mechanisms, and gradient-based attribution to elucidate prediction logic across temporal, modal, and feature dimensions. This approach supports interactive decision-making while enabling explanation generation within seconds. Experimental evaluations on in vitro fertilization (IVF) outcome prediction and wheat yield estimation tasks demonstrate the framework's superior performance, with ablation studies confirming the critical role of temporal modeling. By effectively bridging fragmented gaps in existing literature, this work provides a unified interpretability solution for complex temporal multimodal systems.
π Abstract
Artificial Intelligence (AI) models for temporal multimodal data have potential in healthcare and agriculture, but their opacity can limit trust and adoption. We introduce T-MoXAI (Temporal Multimodal eXplainable AI), a hierarchical framework explaining (1) when timepoints influence predictions, using temporal Shapley values; (2) which modalities contribute at those moments, using attention analysis; and (3) what features or image regions drive decisions, using gradient based attribution. A transformer based architecture handles irregular temporal sequences and heterogeneous data, generating all three explanation levels in under one second for interactive decision support. We evaluate the framework on two real world tasks: predicting IVF treatment outcomes from ultrasound sequences and clinical measurements (AUC 0.660 despite significant class imbalance), and forecasting wheat yield from temporal RGB imagery and phenotypic traits ($R^2$ 0.265 amid substantial environmental variability). Ablation studies indicate that temporal modelling is critical in both domains: removing it reduces performance to the equivalent of random guessing. Temporal ROAR experiments provide evidence that the explanations reflect the model's reasoning process. With a unified, domain agnostic architecture and open source implementation, T-MoXAI provides a baseline for temporal multimodal XAI, addressing fragmentation in the field and supporting applications where understanding decisions is as important as predictive accuracy.