🤖 AI Summary
Cross-modal learning suffers from modality dominance, redundant coupling, and spurious correlations, leading to poor generalization, weak interpretability, and insufficient robustness to noise or missing modalities. To address these issues, we propose the Dual-Stream Residual Semantic Disentanglement (DRSD) framework, which explicitly separates modality-specific representations from shared semantic representations via residual decomposition and orthogonal regularization—thereby mitigating cross-modal redundancy and enhancing weak-signal modeling. DRSD integrates a dual-stream architecture, a residual semantic alignment head, contrastive-regressive joint optimization, and covariance-based regularization. Evaluated on two large-scale educational benchmarks, DRSD significantly outperforms unimodal, early-fusion, late-fusion, and co-attention baselines in both next-step and final outcome prediction. It achieves superior generalization, robustness to missing modalities, and enhanced interpretability through disentangled, semantically grounded representations.
📝 Abstract
Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dominance, redundant information coupling, and spurious cross-modal correlations, leading to suboptimal generalization and limited interpretability. In particular, high-variance modalities tend to overshadow weaker but semantically important signals, while naïve fusion strategies entangle modality-shared and modality-specific factors in an uncontrolled manner. This makes it difficult to understand which modality actually drives a prediction and to maintain robustness when some modalities are noisy or missing. To address these challenges, we propose a Dual-Stream Residual Semantic Decorrelation Network (DSRSD-Net), a simple yet effective framework that disentangles modality-specific and modality-shared information through residual decomposition and explicit semantic decorrelation constraints. DSRSD-Net introduces: (1) a dual-stream representation learning module that separates intra-modal (private) and inter-modal (shared) latent factors via residual projection; (2) a residual semantic alignment head that maps shared factors from different modalities into a common space using a combination of contrastive and regression-style objectives; and (3) a decorrelation and orthogonality loss that regularizes the covariance structure of the shared space while enforcing orthogonality between shared and private streams, thereby suppressing cross-modal redundancy and preventing feature collapse. Experimental results on two large-scale educational benchmarks demonstrate that DSRSD-Net consistently improves next-step prediction and final outcome prediction over strong single-modality, early-fusion, late-fusion, and co-attention baselines.