HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
This study addresses the challenge of capturing synergistic information in multimodal learning by proposing the HRIL framework. It reveals that synergy originates from higher-order statistical dependencies and explicitly models the multimodal joint distribution through empirical cross-moment tensor construction and Tucker decomposition. Furthermore, a synergy-aware regularizer integrated with self-supervised contrastive learning is designed to prevent energy concentration and preserve higher-order coupling capabilities. Experimental results demonstrate that the proposed method outperforms existing approaches on both controlled tasks and real-world benchmarks, significantly enhancing model performance in scenarios dominated by synergistic interactions.