🤖 AI Summary
This study addresses the challenge of capturing synergistic information in multimodal learning by proposing the HRIL framework. It reveals that synergy originates from higher-order statistical dependencies and explicitly models the multimodal joint distribution through empirical cross-moment tensor construction and Tucker decomposition. Furthermore, a synergy-aware regularizer integrated with self-supervised contrastive learning is designed to prevent energy concentration and preserve higher-order coupling capabilities. Experimental results demonstrate that the proposed method outperforms existing approaches on both controlled tasks and real-world benchmarks, significantly enhancing model performance in scenarios dominated by synergistic interactions.
📝 Abstract
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.