HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of capturing synergistic information in multimodal learning by proposing the HRIL framework. It reveals that synergy originates from higher-order statistical dependencies and explicitly models the multimodal joint distribution through empirical cross-moment tensor construction and Tucker decomposition. Furthermore, a synergy-aware regularizer integrated with self-supervised contrastive learning is designed to prevent energy concentration and preserve higher-order coupling capabilities. Experimental results demonstrate that the proposed method outperforms existing approaches on both controlled tasks and real-world benchmarks, significantly enhancing model performance in scenarios dominated by synergistic interactions.
📝 Abstract
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
Problem

Research questions and friction points this paper is trying to address.

multimodal representation learning
synergistic information
self-supervised learning
cross-modal interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Synergy
Higher-Order Tensor Modeling
Tucker Decomposition
Self-Supervised Learning
Synergy-Aware Regularizer
💼 Related Jobs
No related jobs found.
Q
Qun Dai
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics
Liangjian Wen
Liangjian Wen
Southwestern University of Finance and Economics Chengdu, China
J
Jiang Duan
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics; Chengdu Everimaging Science and Technology Co., Ltd
Y
Yong Dai
X-Humanoid
D
Dongkai Wang
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics
Maolin Wang
Maolin Wang
City University of Hong Kong
Tensor DecompositionMachine LearningModel Compression
Mingjie Wang
Mingjie Wang
Zhejiang Sci-Tech University, University of Guelph, Memorial University of Newfoundland
Computer VisionDeep Learning
Jianzhuang Liu
Jianzhuang Liu
Shenzhen Institutes of Advanced Technology, University of Chinese Academy of Sciences
Computer VisionImage ProcessingAIGCMachine Learning
He Yan
He Yan
The Hong Kong University of Science and Technology
Organic solar cellspolymer solar cellsorganic transistorsorganic electronics
Z
Zhao Kang
University of Electronic Science and Technology of China