SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the under-training of cross-modal synergistic information in multimodal contrastive learning. Grounded in partial information decomposition theory, this work proposes a plug-and-play framework that predicts fused representations via linear projection and extracts interaction residuals to apply targeted contrastive supervision, thereby efficiently capturing synergistic information with minimal computational overhead. Empirically, the proposed method improves synergistic information capture by 5.98% on the Trifeature benchmark and achieves state-of-the-art performance across multiple real-world datasets, offering an efficient and generalizable enhancement strategy for multimodal representation learning.
📝 Abstract
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes into three components: redundancy shared between modalities, uniqueness specific to each modality, and synergy available only from their joint observation. Recent frameworks extend contrastive learning to capture all three components, yet synergy remains undertrained in practice. We propose SynCo (Synergy Contrastive Learning), a method that directly addresses synergy undertraining through dedicated supervision on an interaction residual. SynCo fits a linear projector to predict the fused representation from independently computed unimodal features, and the resulting interaction residual, which removes the linearly unimodal-predictable component, receives dedicated contrastive supervision at negligible computational cost. On the controlled Trifeature benchmark, SynCo achieves state-of-the-art synergy capture with a $+5.98\%$ gain over the baseline, and on real-world benchmarks from MultiBench, DARai, and MM-IMDb, SynCo consistently outperforms or matches prior methods across diverse modality combinations and task types. The method operates as a plug-in to existing contrastive multimodal frameworks without modifying the underlying fusion architecture and can further improve synergy capture when combined with other methods.
Problem

Research questions and friction points this paper is trying to address.

multimodal contrastive learning
synergy
partial information decomposition
cross-modal representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synergy Contrastive Learning
Interaction Residuals
Partial Information Decomposition
Multimodal Representation
Plug-in Framework