Score
Designs, implements, or evaluates pipelines, networks, and methods that combine heterogeneous inputs — different modalities, feature sets, temporal frames, spatial scales, or separate models — into unified feature representations or fused outputs (including feature-level, model-level, multi-scale, multi-frame, and multi-source fusion). This work covers aligning and normalizing disparate formats and schemas, grouping and matching features, applying feature perturbation and weighting schemes, handling real and complex-valued features, and developing fusion strategies and techniques that preserve modality-specific signals while supporting downstream analysis or prediction.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing approaches to enhancing the interpretability of multimodal unified representations rely on discretized representations but suffer from two key limitations: (1) Euclidean distance-based quantification ignores dimensional heterogeneity, inducing representation redundancy; and (2) uniform cross-modal alignment neglects modality-specific characteristics. To address these issues, we propose Training-Free Codebook Optimization (TOC) and Fine-Grained/Coarse-Grained Inter-Modal Information Decoupling (FCID)—the first framework enabling post-pretraining, gradient-free representation refinement and modality-adaptive information decoupling. TOC mitigates quantization redundancy via unsupervised codebook refinement, while FCID explicitly models modality-specific properties and disentangles shared versus private cross-modal information. Evaluated on cross-modal retrieval and zero-shot transfer tasks, our method achieves significant improvements over state-of-the-art baselines: representation redundancy is reduced by 37%, and modality specificity is enhanced by 21%.
To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.
Current biomedical multimodal models face critical bottlenecks: reliance on end-to-end training, exponential growth in computational complexity with modality count, severe performance degradation under extreme modality imbalance, and rigid topological coupling. To address these, we propose MM-Lego—a tuning-free, universal multimodal fusion framework. It introduces a novel frequency-domain feature harmonization mechanism that achieves shape alignment and interference-free merging of arbitrary unimodal encoders. We further design modality-agnostic wrappers and zero-/few-shot model merging strategies, enabling topology-agnostic fusion and robust modeling under highly imbalanced modalities. Crucially, MM-Lego requires no fine-tuning yet matches or surpasses end-to-end models in performance, while maintaining full encoder compatibility. Evaluated across seven biomedical benchmark datasets, it achieves state-of-the-art results on five—demonstrating unprecedented flexibility, efficiency, and generalizability in biomedical multimodal learning.
To address weak structural modeling, shallow cross-modal interactions, difficult alignment, and poor interpretability in fusing heterogeneous multimodal features—spanning domains, granularities (e.g., token, patch, frame, clip), and modalities—this paper proposes a relation-centered, learnable graph-power fusion paradigm. It maps high-dimensional features into an interpretable graph space and constructs cross-granularity relational graphs. A learnable graph-power operator is introduced to aggregate element-wise relational scores via multivariate polynomials over homogeneous graphs, enabling structural-aware deep interaction. The method balances expressive power and interpretability, achieving multimodal fusion (text, image, video) without explicit alignment. Evaluated on video anomaly detection, it significantly outperforms concatenation, attention-based, and conventional nonlinear fusion baselines, demonstrating strong generalization and effectiveness.
RGB-D multimodal fusion mechanisms have long suffered from poor interpretability, and the fundamental nature of cross-modal complementarity remains unclear. Method: This paper establishes the first interpretability analysis framework specifically targeting the fusion process, introducing a joint metric of semantic variance and feature similarity to systematically characterize cross-modal representation consistency, intra-modal evolutionary patterns, and collaborative optimization logic. Through cross-layer feature comparison and quantitative semantic analysis, we identify a prevalent imbalance between consistency and specificity in mainstream fusion strategies. Contribution/Results: We formalize a “specificity-driven inference under consistency constraints” principle that explicates cross-modal complementarity. Our framework provides both theoretical foundations and a verifiable evaluation paradigm for designing trustworthy, generalizable multimodal fusion models.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.
This work addresses the limitations of existing approaches that predominantly rely on non-native late-fusion architectures and lack a systematic definition or unified framework for native multimodal modeling. The paper proposes a formal roadmap for Native Multimodal Modeling (NMM), introducing, for the first time, a rigorous formulation of “architectural nativeness.” Leveraging an input–output duality perspective, it categorizes models into three types: Multi-to-Text, Multi-to-Target, and Multi-to-Multi. A full-stack, industrial-grade NMM framework is developed, encompassing data governance, early/mid-stage fusion, end-to-end training, and deployment, all realized through a unified Transformer architecture that enables symbiotic cross-modal understanding and generation. This paradigm demonstrates superior performance and strong scalability across multiple tasks, offering a clear pathway toward truly native multimodal models.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
This work addresses the challenge in multimodal sentiment analysis where modality-specific signal refinement and cross-modal interaction modeling often interfere with each other due to conflicting optimization objectives. To resolve this, the authors propose SeRIn, a novel architecture that decouples modality separation and cross-modal interaction into structured priors through a three-stage pipeline—separation, refinement, and integration—processing unimodal representations and cross-modal interactions via independent pathways before fusing them at the prediction stage. Notably, SeRIn adaptively adjusts modality weights without requiring explicit supervision. The method achieves state-of-the-art performance on the CH-SIMS and CMU-MOSEI benchmarks, yielding significant improvements across all evaluation metrics.
This work addresses the performance degradation in multimodal classification caused by modality imbalance by proposing deep ensembling as an alternative to explicit modality fusion, achieving effective multimodal classification through the combination of unimodal networks. The key contributions include the first demonstration that superior performance can be attained without explicit fusion, a heuristic strategy for allocating the number of ensemble models based on each modality’s predictive capability, and the construction of a controllable synthetic multimodal data framework with fitted scaling laws. Experiments show that, under identical parameter budgets, the proposed method significantly outperforms state-of-the-art late-fusion and intermediate-fusion approaches on both real-world and synthetic datasets, while the derived scaling laws reveal an asymptotic upper bound on ensemble performance.