Score
Designs models and algorithms that decompose inputs from multiple sources into representation components that are shared across sources and components that are private to each source, and that train to align and separate those components. This includes methods for learning modality-shared and modality-specific features, aligning shared components across patterns, handling missing modalities, and reducing representation-learning error for rare or infrequent patterns.
Existing approaches struggle to effectively model the dynamic interplay among redundant, unique, and synergistic information at the sample level in multimodal learning. This work proposes a novel information-theoretic decomposition–based paradigm for multimodal interaction learning, offering the first systematic analysis of the importance of sample-level interactions. By employing a variational architecture, the method explicitly disentangles these three interaction components and integrates a component-aware fine-tuning strategy to adaptively leverage them. Extensive experiments across diverse tasks and architectures consistently demonstrate the superiority of the proposed approach over current state-of-the-art methods, confirming its effectiveness and generality in finely modeling sample-level multimodal interactions.
This paper addresses modality collapse in multimodal fusion—a systemic failure where models neglect certain modalities during training. We identify its root cause: noisy features induce representation entanglement via shared neurons, compounded by rank deficiency in the fusion head, leading to degraded modality-specific representations. We provide the first theoretical explanation grounded jointly in representation entanglement and low-rank constraints. Building on this insight, we propose an explicit basis reallocation algorithm that enforces cross-modal disentanglement and dynamic weight calibration, while ensuring robust inference under modality missingness. Our method integrates cross-modal knowledge distillation, neuron-level attribution, and rank-constrained fusion head modeling. Evaluated on multiple benchmarks, it significantly mitigates modality collapse, substantially improving model generalization and stability under partial modality absence.
In multimodal data, modality-specific and cross-modal shared information are deeply entangled, making disentanglement challenging. Method: This paper proposes DisentangledSSL, a controllable disentangled representation learning framework that—uniquely under non-minimal necessary information (MNI)-inaccessible conditions—systematically analyzes disentanglement optimality. It integrates the information bottleneck principle, self-supervised contrastive learning, mutual information estimation, variational inference, and modality-adversarial constraints to achieve theoretically grounded, controllable disentanglement. Contribution/Results: Evaluated on synthetic and real-world benchmarks—including vision-language and molecular-phenotype datasets—DisentangledSSL achieves average improvements of 5.2%–9.8% over baselines in prediction and cross-modal retrieval tasks. It significantly enhances model interpretability, robustness to distribution shifts, and counterfactual generation capability, offering principled solutions for multimodal representation learning.
Existing feature disentanglement (FD)-based multimodal MRI methods suffer from loss of cross-modal shared information when handling ≥3 modalities and lack explicit modeling of relationships among disentangled features during fusion. To address these limitations, we propose a Complete Feature Disentanglement (CFD) strategy and a Dynamic Mixture-of-Experts Fusion (DMF) module. CFD introduces the novel concept of “modality-partially shared features,” enabling full-order disentanglement that preserves both modality-specific and hierarchically shared representations. The DMF module employs a dynamic gating mechanism to jointly model local–global feature interactions, thereby enhancing both interpretability and fusion accuracy. Evaluated on three multimodal MRI classification tasks, our approach consistently outperforms state-of-the-art methods, demonstrating the effectiveness of restoring shared-information integrity and dynamically modeling inter-feature relationships.
This work addresses the modality gap in multimodal representation learning induced by the InfoNCE objective, which manifests as a conflict between inter-modal alignment and uniformity, as well as intra-modal alignment inconsistencies. The paper proposes the first framework that decouples alignment and uniformity in multimodal learning, employing Hölder divergence–based alignment optimization alongside a dedicated uniformity loss to effectively mitigate these conflicts. Theoretically, the proposed objective is shown to serve as a valid proxy for the global Hölder divergence between multimodal distributions. Notably, the method requires no task-specific components and consistently improves performance across both discriminative tasks (e.g., retrieval) and generative tasks (e.g., UnCLIP), demonstrating its generality and effectiveness.
This study addresses the limitation of existing multimodal joint decomposition methods that necessitate full-dimensional sharing. To overcome this constraint, we propose PathFinder, a generalized joint low-rank decomposition framework utilizing a path connection mechanism to support non-fully-shared dimensions and unify diverse matrix factorization paradigms. The core contribution of this work lies in enabling cross-modal pattern discovery and missing data prediction across heterogeneous datasets. By effectively resolving structural alignment challenges inherent in complex multimodal fusion, PathFinder provides a more generalizable theoretical foundation and practical solution for the joint analysis of heterogeneous data. This approach significantly extends the applicability of tensor decomposition techniques to real-world scenarios where strict dimensional correspondence is unavailable, thereby advancing robust multimodal learning under incomplete or misaligned data conditions.
This study addresses the challenge of effectively disentangling shared, modality-specific, and cross-modal information from continuous representations in multimodal prediction. We propose an explicit disentanglement framework based on latent variables, which decomposes the joint distribution by leveraging invertible normalizing flows and low-rank latent variable models. Furthermore, this approach integrates intermediate-layer contrastive learning with masked autoencoding objectives to achieve unified representation learning guided by structured likelihoods. By analyzing continuous multimodal interaction mechanisms from a latent variable perspective, our method yields robust multimodal information disentanglement and significantly improves predictive performance across multiple benchmarks.
This work addresses the challenges of modality missingness, task interference, and catastrophic forgetting faced by large-scale multimodal models in continuous data streams. To this end, we propose Dual-Decomposed LoRA Experts (DD-LoRA), a novel architecture that dynamically constructs LoRA update matrices through a decoupled pool of modality-specific factors. By integrating a task-partitioning framework, cross-modal guided routing, and a task-key memory mechanism, our approach enables efficient and stable continual learning. As the first to introduce dual-decomposed low-rank structures into continual learning under missing modalities, DD-LoRA significantly mitigates interference between modalities and tasks, supports task-agnostic inference, and effectively prevents forgetting. Extensive experiments on mainstream CMML benchmarks demonstrate substantial performance gains over current state-of-the-art methods, validating the efficacy of architecture-aware LoRA design in real-world multimodal scenarios.
This work addresses the reliance of conventional multimodal large language models on costly and hard-to-scale fully aligned multimodal data. The authors propose a two-stage framework that trains such models using only pairwise modality data. In the first stage, a shared latent space is constructed through within-modality reconstruction and pairwise contrastive learning. In the second stage, new modality encoders are integrated with a pretrained decoder to enable cross-modal transfer and generation. Theoretical analysis establishes conditions under which aligned representations can be achieved using only pairwise data, introducing inductive biases based on partial alignment and minimal latent norms to eliminate the need for complete joint multimodal observations. The approach successfully incorporates 3D point clouds and tactile modalities into a pretrained model, achieving strong cross-modal performance across three pairs of modalities.
This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.