shared-private representation learning

Designs models and algorithms that decompose inputs from multiple sources into representation components that are shared across sources and components that are private to each source, and that train to align and separate those components. This includes methods for learning modality-shared and modality-specific features, aligning shared components across patterns, handling missing modalities, and reducing representation-learning error for rare or infrequent patterns.

shared-privaterepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing approaches struggle to effectively model the dynamic interplay among redundant, unique, and synergistic information at the sample level in multimodal learning. This work proposes a novel information-theoretic decomposition–based paradigm for multimodal interaction learning, offering the first systematic analysis of the importance of sample-level interactions. By employing a variational architecture, the method explicitly disentangles these three interaction components and integrates a component-aware fine-tuning strategy to adaptively leverage them. Extensive experiments across diverse tasks and architectures consistently demonstrate the superiority of the proposed approach over current state-of-the-art methods, confirming its effectiveness and generality in finely modeling sample-level multimodal interactions.

dynamic interactioninformation decompositionmultimodal interaction

A Closer Look at Multimodal Representation Collapse

May 28, 2025
AC
Abhra Chaudhuri
🏛️ Fujitsu Research of Europe | University of Surrey

This paper addresses modality collapse in multimodal fusion—a systemic failure where models neglect certain modalities during training. We identify its root cause: noisy features induce representation entanglement via shared neurons, compounded by rank deficiency in the fusion head, leading to degraded modality-specific representations. We provide the first theoretical explanation grounded jointly in representation entanglement and low-rank constraints. Building on this insight, we propose an explicit basis reallocation algorithm that enforces cross-modal disentanglement and dynamic weight calibration, while ensuring robust inference under modality missingness. Our method integrates cross-modal knowledge distillation, neuron-level attribution, and rank-constrained fusion head modeling. Evaluated on multiple benchmarks, it significantly mitigates modality collapse, substantially improving model generalization and stability under partial modality absence.

Analyzing noisy feature entanglement causing modality collapseProposing cross-modal distillation to prevent modality collapseUnderstanding modality collapse in multimodal fusion models

An Information Criterion for Controlled Disentanglement of Multimodal Data

Oct 31, 2024
CW
Chenyu Wang
🏛️ MIT | Broad Institute of MIT and Harvard | TU Munich

In multimodal data, modality-specific and cross-modal shared information are deeply entangled, making disentanglement challenging. Method: This paper proposes DisentangledSSL, a controllable disentangled representation learning framework that—uniquely under non-minimal necessary information (MNI)-inaccessible conditions—systematically analyzes disentanglement optimality. It integrates the information bottleneck principle, self-supervised contrastive learning, mutual information estimation, variational inference, and modality-adversarial constraints to achieve theoretically grounded, controllable disentanglement. Contribution/Results: Evaluated on synthetic and real-world benchmarks—including vision-language and molecular-phenotype datasets—DisentangledSSL achieves average improvements of 5.2%–9.8% over baselines in prediction and cross-modal retrieval tasks. It significantly enhances model interpretability, robustness to distribution shifts, and counterfactual generation capability, offering principled solutions for multimodal representation learning.

Disentangling shared and modality-specific information in multimodal data.Enabling downstream tasks like counterfactual outcome generation.Improving interpretability and robustness in multimodal representation learning.

Completed Feature Disentanglement Learning for Multimodal MRIs Analysis

Jul 06, 2024
TL
Tianling Liu
🏛️ Tianjin University | The University of Hong Kong

Existing feature disentanglement (FD)-based multimodal MRI methods suffer from loss of cross-modal shared information when handling ≥3 modalities and lack explicit modeling of relationships among disentangled features during fusion. To address these limitations, we propose a Complete Feature Disentanglement (CFD) strategy and a Dynamic Mixture-of-Experts Fusion (DMF) module. CFD introduces the novel concept of “modality-partially shared features,” enabling full-order disentanglement that preserves both modality-specific and hierarchically shared representations. The DMF module employs a dynamic gating mechanism to jointly model local–global feature interactions, thereby enhancing both interpretability and fusion accuracy. Evaluated on three multimodal MRI classification tasks, our approach consistently outperforms state-of-the-art methods, demonstrating the effectiveness of restoring shared-information integrity and dynamically modeling inter-feature relationships.

Improved feature disentanglement and fusion strategiesInadequate feature relationship interpretationLoss of shared information in multimodal MRIs

This work addresses the modality gap in multimodal representation learning induced by the InfoNCE objective, which manifests as a conflict between inter-modal alignment and uniformity, as well as intra-modal alignment inconsistencies. The paper proposes the first framework that decouples alignment and uniformity in multimodal learning, employing Hölder divergence–based alignment optimization alongside a dedicated uniformity loss to effectively mitigate these conflicts. Theoretically, the proposed objective is shown to serve as a valid proxy for the global Hölder divergence between multimodal distributions. Notably, the method requires no task-specific components and consistently improves performance across both discriminative tasks (e.g., retrieval) and generative tasks (e.g., UnCLIP), demonstrating its generality and effectiveness.

alignment-uniformity conflictdistribution gapInfoNCE

Latest Papers

What's happening recently
View more

This study addresses the limitation of existing multimodal joint decomposition methods that necessitate full-dimensional sharing. To overcome this constraint, we propose PathFinder, a generalized joint low-rank decomposition framework utilizing a path connection mechanism to support non-fully-shared dimensions and unify diverse matrix factorization paradigms. The core contribution of this work lies in enabling cross-modal pattern discovery and missing data prediction across heterogeneous datasets. By effectively resolving structural alignment challenges inherent in complex multimodal fusion, PathFinder provides a more generalizable theoretical foundation and practical solution for the joint analysis of heterogeneous data. This approach significantly extends the applicability of tensor decomposition techniques to real-world scenarios where strict dimensional correspondence is unavailable, thereby advancing robust multimodal learning under incomplete or misaligned data conditions.

Common pattern discoveryJoint low-rank decompositionLinked datasets

This study addresses the challenge of effectively disentangling shared, modality-specific, and cross-modal information from continuous representations in multimodal prediction. We propose an explicit disentanglement framework based on latent variables, which decomposes the joint distribution by leveraging invertible normalizing flows and low-rank latent variable models. Furthermore, this approach integrates intermediate-layer contrastive learning with masked autoencoding objectives to achieve unified representation learning guided by structured likelihoods. By analyzing continuous multimodal interaction mechanisms from a latent variable perspective, our method yields robust multimodal information disentanglement and significantly improves predictive performance across multiple benchmarks.

Cross-modal dependenciesInformation decompositionLatent variable model

This work addresses the challenges of modality missingness, task interference, and catastrophic forgetting faced by large-scale multimodal models in continuous data streams. To this end, we propose Dual-Decomposed LoRA Experts (DD-LoRA), a novel architecture that dynamically constructs LoRA update matrices through a decoupled pool of modality-specific factors. By integrating a task-partitioning framework, cross-modal guided routing, and a task-key memory mechanism, our approach enables efficient and stable continual learning. As the first to introduce dual-decomposed low-rank structures into continual learning under missing modalities, DD-LoRA significantly mitigates interference between modalities and tasks, supports task-agnostic inference, and effectively prevents forgetting. Extensive experiments on mainstream CMML benchmarks demonstrate substantial performance gains over current state-of-the-art methods, validating the efficacy of architecture-aware LoRA design in real-world multimodal scenarios.

Continual LearningContinual Missing Modality LearningLarge Multimodal Models

This work addresses the reliance of conventional multimodal large language models on costly and hard-to-scale fully aligned multimodal data. The authors propose a two-stage framework that trains such models using only pairwise modality data. In the first stage, a shared latent space is constructed through within-modality reconstruction and pairwise contrastive learning. In the second stage, new modality encoders are integrated with a pretrained decoder to enable cross-modal transfer and generation. Theoretical analysis establishes conditions under which aligned representations can be achieved using only pairwise data, introducing inductive biases based on partial alignment and minimal latent norms to eliminate the need for complete joint multimodal observations. The approach successfully incorporates 3D point clouds and tactile modalities into a pretrained model, achieving strong cross-modal performance across three pairs of modalities.

aligned datasetsmultimodal LLMspairwise modalities

This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.

feature imputationincomplete observationsmissing modalities

Hot Scholars

ZC

Zhiguang Cao

Singapore Management University
Learning to OptimizeNeural Combinatorial OptimizationComputational Intelligence
SJ

Shengqin Jiang

Nanjing University of Information Science and Technology
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
TE

Thomas Ebner

Fraunhofer Heinrich Hertz Institute, HHI
Volumetric VideoVirtual RealityAugmented RealityMixed Reality
YZ

Yifan Zhan

The University of Tokyo
3D Vision