text-guided disentanglement

Designs and implements representation-learning models and analysis pipelines that decompose multimodal (text and visual) embeddings into disentangled components, using text as guidance to separate shared semantic factors from modality-specific features. This includes methods to select top-k matching textual descriptions, derive modality-shared features (for example via residuals), reconstruct modality-specific representations, align prompt and visual embeddings, and impose constraints such as orthogonality between feature subspaces to enforce semantic disentanglement.

text-guideddisentanglement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

An Information Criterion for Controlled Disentanglement of Multimodal Data

Oct 31, 2024
CW
Chenyu Wang
🏛️ MIT | Broad Institute of MIT and Harvard | TU Munich

In multimodal data, modality-specific and cross-modal shared information are deeply entangled, making disentanglement challenging. Method: This paper proposes DisentangledSSL, a controllable disentangled representation learning framework that—uniquely under non-minimal necessary information (MNI)-inaccessible conditions—systematically analyzes disentanglement optimality. It integrates the information bottleneck principle, self-supervised contrastive learning, mutual information estimation, variational inference, and modality-adversarial constraints to achieve theoretically grounded, controllable disentanglement. Contribution/Results: Evaluated on synthetic and real-world benchmarks—including vision-language and molecular-phenotype datasets—DisentangledSSL achieves average improvements of 5.2%–9.8% over baselines in prediction and cross-modal retrieval tasks. It significantly enhances model interpretability, robustness to distribution shifts, and counterfactual generation capability, offering principled solutions for multimodal representation learning.

Disentangling shared and modality-specific information in multimodal data.Enabling downstream tasks like counterfactual outcome generation.Improving interpretability and robustness in multimodal representation learning.

Standard sparse autoencoders tend to learn “split dictionaries” in multimodal embedding spaces, where features activate exclusively for a single modality, thereby disrupting cross-modal semantic alignment. To address this issue, this work proposes the first autoencoder framework that integrates group sparsity regularization with cross-modal random masking, explicitly promoting cross-modal consistency within multimodal embedding spaces such as those of CLIP or CLAP. The proposed approach effectively mitigates modality splitting, substantially reduces the occurrence of dead neurons, and enhances the semantic meaningfulness, cross-modal alignment, interpretability, and controllability of the learned features in multimodal tasks.

interpretable semanticsmodality alignmentmultimodal embeddings

Language-Guided Visual Perception Disentanglement for Image Quality Assessment and Conditional Image Generation

Mar 04, 2025
ZY
Zhichao Yang
🏛️ Xidian University | Chongqing Three Gorges University | Paris-Saclay University

Vision-language models like CLIP prioritize semantic understanding while tightly coupling perceptual features, hindering fine-grained separation of perception and semantics—critical for image quality assessment (IQA) and conditional image generation (CIG). Method: We propose a language-guided visual perception disentanglement paradigm. First, we introduce I&2T—the first dual-text-annotated dataset for IQA/CIG, providing disentangled perceptual and semantic descriptions for each image. Building upon it, we design DeCLIP: a framework leveraging multimodal contrastive learning, dual-branch text supervision, and CLIP feature-space remapping to explicitly disentangle perceptual and semantic representations. Contribution/Results: DeCLIP preserves CLIP’s zero-shot transfer capability while significantly improving technical and aesthetic quality estimation accuracy in IQA and enabling fine-grained perceptual controllability in CIG. It outperforms state-of-the-art methods across multiple benchmarks. Code, models, and the I&2T dataset are fully open-sourced.

Disentangle semantic and perceptual features in imagesEnhance conditional image generation with guided text descriptionsImprove image quality assessment using disentangled representations

Multimodal extension is often hindered by the high annotation cost of large-scale paired data, particularly in specialized domains such as medical imaging and molecular analysis. This work proposes TextME, a framework that, for the first time, maps diverse modalities—including images, audio, 3D, X-rays, and molecular data—into the embedding space of large language models using only textual descriptions, without any modality-paired supervision. By leveraging the geometric structure of pretrained contrastive encoders, TextME enables zero-shot cross-modal transfer purely through text-driven alignment. This approach establishes a novel paradigm for modality extension, achieving effective zero-shot retrieval across heterogeneous, unaligned modalities—such as audio-to-image or 3D-to-X-ray—while preserving the representational capacity of the pretrained encoders.

modality expansionmultimodal representationpaired datasets

The evolution of embedding techniques from word vectors to multimodal representations remains fragmented, lacking a unified framework that integrates advances across linguistic, cross-lingual, personalized, and multimodal domains—particularly for embodied multimodal learning in large language models. Method: We systematically survey static and contextual language representations, cross-lingual and personalized modeling, sentence/document embeddings, and multimodal fusion in vision, robotics, and cognitive science. We synthesize recent progress in interpretability, model compression, numerical encoding, and bias mitigation, and propose a novel paradigm emphasizing strong alignment across non-textual modalities and scalable training. Contributions: We construct a comprehensive knowledge graph of end-to-end embedding technologies—from Word2Vec and BERT to GPT, generative topic models, and multimodal alignment/distillation methods—identifying key technical bottlenecks and ethical challenges. This work delivers the first systematic roadmap for multimodal, embodied learning in foundation models.

Addressing compression, interpretability and bias challengesEvolving from sparse to dense word embeddingsExtending embeddings to multimodal domains

Latest Papers

What's happening recently
View more

Existing multimodal sentiment analysis methods overlook signals shared exclusively by subsets of modalities, limiting the expressiveness and discriminative power of learned representations. To address this, this work proposes a tri-subspace disentanglement framework that explicitly decomposes features into three complementary subspaces: globally shared, pairwise modality-shared, and modality-private. Subspace independence is enforced through disentanglement supervision and structural regularization. Furthermore, a Subspace-Aware Cross-Attention (SACA) module is introduced to enable fine-grained, adaptive fusion. This approach is the first to model multi-granularity cross-modal affective cues, achieving state-of-the-art performance on CMU-MOSI and CMU-MOSEI with MAE = 0.691 and ACC-7 = 54.9%, respectively. The framework also demonstrates successful transferability to multimodal intent recognition tasks.

cross-modal synergiesmodality-specific featuresMultimodal Sentiment Analysis

This work addresses the limitation of existing cross-modal alignment methods, which often conflate semantic and non-semantic information, leading to insufficient semantic consistency and alignment bias caused by modality gaps. To overcome this, we propose a semantic alignment framework based on constrained disentanglement and distribution sampling. Specifically, a dual-path UNet architecture adaptively disentangles visual and linguistic representations into semantic and modality-specific components, aligning only the extracted semantic factors. Furthermore, a multi-constraint optimization strategy combined with distribution-aware sampling is introduced to effectively bridge inter-modality discrepancies, thereby enhancing the reasonableness and robustness of alignment. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches across multiple benchmarks and backbone architectures, achieving performance gains of 6.6% to 14.2%.

cross-modal alignmentembedding decouplingmodality gap

Text-to-image diffusion models struggle to simultaneously maintain subject consistency and text alignment in multi-image generation; existing approaches rely on fine-tuning or image conditioning, incurring high computational costs and poor generalization. This paper proposes a training-free geometric disentanglement method: for the first time, it leverages the geometric structure of the text embedding space, explicitly decoupling shared subject representations from scene descriptions via token-level embedding rescaling and semantic suppression—thereby mitigating cross-frame semantic leakage. The method is plug-and-play and requires only a single text prompt. Experiments demonstrate substantial improvements across multiple benchmarks: subject consistency increases by 32.7% (ID preservation rate) and text alignment accuracy improves by +0.18 in CLIP-Score, surpassing state-of-the-art methods including 1Prompt1Story.

Achieving training-free subject-consistent generation without fine-tuningEliminating semantic leakage and text misalignment in embeddingsPreserving subject consistency across multiple text-to-image outputs

This work addresses the instability in multimodal sentiment analysis caused by representation misalignment when using independently pretrained modality encoders, which often fail to surpass strong text-only baselines. To overcome this limitation, the authors propose a multimodal framework centered on explicit alignment: visual inputs are first transformed into structured textual descriptions via a vision-language model, thereby unifying heterogeneous modalities into a shared linguistic space. Cross-modal alignment is further enhanced through semantic token selection and batch uniformity regularization, prioritizing alignment over complex fusion mechanisms. The proposed approach significantly outperforms existing unimodal and multimodal methods across multiple sentiment and emotion benchmarks, demonstrating the effectiveness and robustness of an alignment-first strategy.

affective computingmodality encodersmultimodal sentiment analysis

This work addresses the limited controllability in medical image generation caused by the large modality gap and semantic entanglement between text and images. To this end, the authors propose a vision-guided textual disentanglement framework that introduces, for the first time, a cross-modal latent alignment mechanism to decompose unstructured clinical text into disentangled semantic representations—such as anatomical structure and imaging style. These disentangled features are then integrated into a Diffusion Transformer (DiT) architecture via a Hybrid Feature Fusion Module (HFFM), enabling fine-grained structural control during image synthesis. Experimental results on three medical imaging datasets demonstrate that the proposed method significantly outperforms existing approaches, achieving not only higher image generation quality but also improved performance on downstream classification tasks.

fine-grained controlmedical image generationmodality gap

Hot Scholars

GP

Georgios Papoudakis

Unknown affiliation
reinforcement learningmulti-agent systemslarge language models
KC

Kangning Cui

Research Assistant Professor of Computer Science, Wake Forest University
Applied MathematicsComputational SustainabilityMedical Imaging
YL

Yuyang Li

Institute for AI, Peking University
Robotic ManipulationTactile SensingHuman-Object Interaction
SW

Shuo Wang

University of Science and Technology of China
Computer VisionMultimedia
VS

Vera Schmitt

Head of XplaiNLP Research Group at TU Berlin
NLP/LLMsXAIHCIDisinformation