multimodal latent diffusion

Design and implement diffusion-based generative models that learn and operate in a shared stochastic latent space linking multiple modalities by jointly training modality-specific encoders and decoders (including unified or tri-branch architectures). Build and fit priors over those latents (e.g., flow-based priors) and engineering subset-conditioned any-to-any generation mechanisms to enable generation or translation between arbitrary combinations of input and output modalities.

multimodallatentdiffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge

Oct 23, 2025
NB
Nimrod Berman
🏛️ Bosch AI Center | Ben-Gurion University of the Negev | Technical University of Munich

This work addresses key limitations in cross-modal translation (MT)—including reliance on aligned dimensions, Gaussian prior assumptions, and modality-specific architectures—by proposing a universal, theoretically grounded solution. We introduce the Latent Denoising Diffusion Bridging Model (LDDBM), a framework enabling bidirectional translation between arbitrary modalities without requiring dimension-wise alignment or shared prior assumptions. LDDBM employs a domain-agnostic encoder-decoder architecture that jointly optimizes contrastive alignment loss and predictive loss within a shared latent space, augmented by a latent-space noise prediction mechanism to enhance training stability. Experiments demonstrate that LDDBM significantly outperforms state-of-the-art methods on diverse tasks—including multi-view-to-3D reconstruction, image super-resolution, and multi-view scene synthesis—establishing a new strong baseline for general-purpose cross-modal translation.

Enabling arbitrary modality pairs without requiring aligned dimensionsOvercoming restrictive assumptions in existing modality translation methodsTranslating information across different sensory modalities

Generator Matching: Generative modeling with arbitrary Markov processes

Oct 27, 2024
PH
P. Holderrieth
🏛️ MIT | Meta | Weizmann Institute of Science

This work addresses the challenge of developing a modality-agnostic, unified generative modeling framework that generalizes and unifies Markovian generative approaches. We propose Generator Matching—a principled framework grounded in arbitrary Markov processes (including continuous diffusion, flow, discrete transition, and jump processes)—which models data distributions by rigorously aligning conditional and marginal generators. Our contributions are threefold: (i) the first unified treatment of diffusion models, flow matching, and discrete diffusion under a single theoretical umbrella; (ii) the first systematic extension of generative modeling to non-standard jump processes; and (iii) support for rigorous superposition of Markov generators and joint multimodal modeling. Experiments demonstrate substantial performance gains on image and multimodal generation tasks, with superposed jump processes delivering significant empirical improvements.

Enables rigorous multimodal model constructionExpands design to new Markov processesUnifies various generative modeling methods

This study addresses the fundamental disconnect between understanding and generation capabilities in multimodal generative AI. We propose a unified modeling paradigm that systematically characterizes the intrinsic trade-offs between autoregressive and diffusion-based modeling, as well as between dense and Mixture-of-Experts (MoE) architectures. Our approach integrates multimodal large language models (MLLMs), diffusion probabilistic modeling, MoE-based sparse computation, and cross-modal alignment mechanisms, grounded in large-scale multimodal pretraining data analysis. This enables precise delineation of modeling differences and complementary boundaries between the two dominant paradigms. The work yields an extensible “understanding–generation” joint modeling decision atlas, offering both theoretical foundations and practical design principles for efficient, unified multimodal generative AI systems.

Exploring architectures for understanding and generation unificationSummarizing datasets for multi-modal generative AI pretrainingUnifying multi-modal LLMs and diffusion models for AI

Diffusion Models: A Comprehensive Survey of Methods and Applications

Sep 02, 2022
LY
Ling Yang
🏛️ Peking University | University of California, Los Angeles | Carnegie Mellon University | University of California, Merced

Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.

Categorizing research into efficient sampling and likelihood estimationReviewing interdisciplinary applications across scientific fieldsSurveying diffusion models' methods and applications comprehensively

Conditional sampling within generative diffusion models

Sep 15, 2024
ZZ
Zheng Zhao
🏛️ Uppsala University

This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.

addressing Bayesian inverse problemsconditional sampling in generative modelsleveraging joint and marginal distributions

Latest Papers

What's happening recently
View more

Existing cross-modal generation methods are limited by reliance on text-aligned data, fully paired training setups, or deterministic mappings, hindering flexible and consistent arbitrary-to-arbitrary modality synthesis. This work proposes an end-to-end unified multimodal latent diffusion framework that jointly trains modality-specific encoders and decoders with a streaming prior within a shared stochastic latent space. By introducing a variational inference–based routing objective, the approach balances consistency, predictive adequacy, and content minimality. It is the first to extend latent diffusion models to multimodal arbitrary-to-arbitrary generation, supporting both conditional synthesis and unconditional joint sampling. Experiments demonstrate that the method matches or surpasses state-of-the-art baselines in conditional generation on PolyMNIST-Quadrant-Labels and large-scale image–text–audio benchmarks, while achieving significantly improved consistency in unconditional generation compared to existing approaches.

any-to-any generationcross-modal generationlatent coherence

Existing mask generation methods struggle with alignment difficulties and training instability in multimodal settings, hindering unified generative modeling of discrete (e.g., text) and continuous (e.g., image) data. To address this, this work proposes the CoM-DAD framework, which introduces a hierarchical dual-process generative mechanism: it first models the cross-modal semantic manifold via continuous latent diffusion, then leverages this semantic representation as a prior to generate concrete tokens through a discrete absorbing diffusion process with variable-rate noise scheduling. The approach innovatively integrates coupled manifold-aware discrete absorbing diffusion, adaptive noise scheduling, and a stochastic mixed-modality transfer strategy, achieving efficient cross-modal alignment without relying on heavy contrastive dual encoders. Experiments demonstrate that the method significantly enhances training stability and achieves superior generation quality and semantic coherence in unified text-to-image synthesis, offering a scalable new paradigm for multimodal generation.

discrete-continuous gapmasked language modelsmodality alignment

This work addresses the limitation of existing diffusion models, which are largely confined to unimodal or bimodal settings and struggle to jointly model text, images, and audio. We propose the first trilingual masked discrete diffusion model, pretrained from scratch with 3 billion parameters on a dataset comprising 6.4 trillion tokens, enabling unified generation across all three modalities. Through systematic investigation of multimodal scaling laws, modality mixing ratios, noise scheduling, and batch size effects, we introduce an SDE-based reparameterization method that decouples physical and logical batch sizes, significantly simplifying hyperparameter tuning. Additionally, we design an efficient inference sampling strategy that achieves strong performance across text generation, text-to-image synthesis, and text-to-speech tasks, establishing the first comprehensive open benchmark for multimodal diffusion models.

discrete diffusionmasked diffusion modelsmultimodal scaling laws

This work addresses the limitation of conventional variational autoencoders (VAEs), whose encoders—constrained by the reparameterization trick—struggle to model complex posterior distributions. To overcome this, the authors propose a novel encoder that, for the first time, integrates a diffusion model into the VAE encoding process. They further introduce an alternating training strategy inspired by the Expectation-Maximization (EM) algorithm, which effectively aligns the optimization objectives of the encoder and decoder, thereby ensuring reliable synchronization in the latent space. The proposed approach preserves the simplicity and efficiency of standard diffusion model training while substantially enhancing the model’s capacity to capture complex data distributions and improving reconstruction quality.

decoder pressurediffusion encoderencoder-decoder synchronization

Hot Scholars

YL

Yuebing Liang

Massachusetts Institute of Technology, University of Hong Kong
urban computingintelligent transportation
ZM

Zhendong Mao

University of Science and Technology of China
CV,NLP
CY

Chenxiao Yang

Toyota Technological Institute at Chicago
Machine Learning
ZH

Zhenliang He

Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionAIGC