Score
Design and implement diffusion-based generative models that learn and operate in a shared stochastic latent space linking multiple modalities by jointly training modality-specific encoders and decoders (including unified or tri-branch architectures). Build and fit priors over those latents (e.g., flow-based priors) and engineering subset-conditioned any-to-any generation mechanisms to enable generation or translation between arbitrary combinations of input and output modalities.
This work addresses key limitations in cross-modal translation (MT)—including reliance on aligned dimensions, Gaussian prior assumptions, and modality-specific architectures—by proposing a universal, theoretically grounded solution. We introduce the Latent Denoising Diffusion Bridging Model (LDDBM), a framework enabling bidirectional translation between arbitrary modalities without requiring dimension-wise alignment or shared prior assumptions. LDDBM employs a domain-agnostic encoder-decoder architecture that jointly optimizes contrastive alignment loss and predictive loss within a shared latent space, augmented by a latent-space noise prediction mechanism to enhance training stability. Experiments demonstrate that LDDBM significantly outperforms state-of-the-art methods on diverse tasks—including multi-view-to-3D reconstruction, image super-resolution, and multi-view scene synthesis—establishing a new strong baseline for general-purpose cross-modal translation.
This work addresses the challenge of developing a modality-agnostic, unified generative modeling framework that generalizes and unifies Markovian generative approaches. We propose Generator Matching—a principled framework grounded in arbitrary Markov processes (including continuous diffusion, flow, discrete transition, and jump processes)—which models data distributions by rigorously aligning conditional and marginal generators. Our contributions are threefold: (i) the first unified treatment of diffusion models, flow matching, and discrete diffusion under a single theoretical umbrella; (ii) the first systematic extension of generative modeling to non-standard jump processes; and (iii) support for rigorous superposition of Markov generators and joint multimodal modeling. Experiments demonstrate substantial performance gains on image and multimodal generation tasks, with superposed jump processes delivering significant empirical improvements.
This study addresses the fundamental disconnect between understanding and generation capabilities in multimodal generative AI. We propose a unified modeling paradigm that systematically characterizes the intrinsic trade-offs between autoregressive and diffusion-based modeling, as well as between dense and Mixture-of-Experts (MoE) architectures. Our approach integrates multimodal large language models (MLLMs), diffusion probabilistic modeling, MoE-based sparse computation, and cross-modal alignment mechanisms, grounded in large-scale multimodal pretraining data analysis. This enables precise delineation of modeling differences and complementary boundaries between the two dominant paradigms. The work yields an extensible “understanding–generation” joint modeling decision atlas, offering both theoretical foundations and practical design principles for efficient, unified multimodal generative AI systems.
Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
Existing cross-modal generation methods are limited by reliance on text-aligned data, fully paired training setups, or deterministic mappings, hindering flexible and consistent arbitrary-to-arbitrary modality synthesis. This work proposes an end-to-end unified multimodal latent diffusion framework that jointly trains modality-specific encoders and decoders with a streaming prior within a shared stochastic latent space. By introducing a variational inference–based routing objective, the approach balances consistency, predictive adequacy, and content minimality. It is the first to extend latent diffusion models to multimodal arbitrary-to-arbitrary generation, supporting both conditional synthesis and unconditional joint sampling. Experiments demonstrate that the method matches or surpasses state-of-the-art baselines in conditional generation on PolyMNIST-Quadrant-Labels and large-scale image–text–audio benchmarks, while achieving significantly improved consistency in unconditional generation compared to existing approaches.
本文探讨了扩散模型和流模型在表示学习中的应用,提出一个三层框架来组织现有工作,并分类方法以解决图像分类等任务中的挑战。
Existing mask generation methods struggle with alignment difficulties and training instability in multimodal settings, hindering unified generative modeling of discrete (e.g., text) and continuous (e.g., image) data. To address this, this work proposes the CoM-DAD framework, which introduces a hierarchical dual-process generative mechanism: it first models the cross-modal semantic manifold via continuous latent diffusion, then leverages this semantic representation as a prior to generate concrete tokens through a discrete absorbing diffusion process with variable-rate noise scheduling. The approach innovatively integrates coupled manifold-aware discrete absorbing diffusion, adaptive noise scheduling, and a stochastic mixed-modality transfer strategy, achieving efficient cross-modal alignment without relying on heavy contrastive dual encoders. Experiments demonstrate that the method significantly enhances training stability and achieves superior generation quality and semantic coherence in unified text-to-image synthesis, offering a scalable new paradigm for multimodal generation.
This work addresses the limitation of existing diffusion models, which are largely confined to unimodal or bimodal settings and struggle to jointly model text, images, and audio. We propose the first trilingual masked discrete diffusion model, pretrained from scratch with 3 billion parameters on a dataset comprising 6.4 trillion tokens, enabling unified generation across all three modalities. Through systematic investigation of multimodal scaling laws, modality mixing ratios, noise scheduling, and batch size effects, we introduce an SDE-based reparameterization method that decouples physical and logical batch sizes, significantly simplifying hyperparameter tuning. Additionally, we design an efficient inference sampling strategy that achieves strong performance across text generation, text-to-image synthesis, and text-to-speech tasks, establishing the first comprehensive open benchmark for multimodal diffusion models.
This work addresses the limitation of conventional variational autoencoders (VAEs), whose encoders—constrained by the reparameterization trick—struggle to model complex posterior distributions. To overcome this, the authors propose a novel encoder that, for the first time, integrates a diffusion model into the VAE encoding process. They further introduce an alternating training strategy inspired by the Expectation-Maximization (EM) algorithm, which effectively aligns the optimization objectives of the encoder and decoder, thereby ensuring reliable synchronization in the latent space. The proposed approach preserves the simplicity and efficiency of standard diffusion model training while substantially enhancing the model’s capacity to capture complex data distributions and improving reconstruction quality.