unified diffusion modeling

Design, build, and analyze diffusion-based generative systems that jointly represent and iteratively denoise shared latent states to produce multiple outputs or tasks (e.g., perception and planning) and/or temporally-extended sequences. This includes specifying diffusion processes and noise schedules, conditioning and bidirectional information-exchange mechanisms between task branches, integration of diffusion modules into multi-task pipelines, and training/evaluation procedures for noise-conditioned multi-task objectives.

unifieddiffusionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.74
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$243K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Denoising Task Difficulty-based Curriculum for Training Diffusion Models

Mar 15, 2024
JK
Jin-Young Kim
🏛️ TwelveLabs | Ajou University

The relative difficulty of denoising tasks across timesteps in diffusion models remains controversial. Method: This work systematically quantifies denoising difficulty per timestep, leveraging both the convergence behavior of denoising error and the relative entropy between true and predicted distributions—revealing that early (low-timestep) denoising is significantly more challenging. Building on this insight, we propose a “curriculum learning” paradigm: timesteps are clustered by difficulty and trained progressively in stages, with joint optimization of the noise schedule. Contribution/Results: Our approach departs from conventional parallel full-timestep training, requiring no architectural or loss-function modifications and remaining compatible with diverse diffusion model enhancements. Extensive experiments on unconditional generation, class-conditional generation, and text-to-image synthesis demonstrate substantial improvements in both model performance and convergence speed.

Curriculum learning for diffusion modelsDenoising task difficulty conflictImproved convergence in image generation

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

Apr 15, 2025
ZP
Ziqi Pang
🏛️ University of Illinois Urbana-Champaign

Generative diffusion models face a fundamental mismatch in discriminative tasks (e.g., referring image segmentation): their inherently tolerant generation mechanism conflicts with the strict accuracy requirements imposed on intermediate inference steps. This work systematically characterizes the non-uniform impact of denoising timesteps on perceptual quality—a previously unexplored insight. We propose three key innovations: (1) a timestep-aware loss that dynamically weights supervision across diffusion steps; (2) distribution-shift-robust data augmentation to enhance generalization; and (3) a prompt-driven interactive backward sampling mechanism enabling iterative refinement. All techniques operate via discriminative fine-tuning—no architectural modifications are required. Evaluated on depth estimation, referring image segmentation, and general perception benchmarks, our approach achieves state-of-the-art performance, significantly improving multimodal discriminative accuracy and interaction robustness.

Address perception degradation in later denoising stepsAlign generative denoising with discriminative tasks for accuracyEnable interactive correction via generative diffusion processes

Conditional Image Synthesis with Diffusion Models: A Survey

Sep 28, 2024
ZZ
Zheyuan Zhan
🏛️ Zhejiang University | University at Buffalo, State University of New York | Zhejiang University of Technology

Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.

Categorizing conditioning approaches in denoising networks and samplingIdentifying unsolved problems in diffusion-based conditional image generationSurveying diffusion models for conditional image synthesis challenges

This work addresses the lack of a general framework for modeling time-varying latent states in existing generative models, which often rely on auxiliary stochastic processes that are difficult to sample. The authors propose a novel approach that treats observation generation as a deterministic mapping of a tractable Markov process, employing an image-space stochastic process generator whose one-time marginal distribution matches that of a target projected process. The key innovation lies in extending Generator Matching—previously limited to static latent variables—to time-varying latent processes for the first time. By integrating stochastic process theory, Markov projections, and flow matching techniques, the method establishes a unified generative modeling framework. This framework not only subsumes existing models with discrete latent processes as special cases but also accommodates a broader class of time-varying latent conditions while rigorously ensuring consistency between the generated and target marginal distributions.

conditional processesgenerative modelsgenerator matching

Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task

Oct 13, 2023
MO
Maya Okawa
🏛️ Harvard University | NTT Research, Inc. | University of Michigan

While generative models produce high-fidelity samples, they remain unreliable for compositional generalization—i.e., synthesizing novel concept combinations outside the training distribution. This work systematically investigates the compositional generalization of conditional diffusion models in a controlled synthetic setting. We decouple and independently manipulate three key data properties—frequency, structural composition, and factor disentanglement—to construct a multi-dimensional out-of-distribution (OOD) evaluation framework. Our core findings reveal a multiplicative emergence of compositional capability: overall performance equals the product of subtask accuracies; the order of capability emergence is governed by data generative structure; and low-frequency compositions incur substantially higher optimization costs. This is the first empirical demonstration of nonlinear, emergent compositional generalization in diffusion models, establishing quantitative links among data structure, capability emergence, and optimization cost—providing a data-centric theoretical foundation and design principles for trustworthy generative compositional reasoning.

Analyzing data requirements for out-of-distribution generationExploring emergence of multiplicative abilities in generative modelsUnderstanding compositional generalization in diffusion models

Latest Papers

What's happening recently
View more

Existing diffusion synchronization methods rely on heuristic designs and require task-specific tuning, limiting their generalization. This work proposes a unified diffusion synchronization framework grounded in optimal control theory. At test time, it employs variational optimization to infer control variables that steer multiple diffusion trajectories toward a consistent output while remaining faithful to the pretrained prior. The approach provides the first principled theoretical foundation for diffusion synchronization and enables cross-modal cooperative generation without any additional training. Evaluated across three distinct tasks, the method substantially outperforms existing baselines, significantly improving both generation consistency and applicability.

collaborative generationdiffusion synchronizationgeneralizability

Deterministic Discrete Denoising

Sep 25, 2025
HS
Hideyuki Suzuki
🏛️ The University of Osaka | The University of Tokyo

To address the inefficiency and instability in sampling caused by stochastic denoising in discrete diffusion models, this paper proposes the first deterministic denoising framework that requires neither model retraining nor continuous embeddings. Methodologically, it constructs a deterministic reverse transition process based on a Markov chain, integrating an enhanced herding algorithm and weak chaotic dynamics to enable deterministic trajectory evolution over discrete state spaces. The core contribution is the systematic introduction of deterministic reverse processes into discrete diffusion modeling—eliminating the inherent randomness and redundant iterations of conventional sampling. Experiments demonstrate significant improvements: up to 3.2× faster sampling and enhanced sample quality (FID reduced by 18.7%, BLEU increased by 2.4%) on both text and image generation tasks. Performance matches that of continuous diffusion models, establishing a novel paradigm for discrete generative modeling.

Aims to improve efficiency and quality in generation tasksProposes deterministic denoising for discrete diffusion modelsReplaces stochastic process without retraining or embeddings

Diffusion Models: A Mathematical Introduction

Nov 13, 2025
SM
Sepehr Maleki
🏛️ University of Lincoln | Trainline

Existing theoretical analyses of diffusion generative models are fragmented and suffer from inconsistent notation, hindering unified understanding and principled development. Method: This work establishes a rigorous, unified mathematical framework grounded in fundamental properties of Gaussian distributions. It systematically derives the closed-form marginal distribution of the forward noising process, the analytical form of the reverse posterior, and the variational lower bound, ultimately yielding an optimization objective equivalent to noise prediction. Contribution/Results: The framework reveals the intrinsic equivalence between DDIM and rectified flow; provides a unified probabilistic interpretation of classifier-guided and classifier-free guidance; and integrates SDE/ODE formulations, the Fokker–Planck equation, flow matching, and multi-scale modeling—ensuring both theoretical coherence and practical implementability. Validated on mainstream models including Stable Diffusion, the framework enables efficient sampling and precise modeling while unifying disparate theoretical perspectives.

Analyzing likelihood estimation and accelerated sampling techniquesDeriving diffusion models from Gaussian distribution fundamentalsExplaining guided diffusion through score correction methods

This work addresses the desynchronization artifacts commonly observed in multimodal diffusion models due to their limited understanding of dynamic generative mechanisms. By formulating an analytically tractable framework based on coupled Ornstein-Uhlenbeck processes, the study integrates non-equilibrium statistical physics and spectral analysis to reveal that modalities stabilize sequentially—each governed by distinct characteristic timescales—rather than synchronously. The authors introduce the concept of a “synchronization gap” to quantify disparities in stabilization rates across modalities, and derive rigorous bounds linking coupling strength to symmetry-breaking stability, interpreting the system as a tunable temporal spectral filter. Experiments on MNIST demonstrate that time-dependent coupling schedules can precisely control modality-specific timescales, offering a theoretically grounded alternative to heuristic guidance strategies.

coupling strengthdiffusion modelsdynamical phase transitions

Hot Scholars

YS

Ying Shan

Distinguished Scientist at Tencent, Director of ARC Lab & AI Lab CVC
Deep learningcomputer visionmachine learningpaid search
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning
EX

Enze Xie

NVIDIA Research, MMLab@HKU
computer visiongenerative AI
MT

Molei Tao

Associate Professor, Georgia Institute of Technology
foundation of machine learningapplied & computational mathstochastic/nonlinear dynamics
BK

Bernhard Kainz

FAU Erlangen-Nürnberg, Imperial College London
human-in-the-loop computingmachine learningmedical image analysis