Score
Designs, implements, and evaluates conditional generative samplers and training procedures built on diffusion and score-based frameworks—e.g., conditional DDPMs, conditional denoising/score models, latent diffusion, and Poisson-diffusion bridges—that produce samples consistent with specified conditions or partial observations. Work covers constructing conditioning mechanisms and bridge samplers, deriving training losses and score/estimator updates, building latent-space and diffusion-based embeddings and initialization strategies, applying domain-decomposed or specialized variants (e.g., graph-latent or motion-prior samplers), and analyzing sampling quality, coverage, and scalability.
Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
This work addresses the challenge of zero-shot conditional generation from pretrained unconditional diffusion models—specifically, generating samples satisfying complex logical constraints (e.g., structural conditions on tables, images, or time series) without fine-tuning. We propose a neural-symbolic soft-constraint embedding method that encodes first-order logic constraints as differentiable soft penalties and directly perturbs the score function to achieve theoretically consistent approximation of the conditional distribution—bypassing classifier-guided sampling or costly retraining. Our approach integrates score-based modeling, symbolic logic encoding, score correction, and stabilized sampling. Experiments across diverse data modalities demonstrate that our method achieves high-fidelity approximation of the true conditional distribution, significantly outperforming existing zero-shot conditional generation baselines.
This work investigates theoretical performance guarantees for denoising diffusion models under score function mismatch, specifically in the zero-shot conditional sampling setting—where the target conditional distribution differs from the unconditional training distribution. Methodologically, it establishes the first explicit convergence bounds dependent on data dimensionality and conditional structure; quantitatively characterizes the asymptotic sampling bias in terms of accumulated score mismatch; and proposes a bias-optimal linear zero-shot conditional sampler. The theoretical analysis rigorously derives these results using probability measure convergence theory and linear inverse problem modeling, covering common target distributions such as those with bounded support and Gaussian mixtures. Numerical experiments demonstrate that the proposed sampler effectively suppresses distributional bias induced by score mismatch.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work addresses the theoretical complexity and inconsistent interpretations of diffusion models by proposing a concise, self-contained unifying framework grounded in signal processing. Methodologically, it abandons conventional Markov chain and reverse stochastic differential equation (SDE) formulations, instead modeling generation as a stochastic walk coupled with the Tweedie formula—thereby decoupling score estimation, noise scheduling, and sampling, and enabling likelihood-free conditional generation. Key contributions include: (1) the first self-contained theoretical interpretation independent of reverse SDEs or probability flows; (2) full decoupling of noise scheduling between training and sampling, enhancing flexibility in conditional synthesis; and (3) faithful reproduction and unification of major models—including DDPM, DDIM, and Score SDE—while maintaining state-of-the-art performance on image generation and inverse problems, significantly improving both theoretical parsimony and practical interpretability.
This work addresses the lack of a systematic theoretical understanding of how finite-sample learning, neural network parameterization, and numerical discretization jointly affect generation quality in diffusion models. The authors develop a unified framework for convergence and generalization analysis, decomposing the overall generation error— for the first time—into four interpretable components: forward truncation error, backward discretization error, generalization error (accounting for both data finiteness and forward discretization), and optimization gap. Leveraging a ResNet-type score estimator and combining tools from numerical analysis of stochastic differential equations with total variation distance bounds, they quantitatively characterize the joint influence of training sample size, temporal grid density, and optimization accuracy on generation fidelity, thereby establishing end-to-end theoretical guarantees.
This work addresses the longstanding challenge in diffusion modeling of simultaneously enabling simulation-free training and finite-time generation. The authors propose a novel reference diffusion process whose marginal distributions exactly match the target distribution, and whose time-varying conditional distributions facilitate a well-defined reversal. This formulation reveals that score matching naturally arises as the consequence of reversing the reference process and further shows that conditional flow matching corresponds to its small-noise limiting case. The resulting framework is the first to jointly support training without requiring forward simulations and generation within a finite time horizon, thereby not only broadening the theoretical foundations of diffusion models but also enhancing their practical flexibility.
This work demonstrates that diffusion models, score-based generative models, and flow matching methods—despite their apparent formal differences—share a unified continuous-time generative mechanism. By constructing a measure-theoretic framework, the paper unifies these approaches as learning time-dependent vector fields that transport a reference distribution to the data distribution, with distributional evolution governed by the continuity equation and the Fokker–Planck equation. It establishes, for the first time under a common perspective, the equivalence and distinctions among the three paradigms, clarifies the relationship between probability flow ODEs and stochastic backward dynamics, and identifies flow matching as essentially a velocity field regression problem. The study further provides a systematic comparison of objective functions, sampling strategies, and discretization errors, links the framework to Schrödinger bridges and entropy-regularized optimal transport, and summarizes theoretical guarantees and open challenges regarding approximation capacity, stability, and scalability.
Standard diffusion models rely on classifier-free guidance (CFG) during inference to generate high-quality conditional samples, indicating that their training objective lacks explicit modeling of inter-class discriminability. This work proposes the Maximum Conditional-to-Unconditional Likelihood Ratio (MCLR) alignment objective, which explicitly maximizes the ratio between conditional and unconditional likelihoods during training, enabling the model to achieve CFG-level generation quality under standard reverse sampling without requiring inference-time guidance. We provide the first theoretical proof that CFG corresponds to the optimal solution of a weighted MCLR objective, offering a mechanistic explanation for CFG. Experiments demonstrate that models fine-tuned with MCLR match CFG in both qualitative and quantitative metrics while eliminating the need for guidance during inference.