Score
Designs, implements, and evaluates generative models that perform diffusion-based sampling and denoising in a learned latent space, including their architectures, training objectives, and inference procedures. This work covers modeling non-Gaussian and multimodal latent transitions, generating latent trajectories for forecasting, and improving model fit to complex sequential latent dynamics.
Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
In diffusion model training, uniform timestep sampling ignores the variance heterogeneity of gradients across timesteps, rendering high-variance timesteps convergence bottlenecks. To address this, we propose the first online evaluation mechanism that dynamically assesses—per iteration—the impact of gradient updates on the objective function, enabling adaptive, non-uniform timestep sampling focused on optimization-sensitive timesteps. Our method transcends conventional static weighting or heuristic sampling by unifying gradient variance analysis, objective impact tracking, and importance sampling. It achieves principled, real-time timestep prioritization without requiring precomputed statistics or architectural modifications. Experiments across diverse datasets, noise schedules, and network architectures demonstrate consistent improvements: our approach accelerates convergence and enhances final model performance compared to state-of-the-art timestep sampling and weighting strategies.
This work addresses the theoretical complexity and inconsistent interpretations of diffusion models by proposing a concise, self-contained unifying framework grounded in signal processing. Methodologically, it abandons conventional Markov chain and reverse stochastic differential equation (SDE) formulations, instead modeling generation as a stochastic walk coupled with the Tweedie formula—thereby decoupling score estimation, noise scheduling, and sampling, and enabling likelihood-free conditional generation. Key contributions include: (1) the first self-contained theoretical interpretation independent of reverse SDEs or probability flows; (2) full decoupling of noise scheduling between training and sampling, enhancing flexibility in conditional synthesis; and (3) faithful reproduction and unification of major models—including DDPM, DDIM, and Score SDE—while maintaining state-of-the-art performance on image generation and inverse problems, significantly improving both theoretical parsimony and practical interpretability.
This paper investigates the fundamental differences in representation learning mechanisms between diffusion models and classification models, specifically addressing whether diffusion models learn more balanced and comprehensive data representations. Method: We establish the first theoretical framework to comparatively analyze their feature learning dynamics, proving that the denoising objective implicitly optimizes feature balance and completeness. Combining theoretical analysis with experiments on both synthetic and real-world datasets—under standard U-Net architectures—we validate the mechanism via feature visualization and interpretability evaluation. Results: Diffusion models exhibit significantly enhanced inter-class separability and intra-class consistency in intermediate-layer features, outperforming structurally identical classification models. Our core contribution is uncovering the intrinsic regularizing effect of the denoising objective on representation quality, offering a novel theoretical perspective and empirical evidence for understanding the superior representational capacity of diffusion models.
It remains unclear whether the initial noise latent space of diffusion models inherently encodes structured information predictive of generated sample attributes such as class labels. This work addresses this question by leveraging a pretrained classifier to assign confidence scores to unconditionally generated samples and subsequently isolating the corresponding initial noise vectors associated with high-confidence outputs. The study reveals, for the first time, that a class-discriminative structure emerges in the latent space only after filtering by classifier confidence. Experimental results demonstrate that this high-confidence noise subset substantially enhances class separability in the latent space, establishing a novel paradigm for guidance-free conditional generation that requires neither additional annotations nor architectural modifications to the underlying model.
Generative models suffer from a “self-consumption loop” when synthetic data is reused across multiple generations, leading to training instability and model collapse. Method: We propose a latent-space filtering approach that requires no additional data or human annotations. Our method identifies progressive degradation of low-dimensional structure in the latent space across generations, establishes a theoretical analysis framework grounded in this observation, and leverages diffusion models to dynamically characterize the degradation process. We further design a quantitative metric for low-dimensional structural quality and an adaptive filtering strategy to precisely identify and remove low-fidelity synthetic samples from mixed datasets. Contribution/Results: Extensive experiments on multiple real-world benchmarks demonstrate that our method significantly outperforms existing baselines. It effectively mitigates model collapse without any extra training overhead and consistently improves multi-generation training stability and performance.
This work addresses the limitations of existing deep state-space models, which rely on Gaussian assumptions and struggle to capture complex, multimodal latent dynamics. While diffusion models offer strong expressive power, they lack structured mechanisms for temporal reasoning. To bridge this gap, we propose a novel latent variable state-space model that, for the first time, integrates a non-Gaussian diffusion process into the latent state transition mechanism and enables joint training of an autoencoder and a diffusion model on sequential data. This approach overcomes the restrictive Gaussian assumption and supports unified modeling and inference of multimodal temporal dynamics. Experiments demonstrate that our model significantly outperforms state-of-the-art deep state-space models in both fitting accuracy and predictive performance on synthetic time series with complex transition characteristics.
This work addresses the sensitivity of latent diffusion models to sampling perturbations, which often arises from variance collapse in the latent space and leads to degraded generation quality. For the first time, this study explicitly identifies the critical role of sampling perturbation robustness in generative performance and proposes a variance-expansion loss to learn perturbation-robust latent representations while preserving high reconstruction fidelity. The method achieves an adaptive balance through an adversarial trade-off between reconstruction accuracy and latent variance, optimized within a β-VAE encoder framework coupled with diffusion-based sampling. Extensive experiments demonstrate consistent improvements in generation quality across diverse latent diffusion architectures, confirming that enhancing robustness in the latent space effectively stabilizes and elevates image synthesis performance.
This work investigates the generalization mechanisms of diffusion models when they do not memorize training data, revealing the geometric structure and dynamic evolution of their generated distributions. By introducing a data-dependent log-density ridge manifold, the authors characterize a three-stage behavior of generation trajectories—approach, alignment, and sliding—and quantitatively analyze how normal and tangential motions, influenced by training error, govern cross-modal generation capabilities. Integrating manifold analysis, random feature models, and diffusion dynamics, the study establishes the first explicit link between the inductive bias of diffusion models and the geometry of ridge manifolds, elucidating how architectural bias and training accuracy jointly shape generative behavior. The predicted directional effects are validated on synthetic multimodal distributions and latent-space MNIST experiments, demonstrating applicability across both low- and high-dimensional settings.