Score
Design and implement transformer-based modulation modules for diffusion generative models that adaptively scale and shift intermediate feature maps and inject conditioning signals. Build and evaluate DIT-style components that incorporate elapsed-time or auxiliary context into the denoising trajectory and calibrate generation to heteroscedastic noise.
This survey addresses key challenges in applying denoising diffusion models to computer vision—namely, fragmented applications, unclear theoretical connections, and low sampling efficiency—by establishing the first comprehensive, CV-oriented diffusion model taxonomy. Methodologically, it unifies the three dominant paradigms—Denoising Diffusion Probabilistic Models (DDPM), Noise Conditional Score Networks (NCSN), and Stochastic Differential Equations (SDE)—and introduces a multi-perspective classification framework. It rigorously clarifies their mathematical relationships with VAEs and GANs in terms of probabilistic modeling, score matching, and variational inference. The work innovatively elucidates the generative mechanisms underlying forward noising and reverse denoising processes, and precisely characterizes core challenges including scalability, sampling efficiency, and conditional control. As a foundational reference, this survey has been widely cited in subsequent research on efficient sampling, cross-modal diffusion, and theoretical generalizations, significantly shaping the development of diffusion-based vision methodologies.
Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.
This work addresses the unclear mechanism of massive activations (MAs) in Diffusion Transformers (DiTs), which leads to insufficient generation detail and weak representational discriminability. The study reveals for the first time that MAs are spatially concentrated on image tokens and channel-wise focused on fixed dimensions, predominantly governed by the denoising timestep. Building on this insight, the authors propose EMA, a unified, training-free modulation framework that enhances generation quality through MA-driven detail guidance and boosts dense feature discriminability via MA-based modulation. Experiments demonstrate that EMA consistently improves both image generation fidelity and visual representation performance across diverse DiT architectures, offering strong local refinement capabilities while maintaining computational efficiency during inference.
Existing diffusion models often lack explicit awareness of image quality during the denoising process, leading to misaligned outputs, visual inconsistencies, and insufficient fidelity. To address this limitation, this work proposes a Quality Representation Module (QRM), which employs a lightweight Transformer to learn quality-aware representations conditioned on both textual prompts and timesteps. These representations modulate the adaptive LayerNorm layers within a Diffusion Transformer (DiT), thereby injecting quality-sensitive signals into the denoising dynamics. Notably, QRM introduces—for the first time—a lightweight quality-aware modulation mechanism into DiT architectures without altering the sampling strategy or backbone structure. Extensive experiments demonstrate that QRM consistently enhances image quality across multiple DiT baselines, and ablation studies confirm the effectiveness of its loss formulation and architectural design.
To address high gradient variance, slow convergence, and reliance on normalization layers (e.g., AdaLN) in Diffusion Transformers (DiTs), this work proposes a magnitude-preserving network design and Rotation Modulation—a novel conditional modulation mechanism. The magnitude-preserving design replaces conventional normalization layers by constraining activation magnitudes, thereby enhancing training stability. Rotation Modulation parameterizes conditional transformations on the SO(2) group, substituting AdaLN’s scale-and-shift operations with lightweight, learnable 2D rotations. This is the first introduction of magnitude preservation into DiT architectures and the first rotation-based conditional modulation paradigm for diffusion models. Experiments demonstrate a 12.8% reduction in FID score; combining rotation modulation with scaling matches AdaLN’s performance while reducing parameter count by 5.4%. The implementation is open-sourced.
Diffusion models suffer from inefficient representation learning and limited generation quality due to semantically impoverished latent spaces. To address this, we propose REPA (Representation Alignment), a novel regularization method that explicitly aligns denoising latent states—corrupted by noise—with clean-image representations extracted from high-quality external vision encoders (e.g., CLIP or DINO) within diffusion Transformers (DiT/SiT). This alignment is enforced via a projection-based loss, optimized end-to-end to enhance semantic consistency in the latent space. Experiments demonstrate that REPA accelerates SiT training by over 17.5×, enabling a SiT model to match the performance of a 7M-step SiT-XL within fewer than 400K steps. With classifier-free guidance (CFG), the method achieves an FID of 1.42—setting a new state-of-the-art at the time. REPA establishes a principled paradigm for improving representation learning in diffusion models through explicit cross-architecture semantic alignment.
Although Diffusion Transformers (DiTs) achieve impressive performance in image generation, their high computational and memory demands hinder deployment on edge devices. To address this challenge, this work proposes an efficient DiT framework featuring three key innovations: an adaptive global-local sparse attention mechanism to reduce computational complexity, an elastic training strategy within a unified hypernetwork enabling dynamic model scaling, and a four-step generative approach—KG-DMD—that integrates distribution matching with knowledge distillation. The resulting framework enables high-fidelity image generation in just four steps across diverse edge hardware platforms, significantly improving inference efficiency while maintaining visual quality and achieving a favorable balance between real-time performance and generation fidelity.
This work addresses the inefficiency in parameter utilization of Diffusion Transformers for generative tasks, which hampers their denoising performance. To overcome this limitation, the authors propose a lightweight calibration method that introduces approximately 100 learnable scaling parameters and formulates calibration as a black-box reward optimization problem, efficiently solved via an evolutionary algorithm. Requiring only minimal parameter fine-tuning, the approach significantly enhances generation quality and reduces inference steps across various text-to-image diffusion models while preserving high fidelity. This strategy achieves both parameter efficiency and computational efficiency in optimizing Diffusion Transformers, offering a practical and scalable solution for improving generative performance without extensive retraining or architectural modifications.
This work addresses the lack of theoretical understanding regarding why Transformers effectively learn optimal denoisers in diffusion models, particularly the mechanism by which they converge to the Bayes-optimal solution under non-convex loss. We establish, for the first time, a global convergence theory for Transformers trained on denoising diffusion probabilistic models (DDPMs) within a multi-label Gaussian mixture setting. By analyzing the population DDPM objective, modeling the multi-label Gaussian mixture distribution, and dissecting the self-attention architecture, we reveal how self-attention implements mean-field denoising and asymptotically approaches the minimum mean squared error (MMSE) estimator. Our analysis quantifies the required number of tokens per sample and training iterations to achieve a prescribed score-matching error. Numerical experiments corroborate the theoretical predictions and demonstrate alignment with MMSE estimation.
This work challenges the conventional assumption that diffusion models inherently require explicit timestep embeddings, investigating their necessity in the denoising process. Through theoretical analysis and empirical validation, the study demonstrates for the first time that under certain conditions, both U-Net and Diffusion Transformer architectures can converge to a global optimum without explicit timestep conditioning, implicitly inferring the noise scale. Ablation studies and generative evaluations on CelebA and CIFAR-10 show that such timestep-agnostic models achieve competitive or superior performance compared to standard timestep-conditioned counterparts in terms of FID, precision, and recall, while preserving high structural fidelity.
Existing methods struggle to accurately recover the initial noise latent variable from images generated by DDIM, achieving reasonable reconstruction quality but insufficient latent prediction accuracy. This work proposes a hybrid inversion approach that first employs gradient descent for direct inversion and subsequently refines the estimate through fixed-point iteration to more precisely recover the initial latent variable. The study introduces, innovatively, a “self-interpolation test” as a novel evaluation metric to comprehensively assess latent prediction fidelity. Experimental results demonstrate that the proposed method significantly improves both latent prediction accuracy and image reconstruction quality across three benchmark datasets, consistently outperforming existing approaches in self-interpolation test performance.