Score
Designs, builds, or analyzes generative and denoising diffusion models that operate directly in image/pixel/feature and point‑map representations (including image‑space, pixel‑space, feature‑space, point‑map, and spatio‑temporal/4D variants) rather than in compressed latents. This work covers training, sampling, and evaluation methods that produce or refine raw pixels and point maps in a single stage—bypassing lossy latent compression—to recover sharper geometric and temporal structure and remain robust in ambiguous regions.
This survey addresses key challenges in applying denoising diffusion models to computer vision—namely, fragmented applications, unclear theoretical connections, and low sampling efficiency—by establishing the first comprehensive, CV-oriented diffusion model taxonomy. Methodologically, it unifies the three dominant paradigms—Denoising Diffusion Probabilistic Models (DDPM), Noise Conditional Score Networks (NCSN), and Stochastic Differential Equations (SDE)—and introduces a multi-perspective classification framework. It rigorously clarifies their mathematical relationships with VAEs and GANs in terms of probabilistic modeling, score matching, and variational inference. The work innovatively elucidates the generative mechanisms underlying forward noising and reverse denoising processes, and precisely characterizes core challenges including scalability, sampling efficiency, and conditional control. As a foundational reference, this survey has been widely cited in subsequent research on efficient sampling, cross-modal diffusion, and theoretical generalizations, significantly shaping the development of diffusion-based vision methodologies.
Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.
This study addresses the low training efficiency of pixel-space diffusion models by proposing a latent-to-pixel space transfer strategy. Through systematic optimization of weight initialization, data composition, noise scheduling, and decoder architecture, we establish an efficient pixel-space training paradigm. Experiments demonstrate that this approach effectively resolves slow pre-training convergence, achieving performance that matches or surpasses latent-space models while improving end-to-end inference speed by 3.18× to 4.75×. Consequently, this work provides a comprehensive guideline for high-performance and practical pixel-level training in high-resolution image generation, bridging the gap between theoretical efficacy and computational efficiency in generative modeling.
Current image generation models often produce spatially inconsistent outputs and geometric distortions due to the lack of explicit scene-structure modeling. To address this, we propose a collaborative denoising framework that jointly generates images and their intrinsic attributes—namely depth maps and semantic segmentation masks—thereby implicitly learning geometric and layout constraints through a shared latent space. Our method builds upon a pre-trained latent diffusion model and employs a lightweight autoencoder to fuse multi-source intrinsic attributes as structural priors. A cross-domain information-sharing mechanism enables synchronized denoising across the image and attribute domains, requiring neither 3D supervision nor additional annotations. Experiments demonstrate that our approach preserves text-image alignment and visual fidelity while significantly improving spatial plausibility. It achieves state-of-the-art performance across quantitative metrics—including layout consistency and depth fidelity—outperforming all baseline methods.
Diffusion models achieve high-quality generation but suffer from severe deviation of inverted latent representations—e.g., those obtained via DDIM inversion—from the Gaussian prior, leading to ill-posed latent-to-image mapping and poor diversity in interpolation and editing. This work is the first to systematically identify the root cause: amplified noise prediction errors in smooth image regions, which distort latent-space structure. Through noise prediction error visualization, quantitative evaluation of latent-space editability, and statistical analysis of structural patterns, we empirically confirm systematic structural biases in inverted latents. Our analysis provides an interpretable diagnostic framework for latent-space distortion and, theoretically, establishes essential constraints for constructing semantically consistent and highly controllable diffusion latent spaces. These insights lay a principled foundation for designing next-generation controllable generation and editing methods.
Diffusion models unexpectedly generate cartoonized or blurry images—nonexistent in training data—within high-density regions of the learned distribution. Method: We propose Mode Tracking Theory to precisely localize modes in the diffusion denoising distribution; design a zero-overhead SDE likelihood tracking method that estimates and optimizes sample likelihood without additional computation; and develop an efficient high-density sampler that targets atypical, high-likelihood samples overlooked by conventional samplers. Results: Experiments demonstrate substantial improvement in sampling likelihood, stable generation of cartoon/blurry high-density images, and faithful reproduction of this phenomenon on purely real-image datasets—revealing an intrinsic, implicit structural bias inherent to diffusion models.
To address the lack of systematic surveys and unified modeling frameworks for diffusion models in low-level vision, this paper presents the first comprehensive survey covering over 20 tasks—including image restoration, enhancement, and generation. We propose three general-purpose diffusion modeling paradigms, theoretically unify them with GANs and VAEs, and rigorously delineate their boundaries. A dual-perspective classification scheme—structured by both architecture and task—is introduced and extended to cross-domain applications (e.g., medical imaging, remote sensing, video). We conduct benchmarking with joint efficiency–performance evaluation and open-source a resource repository featuring 20+ models and standardized evaluation metrics. Key contributions include: (1) the first structured taxonomy for diffusion-based low-level vision; (2) a cross-task transferability analysis framework; and (3) identification of seven critical future research directions—collectively advancing both theoretical foundations and practical deployment of diffusion models in low-level vision.
Existing methods struggle to accurately recover the initial noise latent variable from images generated by DDIM, achieving reasonable reconstruction quality but insufficient latent prediction accuracy. This work proposes a hybrid inversion approach that first employs gradient descent for direct inversion and subsequently refines the estimate through fixed-point iteration to more precisely recover the initial latent variable. The study introduces, innovatively, a “self-interpolation test” as a novel evaluation metric to comprehensively assess latent prediction fidelity. Experimental results demonstrate that the proposed method significantly improves both latent prediction accuracy and image reconstruction quality across three benchmark datasets, consistently outperforming existing approaches in self-interpolation test performance.
This work addresses the limitations of latent diffusion models, which suffer from detail loss and misalignment between representation and generation objectives due to fixed visual encoders. To overcome these issues, the authors propose an end-to-end pixel-space diffusion Transformer framework that operates directly in the pixel domain without relying on VAE compression. The approach integrates a continuous generation mechanism within a unified multimodal Transformer architecture, sharing a common token space for both images and text. By carefully optimizing noise scheduling, loss weighting, and model scaling strategies, the method significantly enhances fine-grained detail fidelity in high-resolution image synthesis. This paradigm offers a promising direction toward building integrated multimodal vision foundation models capable of both generative and perceptual tasks.
This work addresses the structural ambiguity and training complexity in single-image 3D geometry reconstruction caused by reliance on latent-space compression or hybrid architectures. We propose an extremely minimalist pixel-space diffusion Transformer that trains an end-to-end diffusion model directly on raw 3D point patches, eliminating implicit encoding, complex loss functions, and the need for a point-patch tokenizer. Leveraging pretrained DINOv2 image features as conditioning guidance for geometry generation and built upon a standard ViT backbone, our method achieves the first purely pixel-space diffusion-based geometric reconstruction. It significantly enhances geometric sharpness and robustness—particularly in transparent and highly ambiguous regions—and outperforms existing implicit diffusion approaches.
This work addresses the high inference latency and computational cost of diffusion models arising from their iterative denoising process. The authors propose a timestep-aware dynamic inference acceleration method that learns dedicated masks for each denoising step to dynamically skip redundant network blocks and reuse intermediate features, thereby reducing computation. To mitigate the high memory overhead of global backpropagation, mask optimization is performed independently per timestep. Stability is further enhanced through timestep-aware loss scaling and a knowledge-guided mask refinement strategy. The approach achieves significant inference speedups across diverse architectures—including DDPM, LDM, DiT, and PixArt—while preserving generation quality.
This work investigates how to construct a diffusion-friendly latent space to enhance generation quality, moving beyond the sole optimization of reconstruction fidelity. The authors systematically evaluate diverse visual tokenizer architectures, regularization strategies, and latent configurations across multiple diffusion backbones. They introduce a novel metric, Velocity Irreducible Variance (VIV), to quantify velocity ambiguity in the latent space arising from trajectory intersections. Experimental results demonstrate that VIV serves as a robust predictor of generation quality, consistently outperforming other latent-space attributes across various settings. The study further uncovers several key characteristics of latent representations that exhibit strong generalization capabilities, offering actionable insights for designing better latent spaces tailored to diffusion models.