Score
Designs and implements denoising diffusion probabilistic models that iteratively remove noise from noisy boundary or span representations to recover precise token-span boundaries and sequence labels. Builds the denoising networks, noise schedules, and sampling/decoding procedures (including non‑autoregressive parallel decoding) that progressively refine multi-token spans and boundary predictions.
This survey addresses key challenges in applying denoising diffusion models to computer vision—namely, fragmented applications, unclear theoretical connections, and low sampling efficiency—by establishing the first comprehensive, CV-oriented diffusion model taxonomy. Methodologically, it unifies the three dominant paradigms—Denoising Diffusion Probabilistic Models (DDPM), Noise Conditional Score Networks (NCSN), and Stochastic Differential Equations (SDE)—and introduces a multi-perspective classification framework. It rigorously clarifies their mathematical relationships with VAEs and GANs in terms of probabilistic modeling, score matching, and variational inference. The work innovatively elucidates the generative mechanisms underlying forward noising and reverse denoising processes, and precisely characterizes core challenges including scalability, sampling efficiency, and conditional control. As a foundational reference, this survey has been widely cited in subsequent research on efficient sampling, cross-modal diffusion, and theoretical generalizations, significantly shaping the development of diffusion-based vision methodologies.
This paper investigates the intrinsic mechanisms underlying the generalization capability of diffusion models, particularly their robust cross-architecture performance. Method: We propose a training-free, interpretable analytical framework that establishes, for the first time, a theoretical connection between diffusion model generalization and local empirical denoisers. Our approach models local inductive biases, designs a multi-scale local denoiser aggregation algorithm, and evaluates behavioral consistency across forward and reverse processes. Contribution/Results: We demonstrate that generalization primarily arises from local denoising operations providing a high-fidelity approximation to the training objective. Experiments show our method achieves superior visual fidelity and lower mean squared error in reproducing neural denoising behavior compared to existing approaches. Furthermore, we validate the universality of the “local denoising dominates generalization” principle across diverse network architectures, thereby overcoming the black-box attribution bottleneck in diffusion model analysis.
Existing discrete diffusion language models predominantly adopt full-decoder architectures, where each denoising step requires executing the entire network, resulting in high computational overhead and inefficient inference. Method: We propose the first encoder-decoder-based discrete diffusion model: a dedicated encoder learns clean-text representations, while a lightweight decoder performs iterative denoising; combined with block-wise sequence partitioning and specialized training/sampling algorithms, this design decouples representation learning from noise removal. Contribution/Results: Our architecture significantly improves training stability and inference throughput. Empirical evaluation on summarization, machine translation, and mathematical reasoning demonstrates superior quality–latency trade-offs at reduced computational cost. This work establishes a new paradigm for efficient discrete diffusion modeling.
To address the degradation in generation quality caused by misalignment between noise levels and distances to the data manifold during diffusion model denoising, this paper proposes a noise-level calibration mechanism. It explicitly models noise level as a proxy for the distance from a sample to the data manifold and introduces a lightweight, plug-and-play auxiliary correction network that dynamically refines denoising estimates at each step. The method requires no retraining of the backbone diffusion model and uniformly supports diverse restoration tasks—including image inpainting, super-resolution, deblurring, color enhancement, and compression artifact removal—while remaining compatible with mainstream samplers (e.g., DDIM). Correction subnetworks are constructed solely from pre-trained denoising networks, incorporating manifold geometric priors and task-specific constraints (e.g., masks, degradation kernels, frequency-domain restrictions). Experiments demonstrate consistent improvements in FID and LPIPS across both unconditional generation and restoration tasks, with minimal computational overhead and stable performance gains.
Discrete diffusion models currently lack a unified theoretical framework, hindering systematic comparison across existing approaches. This work proposes a cohesive perspective grounded in discrete state-space formulations, integrating diverse methodologies—such as transition matrices, masking/absorbing states, and score/ratio-based formulations—into a common design space. By elucidating the intrinsic connections and trade-offs among training objectives, inference algorithms, and evaluation protocols, the study clarifies the design spectrum of discrete diffusion models. Furthermore, through the unification of multiple diffusion paradigms alongside systematic optimization and scalability analyses, this research establishes a foundation for principled method comparison, guides future investigations, and fosters synergistic development across the field.
To address the low denoising efficiency and reconstruction quality degradation caused by sequential randomness in discrete diffusion models, this paper proposes a “planning-based denoising” framework that decouples the process into a learnable position planner and a local denoiser, enabling on-demand identification of critical positions and precise token/image restoration. This two-stage paradigm overcomes inherent limitations of conventional mask diffusion—namely, uniform or random denoising—by introducing iterative adaptive mask selection and joint token/image training. Evaluated on text8, OpenWebText, and ImageNet 256×256, our method significantly outperforms state-of-the-art mask diffusion models: it achieves language modeling perplexity closely approaching autoregressive baselines while simultaneously improving both inference efficiency and fidelity in image generation.
This work addresses a fundamental mismatch in Unified Diffusion Models (UDMs), where the standard plug-in bridge parameterization fails to align with the true denoising posterior, leading to inconsistencies between the training objective and generative dynamics. To resolve this, the authors propose a leave-one-out posterior–based denoiser parameterization and introduce an absorbing-state Markov chain reconstruction framework that reformulates UDMs as a mask-like diffusion sampling process. This formulation exposes a theoretical inconsistency between the plug-in evidence lower bound (ELBO) and cross-entropy denoising objectives, yielding an exact transformation relationship. Building upon this insight, they devise a prediction-correction sampling scheme and a temperature optimization strategy that require no additional training. Experiments demonstrate that the proposed leave-one-out parameterization substantially improves language generation quality, with the absorbing-state construction matching or surpassing state-of-the-art mask-based diffusion models in performance.
Discrete diffusion models have demonstrated strong performance in language modeling, yet the learning order of data support structure and frequency information during denoising remains unclear. This work theoretically and empirically reveals that such models first learn the support set—such as grammatical validity—and subsequently refine internal frequency distributions. We establish, for the first time, that under low-noise conditions, a single-step reverse edit decomposes into a leading-order term governing support membership and a finer coefficient capturing frequency characteristics. Furthermore, we elucidate a hierarchical mechanistic distinction between uniform and absorbing diffusion processes. Through asymptotic analysis, masked language diffusion models, and experiments on regular language tasks, we confirm that support identification precedes frequency ranking, and that both diffusion types exhibit the predicted rate separation phenomenon.
Diffusion language models struggle to achieve truly order-agnostic generation in fast parallel decoding due to their sensitivity to denoising order. This work formalizes the resulting “order collapse” problem for the first time as an issue of compatibility among local conditional distributions. Adopting a non-conservative field perspective, it introduces order-induced pseudo-joint distributions and local denoising circulations to uncover the root cause of path dependence. Building upon probabilistic graphical models and circulation decomposition, the paper proposes a theoretical framework that enables diagnosis of order-freedom solely at inference time. This framework effectively disentangles path dependence from two distinct error sources: conditional dependency errors arising from parallel updates and order-specific estimation errors, thereby providing the first quantifiable analytical tool for evaluating the order-agnostic properties of diffusion language models.
Existing methods for discrete data generation operate directly in discrete spaces, often leading to abrupt state transitions that compromise generation quality and stability. This work proposes a non-Markovian denoising framework based on probability simplices, modeling the discrete data generation process within a continuous simplex space. By introducing a conditionally independent noise mechanism, the approach circumvents the redundant constraints inherent in conventional discrete diffusion models, thereby simplifying the model architecture while preserving theoretical rigor. Extensive experiments on multiple synthetic and real-world graph datasets demonstrate that the proposed method significantly outperforms current discrete diffusion and flow-matching approaches, confirming its effectiveness and superiority in discrete generative tasks.