Score
Designs and implements diffusion-based conditional generative models and sampling/denoising procedures that produce outputs conditioned on input images, including inpainting and masked-generation variants that reconstruct or complete specified masked regions or token sequences. Builds mask-aware conditioning mechanisms, loss functions and sampling constraints to enforce consistency with observed image content while supporting diversity and improved temporal and shape fidelity of generated trajectories or token streams.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
To address the challenges of unifying diverse image-conditioned generation tasks—namely, modeling complexity and parameter explosion—this paper proposes the first lightweight, single-stage unified diffusion framework. Our method jointly models the joint distribution of image pairs (e.g., RGB-depth) and supports generation, estimation, joint synthesis, signal-guided synthesis, and coarse-grained control—all within a single model, with only a 15% parameter increase, no architectural modifications, and no auxiliary networks. Key innovations include: (i) native support for non-spatially-aligned and coarse-grained conditioning inputs, breaking away from conventional multi-stage or multi-model paradigms; and (ii) a flexible sampling strategy enabling zero-overhead task switching. Experiments demonstrate that our single model matches or exceeds the performance of task-specific baselines, significantly outperforms existing unified approaches, and seamlessly integrates heterogeneous conditioning signals—advancing the practicality of controllable image generation.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
This work presents a systematic survey of denoising diffusion-based image editing, focusing on inpainting and outpainting, with particular emphasis on text-guided editing. To address the lack of standardized evaluation, we introduce EditEval—the first comprehensive benchmark for text-guided image editing—and propose LMM Score, a novel multimodal evaluation metric leveraging large multimodal models. We further provide the first unified taxonomy and empirical comparison between multimodal conditional editing methods and traditional context-driven approaches. Additionally, we release Awesome-Diffusion-Model-Based-Image-Editing-Methods, an open-source repository curating state-of-the-art techniques. Our study establishes a technical landscape spanning theoretical foundations, methodological frameworks, and evaluation standards. It identifies key limitations—including scalability, controllability, and evaluation consistency—and outlines concrete directions for future research. The work thus bridges critical gaps in both methodology and assessment, advancing the rigor and reproducibility of diffusion-based image editing.
This survey addresses key challenges in applying denoising diffusion models to computer vision—namely, fragmented applications, unclear theoretical connections, and low sampling efficiency—by establishing the first comprehensive, CV-oriented diffusion model taxonomy. Methodologically, it unifies the three dominant paradigms—Denoising Diffusion Probabilistic Models (DDPM), Noise Conditional Score Networks (NCSN), and Stochastic Differential Equations (SDE)—and introduces a multi-perspective classification framework. It rigorously clarifies their mathematical relationships with VAEs and GANs in terms of probabilistic modeling, score matching, and variational inference. The work innovatively elucidates the generative mechanisms underlying forward noising and reverse denoising processes, and precisely characterizes core challenges including scalability, sampling efficiency, and conditional control. As a foundational reference, this survey has been widely cited in subsequent research on efficient sampling, cross-modal diffusion, and theoretical generalizations, significantly shaping the development of diffusion-based vision methodologies.
Existing GAN-based semantic image synthesis methods suffer from inherent trade-offs between generation quality and diversity. To address this, we propose the first semantic image synthesis framework built upon denoising diffusion probabilistic models (DDPMs). Our method fundamentally decouples two key inputs: noisy images are fed into the U-Net encoder, while semantic layouts guide a dedicated decoder path via multi-level Spatially-Adaptive Denormalization (SPADE). Crucially, we introduce classifier-free guidance—the first such application in semantic diffusion synthesis—to substantially improve layout-to-pixel alignment. Evaluated on four standard benchmarks—Cityscapes, ADE20K, COCO-Stuff, and Mapillary Vistas—our approach achieves state-of-the-art performance: FID of 14.3 and LPIPS of 0.52, demonstrating significant gains in both visual fidelity and semantic consistency.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.
This work addresses the challenge of precise control in conditional discrete generative models when confronted with unseen condition combinations. The authors propose a theory-driven, composable discrete generation framework that integrates parallel token prediction with an absorbing diffusion mechanism and a concept-weighted conditional fusion strategy. This approach enables accurate modeling of an arbitrary number and combination of conditions while unifying mask-based generation within the same paradigm. Leveraging compositional vocabularies derived from VQ-VAE/VQ-GAN, the method achieves a 63.4% average reduction in error rate, a 9.58 improvement in FID, and 2.3–12× faster inference across three datasets. Furthermore, it successfully extends to pretrained text-to-image models, enabling fine-grained controllable generation.
This work addresses the limitations of existing conditional image generation methods, which often suffer from task-specific designs or lack training-free guidance, leading to information bottlenecks and error accumulation due to aggressive conditioning signal compression. To overcome these issues, the authors propose a unified, training-free inference framework that injects noisy conditioning signals during early denoising stages and adaptively guides the denoising process to extract task-relevant features. Additionally, they introduce a contrastive trajectory optimization mechanism between adjacent denoising states to refine the generation path. The method demonstrates significant improvements over current baselines across diverse tasks—including style transfer, super-resolution, and deblurring—achieving both high fidelity and superior perceptual quality while exhibiting strong cross-task generalization capabilities.
Standard masked diffusion models neglect the prediction of clean states at masked positions during the reverse denoising process, limiting their step-wise optimization capability. This work proposes a post-training self-conditioning adaptation method that requires no retraining, enabling each denoising step to condition on the model’s own prior predictions of clean states—without resorting to recurrent hidden states or auxiliary models. By overcoming the constraints of conventional partial self-conditioning strategies, the approach substantially enhances generation performance: it outperforms baseline methods across multiple tasks, reducing the generation perplexity of the OWT model by nearly 50% (from 42.89 to 23.72) and achieving higher quality and fidelity in image, molecular, and genomic sequence generation.
This work addresses the scarcity of high-quality labeled data and the high cost of annotation in deep learning by proposing a representation-conditioned latent diffusion model for controllable image synthesis, leveraging self-supervised visual representations such as DINOv2/v3 and CLIP. By integrating pretrained visual representations into the diffusion process, the method substantially enhances both the quality and class coverage of generated samples while enabling efficient data filtering and augmentation. Experiments on ImageNet100 demonstrate that classifiers trained solely on synthetic data generated by this approach achieve a 10.76 percentage point improvement in accuracy over those trained with conventional class-conditional generation methods, and even surpass models trained on real data by 2.0 percentage points.