Score
Designs and implements diffusion-based generative models and transformer architectures that modulate the denoising process via binary or continuous masks to control which parts of an input are reconstructed, generated, or left unchanged. Builds mask‑conditioned transformers and inference pipelines that use mask signals to switch between conditional, unconditional, modality‑specific, and joint rollouts while reusing a shared generative backbone.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work presents a systematic survey of denoising diffusion-based image editing, focusing on inpainting and outpainting, with particular emphasis on text-guided editing. To address the lack of standardized evaluation, we introduce EditEval—the first comprehensive benchmark for text-guided image editing—and propose LMM Score, a novel multimodal evaluation metric leveraging large multimodal models. We further provide the first unified taxonomy and empirical comparison between multimodal conditional editing methods and traditional context-driven approaches. Additionally, we release Awesome-Diffusion-Model-Based-Image-Editing-Methods, an open-source repository curating state-of-the-art techniques. Our study establishes a technical landscape spanning theoretical foundations, methodological frameworks, and evaluation standards. It identifies key limitations—including scalability, controllability, and evaluation consistency—and outlines concrete directions for future research. The work thus bridges critical gaps in both methodology and assessment, advancing the rigor and reproducibility of diffusion-based image editing.
The necessity of noise conditioning in denoising generative models remains unchallenged despite its ubiquitous adoption. Method: We systematically evaluate the impact of removing noise conditioning across diverse denoising architectures via theoretical error analysis, ablation studies, and FID-optimized unconditional sampling. Contribution/Results: Contrary to prevailing assumptions, most denoising models exhibit robust performance without noise conditioning—and in several cases, achieve lower FID scores than their conditioned counterparts. We introduce the first high-performance noise-unconditional diffusion model, attaining a FID of 2.23 on CIFAR-10—narrowing the gap with state-of-the-art conditional models significantly. Our findings demonstrate that the denoising generative paradigm need not rely on explicit noise conditioning, opening new avenues for architectural simplification, computational efficiency gains, and foundational theoretical reexamination.
To address the challenges of few-shot learning, sparse labeling, heterogeneous (numerical and categorical) conditioning variables, and constrained computational resources in engineering applications, this paper proposes a masked conditional generative paradigm. We design a unified learnable embedding to jointly model heterogeneous conditions and introduce a masked conditional scheduling mechanism that explicitly simulates missing conditions during training to enhance robustness to incomplete inputs. Furthermore, we construct a lightweight collaborative architecture integrating a variational autoencoder and a latent diffusion model, coupled with knowledge distillation from pre-trained large models. Experiments on 2D point cloud and engineering image datasets demonstrate that the method enables efficient training with only a small number of labeled samples; achieves a 32% reduction in Fréchet Inception Distance (FID); significantly improves conditional fidelity; and simultaneously ensures strong controllability and high generation quality.
Masked Diffusion Models (MDMs) suffer from redundant computation in discrete sequence generation due to binary masking, which causes tokens to remain unchanged across many sampling steps. To address this, we propose Partial Masking (Prime), the first framework to introduce continuous-interpolated intermediate mask states into discrete diffusion, enabling token-level fine-grained denoising and overcoming the rigid all-or-nothing masking paradigm. Methodologically, we formulate a variational training objective and design a dedicated architecture that eliminates reliance on autoregressive structures. Experiments demonstrate state-of-the-art performance: perplexity of 15.36 on OpenWebText for text generation, and FID scores of 3.26 (CIFAR-10) and 6.98 (ImageNet-32) for image generation—surpassing existing MDMs and hybrid models. Our core contribution is a differentiable intermediate masking mechanism that unifies discrete token representation with continuous denoising dynamics.
Prior work lacks a systematic analysis of the inference mechanisms underlying Masked Generative Transformers (MGTs). Method: This paper introduces the first systematic “design choice set” for MGT inference and proposes an enhanced inference framework for high-resolution image generation, integrating mask reweighting, hierarchical sampling, and diffusion-model-inspired acceleration. Built upon MaskGIT and Meissonic architectures, it unifies masked image modeling, discrete token prediction, diffusion-prior guidance, and adaptive resampling. Contribution/Results: Evaluated on the HPS v2 benchmark, the Meissonic-1024×1024 model achieves ~70% win-rate improvement. All components are modular, plug-and-play, and yield cumulative gains. The work establishes a reproducible, scalable design paradigm and empirical benchmark for efficient MGT inference.
Existing diffusion models suffer from low sampling efficiency, high memory overhead, and limited generation diversity in zero-shot image-to-image (I2I) translation. This paper proposes a training-free, fully black-box filtering guidance method: lightweight, adaptive filtering operations are applied at the input of each diffusion step, enabling model- and sampler-agnostic intervention. Key contributions include: (i) the first architecture- and sampler-agnostic universal filtering guidance; (ii) continuous, tunable guidance strength; and (iii) a novel, general interpretability perspective for self-attention mechanisms. Our method operates via gradient-free, iterative input reweighting—requiring no architectural modification or parameter optimization. Evaluated across multiple I2I tasks, it matches or surpasses task-specific state-of-the-art methods in structural fidelity while incurring negligible inference overhead.
This work investigates the mechanisms by which diffusion models generate highly realistic images that differ from their training data—referred to as “creativity”—and demonstrates that this capability stems from the alignment between the denoiser architecture and the target data distribution. Through theoretical analysis and empirical experiments, the study derives explicit forms of the generated distribution for linear, polynomial, and bottleneck-style denoisers for the first time, and systematically evaluates the behavior of various architectures, including UNet variants, throughout the diffusion process. The findings reveal that minor architectural modifications to the UNet significantly impact generation fidelity, thereby underscoring the critical role of the denoiser’s inductive bias and its alignment with the target distribution in determining model performance.
Standard masked diffusion models neglect the prediction of clean states at masked positions during the reverse denoising process, limiting their step-wise optimization capability. This work proposes a post-training self-conditioning adaptation method that requires no retraining, enabling each denoising step to condition on the model’s own prior predictions of clean states—without resorting to recurrent hidden states or auxiliary models. By overcoming the constraints of conventional partial self-conditioning strategies, the approach substantially enhances generation performance: it outperforms baseline methods across multiple tasks, reducing the generation perplexity of the OWT model by nearly 50% (from 42.89 to 23.72) and achieving higher quality and fidelity in image, molecular, and genomic sequence generation.
This work addresses a fundamental mismatch in Unified Diffusion Models (UDMs), where the standard plug-in bridge parameterization fails to align with the true denoising posterior, leading to inconsistencies between the training objective and generative dynamics. To resolve this, the authors propose a leave-one-out posterior–based denoiser parameterization and introduce an absorbing-state Markov chain reconstruction framework that reformulates UDMs as a mask-like diffusion sampling process. This formulation exposes a theoretical inconsistency between the plug-in evidence lower bound (ELBO) and cross-entropy denoising objectives, yielding an exact transformation relationship. Building upon this insight, they devise a prediction-correction sampling scheme and a temperature optimization strategy that require no additional training. Experiments demonstrate that the proposed leave-one-out parameterization substantially improves language generation quality, with the absorbing-state construction matching or surpassing state-of-the-art mask-based diffusion models in performance.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.