mask-modulated diffusion modeling

Designs and implements diffusion-based generative models and transformer architectures that modulate the denoising process via binary or continuous masks to control which parts of an input are reconstructed, generated, or left unchanged. Builds mask‑conditioned transformers and inference pipelines that use mask signals to switch between conditional, unconditional, modality‑specific, and joint rollouts while reusing a shared generative backbone.

mask-modulateddiffusionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Diffusion Model-Based Image Editing: A Survey

Feb 27, 2024
YH
Yi Huang
🏛️ Shenzhen Institute of Advanced Technology | Chinese Academy of Sciences | University of Chinese Academy of Sciences | Southern University of Science and Technology | Adobe Inc | Apple Inc

This work presents a systematic survey of denoising diffusion-based image editing, focusing on inpainting and outpainting, with particular emphasis on text-guided editing. To address the lack of standardized evaluation, we introduce EditEval—the first comprehensive benchmark for text-guided image editing—and propose LMM Score, a novel multimodal evaluation metric leveraging large multimodal models. We further provide the first unified taxonomy and empirical comparison between multimodal conditional editing methods and traditional context-driven approaches. Additionally, we release Awesome-Diffusion-Model-Based-Image-Editing-Methods, an open-source repository curating state-of-the-art techniques. Our study establishes a technical landscape spanning theoretical foundations, methodological frameworks, and evaluation standards. It identifies key limitations—including scalability, controllability, and evaluation consistency—and outlines concrete directions for future research. The work thus bridges critical gaps in both methodology and assessment, advancing the rigor and reproducibility of diffusion-based image editing.

Analyzing learning strategies and user-input conditions.Proposing EditEval benchmark for text-guided editing evaluation.Surveying diffusion models for image editing tasks.

Must-Read Papers

Most classic and influential ideas
View more

The necessity of noise conditioning in denoising generative models remains unchallenged despite its ubiquitous adoption. Method: We systematically evaluate the impact of removing noise conditioning across diverse denoising architectures via theoretical error analysis, ablation studies, and FID-optimized unconditional sampling. Contribution/Results: Contrary to prevailing assumptions, most denoising models exhibit robust performance without noise conditioning—and in several cases, achieve lower FID scores than their conditioned counterparts. We introduce the first high-performance noise-unconditional diffusion model, attaining a FID of 2.23 on CIFAR-10—narrowing the gap with state-of-the-art conditional models significantly. Our findings demonstrate that the denoising generative paradigm need not rely on explicit noise conditioning, opening new avenues for architectural simplification, computational efficiency gains, and foundational theoretical reexamination.

Challenges necessity of noise conditioningExplores denoising models without noise conditioningIntroduces competitive noise-unconditional model

Masked Conditioning for Deep Generative Models

May 22, 2025
PM
Phillip Mueller
🏛️ BMW Group | University of Augsburg | Ludwig-Maximilians-University Munich

To address the challenges of few-shot learning, sparse labeling, heterogeneous (numerical and categorical) conditioning variables, and constrained computational resources in engineering applications, this paper proposes a masked conditional generative paradigm. We design a unified learnable embedding to jointly model heterogeneous conditions and introduce a masked conditional scheduling mechanism that explicitly simulates missing conditions during training to enhance robustness to incomplete inputs. Furthermore, we construct a lightweight collaborative architecture integrating a variational autoencoder and a latent diffusion model, coupled with knowledge distillation from pre-trained large models. Experiments on 2D point cloud and engineering image datasets demonstrate that the method enables efficient training with only a small number of labeled samples; achieves a 32% reduction in Fréchet Inception Distance (FID); significantly improves conditional fidelity; and simultaneously ensures strong controllability and high generation quality.

Enabling generative models with limited computational resourcesHandling small, sparse, mixed-type datasets in engineeringImproving generation quality with small models and pretrained foundations

Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking

May 24, 2025
CC
Chen-Hao Chao
🏛️ University of Toronto | Vector Institute | NVIDIA AI Technology Center | National Taiwan University

Masked Diffusion Models (MDMs) suffer from redundant computation in discrete sequence generation due to binary masking, which causes tokens to remain unchanged across many sampling steps. To address this, we propose Partial Masking (Prime), the first framework to introduce continuous-interpolated intermediate mask states into discrete diffusion, enabling token-level fine-grained denoising and overcoming the rigid all-or-nothing masking paradigm. Methodologically, we formulate a variational training objective and design a dedicated architecture that eliminates reliance on autoregressive structures. Experiments demonstrate state-of-the-art performance: perplexity of 15.36 on OpenWebText for text generation, and FID scores of 3.26 (CIFAR-10) and 6.98 (ImageNet-32) for image generation—surpassing existing MDMs and hybrid models. Our core contribution is a differentiable intermediate masking mechanism that unifies discrete token representation with continuous denoising dynamics.

Improves performance on text and image generation tasksIntroduces partial masking for intermediate token statesReduces redundant computation in masked diffusion models

Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer

Nov 16, 2024
SS
Shitong Shao
🏛️ Hong Kong University of Science and Technology | Mohamed bin Zayed University of Artificial Intelligence

Prior work lacks a systematic analysis of the inference mechanisms underlying Masked Generative Transformers (MGTs). Method: This paper introduces the first systematic “design choice set” for MGT inference and proposes an enhanced inference framework for high-resolution image generation, integrating mask reweighting, hierarchical sampling, and diffusion-model-inspired acceleration. Built upon MaskGIT and Meissonic architectures, it unifies masked image modeling, discrete token prediction, diffusion-prior guidance, and adaptive resampling. Contribution/Results: Evaluated on the HPS v2 benchmark, the Meissonic-1024×1024 model achieves ~70% win-rate improvement. All components are modular, plug-and-play, and yield cumulative gains. The work establishes a reproducible, scalable design paradigm and empirical benchmark for efficient MGT inference.

Accelerating MGT sampling processBridging DM and ARM discrepanciesEnhancing MGT inference techniques

Filtered-Guided Diffusion: Fast Filter Guidance for Black-Box Diffusion Models

Jun 29, 2023
ZG
Zeqi Gu
🏛️ Cornell Tech | Cornell University

Existing diffusion models suffer from low sampling efficiency, high memory overhead, and limited generation diversity in zero-shot image-to-image (I2I) translation. This paper proposes a training-free, fully black-box filtering guidance method: lightweight, adaptive filtering operations are applied at the input of each diffusion step, enabling model- and sampler-agnostic intervention. Key contributions include: (i) the first architecture- and sampler-agnostic universal filtering guidance; (ii) continuous, tunable guidance strength; and (iii) a novel, general interpretability perspective for self-attention mechanisms. Our method operates via gradient-free, iterative input reweighting—requiring no architectural modification or parameter optimization. Evaluated across multiple I2I tasks, it matches or surpasses task-specific state-of-the-art methods in structural fidelity while incurring negligible inference overhead.

Addressing limited output diversity from deterministic sampling approachesEnhancing control over guidance strength and frequency in image translationOvercoming high computational costs in diffusion-based image generation methods

Latest Papers

What's happening recently
View more

This work investigates the mechanisms by which diffusion models generate highly realistic images that differ from their training data—referred to as “creativity”—and demonstrates that this capability stems from the alignment between the denoiser architecture and the target data distribution. Through theoretical analysis and empirical experiments, the study derives explicit forms of the generated distribution for linear, polynomial, and bottleneck-style denoisers for the first time, and systematically evaluates the behavior of various architectures, including UNet variants, throughout the diffusion process. The findings reveal that minor architectural modifications to the UNet significantly impact generation fidelity, thereby underscoring the critical role of the denoiser’s inductive bias and its alignment with the target distribution in determining model performance.

CreativityDenoiser ArchitectureDiffusion Models

Standard masked diffusion models neglect the prediction of clean states at masked positions during the reverse denoising process, limiting their step-wise optimization capability. This work proposes a post-training self-conditioning adaptation method that requires no retraining, enabling each denoising step to condition on the model’s own prior predictions of clean states—without resorting to recurrent hidden states or auxiliary models. By overcoming the constraints of conventional partial self-conditioning strategies, the approach substantially enhances generation performance: it outperforms baseline methods across multiple tasks, reducing the generation perplexity of the OWT model by nearly 50% (from 42.89 to 23.72) and achieving higher quality and fidelity in image, molecular, and genomic sequence generation.

discrete sequence generationgenerative refinementmasked diffusion models

This work addresses a fundamental mismatch in Unified Diffusion Models (UDMs), where the standard plug-in bridge parameterization fails to align with the true denoising posterior, leading to inconsistencies between the training objective and generative dynamics. To resolve this, the authors propose a leave-one-out posterior–based denoiser parameterization and introduce an absorbing-state Markov chain reconstruction framework that reformulates UDMs as a mask-like diffusion sampling process. This formulation exposes a theoretical inconsistency between the plug-in evidence lower bound (ELBO) and cross-entropy denoising objectives, yielding an exact transformation relationship. Building upon this insight, they devise a prediction-correction sampling scheme and a temperature optimization strategy that require no additional training. Experiments demonstrate that the proposed leave-one-out parameterization substantially improves language generation quality, with the absorbing-state construction matching or surpassing state-of-the-art mask-based diffusion models in performance.

absorbing statedenoising posteriorleave-one-out

This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.

controllable image generationdiffusion modelsgeneration control

Hot Scholars

ZH

Zemin Huang

PhD student, Westlake University, Zhejiang University
Diffusion ModelAutoregressive ModelDiffusion Distillation
MM

Mathilde Mougeot

Full Professor at ENSIIE & Researcher at Borelli Center, ENS Paris-Saclay
Data scienceMachine learning
YG

Yu-Gang Jiang

Professor, Fudan University. IEEE & IAPR Fellow
Video AnalysisEmbodied AITrustworthy AI