mask-conditioned generation

Designs, builds, and evaluates generative models and editing pipelines that take binary or soft masks (attention masks, operation masks, region masks) as explicit conditioning inputs to control where and how content is synthesized or altered, while preserving content outside masked regions. Implements per-region, operation-aware behaviors — e.g., different edit operations or coordination roles (simultaneous, reactive, leader) — trains and augments models across diverse mask configurations for robustness, and integrates mask-aware attention or sub-program execution to ensure localized, coordinated edits.

mask-conditionedgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of region-based image editing—namely precise spatial localization, preservation of background consistency, and seamless boundary blending—by introducing MaskFlow, a novel framework that integrates mask information into the flow-matching generative process for the first time. MaskFlow employs mask-guided probabilistic paths to jointly regulate content generation within editable regions and preservation outside them. Furthermore, it incorporates a Soft-Poisson de-blending module to refine the vector field, enabling natural fusion between foreground and background. Coupled with a mask-driven MEData synthesis strategy, the proposed method consistently outperforms existing approaches on both natural images and infographics, with quantitative and qualitative results demonstrating superior performance in editing accuracy, background fidelity, and boundary seamlessness.

background preservationprecise localizationregional image editing

This work addresses the challenges of precise execution in e-commerce image editing, where multiple operations, localized modifications, and auditability requirements often lead to partial failures in existing methods due to the tight coupling of intent understanding, region localization, and image generation. To overcome this, we propose a decoupled cognitive-generation agent framework that leverages a vision-language model to construct region-anchored editing agendas, guides a diffusion-based editor with operation-aware masks for stepwise execution, and incorporates a reflection-driven iterative mechanism to ensure editing completeness and error correctability. Our approach achieves the first structured decoupling of cognitive reasoning and generative rendering, significantly outperforming open-source alternatives on our newly curated benchmark, EComEditBench, while matching the instruction accuracy and editing fidelity of powerful closed-source models and enabling traceable, recoverable multi-turn editing workflows.

e-commerce image editingedit fidelityinstruction ambiguity

EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing

Dec 12, 2025
WC
Wei Chow
🏛️ ByteDance | National University of Singapore | HKUST(GZ) | Shanghai Jiao Tong University

Diffusion models often distort non-target regions in image editing due to global denoising. This paper proposes the first localized image editing framework based on Masked Generative Transformers (MGT), which avoids global perturbations via token-level mask prediction, enabling precise region editing and strong fidelity preservation. Key contributions include: (1) the first adaptation of MGT to image editing; (2) a cross-layer attention map fusion mechanism to enhance spatial localization accuracy; and (3) region-preserving sampling and attention-injection fine-tuning—parameter-free strategies that adapt pre-trained MGTs without adding parameters. Evaluated on four benchmarks, our method matches diffusion models in editing quality while accelerating inference by 6×; it improves style transfer and domain adaptation by 3.6% and 17.6%, respectively, with <1B parameters. We also introduce CrispEdit-2M, a high-resolution image editing dataset.

Addresses unintended modifications in non-target regions during image editingEnhances editing quality and speed while preserving integrity of surrounding areasIntroduces a localized decoding approach using Masked Generative Transformers for precise edits

Click2Mask: Local Editing with Dynamic Mask Generation

Sep 12, 2024
OR
Omer Regev
🏛️ The Hebrew University of Jerusalem

To address the challenge in local image editing—namely, reliance on precise masks or complex object localization, which hinders usability for non-expert users—this paper proposes a click-driven lightweight editing framework. Given only a single click and a text prompt, the method automatically generates a semantically coherent editing region at the clicked location and seamlessly inserts new content. Technically, it introduces a novel dynamic mask generation mechanism that requires no pre-trained segmentation models or fine-tuning. By integrating a blended latent diffusion (BLD) architecture with a mask-guided CLIP semantic loss, the approach enables progressive, click-triggered mask expansion. Experiments demonstrate state-of-the-art performance across multiple automated metrics; human evaluation further confirms high visual quality and robustness. The method significantly reduces user interaction overhead while maintaining editing fidelity and flexibility.

Image EditingLocal ModificationUser-Friendliness

Existing diffusion-based image editing methods lack a unified theoretical framework that jointly addresses core criteria such as controllability, fidelity to user intent, semantic consistency, locality, and perceptual quality. This work formulates diffusion editing as a guided transport process on the image manifold and introduces, for the first time, a task-agnostic evaluation framework. By leveraging conditional inverse-time generative operators, noise dynamics analysis, mask-guided local modeling, and Lipschitz error propagation theory, the study systematically characterizes the trade-offs among multiple performance dimensions across different editing paradigms. Theoretical analysis reveals bounds on the influence of guidance strength and inversion error on non-target regions, as well as the error accumulation mechanism in multi-step editing. Empirical benchmarking further demonstrates significant disparities among state-of-the-art methods in instruction following, region preservation, and semantic stability.

controllabilitydiffusion-based editinglocality

Latest Papers

What's happening recently
View more

This work proposes an efficient and controllable video generation method that drives subject motion using a single image and a binary mask sequence while preserving realistic environmental interactions. The key innovation lies in identifying critical “motion layers” within a unified MMDiT architecture and applying LoRA fine-tuning exclusively to these layers. Additionally, a lightweight MaskAdapter is introduced to encode the mask sequence into latent residual signals, which are injected into the model via cosine-weighted scheduling. Without modifying the backbone network, the approach achieves precise control over dynamic content and establishes state-of-the-art performance in motion fidelity and perceptual realism across multiple datasets, significantly outperforming existing methods with lower computational overhead.

controllable video generationinteractive video synthesismask-guided generation

Existing one-step image editing methods lack explicit spatial control, making it challenging to achieve strong semantic and structurally consistent modifications within user-specified regions. This work proposes a locally adaptive editing framework that leverages a mask-aware mechanism to automatically identify semantically relevant areas and applies adaptive modulation in the latent space to precisely edit only the target region while preserving the rest of the image unchanged. By integrating internal feature-driven editable region discovery, localized latent modulation, and spatial constraints, the method significantly outperforms current one-step approaches on PIE-Bench, striking an effective balance between editing fidelity and computational efficiency. These results underscore the critical role of explicit spatial reasoning in enabling high-quality image editing.

localized semantic transformationmask-aware editingone-step image editing

Existing multimodal diffusion Transformers lack a unified and effective safety mechanism for image generation and editing, often failing to prevent the synthesis of harmful content. This work proposes a training-free Unified Visual Regulator (UVR), which, for the first time, identifies a task-agnostic harmful semantic priming phase within multimodal attention. Leveraging this insight, UVR dynamically analyzes information flow and proactively modulates relevant attention heads at an early stage to precisely suppress the propagation of unsafe signals. Requiring no additional training, the method seamlessly integrates into both image synthesis and editing pipelines. It achieves 91% and 77% harmful content mitigation rates respectively, substantially outperforming existing approaches while preserving high visual fidelity.

image-to-image editingmultimodal diffusion transformerssafe image generation

This work addresses the lack of a unified theoretical foundation in existing attention mask designs. It establishes, for the first time, a formal connection between attention masks and partially ordered structures, proving that information flow in sufficiently deep multi-layer Transformers converges to a Hasse diagram. The mask design problem is thereby reformulated as finding the minimal common supergraph of such Hasse diagrams, yielding a general framework that derives attention masks directly from task families. Leveraging this framework, the authors propose two novel mechanisms—Block Two-Stream Attention and Butterfly Attention—and derive block-wise causal masks and fully supervised bidirectional masks that guarantee consistency between training and inference. Empirical results validate both the effectiveness and generality of the proposed approach.

attention masksHasse diagramsinformation flow

This work addresses the challenge in diffusion-based image editing where global denoising mechanisms inadvertently alter non-target regions due to tight coupling between edited and unedited areas. To resolve this, the paper introduces EditMGT, the first framework to adapt Masked Generative Transformers (MGT) for image editing. EditMGT leverages MGT’s localized token prediction to strictly confine edits within user-specified masks and incorporates a multi-layer attention fusion module to generate precise spatial guidance signals. Combined with a region-preserving sampling strategy, the approach effectively suppresses unintended modifications outside the target area. Despite using only 0.96 billion parameters, EditMGT achieves state-of-the-art image fidelity across multiple benchmarks and accelerates editing by a factor of six compared to existing methods.

context preservationdiffusion modelsedit localization

Hot Scholars

TN

Trung-Nghia Le

University of Science, VNU-HCM
Applied Deep LearningApplied Computer VisionMultimedia Security
MT

Minh-Triet Tran

University of Science & John von Neumann Institute, VNU-HCM
Cryptography and SecurityMultimedia and InteractionComputer Vision and Machine LearningSoftware Engineering
TV

Tam V. Nguyen

Associate Professor and Graduate Program Director, Dept. of Computer Science, University of Dayton
Artificial IntelligenceMultimedia AnalysisMixed RealityComputer Vision and Deep Learning
QC

Qifeng Chen

HKUST
Computational PhotographyImage SynthesisGenerative AIAutonomous Driving