Score
Design and implement algorithms and pipelines that apply, update, and manage masks over model outputs to iteratively denoise and selectively edit local regions. Build mechanisms for multi-turn masking and mask-guided denoising that enable incremental refinement of outputs and scaling of test‑time reasoning.
Existing masked diffusion models lack the capacity for multi-round iterative reasoning, making it difficult to emulate human-like local refinement processes. This work proposes a lightweight post-training approach that introduces a reflective masking mechanism, enabling multi-step iterative denoising at test time. It further incorporates a history-aware reference mechanism that leverages intermediate denoising states without requiring additional parameters or architectural modifications, thereby enhancing inference performance. As the first method to realize multi-round reflective reasoning within masked diffusion models, it significantly outperforms standard baselines across diverse cross-modal tasks—including text generation, Sudoku solving, and image editing—demonstrating strong generalization and potential as a universal reasoning primitive.
Real-world image matting is hindered by weak semantic understanding, difficulty in distinguishing multiple foreground instances, poor fine-detail recovery, and scarcity of high-quality annotated data. To address these challenges, we propose Mask2Alpha, an iterative refinement framework that introduces a mask-guided feature selection mechanism and a sparse convolution-based progressive optimization strategy, integrated with semantic priors extracted via self-supervised Vision Transformers (ViTs). Our method enables fine-grained instance-aware matting and high-resolution alpha matte reconstruction within a multi-stage iterative architecture—without requiring additional manual annotations—thereby enhancing model generalization. Evaluated on multiple real-world benchmarks, Mask2Alpha achieves state-of-the-art performance, notably improving matting accuracy (e.g., +0.8% F-measure on Composition-1k) and inference efficiency (1.7× faster than baseline models). The framework delivers an efficient, robust, end-to-end solution for content creation and AR applications.
This work addresses the tendency of existing large diffusion Transformers to propagate editing effects beyond intended regions during local image editing, owing to the absence of explicit spatial localization mechanisms. To resolve this, the authors propose REDEdit, a framework that introduces lightweight region-aware block adapters and a SpatialGate routing mechanism atop a frozen DiT backbone, enabling structured injection of positional conditioning signals and effective disentanglement of editing semantics from spatial location. Additionally, a jointly trained MaskPredictor head enables, for the first time, end-to-end mask-free local editing by accurately localizing target regions without requiring user-provided masks. A novel Region-Aware Loss further enhances spatial precision. Experiments demonstrate that REDEdit achieves state-of-the-art performance on both MagicBrush and Emu-Edit Test benchmarks, significantly outperforming existing methods that rely either on ground-truth masks or operate without any mask guidance.
This work addresses the problem of localized 3D mesh editing guided by a single edited image. We propose an interactive editing framework based on mask-conditioned reconstruction: user-specified 3D regions serve as geometric masks, and the edited image acts as a conditioning signal to guide a Large Reconstruction Model (LRM) to reconstruct only the masked regions while preserving high fidelity in the unmasked areas. To our knowledge, this is the first method to adapt LRM for real-time, mask-conditioned mesh editing. Our approach integrates multi-view-consistent mask rendering, stochastic 3D occlusion synthesis, and single-view conditional injection, enabling high-quality geometric updates in a single forward pass. The framework supports diverse semantic edits—including deformation, part replacement, and detail sculpting—achieving state-of-the-art reconstruction quality while accelerating inference by 10× over the best prior baseline.
Existing diffusion-based image editing methods suffer from limited localization control and interactivity. To address this, we propose a training-free, layer-based real-time editing framework. Its core innovation is the layered diffusion brush mechanism: fine-grained spatially masked intervention applied at intermediate denoising layers, decoupling region masks, visibility, and editing order to enable arbitrary, independent, and parallel layer editing. The method integrates prompt guidance with spatial constraints within a lightweight layer editor architecture, achieving efficient GPU-accelerated inference (140 ms per 512×512 image). User studies demonstrate significant improvements over InstructPix2Pix and Stable Diffusion Inpainting across attribute adjustment, error correction, and multi-step object placement tasks. To our knowledge, this is the first approach enabling high-fidelity, contextually consistent, interactive creative editing with real-time responsiveness and precise spatial control.
This work addresses the challenges of region-based image editing—namely precise spatial localization, preservation of background consistency, and seamless boundary blending—by introducing MaskFlow, a novel framework that integrates mask information into the flow-matching generative process for the first time. MaskFlow employs mask-guided probabilistic paths to jointly regulate content generation within editable regions and preservation outside them. Furthermore, it incorporates a Soft-Poisson de-blending module to refine the vector field, enabling natural fusion between foreground and background. Coupled with a mask-driven MEData synthesis strategy, the proposed method consistently outperforms existing approaches on both natural images and infographics, with quantitative and qualitative results demonstrating superior performance in editing accuracy, background fidelity, and boundary seamlessness.
Existing one-step image editing methods lack explicit spatial control, making it challenging to achieve strong semantic and structurally consistent modifications within user-specified regions. This work proposes a locally adaptive editing framework that leverages a mask-aware mechanism to automatically identify semantically relevant areas and applies adaptive modulation in the latent space to precisely edit only the target region while preserving the rest of the image unchanged. By integrating internal feature-driven editable region discovery, localized latent modulation, and spatial constraints, the method significantly outperforms current one-step approaches on PIE-Bench, striking an effective balance between editing fidelity and computational efficiency. These results underscore the critical role of explicit spatial reasoning in enabling high-quality image editing.
This work addresses the scarcity of low-cost, large-scale pixel-level annotations that limits current image manipulation localization (IML) methods. The authors propose a novel automatic annotation framework that obviates manual masks by efficiently extracting high-quality localization supervision from publicly available text-driven edited image pairs. Their approach leverages vision foundation models to compute semantic feature discrepancies, integrates instruction-guided spatial priors, and introduces several key innovations—including bidirectional cross-modal refinement, VAE round-trip noise calibration, EMA-based self-training, and an editing-noise disentanglement loss—to effectively bridge the domain gap between diffusion-based image editing and IML training. Evaluated on five benchmarks, the method substantially outperforms existing approaches (+12.20% F1, +11.16% IoU) and yields a 1.1-million-sample IML training set that boosts the average F1 score of six detectors by 18.34%.
This work addresses the computational redundancy in masked diffusion models, which denoise entire images during inference despite large masked regions. To overcome this inefficiency, the authors propose MASQ, a hardware-software co-designed accelerator architecture that integrates staged multi-precision quantization (MXINT8/4/2), timestep-aware scheduling, mask-aware computation, and customized non-matrix operation optimizations. MASQ introduces a novel strategy that jointly leverages spatial semantic importance and dynamic precision allocation, featuring a block-level multi-precision compute engine and a dedicated mask management unit. Experimental results demonstrate that MASQ achieves up to 16.06× speedup and 4.93× higher energy efficiency compared to NVIDIA A100 and Orin NX platforms, while preserving high-fidelity generation quality.