Mask Consistency Regularization in Object Removal

📅 2025-09-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Diffusion models for image object removal commonly suffer from two critical issues: mask hallucination (generating semantically irrelevant content) and mask shape bias (overfitting to mask contours). To address these, this paper proposes a mask consistency regularization training strategy. Our method introduces a dual-branch mask perturbation mechanism—morphological dilation perturbation to enhance semantic awareness, and elastic deformation perturbation to break geometric dependency on the mask—and enforces output consistency across perturbed variants, thereby compelling the model to rely on contextual cues rather than mask geometry. Integrated into standard diffusion frameworks, this strategy requires no architectural modifications and improves inpainting fidelity and contextual coherence solely through training paradigm optimization. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, with significant improvements in both quantitative metrics (LPIPS, FID) and visual quality over existing methods.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Feature Construction/ReformulationSearch and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Object removal, a challenging task within image inpainting, involves seamlessly filling the removed region with content that matches the surrounding context. Despite advancements in diffusion models, current methods still face two critical challenges. The first is mask hallucination, where the model generates irrelevant or spurious content inside the masked region, and the second is mask-shape bias, where the model fills the masked area with an object that mimics the mask's shape rather than surrounding content. To address these issues, we propose Mask Consistency Regularization (MCR), a novel training strategy designed specifically for object removal tasks. During training, our approach introduces two mask perturbations: dilation and reshape, enforcing consistency between the outputs of these perturbed branches and the original mask. The dilated masks help align the model's output with the surrounding content, while reshaped masks encourage the model to break the mask-shape bias. This combination of strategies enables MCR to produce more robust and contextually coherent inpainting results. Our experiments demonstrate that MCR significantly reduces hallucinations and mask-shape bias, leading to improved performance in object removal.
Problem

Research questions and friction points this paper is trying to address.

Reducing mask hallucination in object removal
Addressing mask-shape bias in image inpainting
Improving contextual coherence for removed regions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mask Consistency Regularization training strategy
Dilation and reshape mask perturbations
Enforcing output consistency across perturbed branches
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hua Yuan
School of Computer Science and Engineering, Southeast University, Nanjing, China
J
Jin Yuan
AI Lab, Lenovo Research, Beijing, China
Y
Yicheng Jiang
School of Computer Science and Engineering, Southeast University, Nanjing, China
Y
Yao Zhang
AI Lab, Lenovo Research, Beijing, China
Xin Geng
Xin Geng
School of Computer Science and Engineering, Southeast University
Artificial IntelligencePattern RecognitionMachine Learning
Y
Yong Rui
School of Computer Science and Engineering, Southeast University, Nanjing, China