Score
Designs, implements, or evaluates models and algorithms that produce pixel- or region-level object masks from visual data (images or video), including mask predictors and segmentation mask generators. Work covers content-adaptive masking and region selection, masked image modeling, and techniques to preserve temporal and perceptual mask coherence (e.g., reduce flicker across frames).
To address three key challenges in text-to-video generation—poor motion consistency, high training costs, and strong data dependency—this paper proposes a lightweight motion control method guided by dynamic mask sequences. The approach integrates foreground mask guidance and dynamic mask sequence modeling within a diffusion-based framework. Its core contributions are: (1) a novel mask-driven motion sequence mechanism that enables precise text-foreground spatial alignment and controllable motion trajectory generation; (2) a hybrid strategy combining first-frame parameter sharing with autoregressive temporal expansion to ensure stability in long-video synthesis; and (3) efficient trainability using only a small amount of annotated data. Extensive experiments on video editing and artistic generation tasks demonstrate substantial improvements in motion consistency and visual fidelity. Both quantitative metrics (e.g., FVD, FID) and qualitative evaluations confirm superior performance over state-of-the-art methods.
This work addresses the challenges of region-based image editing—namely precise spatial localization, preservation of background consistency, and seamless boundary blending—by introducing MaskFlow, a novel framework that integrates mask information into the flow-matching generative process for the first time. MaskFlow employs mask-guided probabilistic paths to jointly regulate content generation within editable regions and preservation outside them. Furthermore, it incorporates a Soft-Poisson de-blending module to refine the vector field, enabling natural fusion between foreground and background. Coupled with a mask-driven MEData synthesis strategy, the proposed method consistently outperforms existing approaches on both natural images and infographics, with quantitative and qualitative results demonstrating superior performance in editing accuracy, background fidelity, and boundary seamlessness.
To address the challenge in local image editing—namely, reliance on precise masks or complex object localization, which hinders usability for non-expert users—this paper proposes a click-driven lightweight editing framework. Given only a single click and a text prompt, the method automatically generates a semantically coherent editing region at the clicked location and seamlessly inserts new content. Technically, it introduces a novel dynamic mask generation mechanism that requires no pre-trained segmentation models or fine-tuning. By integrating a blended latent diffusion (BLD) architecture with a mask-guided CLIP semantic loss, the approach enables progressive, click-triggered mask expansion. Experiments demonstrate state-of-the-art performance across multiple automated metrics; human evaluation further confirms high visual quality and robustness. The method significantly reduces user interaction overhead while maintaining editing fidelity and flexibility.
Existing text removal methods primarily target simple outdoor scenes and struggle with real-world images containing high-density, complex text layouts; moreover, their performance is highly sensitive to mask shape, necessitating costly manual parameter tuning. To address dense text images, this paper proposes an automated mask shape learning framework integrating deformable mask modeling and Bayesian optimization. First, we construct character-level deformable contour masks and empirically demonstrate that minimal covering masks are suboptimal, highlighting the critical role of fine-grained contour adjustment. Second, we formulate mask shape optimization as a black-box problem, using restoration quality as the feedback objective and leveraging Bayesian optimization to automatically determine optimal shape parameters. Experiments show significant improvements in restoration quality on high-density text images, empirically validating the existence of an optimal mask shape. Our approach delivers an interpretable, reusable, and fully automated solution for industrial-grade text removal.
This work addresses zero-shot in-context segmentation by proposing the first latent diffusion model (LDM)-based framework for the task. Methodologically: (1) it introduces an instruction-driven cross-modal alignment mechanism to map reference images to target segmentation masks semantically; (2) it adopts a two-stage mask strategy to prevent information leakage during inference; and (3) it formulates an enhanced pseudo-mask supervision objective that jointly optimizes generation fidelity and segmentation accuracy. Contributions include: (i) the first extension of LDMs to in-context segmentation; (ii) the first fair, unified benchmark covering both image and video segmentation scenarios; and (iii) state-of-the-art performance on this benchmark—outperforming specialized segmentation models and mainstream vision foundation models—thereby demonstrating the feasibility of unifying segmentation and generative modeling within a single diffusion-based architecture.
Diffusion models for image object removal commonly suffer from two critical issues: mask hallucination (generating semantically irrelevant content) and mask shape bias (overfitting to mask contours). To address these, this paper proposes a mask consistency regularization training strategy. Our method introduces a dual-branch mask perturbation mechanism—morphological dilation perturbation to enhance semantic awareness, and elastic deformation perturbation to break geometric dependency on the mask—and enforces output consistency across perturbed variants, thereby compelling the model to rely on contextual cues rather than mask geometry. Integrated into standard diffusion frameworks, this strategy requires no architectural modifications and improves inpainting fidelity and contextual coherence solely through training paradigm optimization. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, with significant improvements in both quantitative metrics (LPIPS, FID) and visual quality over existing methods.
Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.
This work addresses zero-shot, open-vocabulary semantic segmentation without manual annotations, fine-tuning, prompt engineering, or pre-trained segmentation networks. The proposed method leverages Stable Diffusion and introduces a novel joint modeling of cross-attention—enabling coarse-grained semantic localization—and multi-scale self-attention—facilitating fine-grained region propagation—thereby emulating classical seeded segmentation to achieve end-to-end mask generation from text-guided seeds to semantic expansion. Additionally, a background consistency optimization is incorporated to enhance boundary precision. The approach is plug-and-play and requires no task-specific adaptation. Evaluated on PASCAL VOC and COCO, it significantly outperforms existing generative segmentation methods in zero-shot settings. By unifying diffusion-based attention mechanisms with segmentation principles, this work establishes an efficient, interpretable, and annotation-free paradigm for open-vocabulary pixel-level segmentation.
This work addresses the heavy reliance on manual intervention—such as hand-crafted mask resizing and occlusion inpainting—in existing product catalog image synthesis methods. The authors propose a model-agnostic, end-to-end automated framework that generates high-quality composites from only a product image and a background, leveraging a novel dimension-aware masking algorithm and an occlusion-aware hybrid inpainting mechanism. Key contributions include the dimension-aware masking strategy, the occlusion-aware inpainting approach, the CatalogStitch-Eval benchmark comprising 58 complex real-world scenes, and an accompanying visualization toolkit. Experiments demonstrate consistent and significant improvements in both synthesis quality and efficiency across three representative models—ObjectStitch, OmniPaint, and InsertAnything—without requiring any post-processing.