predict object masks

Designs, implements, or evaluates models and algorithms that produce pixel- or region-level object masks from visual data (images or video), including mask predictors and segmentation mask generators. Work covers content-adaptive masking and region selection, masked image modeling, and techniques to preserve temporal and perceptual mask coherence (e.g., reduce flicker across frames).

predictobjectmasks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address three key challenges in text-to-video generation—poor motion consistency, high training costs, and strong data dependency—this paper proposes a lightweight motion control method guided by dynamic mask sequences. The approach integrates foreground mask guidance and dynamic mask sequence modeling within a diffusion-based framework. Its core contributions are: (1) a novel mask-driven motion sequence mechanism that enables precise text-foreground spatial alignment and controllable motion trajectory generation; (2) a hybrid strategy combining first-frame parameter sharing with autoregressive temporal expansion to ensure stability in long-video synthesis; and (3) efficient trainability using only a small amount of annotated data. Extensive experiments on video editing and artistic generation tasks demonstrate substantial improvements in motion consistency and visual fidelity. Both quantitative metrics (e.g., FVD, FID) and qualitative evaluations confirm superior performance over state-of-the-art methods.

Difficulty maintaining text-motion consistencyHigh training costs in text-to-video modelsSubstantial data requirements for video generation

This work addresses the challenges of region-based image editing—namely precise spatial localization, preservation of background consistency, and seamless boundary blending—by introducing MaskFlow, a novel framework that integrates mask information into the flow-matching generative process for the first time. MaskFlow employs mask-guided probabilistic paths to jointly regulate content generation within editable regions and preservation outside them. Furthermore, it incorporates a Soft-Poisson de-blending module to refine the vector field, enabling natural fusion between foreground and background. Coupled with a mask-driven MEData synthesis strategy, the proposed method consistently outperforms existing approaches on both natural images and infographics, with quantitative and qualitative results demonstrating superior performance in editing accuracy, background fidelity, and boundary seamlessness.

background preservationprecise localizationregional image editing

Click2Mask: Local Editing with Dynamic Mask Generation

Sep 12, 2024
OR
Omer Regev
🏛️ The Hebrew University of Jerusalem

To address the challenge in local image editing—namely, reliance on precise masks or complex object localization, which hinders usability for non-expert users—this paper proposes a click-driven lightweight editing framework. Given only a single click and a text prompt, the method automatically generates a semantically coherent editing region at the clicked location and seamlessly inserts new content. Technically, it introduces a novel dynamic mask generation mechanism that requires no pre-trained segmentation models or fine-tuning. By integrating a blended latent diffusion (BLD) architecture with a mask-guided CLIP semantic loss, the approach enables progressive, click-triggered mask expansion. Experiments demonstrate state-of-the-art performance across multiple automated metrics; human evaluation further confirms high visual quality and robustness. The method significantly reduces user interaction overhead while maintaining editing fidelity and flexibility.

Image EditingLocal ModificationUser-Friendliness

Existing text removal methods primarily target simple outdoor scenes and struggle with real-world images containing high-density, complex text layouts; moreover, their performance is highly sensitive to mask shape, necessitating costly manual parameter tuning. To address dense text images, this paper proposes an automated mask shape learning framework integrating deformable mask modeling and Bayesian optimization. First, we construct character-level deformable contour masks and empirically demonstrate that minimal covering masks are suboptimal, highlighting the critical role of fine-grained contour adjustment. Second, we formulate mask shape optimization as a black-box problem, using restoration quality as the feedback objective and leveraging Bayesian optimization to automatically determine optimal shape parameters. Experiments show significant improvements in restoration quality on high-density text images, empirically validating the existence of an optimal mask shape. Our approach delivers an interpretable, reusable, and fully automated solution for industrial-grade text removal.

Addresses complex images with dense text using Bayesian optimizationCreates benchmark for practical text removal tasksDevelops method for optimal mask shapes in text removal

Explore In-Context Segmentation via Latent Diffusion Models

Mar 14, 2024
CW
Chaoyang Wang
🏛️ Peking University | NTU | UC, Merced | ZJU | Skywork AI

This work addresses zero-shot in-context segmentation by proposing the first latent diffusion model (LDM)-based framework for the task. Methodologically: (1) it introduces an instruction-driven cross-modal alignment mechanism to map reference images to target segmentation masks semantically; (2) it adopts a two-stage mask strategy to prevent information leakage during inference; and (3) it formulates an enhanced pseudo-mask supervision objective that jointly optimizes generation fidelity and segmentation accuracy. Contributions include: (i) the first extension of LDMs to in-context segmentation; (ii) the first fair, unified benchmark covering both image and video segmentation scenarios; and (iii) state-of-the-art performance on this benchmark—outperforming specialized segmentation models and mainstream vision foundation models—thereby demonstrating the feasibility of unifying segmentation and generative modeling within a single diffusion-based architecture.

Creates a new benchmark for image and video segmentation.Develops strategies for instruction extraction and output alignment.Explores in-context segmentation using latent diffusion models.

Latest Papers

What's happening recently
View more

Mask Consistency Regularization in Object Removal

Sep 12, 2025
HY
Hua Yuan
🏛️ Southeast University | Lenovo Research | Southeast University | Ministry of Education

Diffusion models for image object removal commonly suffer from two critical issues: mask hallucination (generating semantically irrelevant content) and mask shape bias (overfitting to mask contours). To address these, this paper proposes a mask consistency regularization training strategy. Our method introduces a dual-branch mask perturbation mechanism—morphological dilation perturbation to enhance semantic awareness, and elastic deformation perturbation to break geometric dependency on the mask—and enforces output consistency across perturbed variants, thereby compelling the model to rely on contextual cues rather than mask geometry. Integrated into standard diffusion frameworks, this strategy requires no architectural modifications and improves inpainting fidelity and contextual coherence solely through training paradigm optimization. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, with significant improvements in both quantitative metrics (LPIPS, FID) and visual quality over existing methods.

Addressing mask-shape bias in image inpaintingImproving contextual coherence for removed regionsReducing mask hallucination in object removal

MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning

Nov 16, 2025
JH
Jingshan Hong
🏛️ Zhejiang University of Technology | Zhejiang Normal University

Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.

Addresses underutilization of discarded pixels in supervised image maskingExploits masked regions as semantic diversity sources rather than ignored dataSolves loss of fine-grained features caused by traditional masking methods

SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models

Jul 26, 2025
JH
Joon Hyun Park
🏛️ Hanyang University

This work addresses zero-shot, open-vocabulary semantic segmentation without manual annotations, fine-tuning, prompt engineering, or pre-trained segmentation networks. The proposed method leverages Stable Diffusion and introduces a novel joint modeling of cross-attention—enabling coarse-grained semantic localization—and multi-scale self-attention—facilitating fine-grained region propagation—thereby emulating classical seeded segmentation to achieve end-to-end mask generation from text-guided seeds to semantic expansion. Additionally, a background consistency optimization is incorporated to enhance boundary precision. The approach is plug-and-play and requires no task-specific adaptation. Evaluated on PASCAL VOC and COCO, it significantly outperforms existing generative segmentation methods in zero-shot settings. By unifying diffusion-based attention mechanisms with segmentation principles, this work establishes an efficient, interpretable, and annotation-free paradigm for open-vocabulary pixel-level segmentation.

Generating pixel-level annotation masks without human effortRefining masks using background correspondence for higher accuracyUtilizing Stable Diffusion attention mechanisms for object localization

This work addresses the heavy reliance on manual intervention—such as hand-crafted mask resizing and occlusion inpainting—in existing product catalog image synthesis methods. The authors propose a model-agnostic, end-to-end automated framework that generates high-quality composites from only a product image and a background, leveraging a novel dimension-aware masking algorithm and an occlusion-aware hybrid inpainting mechanism. Key contributions include the dimension-aware masking strategy, the occlusion-aware inpainting approach, the CatalogStitch-Eval benchmark comprising 58 complex real-world scenes, and an accompanying visualization toolkit. Experiments demonstrate consistent and significant improvements in both synthesis quality and efficiency across three representative models—ObjectStitch, OmniPaint, and InsertAnything—without requiring any post-processing.

catalog image generationdimension mismatchmanual intervention

Hot Scholars

MC

Marcella Cornia

Associate Professor, University of Modena and Reggio Emilia
Vision and LanguageGenerative AIComputer VisionDigital Humanities
FS

Fei Shen

National University of Singapore
Controllable GenerationMultimodal Safety
RC

Rita Cucchiara

Università degli Studi di Modena e Reggio Emilia, Italia
Computer VisionPattern RecognitionDeep LearningMultimedia
KR

Kui Ren

Professor and Dean of Computer Science, Zhejiang University, ACM/IEEE Fellow
Data Security & PrivacyAI SecurityIoT & Vehicular Security
CM

Chuofan Ma

PhD student of Electrical and Electronic Engineering, The University of Hong Kong
Computer VisionMachine Learning