adaptive mask refinement

Design and implement algorithms and pipelines that generate, refine, propagate, and post-process binary masks or support masks over data representations, using attention-guided updating, iterative polishing, fusion with mask priors (including prediction-derived and foundation-model priors), and dynamic masking strategies to improve mask precision and resolve attribute confusion among overlapping instances. Evaluate and analyze the refinement procedures for convergence, error modes, and trade-offs between precision, recall, and runtime.

adaptivemaskrefinement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Real-world image matting is hindered by weak semantic understanding, difficulty in distinguishing multiple foreground instances, poor fine-detail recovery, and scarcity of high-quality annotated data. To address these challenges, we propose Mask2Alpha, an iterative refinement framework that introduces a mask-guided feature selection mechanism and a sparse convolution-based progressive optimization strategy, integrated with semantic priors extracted via self-supervised Vision Transformers (ViTs). Our method enables fine-grained instance-aware matting and high-resolution alpha matte reconstruction within a multi-stage iterative architecture—without requiring additional manual annotations—thereby enhancing model generalization. Evaluated on multiple real-world benchmarks, Mask2Alpha achieves state-of-the-art performance, notably improving matting accuracy (e.g., +0.8% F-measure on Composition-1k) and inference efficiency (1.7× faster than baseline models). The framework delivers an efficient, robust, end-to-end solution for content creation and AR applications.

Enhances real-world image matting accuracyImproves semantic and instance awarenessRecovers high-resolution image details

This work addresses the challenge that existing 3D Gaussian Splatting (3DGS) methods struggle to distinguish transient distractors from static backgrounds in scenes with color or semantic ambiguity, leading to contaminated reconstructions. To resolve this, the authors propose RefineSplat, a novel framework that introduces entropy as a key criterion for identifying ambiguous distractors. By integrating entropy-aware adaptive masking with instance segmentation to detect interference regions, and designing an entropy-guided positional gradient optimization mechanism to regulate Gaussian density distribution, RefineSplat significantly improves novel view synthesis quality across multiple datasets while effectively suppressing distractors. Additionally, the authors construct and publicly release Ambiguous Wild—the first dataset specifically targeting ambiguous interference, comprising 18 challenging real-world scenes—to establish a benchmark for future research in this domain.

3D Gaussian Splattingambiguous scenariosdistractor removal

This work addresses common limitations in image segmentation models—such as ambiguous boundaries, semantic inconsistency, and structural errors—by introducing the Phoenix framework. Phoenix generates semantically aware noise through adversarial mask perturbations to simulate realistic segmentation errors and employs a contrastive learning–based tripartite refinement mechanism that simultaneously enhances intra-class feature consistency and inter-class separability. Integrating adversarial learning, embedding attacks, and relational modeling, Phoenix operates as a plug-and-play module without requiring modifications to the backbone architecture. Extensive experiments demonstrate that Phoenix consistently outperforms existing approaches across diverse segmentation tasks, delivering substantial improvements in mask quality and reliably boosting the performance of state-of-the-art models.

adversarial perturbationboundary imperfectionmask refinement

ShadowMaskFormer: Mask Augmented Patch Embeddings for Shadow Removal

Apr 29, 2024
ZL
Zhuohao Li
🏛️ Sun Yat-Sen University | CATL | City University of Hong Kong

Existing Transformer-based shadow removal methods often incorporate shadow priors through complex modifications to the attention mechanism, resulting in bloated architectures and high computational overhead. This paper proposes a lightweight shadow-aware Vision Transformer (ViT) framework. Its core innovation lies in explicitly embedding the shadow mask into the patch embedding layer—specifically at the *front end*, rather than within the attention modules—enabling efficient integration of shadow priors at the very earliest stage of feature extraction. This design avoids redundant structural alterations to self-attention, relying solely on standard multi-head self-attention and supervised mask guidance. Evaluated on ISTD, ISTD+, and SRD benchmarks, our method surpasses state-of-the-art approaches in both reconstruction accuracy—especially within shadow regions—and inference efficiency, while using significantly fewer parameters. The source code is publicly available.

Improves shadow removal using transformer-based modelsReduces computational resources while maintaining effectivenessSimplifies architecture by enhancing early patch embeddings

Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training

Jun 12, 2023
LB
L. Baraldi
🏛️ University of Modena and Reggio Emilia | NVIDIA AI Technology Center | IIT-CNR

To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.

Image UnderstandingPre-trainingVisual Transformers

Latest Papers

What's happening recently
View more

This study addresses the difficulty of disentangling structural contributions from model capacity in mask refinement, as well as its poor cross-generator generalization. We formulate the refinement process as conditional random field (CRF) inference. Specifically, a zero-state mechanism is introduced to handle weakly labeled regions, decoupling boundary from non-boundary errors. Furthermore, DINOv2 features are leveraged to construct image-conditioned latent region consistency constraints, and damped mean-field iterations are unrolled to perform structured corrections. This approach establishes a verification paradigm demonstrating that explicit structure outperforms black-box capacity. Extensive experiments on datasets such as COD10K show that our method significantly surpasses parameter-matched baselines, achieves performance comparable to foundation models with minimal parameters, and effectively enhances generalization to unseen generators.

conditional random fieldgeneralizationimage segmentation

Existing masked diffusion models lack the capacity for multi-round iterative reasoning, making it difficult to emulate human-like local refinement processes. This work proposes a lightweight post-training approach that introduces a reflective masking mechanism, enabling multi-step iterative denoising at test time. It further incorporates a history-aware reference mechanism that leverages intermediate denoising states without requiring additional parameters or architectural modifications, thereby enhancing inference performance. As the first method to realize multi-round reflective reasoning within masked diffusion models, it significantly outperforms standard baselines across diverse cross-modal tasks—including text generation, Sudoku solving, and image editing—demonstrating strong generalization and potential as a universal reasoning primitive.

iterative revisionlocal refinementMask Diffusion Models

This work addresses the scarcity of low-cost, large-scale pixel-level annotations that limits current image manipulation localization (IML) methods. The authors propose a novel automatic annotation framework that obviates manual masks by efficiently extracting high-quality localization supervision from publicly available text-driven edited image pairs. Their approach leverages vision foundation models to compute semantic feature discrepancies, integrates instruction-guided spatial priors, and introduces several key innovations—including bidirectional cross-modal refinement, VAE round-trip noise calibration, EMA-based self-training, and an editing-noise disentanglement loss—to effectively bridge the domain gap between diffusion-based image editing and IML training. Evaluated on five benchmarks, the method substantially outperforms existing approaches (+12.20% F1, +11.16% IoU) and yields a 1.1-million-sample IML training set that boosts the average F1 score of six detectors by 18.34%.

image manipulation localizationmask generationpixel-level annotation

This work addresses the challenge of removing specified objects in dense scenes, where existing methods often suffer from semantic interference caused by visually similar instances, leading to incomplete removal or duplicated artifacts. To overcome this limitation, the authors propose DORS, a diffusion-based object removal framework that introduces a novel dynamic attention routing mechanism comprising Instance Filtering Attention (IFA) and Context-Guided Routing (CGR). This mechanism enables fine-grained control over the attention space, effectively suppressing interference while preserving scene consistency. Furthermore, the study presents DOR-Bench, the first benchmark specifically designed for evaluating dense object removal. Experimental results demonstrate that DORS significantly outperforms current state-of-the-art methods, achieving superior removal completeness and visual coherence.

attention mechanismdense scenesincomplete removal

Hot Scholars

JT

Jie Tang

UW Madison
Computed Tomography
HD

Henghui Ding

Fudan University
Computer VisionMachine LearningSegmentationAIGC
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
PH

Paul Hongsuck Seo

Korea University
Multimodal Interactive IntelligenceVisionSpeech and Language Understanding