Score
Design and implement algorithms and pipelines that generate, refine, propagate, and post-process binary masks or support masks over data representations, using attention-guided updating, iterative polishing, fusion with mask priors (including prediction-derived and foundation-model priors), and dynamic masking strategies to improve mask precision and resolve attribute confusion among overlapping instances. Evaluate and analyze the refinement procedures for convergence, error modes, and trade-offs between precision, recall, and runtime.
Real-world image matting is hindered by weak semantic understanding, difficulty in distinguishing multiple foreground instances, poor fine-detail recovery, and scarcity of high-quality annotated data. To address these challenges, we propose Mask2Alpha, an iterative refinement framework that introduces a mask-guided feature selection mechanism and a sparse convolution-based progressive optimization strategy, integrated with semantic priors extracted via self-supervised Vision Transformers (ViTs). Our method enables fine-grained instance-aware matting and high-resolution alpha matte reconstruction within a multi-stage iterative architecture—without requiring additional manual annotations—thereby enhancing model generalization. Evaluated on multiple real-world benchmarks, Mask2Alpha achieves state-of-the-art performance, notably improving matting accuracy (e.g., +0.8% F-measure on Composition-1k) and inference efficiency (1.7× faster than baseline models). The framework delivers an efficient, robust, end-to-end solution for content creation and AR applications.
This work addresses the challenge that existing 3D Gaussian Splatting (3DGS) methods struggle to distinguish transient distractors from static backgrounds in scenes with color or semantic ambiguity, leading to contaminated reconstructions. To resolve this, the authors propose RefineSplat, a novel framework that introduces entropy as a key criterion for identifying ambiguous distractors. By integrating entropy-aware adaptive masking with instance segmentation to detect interference regions, and designing an entropy-guided positional gradient optimization mechanism to regulate Gaussian density distribution, RefineSplat significantly improves novel view synthesis quality across multiple datasets while effectively suppressing distractors. Additionally, the authors construct and publicly release Ambiguous Wild—the first dataset specifically targeting ambiguous interference, comprising 18 challenging real-world scenes—to establish a benchmark for future research in this domain.
This work addresses common limitations in image segmentation models—such as ambiguous boundaries, semantic inconsistency, and structural errors—by introducing the Phoenix framework. Phoenix generates semantically aware noise through adversarial mask perturbations to simulate realistic segmentation errors and employs a contrastive learning–based tripartite refinement mechanism that simultaneously enhances intra-class feature consistency and inter-class separability. Integrating adversarial learning, embedding attacks, and relational modeling, Phoenix operates as a plug-and-play module without requiring modifications to the backbone architecture. Extensive experiments demonstrate that Phoenix consistently outperforms existing approaches across diverse segmentation tasks, delivering substantial improvements in mask quality and reliably boosting the performance of state-of-the-art models.
Existing Transformer-based shadow removal methods often incorporate shadow priors through complex modifications to the attention mechanism, resulting in bloated architectures and high computational overhead. This paper proposes a lightweight shadow-aware Vision Transformer (ViT) framework. Its core innovation lies in explicitly embedding the shadow mask into the patch embedding layer—specifically at the *front end*, rather than within the attention modules—enabling efficient integration of shadow priors at the very earliest stage of feature extraction. This design avoids redundant structural alterations to self-attention, relying solely on standard multi-head self-attention and supervised mask guidance. Evaluated on ISTD, ISTD+, and SRD benchmarks, our method surpasses state-of-the-art approaches in both reconstruction accuracy—especially within shadow regions—and inference efficiency, while using significantly fewer parameters. The source code is publicly available.
To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.
This study addresses the difficulty of disentangling structural contributions from model capacity in mask refinement, as well as its poor cross-generator generalization. We formulate the refinement process as conditional random field (CRF) inference. Specifically, a zero-state mechanism is introduced to handle weakly labeled regions, decoupling boundary from non-boundary errors. Furthermore, DINOv2 features are leveraged to construct image-conditioned latent region consistency constraints, and damped mean-field iterations are unrolled to perform structured corrections. This approach establishes a verification paradigm demonstrating that explicit structure outperforms black-box capacity. Extensive experiments on datasets such as COD10K show that our method significantly surpasses parameter-matched baselines, achieves performance comparable to foundation models with minimal parameters, and effectively enhances generalization to unseen generators.
Existing masked diffusion models lack the capacity for multi-round iterative reasoning, making it difficult to emulate human-like local refinement processes. This work proposes a lightweight post-training approach that introduces a reflective masking mechanism, enabling multi-step iterative denoising at test time. It further incorporates a history-aware reference mechanism that leverages intermediate denoising states without requiring additional parameters or architectural modifications, thereby enhancing inference performance. As the first method to realize multi-round reflective reasoning within masked diffusion models, it significantly outperforms standard baselines across diverse cross-modal tasks—including text generation, Sudoku solving, and image editing—demonstrating strong generalization and potential as a universal reasoning primitive.
This work addresses the scarcity of low-cost, large-scale pixel-level annotations that limits current image manipulation localization (IML) methods. The authors propose a novel automatic annotation framework that obviates manual masks by efficiently extracting high-quality localization supervision from publicly available text-driven edited image pairs. Their approach leverages vision foundation models to compute semantic feature discrepancies, integrates instruction-guided spatial priors, and introduces several key innovations—including bidirectional cross-modal refinement, VAE round-trip noise calibration, EMA-based self-training, and an editing-noise disentanglement loss—to effectively bridge the domain gap between diffusion-based image editing and IML training. Evaluated on five benchmarks, the method substantially outperforms existing approaches (+12.20% F1, +11.16% IoU) and yields a 1.1-million-sample IML training set that boosts the average F1 score of six detectors by 18.34%.
This work addresses the challenge of removing specified objects in dense scenes, where existing methods often suffer from semantic interference caused by visually similar instances, leading to incomplete removal or duplicated artifacts. To overcome this limitation, the authors propose DORS, a diffusion-based object removal framework that introduces a novel dynamic attention routing mechanism comprising Instance Filtering Attention (IFA) and Context-Guided Routing (CGR). This mechanism enables fine-grained control over the attention space, effectively suppressing interference while preserving scene consistency. Furthermore, the study presents DOR-Bench, the first benchmark specifically designed for evaluating dense object removal. Experimental results demonstrate that DORS significantly outperforms current state-of-the-art methods, achieving superior removal completeness and visual coherence.
本文提出RefineEdit,一种无需训练的图像编辑框架,通过生成精炼网络全局优化二进制图像码,解决基于扩散和因果自回归编辑器在局部控制和解码顺序上的局限。