segmentation mask fusion

Design, build, or analyze algorithms and pipelines that align, merge, and refine multiple segmentation mask proposals into a single coherent, high-quality mask output, performing post-hoc pixel-boundary refinement and optional conversion to polygonal masks. Implement and evaluate specific fusion strategies such as query-guided mask postprocessing, mask-transformer-based fusion, confidence-weighted and size-ordered aggregation of multi-scale proposals, anomaly/gap detection and selection of confident regions, and lightweight parameter-free methods suitable for real-time use.

segmentationmaskfusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Judging from Support-set: A New Way to Utilize Few-Shot Segmentation for Segmentation Refinement Process

Jul 05, 2024
SM
Seonghyeon Moon
🏛️ Brookhaven National Laboratory | Rutgers University | Mohamed Bin Zayed University of Artificial Intelligence

Automatically assessing the success of segmentation refinement is critical for ensuring segmentation reliability and advancing image segmentation techniques. This paper proposes JFS, the first framework to repurpose off-the-shelf few-shot segmentation (FSS) models—such as SegGPT—for segmentation quality assessment. JFS constructs a novel support set from coarse and refined masks and leverages the FSS model to evaluate refinement effectiveness. Crucially, it introduces a mask-driven support set reconstruction mechanism, enabling plug-and-play evaluation without additional training. Experiments on the PASCAL dataset demonstrate that JFS accurately determines the success or failure of mainstream refinement methods (e.g., SEPL), thereby filling a key research gap in automated, reliability-aware refinement assessment. To our knowledge, JFS is the first quantifiable, general-purpose evaluation tool for high-confidence segmentation.

Ensure reliability in high-stakes segmentation applicationsEvaluate success of segmentation refinement methodsLeverage few-shot segmentation for quality assessment

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement

Feb 10, 2025
YL
Yuqi Lin
🏛️ Zhejiang University | Shanghai AI Laboratory | FABU Inc.

To address the low quality of coarse-grain segmentation masks and high annotation costs, this paper proposes SAMRefiner++, a general-purpose and efficient mask refinement framework that is plug-and-play compatible with any pre-trained segmentation model (e.g., SAM) without requiring additional annotations. Methodologically: (i) it introduces a novel noise-robust prompting mechanism integrating distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks; (ii) it designs a “Split–Tackle–Merge” (STM) pipeline to handle multi-object segmentation challenges; and (iii) it incorporates an unsupervised IoU-adaptive optimization module combining iterative calibration with adaptive prompt engineering. Extensive experiments on multiple benchmarks demonstrate significant improvements over state-of-the-art methods in both accuracy and inference efficiency, achieving new SOTA performance. The source code is publicly available.

Enhance quality of coarse masksReduce segmentation annotation costUniversal mask refinement approach

MARS: a Multimodal Alignment and Ranking System for Few-Shot Segmentation

Apr 10, 2025
NC
Nico Catalano
🏛️ Politecnico di Milano

Existing few-shot segmentation methods rely solely on visual similarity for mask selection, leading to suboptimal ranking. To address this, we propose a plug-and-play multimodal alignment and ranking framework that adapts to any mask proposal generator without retraining. Our key innovation is the first introduction of a dual-granularity (local and global) multimodal scoring mechanism, which jointly fuses complementary cues from text, vision, geometry, and semantics for robust mask scoring, selection, and weighted fusion. Extensive experiments demonstrate new state-of-the-art performance on four benchmarks—COCO-20i, Pascal-5i, LVIS-92i, and FSS-1000—yielding significant improvements over prevailing methods. The code is publicly available, facilitating reproducibility and advancing the practical deployment of few-shot segmentation.

Lacks mask selection method beyond visual similarityNeeds robust multimodal ranking for few-shot segmentationRequires plug-and-play system to improve mask proposals

Unitho: A Unified Multi-Task Framework for Computational Lithography

Nov 13, 2025
QJ
Qian Jin
🏛️ Zhejiang University

In computational lithography, tasks such as mask synthesis, design rule violation detection, and layout optimization have traditionally been modeled in isolation, hindered by data scarcity and methodological fragmentation. Method: This paper introduces the first unified multi-task vision large model framework tailored for lithography, built upon a Transformer architecture and trained end-to-end on large-scale industrial-grade simulation data to jointly learn and share knowledge across these three core tasks. Contribution/Results: The framework breaks from conventional single-task paradigms, significantly improving generalization capability and modeling fidelity—outperforming existing academic baselines across multiple benchmarks. It establishes a high-reliability intelligent EDA data foundation and a scalable model infrastructure, providing a novel paradigm for computational lithography.

Enables end-to-end mask generation and design rule verificationSolves data scarcity and limited modeling in lithography simulationUnified framework addresses isolated computational lithography tasks integration

This work addresses common limitations in image segmentation models—such as ambiguous boundaries, semantic inconsistency, and structural errors—by introducing the Phoenix framework. Phoenix generates semantically aware noise through adversarial mask perturbations to simulate realistic segmentation errors and employs a contrastive learning–based tripartite refinement mechanism that simultaneously enhances intra-class feature consistency and inter-class separability. Integrating adversarial learning, embedding attacks, and relational modeling, Phoenix operates as a plug-and-play module without requiring modifications to the backbone architecture. Extensive experiments demonstrate that Phoenix consistently outperforms existing approaches across diverse segmentation tasks, delivering substantial improvements in mask quality and reliably boosting the performance of state-of-the-art models.

adversarial perturbationboundary imperfectionmask refinement

Latest Papers

What's happening recently
View more

This study addresses the difficulty of disentangling structural contributions from model capacity in mask refinement, as well as its poor cross-generator generalization. We formulate the refinement process as conditional random field (CRF) inference. Specifically, a zero-state mechanism is introduced to handle weakly labeled regions, decoupling boundary from non-boundary errors. Furthermore, DINOv2 features are leveraged to construct image-conditioned latent region consistency constraints, and damped mean-field iterations are unrolled to perform structured corrections. This approach establishes a verification paradigm demonstrating that explicit structure outperforms black-box capacity. Extensive experiments on datasets such as COD10K show that our method significantly surpasses parameter-matched baselines, achieves performance comparable to foundation models with minimal parameters, and effectively enhances generalization to unseen generators.

conditional random fieldgeneralizationimage segmentation

This work addresses the challenges of region-based image editing—namely precise spatial localization, preservation of background consistency, and seamless boundary blending—by introducing MaskFlow, a novel framework that integrates mask information into the flow-matching generative process for the first time. MaskFlow employs mask-guided probabilistic paths to jointly regulate content generation within editable regions and preservation outside them. Furthermore, it incorporates a Soft-Poisson de-blending module to refine the vector field, enabling natural fusion between foreground and background. Coupled with a mask-driven MEData synthesis strategy, the proposed method consistently outperforms existing approaches on both natural images and infographics, with quantitative and qualitative results demonstrating superior performance in editing accuracy, background fidelity, and boundary seamlessness.

background preservationprecise localizationregional image editing

This work addresses the limitations of traditional minimal path-based image segmentation methods, which suffer from insufficient accuracy in complex backgrounds and high sensitivity to initialization. To overcome these issues, the authors propose a geodesic framework–based mask proposal voting mechanism that generates diverse and reliable mask candidates through adaptive domain partitioning and constructs the final segmentation via a voting strategy incorporating prior knowledge. The approach innovatively integrates region-based min-cut evolution, geodesic distance computation, and mask scoring maps, effectively eliminating dependence on initial seed points. Experimental results demonstrate that the proposed method significantly outperforms existing minimal path segmentation techniques across multiple challenging scenarios, achieving state-of-the-art performance in both robustness and accuracy.

complex scenariosimage segmentationinitialization sensitivity

Existing 2D instance segmentation models often produce fragmented and inconsistent masks across multiple views, hindering reliable 3D scene understanding. This work proposes a multi-cue guided approach for generating cross-view consistent 2D instance masks by integrating semantic, geometric, and structural information, along with an identity matching mechanism to align instances across viewpoints. These consistent masks are then leveraged to guide the optimization of a 3D Gaussian Splatting feature field, marking the first effective incorporation of instance-level consistency into the Gaussian Splatting framework. Experiments demonstrate that the proposed method significantly improves cross-view mask consistency and 3D instance segmentation stability while preserving high-quality photometric reconstruction, thereby enabling robust downstream editing tasks.

3D Gaussian Splattingcross-view consistencyinstance segmentation

This work addresses the challenge of fairly evaluating the true performance of diverse backbone architectures in image segmentation, which has been hindered by inconsistencies in decoders, training strategies, and pretraining protocols. To this end, the authors propose LUMA (Lightweight Universal Mask Adapter), a lightweight and architecture-agnostic decoder based on cross-attention, enabling unified training recipes across backbones. Using LUMA, they conduct the first systematic benchmark under identical conditions, evaluating 20 backbones—including ViTs, CNNs, and MoEs—across 11 pretraining strategies and multiple resolutions. Their experiments reveal that the choice of pretraining objective exerts a far greater influence on segmentation performance than architectural design. Moreover, LUMA matches the accuracy of the state-of-the-art efficient ViT-based segmenter EoMT at lower computational cost and demonstrates that so-called “efficient” token mixers offer no advantage at high resolutions, with standard ViTs consistently occupying the Pareto frontier in throughput–accuracy trade-offs.

benchmarkingimage segmentationmodel comparison

Hot Scholars

YP

Yulin Pan

Alibaba Group
computer visionmultimedia search
CH

Chen Hui

Harbin Institute of Technology & Nanyang Technological University
image compressionquality assessmentmultimedia securityimage and video processing
DM

Deyu Meng

Professor, Xi'an Jiaotong University
Machine LearningApplied MathematicsComputer VisionArtificial Intelligence