Score
Designs, builds, or analyzes interactive segmentation systems that create, refine, and maintain segmentation masks across multiple co-registered modalities and views; these systems keep a shared segmentation state, propagate or transfer labels between modalities, and support user-guided refinement using multimodal similarity measures (for example spectral similarity).
This paper addresses the Multi-Surface Interactive Segmentation (MMMS) task—precise mask generation for multiple highly adjacent or entangled surfaces within a single image, using only sparse user clicks. Methodologically, we propose the first systematic solution: a click-driven architecture supporting both RGB and non-RGB multimodal inputs, enabling efficient feature-level fusion of interaction signals while remaining compatible with off-the-shelf RGB backbone networks; we also introduce dedicated MMMS evaluation metrics. Key contributions include: (1) formal definition of the novel MMMS task; (2) construction of the first interactive segmentation framework supporting multiple surfaces, multimodal inputs, and minimal user clicks; and (3) state-of-the-art performance on DeLiVER and MFNet benchmarks—reducing mean Number-of-Clicks@90 by 1.28 and 1.19 per surface, respectively—with the RGB-only variant matching or surpassing prior art in single-mask scenarios.
Interactive segmentation models face a trade-off: early fusion incurs high latency, while late fusion degrades fine-grained detail perception within prompt regions. This paper proposes a two-stage lightweight fine-tuning framework that preserves SAM’s efficient image encoding capability while dynamically fusing image and prompt features to enhance target-region detail representation. The core innovation is a plug-and-play Refiner module, which performs prompt-guided feature refinement via cross-modal attention at the feature level and local adaptive reweighting. This enables precise, prompt-aware enhancement without architectural overhaul. The method achieves superior accuracy and efficiency: it significantly outperforms state-of-the-art methods on multiple benchmarks—particularly in mIoU and Boundary F-score—while increasing inference latency by less than 3%, thereby maintaining real-time interactive performance.
Interactive segmentation suffers from “part-object” ambiguity: identical user clicks may correspond either to local regions or entire objects, resulting in unstable predictions and hindering scalable, efficient annotation. To address this, we propose a reference-guided ambiguity-aware segmentation framework. First, we introduce a novel cross-image reference guidance mechanism that aligns features via reference images and their corresponding masks. Second, we construct Target Disassembly, the first benchmark dataset explicitly designed for part-object disambiguation, comprising dedicated part-level and object-level subsets. Third, we design a contrastive learning–driven ambiguity encoding module coupled with a lightweight adapter, enabling zero-shot generalization to unseen object categories and part configurations. Extensive evaluation across multiple benchmarks demonstrates state-of-the-art performance—significantly outperforming RITM, SimpleClick, and other leading methods—while achieving high accuracy, minimal click count, and strong robustness.
Automatically assessing the success of segmentation refinement is critical for ensuring segmentation reliability and advancing image segmentation techniques. This paper proposes JFS, the first framework to repurpose off-the-shelf few-shot segmentation (FSS) models—such as SegGPT—for segmentation quality assessment. JFS constructs a novel support set from coarse and refined masks and leverages the FSS model to evaluate refinement effectiveness. Crucially, it introduces a mask-driven support set reconstruction mechanism, enabling plug-and-play evaluation without additional training. Experiments on the PASCAL dataset demonstrate that JFS accurately determines the success or failure of mainstream refinement methods (e.g., SEPL), thereby filling a key research gap in automated, reliability-aware refinement assessment. To our knowledge, JFS is the first quantifiable, general-purpose evaluation tool for high-confidence segmentation.
Remote sensing referring image segmentation (RRSIS) faces challenges in pixel-level localization due to complex geospatial relationships, highly variable object scales, and weak visual saliency. To address these, we propose CroBIM, a cross-modal bidirectional interaction framework featuring three key innovations: (1) a novel attention-deficiency compensation mechanism, (2) context-aware prompt modulation, and (3) language-guided feature aggregation with a mutual interaction decoder. CroBIM integrates multi-scale language-guided attention and cascaded bidirectional cross-attention to enhance fine-grained alignment between linguistic expressions and remote sensing imagery. We introduce RISBench—a large-scale, manually curated benchmark comprising 52,472 triplets—and evaluate CroBIM on RISBench and two established datasets. Our method achieves significant improvements over state-of-the-art approaches across all benchmarks. The code and RISBench dataset are publicly available.
This work addresses the limitations of traditional interactive image segmentation methods—namely, high interaction burden and parameter sensitivity—and those of deep learning approaches, which rely heavily on large annotated datasets and suffer from unstable iterative refinement. The authors propose a sustainable interactive level set method that decouples user guidance into an independent interaction term and replaces the conventional length regularization with a higher-order regularizer. This formulation yields an evolution equation comprising interaction, regularization, and segmentation components, enabling direct manipulation of the zero-level set. The method supports multi-round dynamic refinement while maintaining strong numerical stability and interaction robustness: competitive performance is achieved in the first segmentation round, and subsequent interactions consistently enhance accuracy. Theoretical analysis and extensive experiments validate its superior constraint capability and convergence properties.
This work addresses the challenge of generating semantic regions for 3D asset segmentation, which traditionally relies on manual intervention and struggles to integrate into interactive content creation pipelines. The authors propose a human-in-the-loop approach for producing editable semantic texture atlases by leveraging multi-view rendering and interactive 2D segmentation—combining SAM² with Label Studio—and back-projecting the results into UV space. A greedy set cover strategy is employed to select key views, enhancing computational efficiency. This method delivers the first unified, editable semantic atlas tailored for XR and game development workflows, enabling downstream tasks such as material assignment and style transfer. Experiments on eight cultural heritage objects demonstrate its effectiveness in handling complex geometries and accurately identifying fine details, cavities, and weak boundaries that require human refinement.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This work addresses the absence of reference-free mask quality assessment methods for language-guided audio-visual segmentation. It introduces, for the first time, the MQA-RefAVS task, which defines and implements reference-free evaluation of segmentation masks in this multimodal setting. The proposed approach leverages a multimodal large language model (MLLM) to jointly integrate audio, video, text, and mask inputs, enabling explicit reasoning to predict IoU scores, identify geometric and semantic error types, and provide actionable suggestions for quality improvement. To support this task, the authors construct MQ-RAVSBench, a comprehensive benchmark featuring diverse mask errors, and propose the MQ-Auditor architecture. Experiments demonstrate that MQ-Auditor outperforms existing open-source and commercial MLLMs in assessment accuracy, effectively detects segmentation failures, and enhances downstream segmentation performance.
This work addresses the challenges of high early-stage uncertainty in manufacturing monitoring system development—leading to redundant modeling and substantial training costs—and the limited transferability of filtering pipelines in cross-domain image segmentation tasks. To tackle these issues, the authors propose a problem-centric design paradigm that constructs an abstract system model to continuously accumulate and retrieve historical segmentation tasks along with their associated filtering pipelines, enabling solution reuse and incremental optimization. The approach integrates similarity-based problem retrieval, abstract modeling, pipeline reuse, and a retrieval-augmented evolutionary learning mechanism. Experimental results demonstrate that the method significantly reduces training costs and late-stage revision risks, provides the first systematic validation of filtering pipeline transferability across similar segmentation tasks, and achieves a favorable balance among complexity, technical requirements, and reliability under lightweight model constraints.