🤖 AI Summary
Interactive segmentation suffers from “part-object” ambiguity: identical user clicks may correspond either to local regions or entire objects, resulting in unstable predictions and hindering scalable, efficient annotation. To address this, we propose a reference-guided ambiguity-aware segmentation framework. First, we introduce a novel cross-image reference guidance mechanism that aligns features via reference images and their corresponding masks. Second, we construct Target Disassembly, the first benchmark dataset explicitly designed for part-object disambiguation, comprising dedicated part-level and object-level subsets. Third, we design a contrastive learning–driven ambiguity encoding module coupled with a lightweight adapter, enabling zero-shot generalization to unseen object categories and part configurations. Extensive evaluation across multiple benchmarks demonstrates state-of-the-art performance—significantly outperforming RITM, SimpleClick, and other leading methods—while achieving high accuracy, minimal click count, and strong robustness.
📝 Abstract
Interactive segmentation aims to segment the specified target on the image with positive and negative clicks from users. Interactive ambiguity is a crucial issue in this field, which refers to the possibility of multiple compliant outcomes with the same clicks, such as selecting a part of an object versus the entire object, a single object versus a combination of multiple objects, and so on. The existing methods cannot provide intuitive guidance to the model, which leads to unstable output results and makes it difficult to meet the large-scale and efficient annotation requirements for specific targets in some scenarios. To bridge this gap, we introduce RefCut, a reference-based interactive segmentation framework designed to address part ambiguity and object ambiguity in segmenting specific targets. Users only need to provide a reference image and corresponding reference masks, and the model will be optimized based on them, which greatly reduces the interactive burden on users when annotating a large number of such targets. In addition, to enrich these two kinds of ambiguous data, we propose a new Target Disassembly Dataset which contains two subsets of part disassembly and object disassembly for evaluation. In the combination evaluation of multiple datasets, our RefCut achieved state-of-the-art performance. Extensive experiments and visualized results demonstrate that RefCut advances the field of intuitive and controllable interactive segmentation. Our code will be publicly available and the demo video is in https://www.lin-zheng.com/refcut.