๐ค AI Summary
This work addresses the integration bias in existing multi-pretrained-model approaches to 3D instance segmentation, which arises from discrepancies in model confidence scores and degrades segmentation accuracy. To overcome this limitation, the authors propose a training-free 3D instance segmentation method that introduces, for the first time, a geometryโvision correspondence mechanism. This mechanism achieves precise alignment between 3D geometric cues and 2D visual cues, effectively eliminating confidence bias. By integrating 3D proposal generation with mask-aware CLIP feature extraction, the method enables high-quality, training-free model ensembling. The approach achieves state-of-the-art performance on multiple established 3D segmentation benchmarks and demonstrates exceptional generalization capability in open-vocabulary semantic segmentation tasks.
๐ Abstract
Accurate 3D instance segmentation in point cloud data is critical for machine vision applications. Recent advancements leverage multiple pre-trained foundation models to generate 3D proposals, followed by the application of proposal aggregation methods, which significantly enhance performance. However, they often produce sub-optimal results due to inherent variations in confidence levels across different segmentation models, resulting in a bias toward the model with higher confidence. This bias is inherently model-dependent and is influenced by factors such as data preprocessing techniques and training strategies. To address this bias, we propose a novel, training-free 3D instance segmentation approach via Geometric Visual Correspondence (GVC-Seg), which exploits the correspondence between 3D geometric cues and 2D visual cues to mitigate the confidence bias. Additionally, a 3D proposal generation module and a mask-aware CLIP feature extraction module are introduced during the instance mask generation and instance semantic reasoning, respectively. In this way, GVC-Seg enhances proposal quality assessment, ensuring unbiased ensemble learning across different models. Extensive experiments demonstrate that our method achieves state-of-the-art performance on several challenging benchmarks, while also exhibiting strong potential in open-vocabulary semantic segmentation settings.