🤖 AI Summary
Current 3D instance segmentation methods rely on predefined categories or explicit manual prompts, limiting their ability to support open-vocabulary and natural-language-instruction-driven end-to-end reasoning. To address this, we propose the first open-vocabulary 3D instance segmentation framework for point clouds that operates without category priors or annotated prompts, directly interpreting implicit semantic intents from free-form language instructions. Our core innovations include a learnable SEG token and an object-identification mechanism, integrated within a unified architecture combining multimodal large language models, point cloud Transformers, and cross-modal alignment techniques to bridge the semantic gap between language and 3D geometry. Extensive experiments on ScanNet demonstrate that our method comprehensively outperforms prior approaches, achieving state-of-the-art performance across zero-shot transfer, few-shot segmentation, and complex instruction understanding tasks.
📝 Abstract
Although perception systems have made remarkable advancements in recent years, particularly in 2D reasoning segmentation, these systems still rely on explicit human instruction or pre-defined categories to identify target objects before executing visual recognition tasks. Such systems have matured significantly, demonstrating the ability to reason and comprehend implicit user intentions in two-dimensional contexts, producing accurate segmentation masks based on complex and implicit query text. However, a comparable framework and structure for 3D reasoning segmentation remain absent. This paper introduces OpenMaskDINO3D, a LLM designed for comprehensive 3D understanding and segmentation. OpenMaskDINO3D processes point cloud data and text prompts to produce instance segmentation masks, excelling in many 3D tasks. By introducing a SEG token and object identifier, we achieve high-precision 3D segmentation mask generation, enabling the model to directly produce accurate point cloud segmentation results from natural language instructions. Experimental results on large-scale ScanNet datasets validate the effectiveness of our OpenMaskDINO3D across various tasks.