OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model

📅 2025-06-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current 3D instance segmentation methods rely on predefined categories or explicit manual prompts, limiting their ability to support open-vocabulary and natural-language-instruction-driven end-to-end reasoning. To address this, we propose the first open-vocabulary 3D instance segmentation framework for point clouds that operates without category priors or annotated prompts, directly interpreting implicit semantic intents from free-form language instructions. Our core innovations include a learnable SEG token and an object-identification mechanism, integrated within a unified architecture combining multimodal large language models, point cloud Transformers, and cross-modal alignment techniques to bridge the semantic gap between language and 3D geometry. Extensive experiments on ScanNet demonstrate that our method comprehensively outperforms prior approaches, achieving state-of-the-art performance across zero-shot transfer, few-shot segmentation, and complex instruction understanding tasks.

Technology Category

Computer Vision: 3D Computer VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Although perception systems have made remarkable advancements in recent years, particularly in 2D reasoning segmentation, these systems still rely on explicit human instruction or pre-defined categories to identify target objects before executing visual recognition tasks. Such systems have matured significantly, demonstrating the ability to reason and comprehend implicit user intentions in two-dimensional contexts, producing accurate segmentation masks based on complex and implicit query text. However, a comparable framework and structure for 3D reasoning segmentation remain absent. This paper introduces OpenMaskDINO3D, a LLM designed for comprehensive 3D understanding and segmentation. OpenMaskDINO3D processes point cloud data and text prompts to produce instance segmentation masks, excelling in many 3D tasks. By introducing a SEG token and object identifier, we achieve high-precision 3D segmentation mask generation, enabling the model to directly produce accurate point cloud segmentation results from natural language instructions. Experimental results on large-scale ScanNet datasets validate the effectiveness of our OpenMaskDINO3D across various tasks.
Problem

Research questions and friction points this paper is trying to address.

Lack of 3D reasoning segmentation framework
Need for implicit query text understanding in 3D
Precision challenges in 3D segmentation mask generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM for 3D understanding and segmentation
SEG token enables precise mask generation
Processes point cloud data with text prompts
🔎 Similar Papers
K
Kunshen Zhang
Wuhan University