OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

📅 2026-06-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D scene graph generation methods are object-centric and struggle to model part-level details and multi-level relationships, limiting fine-grained scene understanding. This work proposes the first open-vocabulary, part-aware unified 3D scene graph framework that jointly represents objects, interactable parts, spatial and functional relationships, and affordances. By integrating object-part knowledge-guided detection, part-aware 3D feature fusion, geometry-prior-initialized relation modeling, and joint optimization with large language models, our approach enables efficient and accurate relational reasoning. We introduce a new benchmark, UniGraph3D, on which our method achieves state-of-the-art performance and significantly enhances perception for a variety of robotic tasks.
📝 Abstract
3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabulary 3DSG generation, existing approaches remain object-centric and encode limited relational information -- restricting their applicability in real-world scenarios that require fine-grained understanding. We propose OP3DSG, an open-vocabulary part-aware 3DSG generation framework that constructs unified graphs that jointly model objects, interactive parts, spatial relations, functional relations, and affordances. OP3DSG integrates object-part knowledge-guided detection with part-aware 3D fusion to preserve small and interaction-relevant components, and employs a geometry-initialized prior graph with LLM-based refinement to reduce spurious relational predictions while enabling efficient graph construction. To systematically evaluate unified 3D scene graph construction, we introduce UniGraph3D, a benchmark designed for part-aware perception and multi-level relational reasoning. Experimental results show that OP3DSG achieves state-of-the-art performance and demonstrates its effectiveness as a perception backbone in diverse real-world robotics tasks.
Problem

Research questions and friction points this paper is trying to address.

3D scene graph
open-vocabulary
part-aware
relational reasoning
real-world environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

open-vocabulary
part-aware
3D scene graph
relational reasoning
foundation models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yirum Kim
Gwangju Institute of Science and Technology (GIST), Gwangju, Republic of Korea
U
Ue-Hwan Kim
Gwangju Institute of Science and Technology (GIST), Gwangju, Republic of Korea