🤖 AI Summary
To address insufficient environmental understanding in indoor intelligent robot navigation, this paper proposes a hierarchical 3D scene graph (3DSG) construction framework. The 3DSG comprises three layers: a metric-semantic foundation layer, an object-level point cloud/visual representation layer, and high-level semantic nodes (e.g., rooms, floors). Innovatively, it introduces large language models (LLMs) for the first time to automate labeling across all hierarchical nodes—including room-level and above—and designs an LLM-based voting mechanism for room classification, significantly improving accuracy and robustness of high-level semantic annotation. The method integrates multi-modal RGB-D and point cloud perception, geometry-semantic joint modeling, prompt-engineered LLM reasoning, and hierarchical graph structure optimization. Experiments demonstrate that the generated 3DSG exhibits rich semantics and structural completeness, achieving substantial improvements over baselines in geometric-semantic fusion accuracy, context-aware navigation, and task planning capability.
📝 Abstract
This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Models (LLMs) to construct hierarchical 3D Scene Graphs (3DSGs) for indoor scenarios. The proposed framework constructs 3DSGs consisting of a fundamental layer with rich metric-semantic information, an object layer featuring precise point-cloud representation of object nodes as well as visual descriptors, and higher layers of room, floor, and building nodes. Thanks to the innovative application of LLMs, not only object nodes but also nodes of higher layers, e.g., room nodes, are annotated in an intelligent and accurate manner. A polling mechanism for room classification using LLMs is proposed to enhance the accuracy and reliability of the room node annotation. Thorough numerical experiments demonstrate the system’s ability to integrate semantic descriptions with geometric data, creating an accurate and comprehensive representation of the environment instrumental for context-aware navigation and task planning.