🤖 AI Summary
Existing 3D scene graph methods rely on object-level, coarse-grained representations, limiting their applicability to functional robot–environment interaction. This work proposes a fine-grained, function-oriented 3D scene graph that explicitly models functionally manipulable parts—such as door handles and light switches—as first-class nodes, enabling a semantic shift from object-level to function-level reasoning. Methodologically, we synthesize multi-source 3D data to generate 2D functional part annotations, train a part-level detector, and integrate it into standard 3D scene graph construction; we further enhance functional grounding via vision-language alignment and task-driven affordance localization. Experiments demonstrate state-of-the-art performance in functional part segmentation and significantly improved accuracy and robustness in mapping natural language instructions to executable robot actions in real-world settings.
📝 Abstract
The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to develop a representation that enables robots to directly interact with their environment by identifying both the location of functional interactive elements and how these can be used. To achieve this, we focus on detecting and storing objects at a finer resolution, focusing on affordance-relevant parts. The primary challenge lies in the scarcity of data that extends beyond instance-level detection and the inherent difficulty of capturing detailed object features using robotic sensors. We leverage currently available 3D resources to generate 2D data and train a detector, which is then used to augment the standard 3D scene graph generation pipeline. Through our experiments, we demonstrate that our approach achieves functional element segmentation comparable to state-of-the-art 3D models and that our augmentation enables task-driven affordance grounding with higher accuracy than the current solutions.