🤖 AI Summary
This study addresses the prohibitive computational overhead of existing open-vocabulary 3D mapping systems caused by frame-wise segmentation and frequent vision-language reasoning. We propose a novel "track-then-fuse" paradigm that directly tracks masks within image streams to maintain short-term identity consistency, subsequently constructing a hierarchical 3D scene graph by fusing sparse keyframe features with dense propagation. By integrating FastSAM, CLIP, and DINOv3 for multi-view feature extraction, our method achieves class-agnostic 3D instance segmentation and efficient retrieval. Experimental results demonstrate highly competitive performance across multiple benchmarks, yielding a 1.7× inference speedup over state-of-the-art approaches while reducing GPU memory consumption by 3.3×. Furthermore, the proposed system is successfully deployed on a quadruped robot, validating its capability for real-time operation in practical scenarios.
📝 Abstract
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.