TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead of existing open-vocabulary 3D mapping systems caused by frame-wise segmentation and frequent vision-language reasoning. We propose a novel "track-then-fuse" paradigm that directly tracks masks within image streams to maintain short-term identity consistency, subsequently constructing a hierarchical 3D scene graph by fusing sparse keyframe features with dense propagation. By integrating FastSAM, CLIP, and DINOv3 for multi-view feature extraction, our method achieves class-agnostic 3D instance segmentation and efficient retrieval. Experimental results demonstrate highly competitive performance across multiple benchmarks, yielding a 1.7× inference speedup over state-of-the-art approaches while reducing GPU memory consumption by 3.3×. Furthermore, the proposed system is successfully deployed on a quadruped robot, validating its capability for real-time operation in practical scenarios.
📝 Abstract
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TRACKGRAPH achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7x faster and uses 3.3x less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.
Problem

Research questions and friction points this paper is trying to address.

Open-vocabulary 3D mapping
3D scene graphs
computational efficiency
Vision-Language inference
online robotics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Vocabulary 3D Scene Graphs
Image-Space Tracking
Mask Propagation
Hierarchical Scene Graph
Multi-view CLIP Embeddings
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
P
Peder Borge Hellesylt
Norwegian University of Science and Technology (NTNU), Trondheim, Norway
A
Albert Gassol Puigjaner
Norwegian University of Science and Technology (NTNU), Trondheim, Norway
Kostas Alexis
Kostas Alexis
NTNU - Norwegian University of Science and Technology
RoboticsUnmanned Aerial VehiclesControlPath PlanningPerception
A
Annette Stahl
Norwegian University of Science and Technology (NTNU), Trondheim, Norway