VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation

📅 2026-02-11
📈 Citations: 0
Influential: 0
📄 PDF

career value

206K/year
🤖 AI Summary
This work proposes a memory- and speed-efficient online semantic SLAM framework to address the challenge of simultaneously achieving geometric consistency, semantic coherence, and computational efficiency in indoor 3D semantic mapping and navigation. By integrating SLAM with semantic mapping, the method constructs a spatiotemporally consistent 3D scene understanding system based on a VGGT tracking head. It processes arbitrarily long video streams via a sliding window, aligns local submaps using camera poses, and lifts 2D instance masks into 3D objects with consistent identity labels. Floor-plane projection is further leveraged to enable efficient navigation support. Evaluated on ScanNet++, the approach achieves competitive point cloud performance while consuming less than 17 GB of GPU memory, enabling real-time interactive assistive navigation with audio feedback.

Technology Category

Application Category

📝 Abstract
We present SceneVGGT, a spatio-temporal 3D scene understanding framework that combines SLAM with semantic mapping for autonomous and assistive navigation. Built on VGGT, our method scales to long video streams via a sliding-window pipeline. We align local submaps using camera-pose transformations, enabling memory- and speed-efficient mapping while preserving geometric consistency. Semantics are lifted from 2D instance masks to 3D objects using the VGGT tracking head, maintaining temporally coherent identities for change detection. As a proof of concept, object locations are projected onto an estimated floor plane for assistive navigation. The pipeline's GPU memory usage remains under 17 GB, irrespectively of the length of the input sequence and achieves competitive point-cloud performance on the ScanNet++ benchmark. Overall, SceneVGGT ensures robust semantic identification and is fast enough to support interactive assistive navigation with audio feedback.
Problem

Research questions and friction points this paper is trying to address.

3D semantic SLAM
indoor scene understanding
assistive navigation
long video streams
temporal coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

VGGT
semantic SLAM
sliding-window mapping
temporal coherence
assistive navigation
A
Anna Gelencsér-Horváth
Pázmány Péter Catholic University, Budapest, Hungary
G
Gergely Dinya
Eötvös Loránd University, Budapest, Hungary
D
Dorka Boglárka Erős
Pázmány Péter Catholic University, Budapest, Hungary
P
Péter Halász
Pázmány Péter Catholic University, Budapest, Hungary
I
Islam Muhammad Muqsit
Pázmány Péter Catholic University, Budapest, Hungary
K
Kristóf Karacs
Pázmány Péter Catholic University, Budapest, Hungary