GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of 3D geometric priors in vision-language navigation (VLN) models and the high inference costs of existing solutions by proposing a training-time geometric supervision framework. The method introduces a depth tokenizer and learnable geometric query tokens, internalizing geometric knowledge into the navigation policy via a self-supervised depth reconstruction loss. Auxiliary modules are discarded after training, achieving zero additional inference overhead. Experiments demonstrate that this framework surpasses current state-of-the-art models on continuous VLN benchmarks. By combining high performance with a lightweight design, the proposed approach facilitates deployment on edge devices.
📝 Abstract
Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inference-time price: depth sensors, 3D encoders, or per-step perception tool calls. We propose GeoScaffold, a geometric supervision framework that pays this price once, at training time, by internalizing geometry into the policy itself. It first learns a compact depth tokenizer on depth maps from the training trajectories and freezes it. It then fine-tunes the policy with a handful of learnable geometry query tokens, training their hidden states to reconstruct navigation-critical geometry such as depth, connectivity, and traversability. This supervision turns the query states into compact geometric latents for action decoding, and through the shared weights also internalizes geometry into the backbone's own representations. Like a scaffold, the tokenizer, target generators, and reconstruction heads are discarded after training, leaving the backbone and action interface unchanged. Extensive experiments show that GeoScaffold consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
3D geometric priors
inference efficiency
embodied navigation
edge deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Navigation
Geometric Latents
Depth Tokenizer
Reconstruction Supervision
Efficient Inference
🔎 Similar Papers
No similar papers found.
Y
Yixuan Jiang
Nanjing University, Nanjing, China
Wentong Li
Wentong Li
Nanjing University of Aeronautics and Astronautics
Computer VisionMachine LearningVision-Language ModelRobotics
A
An Liu
Institute of Automation, Chinese Academy of Sciences, Beijing, China
Z
Zihao Xin
Nanjing University of Aeronautics and Astronautics, Nanjing, China
Fulin Tang
Fulin Tang
Ph.D, University of Chinese Academy of Sciences
SLAM3D reconstructionVLN
C
Cong Leng
MAICRO, Nanjing, China
Y
Yang Gao
Nanjing University, Nanjing, China
Jian Cheng
Jian Cheng
Beijing, China
computational fluid dynamicshigh-order methodsdiscontinuous Galerkin method