GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation
This study addresses the lack of 3D geometric priors in vision-language navigation (VLN) models and the high inference costs of existing solutions by proposing a training-time geometric supervision framework. The method introduces a depth tokenizer and learnable geometric query tokens, internalizing geometric knowledge into the navigation policy via a self-supervised depth reconstruction loss. Auxiliary modules are discarded after training, achieving zero additional inference overhead. Experiments demonstrate that this framework surpasses current state-of-the-art models on continuous VLN benchmarks. By combining high performance with a lightweight design, the proposed approach facilitates deployment on edge devices.