🤖 AI Summary
This study addresses the lack of 3D geometric priors in vision-language navigation (VLN) models and the high inference costs of existing solutions by proposing a training-time geometric supervision framework. The method introduces a depth tokenizer and learnable geometric query tokens, internalizing geometric knowledge into the navigation policy via a self-supervised depth reconstruction loss. Auxiliary modules are discarded after training, achieving zero additional inference overhead. Experiments demonstrate that this framework surpasses current state-of-the-art models on continuous VLN benchmarks. By combining high performance with a lightweight design, the proposed approach facilitates deployment on edge devices.
📝 Abstract
Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inference-time price: depth sensors, 3D encoders, or per-step perception tool calls. We propose GeoScaffold, a geometric supervision framework that pays this price once, at training time, by internalizing geometry into the policy itself. It first learns a compact depth tokenizer on depth maps from the training trajectories and freezes it. It then fine-tunes the policy with a handful of learnable geometry query tokens, training their hidden states to reconstruct navigation-critical geometry such as depth, connectivity, and traversability. This supervision turns the query states into compact geometric latents for action decoding, and through the shared weights also internalizes geometry into the backbone's own representations. Like a scaffold, the tokenizer, target generators, and reconstruction heads are discarded after training, leaving the backbone and action interface unchanged. Extensive experiments show that GeoScaffold consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.