Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between spatial reasoning benchmarks and navigation performance by proposing the Spatial-NPD framework. We first construct the Spatial-Nav-100K dataset and subsequently employ a two-stage fine-tuning strategy combined with large language model training. Through knowledge distillation, a teacher model guides a lightweight student model to achieve efficient navigation without explicit spatial reasoning. The resulting 8B-parameter model attains state-of-the-art performance on benchmarks such as HM3D while maintaining minimal computational overhead, requiring only 45 GPU hours for training and achieving an inference speed of 148 ms per step. This work demonstrates that high-performance embodied navigation can be effectively realized through distilled implicit spatial understanding rather than costly explicit reasoning pipelines.
📝 Abstract
Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation. Guided by these findings, we build \textsc{Spatial-Nav-100K} and fine-tune in two stages, \textit{i.e.} first learning a shared spatial-navigation foundation, and then specializing each phase with the abilities it relies on. We further introduce Spatial-NPD, where a teacher conditioned on spatial priors produces grounded action preferences and distills them into a student policy, so no explicit spatial reasoning is needed at inference. With 45 A100 GPU-hours of policy training, our 8B model reaches SR/SPL of 77.4/35.4 on HM3D-v0.2, 60.2/30.5 on HM3D-v0.1, and 47.9/20.6 on train-unseen MP3D. It outperforms several systems that rely on closed-source models or thousands of GPU-hours of training, at 148 ms per action step. All code and datasets will be publicly available at https://github.com/ylwhxht/Spatial-Nav.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
navigation
benchmark gap
embodied AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Reasoning
Navigation
Knowledge Distillation
Two-stage Fine-tuning
Preference Learning