🤖 AI Summary
This study addresses the challenge of associating local observations with global spatial memory in language-guided navigation by proposing top-down maps as explicit spatial memory and action-conditioned prediction targets. By jointly training navigation and map generation through a shared multimodal backbone, this work pioneers the use of map prediction as an auxiliary supervision signal to correlate historical views with spatial locations, thereby unifying the interfaces for navigation, question answering, and 3D grounding. Experimental results demonstrate that the proposed method achieves success rates of 56.9% and 54.9% on the R2R and RxR benchmarks, respectively, alongside significant improvements in SPL. Furthermore, it exhibits strong performance on ScanQA and surpasses baselines such as NaVid in real-world deployment on a Unitree Go2 robot.
📝 Abstract
Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.