NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant sim-to-real gap and high data acquisition costs in embodied 3D navigation by pioneering the use of high-fidelity visual generative models as a data engine. Specifically, we propose a scalable training framework for vision-language navigation that synthesizes decentralized navigation data via a text-to-video pipeline and a world-action model paradigm, while employing style diversification techniques to augment long-tail distributions. This approach overcomes the conventional trade-offs between simulated and real-world data, enabling low-cost, large-scale synthesis of high-quality data. In total, approximately 400,000 navigation trajectories were generated. Experimental results demonstrate that model performance improves substantially with data scale, achieving a 75% success rate across diverse tasks and environments in real-world flight experiments.
📝 Abstract
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce NavGen, a text-to-video data generation pipeline that produces diverse vision-language navigation (VLN) episodes across indoor and outdoor scenes. We also propose a style-diversification method that scales up long-tail data that are difficult and costly to collect. The resulting dataset contains approximately 400K navigation episodes. We evaluate our dataset against existing UAV navigation datasets across multiple metrics, and find that the model trained on our data generally improves with scale, outperforming those trained on existing datasets. To validate real-world transferability, we deploy the trained model in world-action-model paradigm to real-world flying experiments. The final model achieves a 75\% success rate across different navigation tasks and environments.
Problem

Research questions and friction points this paper is trying to address.

embodied 3D navigation
sim-to-real gap
data scalability
vision-language navigation
long-tail data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Generative Models
Embodied 3D Navigation
Text-to-Video Generation
Style Diversification
World-Action Model