Less Is More: Scalable Visual Navigation from Limited Data

πŸ“… 2026-01-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the data scarcity and generalization bottlenecks in goal-oriented visual navigation caused by reliance on large-scale, high-quality human demonstrations. To this end, we propose LiMoβ€”a data-efficient Transformer-based navigation policy that integrates synthetic SE(2) trajectories generated by a geometric planner with a small set of human demonstrations. By modeling goal-conditioned navigation from a single RGB image, LiMo strategically blends diverse data sources instead of scaling up expert demonstrations, thereby significantly enhancing generalization. Real-robot experiments demonstrate that LiMo achieves superior navigation performance under limited data regimes, validating that data quality and diversity are more critical than sheer data volume for effective policy learning.

Technology Category

Computer Vision: Multi-modal VisionSearch and Optimization: Learning to SearchIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Practical large-scale studies of user experience
πŸ“ Abstract
Imitation learning provides a powerful framework for goal-conditioned visual navigation in mobile robots, enabling obstacle avoidance while respecting human preferences and social norms. However, its effectiveness depends critically on the quality and diversity of training data. In this work, we show how classical geometric planners can be leveraged to generate synthetic trajectories that complement costly human demonstrations. We train Less is More (LiMo), a transformer-based visual navigation policy that predicts goal-conditioned SE(2) trajectories from a single RGB observation, and find that augmenting limited expert demonstrations with planner-generated supervision yields substantial performance gains. Through ablations and complementary qualitative and quantitative analyses, we characterize how dataset scale and diversity affect planning performance. We demonstrate real-robot deployment and argue that robust visual navigation is enabled not by simply collecting more demonstrations, but by strategically curating diverse, high-quality datasets. Our results suggest that scalable, embodiment-specific geometric supervision is a practical path toward data-efficient visual navigation.
Problem

Research questions and friction points this paper is trying to address.

visual navigation
imitation learning
limited data
data efficiency
goal-conditioned navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

imitation learning
visual navigation
synthetic data augmentation
transformer-based policy
geometric planning
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.