🤖 AI Summary
This work addresses the challenge of long-term robot navigation in dynamic environments, where persistent scene understanding and generalization capabilities are often lacking. To this end, the authors propose AllDayNav, a novel framework that introduces implicit, memory-driven reinforcement learning to lifelong navigation for the first time. Departing from conventional explicit maps or scene graphs, AllDayNav employs a self-evolving multimodal memory mechanism that integrates visual keyframes, semantic descriptions, and temporal context. Leveraging large language models, the system autonomously generates open-vocabulary instructions, image-based goals, and structured rewards. Experiments demonstrate that AllDayNav achieves near-perfect success rates in both simulated and real-world environments, significantly outperforming strong baselines—including map-based, vision-language, and reinforcement learning approaches—in terms of path efficiency and robustness, while enabling cross-room, cross-task, and cross-time navigation.
📝 Abstract
Lifelong embodied navigation in dynamic environments requires robots to form persistent scene understanding from fragmentary observations, which remains difficult for existing methods that rely on explicit maps or scene graphs and struggle to generalize beyond structured settings. We propose AllDayNav, a lifelong self-learning navigation framework that implicitly encodes scene dynamics into the billion-scale parameters of a large model via reinforcement learning, powered by a self-evolving multimodal memory that maintains and updates visual keyframes, semantic descriptions, and temporal context while autonomously generating open-vocabulary instructions, image goals, and structured rewards. Experiments in both synthetic and real-world environments across cross-room, cross-episode, and cross-task scenarios show that AllDayNav achieves success rates approaching $100\%$ and consistently surpasses strong map-based, VLM, and RL baselines in path efficiency and robustness, demonstrating implicit, memory-driven reinforcement learning as a scalable alternative to explicit mapping for reliable lifelong navigation.