π€ AI Summary
This study addresses the problem of cross-task memory conflicts that hinder lifelong learning in embodied navigation by proposing a training-free framework that integrates memory processing into the navigation loop. Methodologically, it leverages maps, historical records, and housing knowledge to verify observations and correct actions. A novel structured recovery handoff mechanism replaces conventional summarization, enabling experience reuse across sessions and revealing that lifelong navigation depends on reasoning sessionsβ continuous accumulation of historical experience. By combining multi-turn agent dialogues, SLAM-based pose estimation, and end-of-run summarization, the approach achieves an 18.6β22.6 point improvement in success rate (s-SR) on GOAT-Bench. Specifically, the Astra model attains 83.7 s-SR and 36.9 episode-level SR (e-SR), while IR2R-CE reaches 85.9 s-SR, establishing new state-of-the-art performance.
π Abstract
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new observations. We present NavHarness, a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, its multi-round agentic session draws on maps, task records, and house knowledge, checking them against observations and recording corrections to guide its actions. NavHarness preserves this experience across fresh conversations for new tasks or recovery attempts, while outcome verification and run-end summaries support its later reuse. On GOAT-Bench, NavHarness improves s-SR over context-only independent sessions by 18.6 points with Astra and 22.6 with Opus 5. Using SLAM-estimated poses, NavHarness with GPT-6 Astra achieves state-of-the-art task success of 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE. To understand these gains, we examine how experience is carried between sessions and find that structured recovery handovers outperform length-matched summaries. In extended deployments across houses, consolidation improves navigation beyond retaining maps and task records, with case studies showing how agents use earlier experience to interpret new goals, investigate unresolved questions, and resume failed searches. We suggest that progress towards lifelong navigation depends on how successive reasoning sessions build on prior experience, alongside improvements in single-task capability.