🤖 AI Summary
This work addresses the world-state inconsistency arising from the conventional separation of SLAM and navigation modules. We present the first unified end-to-end navigation framework that integrates SLAM as an intrinsic mechanism within long-horizon world modeling. By leveraging incremental state updates and backend error optimization, the proposed approach maintains a globally consistent world representation while jointly predicting visual, kinematic, and geometric information to enable closed-loop navigation. Experimental results demonstrate that this framework substantially improves navigation performance while preserving high-precision SLAM estimation.
📝 Abstract
We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend refinement of accumulated errors, to maintain a consistent world state during interaction. Given the current observation and a navigation goal, WSM predicts future visual states and jointly estimates their camera motion and dense geometry, grounding visual prediction in an evolving spatial world state. This spatial state is continuously updated as new observations arrive and provides the basis for action generation and closed-loop navigation. WSM is trained end-to-end with a joint navigation--SLAM objective, enabling downstream navigation to benefit directly from SLAM-style state maintenance and refinement while preserving accurate geometric estimation. Experiments demonstrate improved navigation performance together with strong SLAM accuracy, highlighting the potential of SLAM as an intrinsic mechanism for long-horizon world modeling and embodied interaction.