🤖 AI Summary
This work addresses the limited mapping accuracy of conventional RGB-D SLAM systems under sparse depth conditions. We propose a real-time RGB-D SLAM framework built upon Pow3R that innovatively treats depth information as a network prior for two-view 3D reconstruction rather than as geometrically fused data, enabling the inference of missing depth and refinement of point cloud predictions. Efficient execution is achieved through sensor fusion combined with a hybrid variant acceleration algorithm. Experimental evaluations demonstrate that the proposed system operates 1.6× faster than MASt3R-SLAM while reducing trajectory error by 15% and map Chamfer distance by 30%. Furthermore, it consistently outperforms ORB-SLAM3 across multiple datasets, substantially enhancing both pose estimation and dense mapping capabilities in scenarios characterized by sparse depth observations.
📝 Abstract
We present Pow3R-SLAM, a real-time RGB-D simultaneous localization and mapping (SLAM) system that uses Pow3R for tracking and mapping. Inspired by MASt3R-SLAM, a recent work on monocular SLAM using two-view 3D reconstruction priors, we extend the work to incorporate depth as a prior on the network's prediction, rather than as geometry to fuse. Where traditional RGB-D SLAM systems struggle with sparsity in the depth images, Pow3R utilizes the available depth to give a better-conditioned pointmap, while inferring the depths in empty regions from the two-view photometric, depth, and intrinsic data. Evaluated against MASt3R-SLAM following its protocol on 24 sequences from TUM, 7-Scenes, and Replica, Pow3R-SLAM runs 1.6x faster in wall time, has 15% lower mean trajectory error, a 3.1x lower unscaled error, and produces denser maps, with a 30% lower Chamfer distance. We also introduce a hybrid variant that runs 2.1x faster than MASt3R-SLAM at 25.3 frames per second (FPS), while maintaining improved tracking and mapping accuracy. Against ORB-SLAM3 in RGB-D mode, Pow3R-SLAM is more accurate on TUM, 7-Scenes, and ETH3D-SLAM, and completes every TUM sequence. While Pow3R-SLAM can struggle on a small set of self-similar scenes, its overall performance shows that adding depth as a prior for two-view 3D reconstruction SLAM can be beneficial. A project webpage is available at: https://ChrisKolios.github.io/Pow3R-SLAM , and code will be made open-source upon acceptance.