🤖 AI Summary
Traditional monocular visual SLAM (VSLAM) struggles to balance efficiency and accuracy in large-scale environments: sparse point-cloud reconstruction and bundle adjustment incur high computational costs, while the resulting map representation is ill-suited for downstream navigation tasks. This paper proposes 2GO—a lightweight, reconstruction-free monocular VSLAM method. Its core innovation lies in the first integration of high-accuracy two-view loop closure detection with DINOv2-guided monocular depth priors to directly construct and optimize a sparse pose graph. The pipeline leverages SuperPoint+SuperGlue feature matching, two-view geometric modeling, and incremental pose graph optimization (PGO). Evaluated on large-scale benchmarks, 2GO achieves state-of-the-art absolute trajectory error (ATE < 0.15 m), real-time performance (>30 Hz), >90% map storage compression, and robust optimization over trajectories exceeding 10 km.
📝 Abstract
(Visual) Simultaneous Localization and Mapping (SLAM) remains a fundamental challenge in enabling autonomous systems to navigate and understand large-scale environments. Traditional SLAM approaches struggle to balance efficiency and accuracy, particularly in large-scale settings where extensive computational resources are required for scene reconstruction and Bundle Adjustment (BA). However, this scene reconstruction, in the form of sparse pointclouds of visual landmarks, is often only used within the SLAM system because navigation and planning methods require different map representations. In this work, we therefore investigate a more scalable Visual SLAM (VSLAM) approach without reconstruction, mainly based on approaches for two-view loop closures. By restricting the map to a sparse keyframed pose graph without dense geometry representations, our '2GO' system achieves efficient optimization with competitive absolute trajectory accuracy. In particular, we find that recent advancements in image matching and monocular depth priors enable very accurate trajectory optimization from two-view edges. We conduct extensive experiments on diverse datasets, including large-scale scenarios, and provide a detailed analysis of the trade-offs between runtime, accuracy, and map size. Our results demonstrate that this streamlined approach supports real-time performance, scales well in map size and trajectory duration, and effectively broadens the capabilities of VSLAM for long-duration deployments to large environments.