visual slam

Designs, implements, and evaluates algorithms and systems that use visual (camera) data to simultaneously estimate a moving sensor’s pose and build or update a map of the environment. This includes designing visual localization methods, graph-based SLAM formulations and pose-graph optimization, loop-closure and mapping modules, and real-time implementations of visual SLAM pipelines.

visualslam

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Monocular Visual Simultaneous Localization and Mapping: (R)Evolution From Geometry to Deep Learning-Based Pipelines

May 01, 2024
OÁ
Olaya Álvarez-Tuñón
🏛️ Aarhus University | EIVA a/s | Paderborn University

This paper addresses the robustness bottlenecks of monocular visual SLAM in realistic scenarios—namely dynamic scenes, underwater imaging, and high-speed motion. To this end, it proposes the first comprehensive survey framework that jointly ensures classification consistency and quantifiable evaluation. The work systematically traces the evolution of geometric and end-to-end learning paradigms, unifies the modeling of these three canonical challenges, and establishes a reproducible, quantitative evaluation benchmark. It rigorously delineates the boundaries between the two paradigms, integrating multi-view geometry, nonlinear optimization, CNNs/RNNs, domain adaptation, and robust feature learning to enable cross-environment performance analysis. Empirical findings reveal distinct failure modes of existing methods under varying imaging conditions, thereby providing a systematic benchmark and an extensible research roadmap for designing and deploying robust monocular SLAM systems.

Evaluating resilience of SLAM pipelines under varying conditions.Identifying environmental challenges affecting SLAM performance.Surveying geometry-based and learning-based visual SLAM algorithms.

Factor Graph-Based Active SLAM for Spacecraft Proximity Operations

Jan 19, 2025
LT
Lorenzo Ticozzi
🏛️ Georgia Institute of Technology

To address the high uncertainty in monocular visual localization and mapping during spacecraft proximity operations, this paper proposes an autonomous environmental modeling and self-localization method integrating active perception with factor graph smoothing-based SLAM. We introduce information-entropy-driven active perception into the spacecraft SLAM framework for the first time, enabling joint optimization of observation planning and state estimation through online camera pose optimization. The method models the joint posterior distribution of the state trajectory and map landmarks using a factor graph, and dynamically selects observation strategies that maximize information gain based on information-theoretic metrics. Numerical simulations demonstrate that, compared to passive perception, the proposed approach significantly improves pose and map accuracy, accelerating state uncertainty convergence by approximately 40%. These results validate the effectiveness and engineering applicability of the closed-loop “perception–planning–estimation” coordination mechanism.

Environmental MappingMonocular VisionPositioning

This work addresses the high communication overhead and low data efficiency in multi-robot collaborative SLAM, which often stems from reliance on low-level feature matching. The authors propose a distributed SLAM framework based on scene graph matching that leverages RGB-LiDAR fused point clouds for semantic segmentation, extracting discrete objects and their boundaries to construct scene graphs. Notably, inter-robot loop closures are achieved solely by exchanging object labels and centroids, eliminating dependence on raw feature descriptors. Integrated with a multi-stage communication scheme and distributed optimization, the method significantly reduces communication load while preserving localization and mapping accuracy, as demonstrated in both simulated and real-world experiments with legged robots across indoor and outdoor environments.

data efficiencydistributed SLAMmulti-robot

Loop Closure from Two Views: Revisiting PGO for Scalable Trajectory Estimation through Monocular Priors

Mar 20, 2025
TY
Tian Yi Lim
🏛️ ETH Zurich | Microsoft | University of Bonn

Traditional monocular visual SLAM (VSLAM) struggles to balance efficiency and accuracy in large-scale environments: sparse point-cloud reconstruction and bundle adjustment incur high computational costs, while the resulting map representation is ill-suited for downstream navigation tasks. This paper proposes 2GO—a lightweight, reconstruction-free monocular VSLAM method. Its core innovation lies in the first integration of high-accuracy two-view loop closure detection with DINOv2-guided monocular depth priors to directly construct and optimize a sparse pose graph. The pipeline leverages SuperPoint+SuperGlue feature matching, two-view geometric modeling, and incremental pose graph optimization (PGO). Evaluated on large-scale benchmarks, 2GO achieves state-of-the-art absolute trajectory error (ATE < 0.15 m), real-time performance (>30 Hz), >90% map storage compression, and robust optimization over trajectories exceeding 10 km.

Achieve real-time performance in large-scale environments.Enhance Visual SLAM scalability without dense reconstruction.Optimize trajectory accuracy using two-view loop closures.

Multi-cam Multi-map Visual Inertial Localization: System, Validation and Dataset

Dec 05, 2024
YW
Yufei Wei
🏛️ Zhejiang University | Hangzhou Normal University | Hong Kong University of Science and Technology (HKUST)

Real-time robotic control demands causal pose estimation—relying solely on past and current observations—yet visual SLAM violates causality via non-causal loop closure optimization, and visual-inertial odometry (VIO) suffers from unbounded drift. This paper introduces the first causal visual-inertial localization framework: a tightly coupled multi-camera–multi-map architecture that achieves bounded drift through online map selection, cross-map constraint propagation, IMU preintegration, and joint keyframe feature optimization. We establish the first comprehensive evaluation framework for causal localization, including a formal causal error model and a real-time relocalization mechanism. Evaluated on a newly collected long-term campus dataset and public benchmarks, our method guarantees strictly bounded localization error, improves accuracy by 37%, and operates in real time. The system implementation and dataset are publicly released.

Achieves causal pose estimation for robot control loopsAddresses drift accumulation through multi-camera map constraintsProvides real-time localization with bounded error constraints

Latest Papers

What's happening recently
View more

Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions

Oct 01, 2025
TN
Thanh Nguyen Canh
🏛️ Japan Advanced Institute of Science and Technology | Vietnam National University, University of Engineering and Technology

The semantic visual SLAM field lacks a systematic survey, particularly regarding the integration of deep learning and large language models (LLMs). To address this gap, we propose a unified problem formulation that decomposes semantic SLAM into five core modules: visual localization, semantic feature extraction, map construction, data association, and loop closure optimization. We introduce a modular analytical framework that unifies classical geometric approaches with modern semantic understanding techniques—including semantic segmentation, object detection, scene understanding, and LLM-based reasoning—and conduct empirical evaluations on benchmark datasets. Our work provides the first comprehensive taxonomy of technical evolution, critically analyzes limitations of existing methods, and identifies key bottlenecks: semantic consistency, cross-modal alignment, and real-time performance. The study establishes an authoritative knowledge base and a scalable technical roadmap for future research in semantic SLAM.

Identifying future research directions for Semantic SLAM developmentProviding a unified problem formulation and modular solution frameworkSurveying state-of-the-art Semantic SLAM techniques and challenges

This work addresses the challenge of insufficient robustness in real-time localization for high-speed, highly maneuverable drone racing, where motion blur, unstable visual features, and monocular vision constraints degrade performance. To this end, the authors propose a dual-pose-graph optimization framework that tightly integrates visual-inertial odometry with semantic gate detection cues. The approach employs a temporary graph to aggregate multi-view observations across keyframes, generating refined constraints that are subsequently incorporated into the main pose graph. This design effectively controls graph complexity without additional computational overhead, thereby balancing accuracy and real-time performance. Experimental results demonstrate a 56%–74% reduction in absolute trajectory error (ATE) on the TII-RATM dataset, yielding a 10%–12% accuracy improvement over single-graph baselines. In A2RL competition scenarios, the method reduces per-lap localization drift by up to 4.2 meters, enabling real-time onboard deployment.

autonomous drone racingmotion blurreal-time performance

This study addresses the key challenges and synergistic mechanisms in integrating simultaneous localization and mapping (SLAM) with wireless communication, with a focus on the bidirectional enhancement between visual SLAM and radio frequency (RF) signals for state estimation, scale recovery, and perception augmentation. Through a systematic review of wireless channel modeling, RF-aided localization, visual feature extraction, and perception-driven motion control, this work presents the first comprehensive synthesis of their mutual benefits: monocular visual SLAM can resolve scale ambiguity using RF information, while 5G+ communication systems can leverage visual odometry to enhance performance. By integrating geometric channel modeling, Bayesian filtering, and multi-sensor fusion, the study establishes a unified theoretical framework for joint communication, perception, and localization, offering a novel pathway toward high-precision state estimation for autonomous robots.

RF-based LocalizationScale AmbiguitySLAM

This work addresses the world-state inconsistency arising from the conventional separation of SLAM and navigation modules. We present the first unified end-to-end navigation framework that integrates SLAM as an intrinsic mechanism within long-horizon world modeling. By leveraging incremental state updates and backend error optimization, the proposed approach maintains a globally consistent world representation while jointly predicting visual, kinematic, and geometric information to enable closed-loop navigation. Experimental results demonstrate that this framework substantially improves navigation performance while preserving high-precision SLAM estimation.

Embodied AILong-horizonNavigation

This work addresses the inconsistency in state estimation and computational inefficiency arising from the rigid structure of factor graphs under asynchronous multi-sensor measurements. To overcome these limitations, the authors propose an incremental dynamic factor graph construction method that incorporates an external evaluation criterion to select the optimal graph topology in real time. This approach enables synchronous fusion of asynchronous multi-source sensor data and supports on-the-fly graph compression to reduce the number of optimization variables. Experimental results demonstrate that the proposed method maintains map accuracy comparable to conventional approaches while reducing the average number of graph nodes by approximately 30%, thereby significantly lowering computational complexity. The key innovation lies in the evaluation-driven dynamic topology construction and compression mechanism, which effectively balances estimation accuracy and computational efficiency.

asynchronous sensorsfactor graphgraph topology

Hot Scholars

LX

Lihua Xie

Professor of Electrical Engineering, Nanyang Technological University
Robust controlNetworked ControlMult-agent Systems
GK

Giseop Kim

Assistant Professor, Dept of Robotics and Mechatronics Eng, DGIST
Mobile RoboticsField RoboticsSLAM
XY

Xianjia Yu

University of Turku
Sensing and PerceptionRoboticsMachine/Deep Learning
TW

Tomi Westerlund

Professor, University of Turku, Finland - DIWA Flagship (https://digitalwaters.fi/)
Internet of ThingsUAVUGVUSV
YT

Yifu Tao

University of Oxford
3D ReconstructionComputer VisionRobotics