visual-inertial odometry

Designs, implements, and evaluates algorithms and systems that estimate a moving platform's pose, velocity, and trajectory by fusing synchronized visual (camera) and inertial (IMU) measurements. This includes building classical filter- or optimization-based pipelines and learned/neural (neural odometry) methods for state estimation, sensor calibration, and drift mitigation.

visual-inertialodometry

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Multi-IMU Sensor Fusion for Legged Robots

Jul 15, 2025
SY
Shuo Yang
🏛️ Carnegie Mellon University

Legged robots suffer from severe pose and velocity estimation drift during highly dynamic maneuvers—such as impacts, slips, and rapid rotations. To address this, we propose a tightly coupled visual–inertial–legged odometry framework. Our approach innovatively employs a distributed multi-IMU configuration across robot links, jointly leveraging joint encoders and monocular camera data to model and compensate dominant error sources in proprioceptive odometry. We formulate a sliding-window factor graph optimization that incorporates extended Kalman filter (EKF)-based preintegration of inertial and joint measurements, while unifying visual features, IMU preintegrations, and foot motion constraints as factors for joint optimization. Experimental results demonstrate centimeter-level localization accuracy and significantly reduced drift under high-dynamic tasks, markedly improving state estimation robustness. The corresponding C++ implementation and a large-scale real-world dataset are publicly released.

Correcting errors in proprioceptive odometry using multiple IMUsEnhancing state estimation under challenging locomotion conditionsReducing pose and velocity drift in legged robots

Vision-Aided Relative State Estimation for Approach and Landing on a Moving Platform with Inertial Measurements

Dec 22, 2025
TB
Tarek Bouazza
🏛️ I3S | CNRS | Université Côte d’Azur | Université du Québec en Outaouis | Lakehead University | Australian National University | Institut Universitaire de France (IUF)

This work addresses the problem of estimating the relative pose and velocity of an unmanned aerial vehicle (UAV) with respect to an arbitrarily moving planar platform in 3D space, enabling precise approach and landing. The proposed method introduces a tightly coupled estimator that fuses measurements from two inertial measurement units (IMUs) and monocular vision—specifically, the platform’s center-line-of-sight direction and surface normal vector. A novel architecture combines an SO(3) Lie-group complementary filter with a linear Riccati-based cascaded observer; critically, it recovers unobservable attitude angles using only the platform’s linear acceleration under known normal-axis rotational constraints—a first in the literature. Theoretical analysis establishes local exponential convergence and almost-global asymptotic stability of the estimation error. Extensive simulations demonstrate high accuracy and strong robustness against dynamic disturbances, sensor noise, and large initial state errors.

Designing stable cascade observers for position and attitudeEstimating UAV-platform relative state during landingUsing IMU and monocular vision for 3D motion tracking

Observer Design for Optical Flow-Based Visual-Inertial Odometry with Almost-Global Convergence

Aug 28, 2025
TB
Tarek Bouazza
🏛️ Université Côte d’Azur | Université du Québec en Outaouais | Lakehead University | Institut Universitaire de France

Monocular visual-inertial odometry (VIO) suffers from local convergence bottlenecks in velocity and gravity direction estimation due to reliance on initial guesses and local linearization. Method: We propose a cascaded observer architecture with almost global asymptotic stability. First, a Riccati-type joint velocity–gravity observer is constructed in the body frame by fusing optical flow direction measurements with IMU data, ensuring global exponential stability. Second, a complementary observer on SO(3) estimates attitude. Third, spherical-constrained gradient descent optimizes optical flow direction usage. Results: Simulation results demonstrate that, under persistently exciting translational motion, the architecture simultaneously achieves high-accuracy, stable estimation of pose, body-frame velocity, and gravity direction. It overcomes classical VIO limitations—namely sensitivity to initialization and local convergence—and significantly improves robustness and applicability across diverse motion scenarios.

Developing stable observer architecture for visual-inertial odometry with global convergenceEstimating body velocity and gravity direction from optical flow and IMU measurementsSolving constrained minimization for velocity direction extraction from sparse optical flow

This work addresses the issue of cumulative error in learning-based inertial odometry caused by direct regression of absolute position. To mitigate this, the authors propose estimating incremental displacement (Δp) over a 50 ms sliding window and reconstructing the trajectory via numerical integration. The study introduces Kolmogorov–Arnold Networks (KANs) into IMU odometry for the first time, leveraging their learnable B-spline activation functions to effectively suppress long-term error accumulation. Experiments on the EuRoC MAV dataset demonstrate that, compared to conventional multilayer perceptrons (MLPs), KANs reduce cumulative drift error by 44% while using only approximately one-sixth the number of parameters, and exhibit superior stability under both P₅₀ and P₉₀ metrics.

cumulative driftdelta-positionIMU

Latest Papers

What's happening recently
View more

This work addresses the problem of estimating the relative pose and velocity of a moving target using an IMU-equipped observer, assuming availability of either relative position or bearing measurements. By modeling the relative dynamics on the SE₂(3) Lie group and embedding them into ℝ¹⁵ to form a linear time-varying system, the authors propose a hybrid estimation architecture that combines a Riccati observer with an SO(3) nonlinear complementary filter. This approach achieves, for the first time within the SE₂(3) framework, a unified fusion of dual-IMU and relative sensing data, establishing a consistency observability condition that relies solely on target acceleration excitation. The method theoretically guarantees global exponential convergence of the estimation error and almost global asymptotic stability of the orientation estimate. Numerical simulations validate the effectiveness of the proposed technique.

Dual IMUMoving TargetRelative Pose Estimation

Monocular visual-inertial odometry struggles to recover metric scale from vision alone, and the relationship between scale observability and trajectory characteristics remains unclear. This work addresses this gap through an observability analysis, revealing that translational acceleration induced by trajectory curvature is the key factor coupling scale with inertial states. Leveraging the asymmetry between gravity and linear acceleration in the IMU model, the study establishes a theoretical framework and, for the first time, explicitly identifies trajectory curvature as the decisive contributor to scale observability. A lightweight trajectory excitation metric is proposed, computable directly from raw IMU data. Experiments on straight, circular, and figure-eight trajectories yield scale errors of 9.2%, 6.4%, and 4.8%, respectively, with excitation levels spanning four orders of magnitude, thereby validating the effectiveness of trajectory design in enhancing scale recovery accuracy.

IMUmetric scalemonocular visual-inertial odometry

This work addresses the high computational complexity and substantial point correspondence requirements that often hinder practical deployment of relative pose estimation in multi-camera systems. The authors propose two efficient minimal solvers that, for the first time, incorporate IMU-derived priors—either the vertical direction or a known rotation axis—into a minimal solver framework. Requiring only four point correspondences, the approach leverages a novel parametrization and algebraic geometry techniques to reduce the problem to solving a univariate sextic polynomial, a significant simplification over existing octic formulations. Integrated within a RANSAC-based robust estimation pipeline, extensive experiments on synthetic data and the KITTI benchmark demonstrate that the proposed method achieves comparable accuracy while substantially lowering computational cost and data requirements.

computational complexityminimal solversmulti-camera systems

Indoor visual localization is hindered by detection noise, occlusions, and limited camera coverage, leading to multi-stage uncertainties that existing fusion methods fail to explicitly model. This work proposes a component-level error quantification and calibration mechanism that explicitly characterizes the uncertainty in homography calibration, human detection, and motion tracking, and leverages these estimates to optimize multi-camera fusion weights. By transforming the fusion process from a black-box into an interpretable framework, the method significantly enhances trajectory stability and motion smoothness. Experimental results demonstrate that, while yielding only marginal gains in absolute localization accuracy over single-camera baselines, the proposed strategy effectively reduces trajectory variance and substantially improves the continuity and robustness of motion estimation.

detection noiseerror characterizationindoor localization

Hot Scholars

DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics
KA

Kostas Alexis

NTNU - Norwegian University of Science and Technology
RoboticsUnmanned Aerial VehiclesControlPath Planning
MG

Maani Ghaffari

Assistant Professor, University of Michigan
RoboticsMachine LearningRobot PerceptionAutonomous Navigation
MF

Maurice Fallon

Professor, University of Oxford
RoboticsComputer Vision
TY

Tzu-Yuan Lin

Postdoctoral Associate, MIT
RoboticsRobot PerceptionMachine LearningGeometric Deep Learning