Score
Designs, implements, and evaluates algorithms and systems that estimate a moving platform's pose, velocity, and trajectory by fusing synchronized visual (camera) and inertial (IMU) measurements. This includes building classical filter- or optimization-based pipelines and learned/neural (neural odometry) methods for state estimation, sensor calibration, and drift mitigation.
Legged robots suffer from severe pose and velocity estimation drift during highly dynamic maneuvers—such as impacts, slips, and rapid rotations. To address this, we propose a tightly coupled visual–inertial–legged odometry framework. Our approach innovatively employs a distributed multi-IMU configuration across robot links, jointly leveraging joint encoders and monocular camera data to model and compensate dominant error sources in proprioceptive odometry. We formulate a sliding-window factor graph optimization that incorporates extended Kalman filter (EKF)-based preintegration of inertial and joint measurements, while unifying visual features, IMU preintegrations, and foot motion constraints as factors for joint optimization. Experimental results demonstrate centimeter-level localization accuracy and significantly reduced drift under high-dynamic tasks, markedly improving state estimation robustness. The corresponding C++ implementation and a large-scale real-world dataset are publicly released.
This work addresses the problem of estimating the relative pose and velocity of an unmanned aerial vehicle (UAV) with respect to an arbitrarily moving planar platform in 3D space, enabling precise approach and landing. The proposed method introduces a tightly coupled estimator that fuses measurements from two inertial measurement units (IMUs) and monocular vision—specifically, the platform’s center-line-of-sight direction and surface normal vector. A novel architecture combines an SO(3) Lie-group complementary filter with a linear Riccati-based cascaded observer; critically, it recovers unobservable attitude angles using only the platform’s linear acceleration under known normal-axis rotational constraints—a first in the literature. Theoretical analysis establishes local exponential convergence and almost-global asymptotic stability of the estimation error. Extensive simulations demonstrate high accuracy and strong robustness against dynamic disturbances, sensor noise, and large initial state errors.
Monocular visual-inertial odometry (VIO) suffers from local convergence bottlenecks in velocity and gravity direction estimation due to reliance on initial guesses and local linearization. Method: We propose a cascaded observer architecture with almost global asymptotic stability. First, a Riccati-type joint velocity–gravity observer is constructed in the body frame by fusing optical flow direction measurements with IMU data, ensuring global exponential stability. Second, a complementary observer on SO(3) estimates attitude. Third, spherical-constrained gradient descent optimizes optical flow direction usage. Results: Simulation results demonstrate that, under persistently exciting translational motion, the architecture simultaneously achieves high-accuracy, stable estimation of pose, body-frame velocity, and gravity direction. It overcomes classical VIO limitations—namely sensitivity to initialization and local convergence—and significantly improves robustness and applicability across diverse motion scenarios.
This work addresses the issue of cumulative error in learning-based inertial odometry caused by direct regression of absolute position. To mitigate this, the authors propose estimating incremental displacement (Δp) over a 50 ms sliding window and reconstructing the trajectory via numerical integration. The study introduces Kolmogorov–Arnold Networks (KANs) into IMU odometry for the first time, leveraging their learnable B-spline activation functions to effectively suppress long-term error accumulation. Experiments on the EuRoC MAV dataset demonstrate that, compared to conventional multilayer perceptrons (MLPs), KANs reduce cumulative drift error by 44% while using only approximately one-sixth the number of parameters, and exhibit superior stability under both P₅₀ and P₉₀ metrics.
本文提出PLATO方法,通过精确轨迹观测学习IMU偏差动力学及噪声协方差,以提高神经惯性里程计在复杂环境下的运动估计精度。
This work addresses the problem of estimating the relative pose and velocity of a moving target using an IMU-equipped observer, assuming availability of either relative position or bearing measurements. By modeling the relative dynamics on the SE₂(3) Lie group and embedding them into ℝ¹⁵ to form a linear time-varying system, the authors propose a hybrid estimation architecture that combines a Riccati observer with an SO(3) nonlinear complementary filter. This approach achieves, for the first time within the SE₂(3) framework, a unified fusion of dual-IMU and relative sensing data, establishing a consistency observability condition that relies solely on target acceleration excitation. The method theoretically guarantees global exponential convergence of the estimation error and almost global asymptotic stability of the orientation estimate. Numerical simulations validate the effectiveness of the proposed technique.
Monocular visual-inertial odometry struggles to recover metric scale from vision alone, and the relationship between scale observability and trajectory characteristics remains unclear. This work addresses this gap through an observability analysis, revealing that translational acceleration induced by trajectory curvature is the key factor coupling scale with inertial states. Leveraging the asymmetry between gravity and linear acceleration in the IMU model, the study establishes a theoretical framework and, for the first time, explicitly identifies trajectory curvature as the decisive contributor to scale observability. A lightweight trajectory excitation metric is proposed, computable directly from raw IMU data. Experiments on straight, circular, and figure-eight trajectories yield scale errors of 9.2%, 6.4%, and 4.8%, respectively, with excitation levels spanning four orders of magnitude, thereby validating the effectiveness of trajectory design in enhancing scale recovery accuracy.
This work addresses the high computational complexity and substantial point correspondence requirements that often hinder practical deployment of relative pose estimation in multi-camera systems. The authors propose two efficient minimal solvers that, for the first time, incorporate IMU-derived priors—either the vertical direction or a known rotation axis—into a minimal solver framework. Requiring only four point correspondences, the approach leverages a novel parametrization and algebraic geometry techniques to reduce the problem to solving a univariate sextic polynomial, a significant simplification over existing octic formulations. Integrated within a RANSAC-based robust estimation pipeline, extensive experiments on synthetic data and the KITTI benchmark demonstrate that the proposed method achieves comparable accuracy while substantially lowering computational cost and data requirements.
Indoor visual localization is hindered by detection noise, occlusions, and limited camera coverage, leading to multi-stage uncertainties that existing fusion methods fail to explicitly model. This work proposes a component-level error quantification and calibration mechanism that explicitly characterizes the uncertainty in homography calibration, human detection, and motion tracking, and leverages these estimates to optimize multi-camera fusion weights. By transforming the fusion process from a black-box into an interpretable framework, the method significantly enhances trajectory stability and motion smoothness. Experimental results demonstrate that, while yielding only marginal gains in absolute localization accuracy over single-camera baselines, the proposed strategy effectively reduces trajectory variance and substantially improves the continuity and robustness of motion estimation.
本文评估了多种图像匹配方法在无人机视觉里程计中的应用,旨在解决GNSS信号不可用时的导航问题,发现RoMa匹配器表现最佳。