Score
Design and implement visual odometry systems that estimate metric-scale camera motion from a single plenoptic (light-field) sensor by computing depth priors from plenoptic images and integrating them into sliding-window optimization for drift resilience.
This work addresses the scale ambiguity inherent in monocular visual odometry and overcomes limitations of existing light-field methods—such as insufficient accuracy and complex initialization—by proposing a purely optimization-based sparse photometric visual odometry framework. The approach uniquely integrates geometric priors from light-field imaging directly into photometric bundle adjustment, jointly optimizing camera poses and inverse depths within a sliding window while fusing single-frame light-field geometric depth with multi-view temporal constraints. Notably, it achieves metric-scale reconstruction without requiring elaborate initialization and demonstrates significantly improved performance over current light-field visual odometry techniques in both indoor and outdoor environments. Its accuracy rivals that of state-of-the-art optimization- and learning-based methods, effectively mitigating trajectory drift.
To address the challenges of unobservable scale, unstable initialization, and unreliable loop closure in multi-camera SLAM—stemming from arbitrary camera configurations—this paper proposes the first end-to-end visual odometry framework designed for generic multi-camera setups. Methodologically, it introduces a novel learning-driven framework that jointly models multi-stream feature extraction and inter-camera rigid-motion constraints, enabling online scale initialization and refinement. The approach integrates learned feature tracking, multi-camera rigid-body motion priors, multi-source feature map optimization, and multi-view loop closure detection. Evaluated on KITTI-360 and our newly introduced MultiCamData benchmark, the method significantly outperforms existing stereo and multi-camera SLAM systems in pose accuracy, robustness to wide-field-of-view and texture-deprived scenes, and configurational flexibility—requiring no predefined camera geometry. Code and an interactive online demo are publicly available.
Monocular depth estimation suffers from scale ambiguity, while light-field depth estimation lacks public benchmarks and dedicated models. To address these issues, this paper introduces a novel paradigm for metric-scale dense depth estimation from single-shot focused light-field imagery. Methodologically, we propose a two-stage framework: “sparse point cloud guidance → dense relative depth calibration.” First, a learning-based regressor estimates a sparse 3D point cloud; then, geometric calibration combined with multi-source depth maps (DINOv2 and DepthAnything) is fused and aligned to enable end-to-end metric depth prediction. Our contributions include: (i) the Light-Field Stereo (LFS) benchmark—the first light-field depth dataset with stereo-matching ground truth—and (ii) a lightweight, calibratable geometric-learning joint modeling approach. On LFS, our method achieves significantly lower absolute error than state-of-the-art monocular methods, offering a cost-effective solution for robotic perception.
This work addresses the significant performance degradation of existing visual odometry methods under uncalibrated cameras or variable frame rates, which limits their applicability in open-world scenarios such as dashcam recordings. To overcome these challenges, we propose OpenVO, a novel framework that, for the first time, explicitly models temporal dynamics within a two-frame pose regression paradigm and integrates 3D geometric priors from foundation models to enable scale-aware and robust monocular pose estimation. By jointly leveraging temporal information and rich geometric knowledge, OpenVO effectively handles camera miscalibration and frame-rate variability, substantially improving generalization in unconstrained environments. Experimental results demonstrate that our method outperforms state-of-the-art approaches by over 20% on average across KITTI, nuScenes, and Argoverse 2 benchmarks, and reduces pose error by 46%–92% under variable observation frequencies.
Existing visual odometry (VO) methods predominantly rely on two-frame tracking, neglecting temporal context across image sequences. This limitation hinders global motion modeling and trajectory reliability estimation, leading to significant performance degradation under occlusion, dynamic objects, and low-texture conditions. To address this, we propose the first long-horizon, arbitrary-point tracking frontend that jointly exploits visual features, inter-trajectory associations, and temporal evolution cues. Our method introduces a temporal probabilistic modeling framework coupled with a learnable iterative optimization module for per-point uncertainty inference. Key components include multi-cue deep tracking, temporal Bayesian distribution updating, differentiable iterative refinement, and a dynamic anchor selection mechanism. Evaluated on mainstream VO benchmarks, our approach consistently outperforms state-of-the-art methods, achieving substantial improvements in localization robustness and accuracy—particularly in challenging occluded, dynamic, and texture-deprived scenarios.
This work proposes a fully parallelizable, pixel-level distributed visual odometry and depth estimation algorithm that overcomes the inefficiencies of traditional approaches, which rely on transmitting redundant and noisy raw pixel data and struggle with on-sensor deployment. The method introduces, for the first time, an on-chip sensor architecture based on Gaussian Belief Propagation (GBP), where pixels exchange photometric observations and surface normal priors in parallel to reach consensus on camera motion. A keyframe-like anchoring mechanism is incorporated to effectively constrain inter-frame baselines and preserve geometric consistency. Experimental results demonstrate that the proposed approach achieves efficient and robust on-chip visual odometry and depth estimation on real-world datasets.
This work addresses the performance degradation of RGB-D direct sparse visual odometry in complex scenarios involving dynamic objects, occlusions, illumination changes, and unreliable depth measurements, where short-term photometric and geometric consistency assumptions break down. The authors propose Con-DSO, a novel framework that, for the first time, integrates data-driven modeling of consistency uncertainty into RGB-D direct methods. By learning dense photometric and depth-based geometric consistency errors between consecutive frames, Con-DSO generates quality priors that guide pixel selection and enable decoupled photometric–geometric weighting during keyframe tracking, thereby continuously attenuating unreliable observations. This unified approach handles multiple failure modes without relying on external modules or handcrafted rules. Evaluated on five public benchmarks, Con-DSO significantly outperforms baseline methods, reducing absolute trajectory error by over 20% on ICL-NUIM and by 50%–80% on RGB-D Scenes V2, TUM/Bonn Dynamic, and OpenLORIS sequences.
This work addresses the scale ambiguity and drift inherent in monocular depth estimation, which hinder its applicability to robot navigation requiring metric-scale accuracy. The authors propose Metric-DROID, an end-to-end recurrent visual SLAM architecture that anchors depth estimates to physical scale by fusing egomotion from proprioceptive odometry. Key innovations include LSTM-based modeling of high-frequency odometry sequences, an uncertainty-aware bundle adjustment backend (BA_odom) that integrates odometry measurements into geometric constraints via covariance-weighted anchoring, and a selective residual fine-tuning strategy enabling zero-shot metric alignment. This approach effectively mitigates wheel slippage and sensor noise while preserving visual geometric consistency, thereby significantly improving both accuracy and robustness in metric depth estimation.
This work addresses the challenge of achieving robust real-time visual odometry for planetary rovers under extreme illumination conditions and stringent computational constraints. The authors propose a novel monocular visual odometry method leveraging an event camera, which uniquely integrates asynchronous event streams with an Error-State Kalman Filter (ESKF) and incorporates the real-time asynchronous feature tracker RATE. This integration enables low-bandwidth, highly efficient ego-motion estimation. The approach maintains stable performance across high dynamic range lighting and complex terrains, significantly enhancing both the robustness and real-time capability of pose estimation while meeting the practical resource limitations inherent to planetary rover platforms.
为解决长视频中姿态估计的计算成本、上下文模糊性和时间不稳定性问题,提出FFVO方法,通过紧凑表示、层次解码器及轨迹监督实现高效稳定姿态估计。