🤖 AI Summary
This work addresses the significant performance degradation of existing visual odometry methods under uncalibrated cameras or variable frame rates, which limits their applicability in open-world scenarios such as dashcam recordings. To overcome these challenges, we propose OpenVO, a novel framework that, for the first time, explicitly models temporal dynamics within a two-frame pose regression paradigm and integrates 3D geometric priors from foundation models to enable scale-aware and robust monocular pose estimation. By jointly leveraging temporal information and rich geometric knowledge, OpenVO effectively handles camera miscalibration and frame-rate variability, substantially improving generalization in unconstrained environments. Experimental results demonstrate that our method outperforms state-of-the-art approaches by over 20% on average across KITTI, nuScenes, and Argoverse 2 benchmarks, and reduces pose error by 46%–92% under variable observation frequencies.
📝 Abstract
We introduce OpenVO, a novel framework for Open-world Visual Odometry (VO) with temporal awareness under limited input conditions. OpenVO effectively estimates real-world-scale ego-motion from monocular dashcam footage with varying observation rates and uncalibrated cameras, enabling robust trajectory dataset construction from rare driving events recorded in dashcam. Existing VO methods are trained on fixed observation frequency (e.g., 10Hz or 12Hz), completely overlooking temporal dynamics information. Many prior methods also require calibrated cameras with known intrinsic parameters. Consequently, their performance degrades when (1) deployed under unseen observation frequencies or (2) applied to uncalibrated cameras. These significantly limit their generalizability to many downstream tasks, such as extracting trajectories from dashcam footage. To address these challenges, OpenVO (1) explicitly encodes temporal dynamics information within a two-frame pose regression framework and (2) leverages 3D geometric priors derived from foundation models. We validate our method on three major autonomous-driving benchmarks - KITTI, nuScenes, and Argoverse 2 - achieving more than 20 performance improvement over state-of-the-art approaches. Under varying observation rate settings, our method is significantly more robust, achieving 46%-92% lower errors across all metrics. These results demonstrate the versatility of OpenVO for real-world 3D reconstruction and diverse downstream applications.