VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of metric pose recovery in uncalibrated video sequences, where existing methods are often sensitive to initialization or fail to exploit temporal information. The authors propose a novel approach that integrates SLAM-style temporal constraints with global Structure-from-Motion (SfM) optimization. The key innovation lies in explicitly modeling temporal structure as a core component of video-based SfM, combined with wide-baseline dense matching, temporal-aware loop closure detection, and monocular depth priors to guide global bundle adjustment. This enables highly robust metric reconstruction without requiring known camera intrinsics. Evaluated on diverse datasets featuring extreme camera motion and visual symmetries, the method consistently outperforms state-of-the-art SLAM and SfM techniques—both with and without known calibration—achieving significant improvements in accuracy and robustness.
📝 Abstract
Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.
Problem

Research questions and friction points this paper is trying to address.

Structure-from-Motion
SLAM
camera calibration
metric reconstruction
temporal structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structure-from-Motion
SLAM
temporal ordering
metric reconstruction
monocular depth priors
🔎 Similar Papers
2024-03-05IEEE transactions on circuits and systems for video technology (Print)Citations: 0