VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quadratic computational complexity bottleneck of long-sequence visual geometry Transformers and the error accumulation and trajectory drift inherent in chunk-wise alignment. To overcome these limitations, this work proposes a retraining-free coarse-grained skip-edge mechanism that constrains non-adjacent chunks to model long-range dependencies. Furthermore, it leverages reversed input sequences to transform first-frame scale deviations into drift correction signals, and integrates pose graph optimization to eliminate accumulated errors while remaining compatible with loop closure detection strategies. Extensive evaluations on the KITTI dataset and other benchmarks demonstrate that the proposed approach reduces the Absolute Trajectory Error (ATE) by up to 28.3%. The method consistently outperforms the SwiftVGGT baseline, achieving state-of-the-art performance among comparable approaches.
📝 Abstract
Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
Problem

Research questions and friction points this paper is trying to address.

multi-view 3D reconstruction
long sequence scaling
pose graph
accumulated drift
chunk-and-align
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Geometry Transformer
Skip Edges
Drift Correction
Chunk-and-Align
Pose Graph
💼 Related Jobs
No related jobs found.
S
Sungjae Choi
Korea Advanced Institute of Science and Technology, South Korea
H
Hanna Bae
Korea Advanced Institute of Science and Technology, South Korea
S
Sunghyun Baek
Korea Advanced Institute of Science and Technology, South Korea
Junmo Kim
Junmo Kim
School of Electrical Engineering, KAIST
Statistical Signal ProcessingImage ProcessingComputer VisionMachine LearningInformation Theory