HETA++: Global Structure-from-Motion with Hybrid Explicit Translation Averaging

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the susceptibility to collinear degeneracies and outliers in global Structure-from-Motion caused by reliance solely on relative translations or feature tracks. To overcome this limitation, we propose a hybrid explicit translation averaging framework that jointly leverages both relative translation and feature trajectory constraints for the first time. Our approach achieves robust joint estimation of camera poses and 3D points through convex distance-based initialization, non-bilinear angular refinement, spatially balanced feature selection, and coupled rotation–translation optimization. By integrating globally consistent relative translation filtering with reprojection-based bundle adjustment, the method significantly outperforms existing approaches across multiple real-world and unordered datasets, achieving state-of-the-art performance in accuracy, robustness, and computational efficiency.
📝 Abstract
Global Structure-from-Motion (SfM) offers advantages over incremental methods in terms of efficiency and error distribution. However, the task of translation averaging remains challenging. Many existing methods rely solely on relative translations or feature tracks, which either degrade under collinear camera motion or are susceptible to outliers. In this paper, we propose a novel hybrid explicit translation averaging framework that incorporates both relative translations and feature tracks. Specifically, we first refine the relative translations using global camera rotations and remove globally inconsistent relative translations. Next, we employ convex distance-based objective functions to estimate the initial camera positions and 3D points, followed by refinement using a non-bilinear angle-based objective function. Furthermore, since camera rotations are fixed during translation averaging, inaccurate camera rotations can severely limit the accuracy of camera positions. To address this issue, we then robustly refine both camera rotations and camera positions with selected feature tracks through bounded angle-based refinement and subsequent reprojection-based bundle adjustment. In this step, feature tracks are selected to maintain a balanced spatial distribution and improve optimization efficiency. Finally, we perform a complete bundle adjustment using all reliable feature tracks to refine the camera parameters and 3D points. Extensive experiments on various sequential and unordered real-world datasets demonstrate the superior accuracy, robustness, and scalability of our approach, outperforming state-of-the-art methods in both accuracy and computational efficiency.
Problem

Research questions and friction points this paper is trying to address.

Structure-from-Motion
translation averaging
camera rotation
outliers
global SfM
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Translation Averaging
Global Structure-from-Motion
Angle-based Refinement
Robust Bundle Adjustment
Feature Track Selection
🔎 Similar Papers
No similar papers found.