Score
Designs and implements optimization algorithms and pipelines that jointly refine camera and projector poses and scene structure (3D point positions) by minimizing reprojection or reconstruction errors, while incorporating calibration constraints and regularization to improve stability and accuracy. Builds solvers, parameterizations, and error models for stable convergence and precise multi-view pose-and-structure estimation.
This paper addresses the Structure from Motion (SfM) problem—reconstructing 3D scene structure and camera motion from multi-view point correspondences. To overcome the lack of a unifying theoretical perspective in existing approaches, we propose the first taxonomy of SfM methods based on *problem emphasis*: whether the formulation prioritizes motion estimation, structure reconstruction, or their coupled optimization. Grounded in geometric computer vision and multi-view geometry, we formally model diverse problem settings and, for the first time, establish rigorous connections between SfM’s well-posedness conditions and its underlying modeling paradigms. Our framework enhances method interpretability and extensibility, clarifies fundamental theoretical bottlenecks and open challenges, and provides a unified conceptual foundation for algorithm design, theoretical analysis, and pedagogy.
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
This work proposes PnP-ProCay78, a novel algorithm addressing the accuracy and efficiency challenges in initial pose estimation for the planar Perspective-n-Point (PnP) problem. The method introduces a geometrically transparent and computationally efficient hybrid cost function that innovatively combines projection error minimization with a proxy term for translation reconstruction error derived via analytical elimination. Rotation is parameterized using the Cayley representation, and deterministic initialization coupled with least-squares optimization eliminates the need for costly search procedures. Theoretical analysis reveals convergence properties of the optimization trajectory within Cayley space. Experimental results demonstrate that the algorithm achieves projection accuracy comparable to SQPnP and slightly superior to IPPE on both RGB and low-resolution thermal imaging data, while maintaining a more streamlined structure.
Traditional novel view synthesis relies on external Structure-from-Motion (SfM) tools (e.g., COLMAP), which introduce pose estimation errors and computational bottlenecks. To address this, we propose the first end-to-end framework that jointly optimizes 3D Gaussian splatting radiance fields and camera poses. Our method decouples optimization into two alternating stages—Gaussian parameter refinement and pose fine-tuning—guided by differentiable rendering and a geometric-photometric consistency constraint enforced via 3D optical flow. This design significantly suppresses projection errors, especially under large parallax and sparse-texture conditions. Evaluated on multiple standard benchmarks, our approach not only substantially outperforms existing COLMAP-free methods but also surpasses the COLMAP+3DGS baseline across both reconstruction quality and camera pose accuracy.
Existing initialization-free calibration-aware bundle adjustment methods rely on pseudo-object-space errors (pOSE), which are inherently constrained by projective invariance and cannot exploit known camera intrinsics—resulting in reconstructions defined only up to a projective equivalence class, lacking metric scale and Euclidean structure. Method: We propose a novel initialization-free calibrated Structure-from-Motion (SfM) framework that incorporates known camera calibration into pairwise relative rotation estimation, then jointly optimizes rotations via rotation averaging and pOSE refinement, directly solving for a near-metric reconstruction under the similarity transformation group. Contribution/Results: Our method breaks free from projective ambiguity, explicitly preserving scale and rigid-body geometry. It converges with high probability to the global optimum even from completely random initializations. Experiments demonstrate significantly higher reconstruction accuracy and robustness compared to conventional pOSE-based approaches, establishing a new paradigm for unsupervised calibrated SfM.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
This work proposes a ray-based camera calibration framework tailored for 3D reconstruction, addressing the limitations of traditional reprojection error–based methods that rely on 2D calibration boards and inadequately reflect 3D geometric accuracy. Instead of reprojection error, the approach introduces reconstruction error and intersection error as more representative metrics. It employs a novel icosahedral 3D calibration target and a ring-shaped feature detector, integrated with a generalized distortion model and bootstrapping to refine both intrinsic and extrinsic parameter estimates. Experimental results on synthetic data demonstrate that the proposed method reduces average intersection error by approximately 40%, significantly enhances calibration stability, and validates that ray-level metrics provide a more faithful assessment of 3D reconstruction fidelity compared to conventional approaches.
This work addresses the computational inefficiency of conventional bundle adjustment in multi-camera systems, which jointly optimizes camera poses and 3D points. By leveraging the generalized camera model, the authors introduce, for the first time, a purely pose-based geometric constraint that implicitly encodes 3D points through two reference observations and their associated poses, thereby entirely eliminating 3D point parameters from the optimization. The resulting formulation optimizes only camera poses, significantly enhancing computational efficiency while maintaining or even improving pose estimation accuracy. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method consistently outperforms traditional bundle adjustment in terms of speed and accuracy.
本文提出一种统一框架,通过2D重投影测量优化几何特征和高阶关系,解决了在保持计算效率的同时丰富3D结构的问题。
This work addresses the common trade-off in efficient camera pose estimation, where speed is often achieved at the expense of accuracy, and proposes a hybrid paradigm that integrates neural network-based initializations with classical Structure-from-Motion (SfM) optimization. The approach maintains high reconstruction accuracy while significantly reducing the number of required feature points. To systematically evaluate the efficacy of sparse matching and neural initialization in guiding bundle adjustment, the authors construct a novel SfM benchmark tailored for novel view synthesis. Experiments demonstrate that merely lowering feature density can accelerate conventional SfM pipelines, yet the combination of neural priors with traditional optimization achieves the best balance between efficiency and accuracy. The publicly released benchmark aims to advance research in high-precision, efficient SfM methods.
This work addresses the challenge that existing Perspective-n-Point (PnP) solvers struggle with generalized absolute pose problems involving multiple projection centers. To bridge this gap, the authors propose a virtual point modeling approach that unifies standard PnP and generalized pose estimation within a single framework, enabling conventional PnP solvers to be extended to the generalized setting. Building upon this formulation, they derive three novel solvers based on distinct rotation parametrizations—Cayley parameters, quaternions, and rotation matrices—each balancing accuracy and computational efficiency while preserving global optimality. Experimental results demonstrate that the proposed methods significantly outperform current state-of-the-art generalized pose solvers under heteroscedastic noise conditions.