Score
Design and implement algorithms that jointly estimate and iteratively refine camera poses and 3D point clouds or scene geometry by alternating optimization of pose parameters and point positions, including pose-graph optimization and edge-based pose optimization. Build and analyze components such as nearest-neighbor correspondence solvers, edge-map alignment losses, pose error metrics, and pose-conditioned rendering/generation that converge to consistent pose and geometry (optionally without explicit feature tracks) while managing runtime and memory trade-offs.
Existing camera pose estimation methods suffer from low initial pose accuracy, reliance on task-specific descriptors, or the need for network fine-tuning or scene reconstruction. To address these limitations, we propose GS-SMC—the first lightweight pose optimization framework based on 3D Gaussian Splatting (3DGS). GS-SMC leverages a pre-trained 3DGS model to render multi-view images, establishes 2D–2D correspondences via feature matching, and iteratively refines poses using epipolar geometry constraints—without requiring network adaptation or scene re-reconstruction. Its modular architecture enables flexible substitution of feature extraction and matching components, eliminating dependence on specialized descriptors or networks. Evaluated on the 7-Scenes and Cambridge Landmarks benchmarks, GS-SMC reduces translational error by 53.3% and 40.7%, and rotational error by 56.9% and 53.2%, respectively—outperforming state-of-the-art approaches significantly.
Traditional novel view synthesis relies on external Structure-from-Motion (SfM) tools (e.g., COLMAP), which introduce pose estimation errors and computational bottlenecks. To address this, we propose the first end-to-end framework that jointly optimizes 3D Gaussian splatting radiance fields and camera poses. Our method decouples optimization into two alternating stages—Gaussian parameter refinement and pose fine-tuning—guided by differentiable rendering and a geometric-photometric consistency constraint enforced via 3D optical flow. This design significantly suppresses projection errors, especially under large parallax and sparse-texture conditions. Evaluated on multiple standard benchmarks, our approach not only substantially outperforms existing COLMAP-free methods but also surpasses the COLMAP+3DGS baseline across both reconstruction quality and camera pose accuracy.
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
This work addresses the challenge of simultaneously achieving high accuracy, robustness, and loop-closure capability in two-frame pose optimization by proposing a unified framework that integrates geometric and photometric information. For the first time, dense geometric feature descriptors are incorporated into differential photometric optimization, replacing conventional photometric residuals with descriptor-based residuals to enable subpixel-level pose estimation in descriptor space. By synergistically combining the strengths of both geometric and photometric paradigms, this approach explores a novel trajectory for pose optimization grounded in descriptor similarity. Experimental results demonstrate a significant improvement in tracking accuracy; however, overall performance remains slightly inferior to reprojection error–based methods, with the primary bottleneck identified as the relatively flat landscape of descriptor similarity, which limits optimization efficacy.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
本文提出GRIP,通过高斯渲染解决图像与点云间结构不匹配问题,实现精准的2D到3D配准。
本文提出一种基于几何驱动的数据优化方法,通过主轴对齐解决物体姿态估计中的噪声敏感、对称性混淆问题,且无需修改现有网络架构。
This study addresses the limitation of existing 6D pose tracking methods that rely on costly initialization or real-time reconstruction, thereby struggling to meet the real-time demands of robotic manipulation and augmented reality. To this end, this work proposes a lightweight framework for long-term rigid object tracking. Departing from conventional point-based optimization paradigms, the proposed method constructs a compact pose graph modeled exclusively with relative pose constraints weighted by geometrically aligned uncertainties, effectively decoupling computational complexity from the number of correspondences. Evaluated across four real-world benchmarks, the approach achieves accuracy comparable to reconstruction-based trackers at minimal optimization cost, offering an efficient and robust solution for real-time applications.
本文提出一种统一框架,通过2D重投影测量优化几何特征和高阶关系,解决了在保持计算效率的同时丰富3D结构的问题。
This work proposes a novel method for establishing point correspondences across image sequences in real time under unknown 3D scene structure and imaging geometry. The approach introduces a channel-vector-based uncertainty density model and employs an online optimization mechanism driven by Neyman chi-square divergence to iteratively learn mappings between image point sets. By representing channel vectors with basis functions and integrating a density divergence criterion, the algorithm achieves rapid convergence and high-accuracy correspondence estimation under general imaging geometries. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques across multiple metrics, offering a compelling combination of real-time performance, robustness, and accuracy.