Score
Designs and implements methods to compute projective transformation (homography) matrices—typically 3×3 maps—from point or feature correspondences (or other measurements), including robust estimation and iterative refinement; and uses those maps to warp or register images and map planar views into canonical coordinate frames.
本文提出了一种快速贝叶斯方法,通过点对应关系估计单应性矩阵,并明确考虑测量不确定性和先验知识,提高了在噪声条件下的估计精度。
Planar homographies possess eight degrees of freedom, yet conventional four-corner offset parameterizations lack geometric interpretability and require solving an 8×9 linear system to recover the homography matrix. To address this, we propose a decoupled geometric parameterization based on the Similarity–Kernel–Similarity (SKS) decomposition, explicitly factoring the homography into two orthogonal four-dimensional parameter groups: similarity transformations and kernel transformations. Crucially, we establish, for the first time, an analytical linear mapping between kernel parameters and angular offsets, enabling direct, closed-form generation of the homography matrix without linear system solving. Evaluated on deep homography estimation tasks, our method achieves accuracy comparable to four-corner regression while significantly enhancing parameter interpretability and inference efficiency. This work introduces a novel paradigm for homography modeling that unifies geometric meaning with computational advantages.
本文通过多视角结构从运动技术解决了三维定向投影形状分析问题,提高了重建精度和统计推断能力。
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
In point cloud registration, existing methods suffer from fixed iterative optimization paths, implicit correspondence refinement, and single-projection updates prone to local optima. This work introduces, for the first time, denoising diffusion models into the space of doubly stochastic matrices to explicitly model and optimize the distribution of matching matrices. Instead of fixed iterations, it employs the diffusion reverse process—enabling initialization from arbitrary inputs (e.g., white noise)—and integrates Sinkhorn regularization with differentiable geometric feature encoding to enable gradient-guided global matching search. Evaluated on 3DMatch/3DLoMatch and 4DMatch/4DLoMatch benchmarks, our approach achieves significant improvements in both rigid and non-rigid registration accuracy, correspondence quality, and robustness over RAFT-style methods and conventional feature-distance-based approaches.
This study addresses the challenge of achieving high-precision camera guidance and alignment for multiple rectangular planar regions under extremely limited annotation—requiring only a single labeled image. To this end, the authors propose a geometry-centric intra-image navigation framework that leverages homography as the central organizing variable to unify modeling, alignment, and evaluation. The method integrates intra-image augmentation to generate synthetic training data and employs a two-stage inference mechanism—comprising global detection followed by local refinement—alongside a Stable Warp training strategy. This approach substantially improves alignment accuracy even with low-resolution inputs and enables sparse keypoint localization together with sample-level confidence estimation. The work establishes a robust foundation for geometry-driven camera guidance and self-supervised learning in unconstrained video settings.
This work addresses the failure of traditional Structure-from-Motion (SfM) methods in planar scenes due to degeneracy in epipolar geometry. The authors propose a unified framework that leverages homography graph embedding to actively exploit planar structures as geometric constraints. By decomposing multi-view homographies induced by dominant planes, relative camera poses are estimated and organized into a pose graph. Global optimization is then achieved through spectral embedding combined with a maximum spanning tree for consensus maximization. The method significantly outperforms existing SfM approaches in strongly planar environments—such as basketball courts—and achieves state-of-the-art or comparable performance on the IMC Phototourism outdoor benchmark, demonstrating robust camera pose recovery in both degenerate and non-degenerate scenarios.
We present an efficient and robust method for 3D geometric reconstruction that is based solely on the camera-independent linear relationships among a given set of points, which are stable over time and robustly estimated using multiple point matches. We essentially learn, from correspondences between points across several frames, a linear geometric auto-regression matrix $\mathbf{W}$, which establishes how a point in 3D can be expressed as a linear combination of all the others. This matrix is constant and does not depend on the world coordinate system or the camera pose---it is an intrinsic property of the point set. We also show that the principal eigenvectors of $\mathbf{W}$, which all have eigenvalue $1$, provide a homogeneous representation of the 3D point configuration. The first version of our method takes advantage of noisy monocular depth maps in order to obtain, from multiple frames, a robust geometric auto-regression matrix $\mathbf{W}$ of linear relationships between the 3D points. Thus, we build on recent advances in deep learning, which now provide monocular depth estimation models that are fast but very often noisy. Our approach handles noise through robust linear estimation over several frames. The second version of our method does not need monocular depth estimation maps. It applies in cases of weak-perspective projection, when the linear combinations between the 3D points can be robustly estimated from their 2D projections in the image. Note that the camera projection matrix is never used in our derivations. Consequently, our method does not recover camera pose, but only 3D structure. This is a key difference between our method and the related literature on 3D geometric reconstruction.
This work proposes a novel method for establishing point correspondences across image sequences in real time under unknown 3D scene structure and imaging geometry. The approach introduces a channel-vector-based uncertainty density model and employs an online optimization mechanism driven by Neyman chi-square divergence to iteratively learn mappings between image point sets. By representing channel vectors with basis functions and integrating a density divergence criterion, the algorithm achieves rapid convergence and high-accuracy correspondence estimation under general imaging geometries. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques across multiple metrics, offering a compelling combination of real-time performance, robustness, and accuracy.
为解决大规模多摄像头跟踪中的校准难题,提出了一种基于实时检测流的在线单应性校准方法,通过异步优化地面平面几何,实现自动适应环境变化的隐私安全跟踪。