Score
Designs and implements algorithms and pipelines to establish matches between features detected in 2D images and points, vertices, or surfaces in 3D models or scenes. This work covers feature detection and description, descriptor-based matching, outlier rejection and geometric verification to produce reliable correspondences used for camera pose estimation, registration, alignment, or related analyses.
Cross-view image matching remains challenging due to significant viewpoint discrepancies, compounded by the absence of a unified problem formulation, model architecture, and evaluation protocol in the field. This work presents a systematic survey of the area and introduces the first structured taxonomy encompassing feature extraction, uni- and multimodal matchers, the integration of vision foundation models, and robust training strategies. Under a consistent experimental protocol, the study conducts a fair benchmark evaluation of prominent methods, revealing key design principles underlying the evolution from task-specific models toward generalizable correspondence frameworks. Furthermore, it establishes a reproducible evaluation platform and identifies critical future directions, including computational efficiency, robustness under extreme conditions, and cross-domain generalization.
Image matching faces challenges in robustness and accuracy under complex scenes, particularly for visual localization and 3D reconstruction. To address this, we systematically reformulate the conventional multi-stage pipeline and propose the first dual-dimensional taxonomy—aligned with the detection-description-matching-geometric-estimation workflow—to uniformly evaluate twelve deep matching paradigms across pose estimation, homography estimation, and visual localization. Our method integrates differentiable geometric solvers, end-to-end trainable architectures, contrastive/self-supervised feature learning, and robust optimization modules into a standardized benchmark. Extensive experiments reveal fundamental trade-offs among sparse, semi-dense, and dense matching strategies—as well as pose regression paradigms—in terms of accuracy, robustness, and efficiency. The study identifies key open challenges and delineates principled directions for next-generation matching frameworks.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
This study addresses the insufficient robustness and accuracy of local feature matching in overlapping regions of satellite imagery. To this end, the authors construct a manually curated satellite image dataset annotated with GPS coordinates and conduct a systematic evaluation of SIFT and ORB algorithms across the entire matching pipeline—including keypoint detection, descriptor extraction, feature matching, and RANSAC-based geometric verification. Using the inlier ratio as the primary metric for matching quality, the work quantitatively analyzes the impact of keypoint quantity on matching performance. The results reveal a nonlinear relationship between the number of detected keypoints and the inlier ratio, offering empirical evidence and theoretical guidance for algorithm selection and parameter tuning in remote sensing image matching tasks.
Addressing the challenge of simultaneously achieving robustness, efficiency, and generalization in global point cloud registration, this paper introduces an open-source C++ library. The proposed end-to-end pipeline integrates a lightweight Faster-PFH feature descriptor, a k-core graph-theoretic outlier pruning strategy, and robust pose solvers (e.g., RANSAC and TEASER+). Key contributions include: (i) Faster-PFH, which drastically reduces feature computation overhead while preserving discriminability; (ii) k-core pruning, lowering outlier rejection complexity from O(n²) to near-linear time; and (iii) a modular, highly extensible architecture that maintains high accuracy. Extensive experiments on standard benchmarks—including 3DMatch and KITTI—demonstrate that our method achieves 2–5× speedup over state-of-the-art robust registration approaches, with comparable registration accuracy, while supporting large-scale point clouds and cross-scenario generalization.
This paper addresses robust point-set matching under outliers and noise. We propose an invariant matching mechanism based on distance profiles, constructing noise- and outlier-robust distance-based feature representations in abstract metric spaces. Theoretically, we establish, for the first time, high-probability guarantees for successful matching, linking our approach to the Gromov–Wasserstein distance and deriving a novel sample complexity upper bound; we further prove that matching success probability grows exponentially with sample size. Experimentally, our method significantly outperforms baselines—including ICP and RANSAC—on synthetic data and diverse real-world benchmarks featuring structural noise, outliers, and non-rigid deformations. Key contributions include: (i) a unified distance-profile modeling framework; (ii) the first theoretical guarantee of joint robustness to both noise and outliers; and (iii) a rigorous analysis of scalability to general metric spaces.
PARTE方法通过利用平面结构作为补充注册证据,结合点和平面对应关系进行全局点云配准,解决了因重叠区域有限、重复几何和传感器噪声导致的外点问题。
This work addresses the challenge of simultaneously achieving high accuracy, robustness, and loop-closure capability in two-frame pose optimization by proposing a unified framework that integrates geometric and photometric information. For the first time, dense geometric feature descriptors are incorporated into differential photometric optimization, replacing conventional photometric residuals with descriptor-based residuals to enable subpixel-level pose estimation in descriptor space. By synergistically combining the strengths of both geometric and photometric paradigms, this approach explores a novel trajectory for pose optimization grounded in descriptor similarity. Experimental results demonstrate a significant improvement in tracking accuracy; however, overall performance remains slightly inferior to reprojection error–based methods, with the primary bottleneck identified as the relatively flat landscape of descriptor similarity, which limits optimization efficacy.
This study addresses the trade-off in image–point cloud registration between insufficient inliers and an excessively high outlier ratio caused by suboptimal point cloud density, which limits registration accuracy. It presents the first systematic analysis of how point cloud density affects cross-modal registration and introduces a cross-coordinate correspondence pruning mechanism. Specifically, coarse correspondences are projected into the image coordinate system, where a lightweight network fuses geometric and feature information to predict inlier confidence scores for effective outlier rejection. Furthermore, a multi-density point cloud ensemble strategy is employed to enhance inlier recall. The proposed method consistently outperforms existing approaches across multiple benchmarks, achieving a registration recall improvement of at least 8.6%.
本文提出PESTO算法,利用四面体作为通用特征解决LiDAR点云配准问题,特别是在重叠区域有限的环境下,并证明了其在最坏情况下的误差界限。
This study addresses the susceptibility of alternating minimization to local optima and its computational inefficiency in correspondence-free point set alignment. We propose a global optimization method based on support vectors derived from the convex hull vertices of permuted polygons. By proving a tight bound of $n(n-1)$ vertices, we resolve an open problem posed by Rote. Integrating the Procrustes-Wasserstein framework with a branch-and-bound algorithm, our approach achieves exact solutions in 2D and extends naturally to 3D. Evaluated on the MPEG-7 benchmark, the method requires only 12ms on average, achieving a 50-fold speedup over grid search while delivering superior accuracy. These improvements substantially enhance shape retrieval performance, demonstrating both theoretical rigor and practical efficiency for robust point set registration.
This study addresses the problem of improving accuracy in 3D reconstruction and surface normal estimation in stereo vision by leveraging affine correspondences. Recognizing that conventional methods often neglect local geometric deformations, the authors propose a novel approach to estimate local affine transformations from oriented image correspondences and integrate them into the fundamental matrix estimation and epipolar geometry framework to enhance surface normal reconstruction. To quantitatively evaluate performance, a specialized calibration object comprising three mutually orthogonal checkerboard planes is constructed, and experiments are conducted on both synthetic and real images. Results demonstrate that, under typical stereo configurations and planar orientations, the proposed method achieves surface normal estimation errors of only a few degrees in real-world scenes, thereby validating the efficacy and practical limits of modeling affine correspondences.