Score
Design and implement algorithms that compute pixel-wise correspondences between image pixels and object-shape representations by iteratively finding nearest matches in pixel or point space (iterative closest pixel/point) and combining two complementary correspondence paths; and formulate shape-guided pixel correspondence loss terms that penalize pixel–shape mismatch to guide, stabilize, and improve pose or alignment optimization.
Image matching faces challenges in robustness and accuracy under complex scenes, particularly for visual localization and 3D reconstruction. To address this, we systematically reformulate the conventional multi-stage pipeline and propose the first dual-dimensional taxonomy—aligned with the detection-description-matching-geometric-estimation workflow—to uniformly evaluate twelve deep matching paradigms across pose estimation, homography estimation, and visual localization. Our method integrates differentiable geometric solvers, end-to-end trainable architectures, contrastive/self-supervised feature learning, and robust optimization modules into a standardized benchmark. Extensive experiments reveal fundamental trade-offs among sparse, semi-dense, and dense matching strategies—as well as pose regression paradigms—in terms of accuracy, robustness, and efficiency. The study identifies key open challenges and delineates principled directions for next-generation matching frameworks.
This work proposes a novel method for establishing point correspondences across image sequences in real time under unknown 3D scene structure and imaging geometry. The approach introduces a channel-vector-based uncertainty density model and employs an online optimization mechanism driven by Neyman chi-square divergence to iteratively learn mappings between image point sets. By representing channel vectors with basis functions and integrating a density divergence criterion, the algorithm achieves rapid convergence and high-accuracy correspondence estimation under general imaging geometries. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques across multiple metrics, offering a compelling combination of real-time performance, robustness, and accuracy.
This work addresses the problem of establishing dense intrinsic correspondences between non-rigid manifolds. We propose a matrix completion framework that jointly incorporates geometric priors—specifically, manifold Laplacian-guided constraints—and sparse functional landmark localization. Methodologically, we are the first to integrate Laplacian-based geometric regularization with ℓ₁-norm sparsity promotion into a unified matrix completion model, effectively mitigating overfitting under limited supervision and enabling precise identification of functionally consistent regions. Optimization is performed via an efficient numerical algorithm. Extensive evaluation on standard non-rigid matching benchmarks (e.g., FAUST, TOSCA) demonstrates state-of-the-art performance: our method achieves the highest overall accuracy and, under highly sparse supervision (fewer than 10 seed points), reduces average correspondence error by 18.7% compared to the best prior approach—substantially improving robustness and generalization capability.
Existing 3D shape matching methods predominantly assume complete input shapes, while robust partial-observation matching—more reflective of real-world scenarios—remains underexplored. Current benchmarks suffer from limited scale, unrealistic partiality, and absence of cross-dataset ground-truth correspondences. Method: We introduce the first large-scale, standardized benchmark for partial-observation matching: (1) a programmable geometric perturbation framework that synthesizes photorealistic partial deformations with infinite scalability; (2) integration of seven mainstream datasets with manually annotated cross-dataset full-shape correspondences (2,543 pairs); and (3) a multi-level difficulty evaluation protocol. Results: Comprehensive evaluation reveals substantial performance degradation of state-of-the-art methods under realistic partiality. We publicly release the benchmark—including data, baselines, and an open-source evaluation platform—to establish a new standard and accelerate research in partial 3D shape matching.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
In point cloud registration, existing methods suffer from fixed iterative optimization paths, implicit correspondence refinement, and single-projection updates prone to local optima. This work introduces, for the first time, denoising diffusion models into the space of doubly stochastic matrices to explicitly model and optimize the distribution of matching matrices. Instead of fixed iterations, it employs the diffusion reverse process—enabling initialization from arbitrary inputs (e.g., white noise)—and integrates Sinkhorn regularization with differentiable geometric feature encoding to enable gradient-guided global matching search. Evaluated on 3DMatch/3DLoMatch and 4DMatch/4DLoMatch benchmarks, our approach achieves significant improvements in both rigid and non-rigid registration accuracy, correspondence quality, and robustness over RAFT-style methods and conventional feature-distance-based approaches.
This study addresses the challenge of establishing cross-modal semantic correspondences between natural images and textureless 3D shapes, overcoming discrepancies in appearance, geometry, and viewpoint. The proposed method distills deep features from a 2D vision model onto the surface of 3D shapes and computes cross-modal feature similarities between image pixels and shape vertices. It introduces a “best segmentation partner” mechanism that identifies 3D vertices whose most similar image pixels fall within coherent image segmentation regions, thereby enabling semantically consistent image-to-shape alignment. Leveraging this correspondence, shape segmentation is performed directly in 3D space through an end-to-end bootstrapped alignment framework. Notably, the approach requires neither surface textures nor manual annotations, and demonstrates strong generality, robustness, and semantic accuracy across diverse image–shape pairs.
This study addresses the susceptibility of alternating minimization to local optima and its computational inefficiency in correspondence-free point set alignment. We propose a global optimization method based on support vectors derived from the convex hull vertices of permuted polygons. By proving a tight bound of $n(n-1)$ vertices, we resolve an open problem posed by Rote. Integrating the Procrustes-Wasserstein framework with a branch-and-bound algorithm, our approach achieves exact solutions in 2D and extends naturally to 3D. Evaluated on the MPEG-7 benchmark, the method requires only 12ms on average, achieving a 50-fold speedup over grid search while delivering superior accuracy. These improvements substantially enhance shape retrieval performance, demonstrating both theoretical rigor and practical efficiency for robust point set registration.
This study addresses the fundamental challenge of estimating correspondences between 3D shape instances under non-rigid deformations by providing a systematic review of existing approaches, which it categorizes into three major paradigms: spectral methods based on functional maps, combinatorial methods incorporating discrete constraints, and deformation-based techniques that directly recover global alignment. For the first time, these three lines of work are unified within a coherent framework, clarifying their historical development, respective strengths, and limitations. A key contribution lies in demonstrating the emerging potential of vision foundation models for zero-shot correspondence tasks. The paper further highlights pressing challenges such as local shape matching, identifies current bottlenecks, and outlines promising future directions, thereby offering both a comprehensive theoretical foundation and practical guidance for advancing research in this domain.
Establishing dense correspondences for 3D shapes in real-world scenarios is challenged by the absence of annotations, high resolution, topological distortions, and heterogeneous shape representations. This work proposes the ATM framework, which adopts a “model-then-match” paradigm by integrating pretrained vision foundation models with parametric shape priors to learn a unified shape representation in a shared parameter space from multi-view renderings. Dense correspondences are achieved in a zero-shot manner through geometric consistency constraints and spectral refinement, eliminating the need for correspondence-labeled training data. The method is inherently robust to topological noise and seamlessly handles diverse representations—including meshes, point clouds, and 3D Gaussians. It significantly outperforms existing approaches on non-isometric benchmarks, reducing correspondence errors by 73% on TOPKIDS and 37% on SMAL, while maintaining high efficiency and accuracy on wild scans with up to 200,000 vertices.
This work addresses the challenge in low-level vision tasks where paired training data often exhibit global photometric inconsistencies that dominate optimization gradients with irrelevant photometric differences, thereby impairing content restoration. The study is the first to reveal, from a gradient energy perspective, that photometric and structural components within the residual between prediction and target are orthogonal, with the photometric component overwhelmingly dominating gradient energy. To mitigate this, the authors propose Photometric Alignment Loss (PAL), which explicitly eliminates global photometric interference through a closed-form affine color alignment mechanism. Leveraging least-squares decomposition and lightweight matrix inversion, PAL achieves effective decoupling of photometric and structural information with negligible computational overhead. Extensive experiments demonstrate consistent performance gains and improved generalization across six task categories, sixteen datasets, and sixteen network architectures.