Score
Designs and implements differentiable algorithms and model layers that estimate rigid (or similarity) transforms aligning two point clouds without computing explicit point-to-point correspondences, using correspondence-free objectives such as maximum mean discrepancy (MMD). Builds and analyzes the optimization and loss formulations, gradient flow, and integration details needed for end-to-end training and robust handling of partial overlap, missing data, and large initial misalignment.
Point cloud registration (PCR) lacks standardized evaluation protocols, hindering fair cross-method comparison of deep learning (DL) approaches under realistic conditions—namely, noise, outliers, and uncertain initial poses. To address this, we propose the first fine-grained taxonomy for DL-based PCR, systematically categorizing methods along four orthogonal dimensions: supervision paradigm (supervised vs. unsupervised), registration pipeline (end-to-end vs. feature-based), optimization strategy (differentiable vs. iterative), and network architecture (e.g., point-wise, graph-based, or transformer-based). We conduct the first standardized quantitative benchmark across 12 representative methods under unified data splits, training configurations, and evaluation metrics. We publicly release our comprehensive evaluation framework and benchmark results. Empirical analysis reveals fundamental performance boundaries and application-specific trade-offs, identifies critical bottlenecks—including sensitivity to large-angle initial misalignments and limited noise robustness—and provides principled guidance for algorithm design and rigorous evaluation.
This work addresses the challenge of point cloud registration by proposing the first differentiable method based on Maximum Mean Discrepancy (MMD) that operates without explicit correspondences and scales efficiently to large-scale data. By approximating MMD via random Fourier features, the registration problem is cast as a nonlinear least-squares optimization with linear computational complexity. Differentiability of the solution is ensured through the integration of the Levenberg–Marquardt algorithm with the implicit function theorem, enabling end-to-end training. The resulting method can be embedded as a differentiable optimization layer within neural networks, outperforming existing learning-based approaches in both supervised and unsupervised settings. When used standalone, it also surpasses state-of-the-art non-learning registration algorithms in terms of accuracy and scalability.
This paper addresses the challenging problem of dense correspondence estimation between non-rigid point clouds—particularly under realistic conditions including near-isometric or heterogeneous shapes, partial overlap, and noise. We propose an end-to-end deep matching framework that requires neither meshing nor manual annotations. Our method uniquely integrates global semantic priors from pretrained vision models into geometric feature learning for point clouds and introduces a deformation-guided module that jointly optimizes extrinsic alignment accuracy and feature discriminability. By enabling cross-modal fusion of visual and geometric features and modeling differentiable deformation constraints, our approach achieves robust and generalizable dense matching. Extensive experiments on standard benchmarks demonstrate state-of-the-art performance, with significant improvements in robustness to non-rigid deformation, occlusion, and noise compared to existing methods.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
Traditional point-to-point or point-to-plane distance metrics in non-rigid point cloud registration suffer from slow convergence and geometric detail loss. To address this, we propose a symmetric point-to-plane distance metric that jointly enforces positional and normal-based geometric constraints, significantly improving geometric fidelity. Methodologically, we introduce the first symmetric distance formulation for non-rigid registration and integrate it with a deformation-graph-based coarse alignment followed by an alternating optimization scheme within the Majorization-Minimization (MM) framework—balancing robustness, accuracy, and efficiency. Extensive experiments on multiple benchmark datasets demonstrate that our approach achieves state-of-the-art registration accuracy while maintaining high computational efficiency. The source code is publicly available.
In point cloud registration, existing methods suffer from fixed iterative optimization paths, implicit correspondence refinement, and single-projection updates prone to local optima. This work introduces, for the first time, denoising diffusion models into the space of doubly stochastic matrices to explicitly model and optimize the distribution of matching matrices. Instead of fixed iterations, it employs the diffusion reverse process—enabling initialization from arbitrary inputs (e.g., white noise)—and integrates Sinkhorn regularization with differentiable geometric feature encoding to enable gradient-guided global matching search. Evaluated on 3DMatch/3DLoMatch and 4DMatch/4DLoMatch benchmarks, our approach achieves significant improvements in both rigid and non-rigid registration accuracy, correspondence quality, and robustness over RAFT-style methods and conventional feature-distance-based approaches.
This study addresses the fundamental challenge of estimating correspondences between 3D shape instances under non-rigid deformations by providing a systematic review of existing approaches, which it categorizes into three major paradigms: spectral methods based on functional maps, combinatorial methods incorporating discrete constraints, and deformation-based techniques that directly recover global alignment. For the first time, these three lines of work are unified within a coherent framework, clarifying their historical development, respective strengths, and limitations. A key contribution lies in demonstrating the emerging potential of vision foundation models for zero-shot correspondence tasks. The paper further highlights pressing challenges such as local shape matching, identifies current bottlenecks, and outlines promising future directions, thereby offering both a comprehensive theoretical foundation and practical guidance for advancing research in this domain.
This study addresses the susceptibility of alternating minimization to local optima and its computational inefficiency in correspondence-free point set alignment. We propose a global optimization method based on support vectors derived from the convex hull vertices of permuted polygons. By proving a tight bound of $n(n-1)$ vertices, we resolve an open problem posed by Rote. Integrating the Procrustes-Wasserstein framework with a branch-and-bound algorithm, our approach achieves exact solutions in 2D and extends naturally to 3D. Evaluated on the MPEG-7 benchmark, the method requires only 12ms on average, achieving a 50-fold speedup over grid search while delivering superior accuracy. These improvements substantially enhance shape retrieval performance, demonstrating both theoretical rigor and practical efficiency for robust point set registration.
Establishing dense correspondences for 3D shapes in real-world scenarios is challenged by the absence of annotations, high resolution, topological distortions, and heterogeneous shape representations. This work proposes the ATM framework, which adopts a “model-then-match” paradigm by integrating pretrained vision foundation models with parametric shape priors to learn a unified shape representation in a shared parameter space from multi-view renderings. Dense correspondences are achieved in a zero-shot manner through geometric consistency constraints and spectral refinement, eliminating the need for correspondence-labeled training data. The method is inherently robust to topological noise and seamlessly handles diverse representations—including meshes, point clouds, and 3D Gaussians. It significantly outperforms existing approaches on non-isometric benchmarks, reducing correspondence errors by 73% on TOPKIDS and 37% on SMAL, while maintaining high efficiency and accuracy on wild scans with up to 200,000 vertices.
This study addresses the challenges of correspondence ambiguity caused by local geometric similarity and the erroneous exclusion of correct solutions during region matching in non-rigid point cloud registration. To overcome these issues, we propose CoCo-Reg, which decouples regional context from dense matching. Specifically, the method enriches point features via regional patches while preserving the global search space to avoid hard constraints, and optimizes patch similarity through identity-corrected overlap supervision. Experimental results demonstrate that CoCo-Reg reduces the mean correspondence error to 0.0547, achieving a 72.6% improvement over the baseline. Furthermore, the proportion of high-error points decreases significantly from 47.3% to 17.3%, indicating substantial improvements in both registration accuracy and robustness.
This work addresses the challenging alignment problem between generative 3D reconstructions and sparse, noisy monocular observations, which is hindered by scale ambiguity, geometric hallucinations, and initial lack of overlap. The authors propose a training-free geometric alignment framework that recovers metric scale and pose via Sim(3) transformation through a coarse-to-fine strategy for robust initialization and precise refinement. Key innovations include an explicit scale factor to resolve scale ambiguity, a geometry-aware descriptor paired with a decoupled closed-form solver, and a hallucination filtering mechanism to suppress spurious geometry generated by neural models. Evaluated on the newly introduced GenPMOAlign–Where2Place benchmark, the method significantly outperforms both classical geometric and state-of-the-art learning-based approaches, achieving stable and highly accurate alignment.