Score
Designs, builds, or analyzes algorithms and models that estimate correspondences while enforcing geometric priors—e.g., epipolar or manifold constraints, surface-anchored loci, or anatomy-informed surfaces—covering both sparse and dense matching. This includes constructing geometry-aware feature encoders and dense matching pipelines, restricting similarity search to geometrically feasible regions, and measuring matchability and localization uniqueness under those constraints.
This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.
In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.
Establishing dense correspondences for 3D shapes in real-world scenarios is challenged by the absence of annotations, high resolution, topological distortions, and heterogeneous shape representations. This work proposes the ATM framework, which adopts a “model-then-match” paradigm by integrating pretrained vision foundation models with parametric shape priors to learn a unified shape representation in a shared parameter space from multi-view renderings. Dense correspondences are achieved in a zero-shot manner through geometric consistency constraints and spectral refinement, eliminating the need for correspondence-labeled training data. The method is inherently robust to topological noise and seamlessly handles diverse representations—including meshes, point clouds, and 3D Gaussians. It significantly outperforms existing approaches on non-isometric benchmarks, reducing correspondence errors by 73% on TOPKIDS and 37% on SMAL, while maintaining high efficiency and accuracy on wild scans with up to 200,000 vertices.
Existing semantic correspondence methods based on 2D foundation models lack explicit 3D awareness, often confusing symmetric structures, repetitive parts, and visually similar regions that differ in spatial location. This work proposes a 3D-aware post-training framework that requires no manual pose annotations: it leverages SAM3D to automatically acquire instance-level 3D geometry and pose, refines geometric accuracy through render-and-compare optimization, and projects PartField descriptors onto the image plane to fuse with DINO and Stable Diffusion features for learning semantic correspondences. By incorporating precise geometric priors, the method eliminates the limitations of prior approaches that relied on coarse spherical assumptions and strong supervision, significantly improving correspondence accuracy and outperforming existing post-training techniques.
This work addresses the challenge in partial-to-partial 3D shape matching where the overlapping region is unknown and correspondences are difficult to estimate accurately and simultaneously. We propose the first joint optimization framework based on integer linear programming (ILP) that unifies the discovery of the overlapping region and the establishment of neighborhood-preserving correspondences through geometric consistency priors. Both components are solved concurrently within a single optimization process. As the first approach to introduce ILP to this task, our method achieves significantly higher matching accuracy and smoothness compared to existing techniques, while also demonstrating superior scalability.
Existing 3D shape matching methods predominantly assume complete input shapes, while robust partial-observation matching—more reflective of real-world scenarios—remains underexplored. Current benchmarks suffer from limited scale, unrealistic partiality, and absence of cross-dataset ground-truth correspondences. Method: We introduce the first large-scale, standardized benchmark for partial-observation matching: (1) a programmable geometric perturbation framework that synthesizes photorealistic partial deformations with infinite scalability; (2) integration of seven mainstream datasets with manually annotated cross-dataset full-shape correspondences (2,543 pairs); and (3) a multi-level difficulty evaluation protocol. Results: Comprehensive evaluation reveals substantial performance degradation of state-of-the-art methods under realistic partiality. We publicly release the benchmark—including data, baselines, and an open-source evaluation platform—to establish a new standard and accelerate research in partial 3D shape matching.
Existing monocular geometry estimation methods, constrained by 2D image-space modeling, struggle to recover fine-grained 3D structures such as thin objects and small targets, often resulting in local distortions and excessive smoothing. This work proposes a Self-supervised Sparse Voxel Refinement (SSR) mechanism that elevates geometric modeling into 3D space: starting from a coarse point map generated by a base model, it initializes sparse voxels and employs sparse 3D convolutions to aggregate features within true 3D neighborhoods, iteratively refining geometric details through a self-guided strategy. SSR introduces, for the first time, a self-guided sparse 3D voxel representation that achieves high-fidelity, metric-scale reconstruction while maintaining computational efficiency. Experiments demonstrate significant performance gains over state-of-the-art methods across multiple datasets, with both quantitative metrics and visual results confirming its superior ability to recover complex geometric structures.
This study addresses the problem of recovering true correspondences from multiple noisy and independently permuted point clouds in high-dimensional space. While single-view observation exhibits an information-theoretic impossibility threshold—where exact matching becomes infeasible when the signal strength parameter $b < 2$—this work demonstrates for the first time that incorporating multiple views circumvents this fundamental limitation. Leveraging a high-dimensional Gaussian model and tools from random matrix theory, the authors devise a polynomial-time algorithm that, given $K$ views, achieves near-perfect recovery with only $o(n)$ mismatches whenever $b > K/(K-1)$. Notably, with three views, the method enables efficient and accurate matching in the regime $3/2 < b < 2$, which is provably impossible under a single view.
This study addresses the trade-off in image–point cloud registration between insufficient inliers and an excessively high outlier ratio caused by suboptimal point cloud density, which limits registration accuracy. It presents the first systematic analysis of how point cloud density affects cross-modal registration and introduces a cross-coordinate correspondence pruning mechanism. Specifically, coarse correspondences are projected into the image coordinate system, where a lightweight network fuses geometric and feature information to predict inlier confidence scores for effective outlier rejection. Furthermore, a multi-density point cloud ensemble strategy is employed to enhance inlier recall. The proposed method consistently outperforms existing approaches across multiple benchmarks, achieving a registration recall improvement of at least 8.6%.