PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of error accumulation and the absence of geometric priors in multi-view RGB-D point cloud registration by proposing a training-free, zero-shot coarse-to-fine framework. First, it leverages visual geometry foundation models to directly recover metrically consistent global poses, thereby circumventing error propagation inherent in pairwise estimation. Subsequently, fine alignment is achieved by integrating voxel hashing with robust motion bundle adjustment, solved via iteratively reweighted least squares (IRLS) and conjugate gradient optimization. Experimental results demonstrate that the proposed method achieves superior zero-shot registration accuracy and computational efficiency across indoor, object-centric, and outdoor scenes.
📝 Abstract
This paper addresses multiview RGB-D point cloud registration, aiming to estimate global rigid poses for unordered RGB-D scans and align them in a metrically consistent coordinate frame. The conventional pairwise-then-global paradigm suffers from locally optimized pairwise registration, severe error propagation and high computational burden. In particular, existing methods typically treat RGB data as a mere auxiliary matching cue and overlook the holistic geometric priors (e.g., camera poses and 3D models) encoded across image sequences. This paper introduces PointVGGT, a zero-shot framework built upon a novel \emph{foundation-then-refinement} paradigm that systematically leverages visual geometry foundation models (e.g., VGGT) as the computational backbone for robust, training-free multiview RGB-D registration. In the foundation stage, we directly recover metrically consistent global poses (without any pairwise estimation) by grounding the scale-ambiguous pose predictions of the foundation model against metric depth observations. In the refinement stage, we introduce an efficient voxelized spatial hashing mechanism that exploits the globally coherent 3D reconstruction (induced by the foundation model) as a shared spatial anchor, enabling dense multiview correspondences in near-linear time. On top of this, an IRLS-based robust motion-only bundle adjustment is performed using a conjugate gradient solver to jointly minimize the correspondence and reprojection residuals for multiview pose refinement. Extensive experiments on indoor/object-centric/outdoor datasets verify the outstanding zero-shot registration accuracy and computational efficiency of our proposed method.
Problem

Research questions and friction points this paper is trying to address.

multiview RGB-D point cloud registration
global rigid pose estimation
error propagation
geometric priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Registration
Visual Geometry Foundation Model
Multiview RGB-D
Voxelized Spatial Hashing
Bundle Adjustment