Score
Designs and implements calibration and registration systems that estimate spatial (and when needed temporal) transforms and intrinsic/extrinsic parameters between multiple sensing modalities, producing co‑registered geometric and radiometric datasets (e.g., aligned point clouds, imagery, thermal maps). Builds and evaluates sensor integration and co‑calibration pipelines to fuse measurements into a common coordinate frame and validate alignment quality for downstream processing.
To address low accuracy, poor real-time performance, and reliance on calibration targets or hand-crafted environmental features in LiDAR point cloud–camera image cross-modal registration, this paper proposes an unsupervised, end-to-end multi-view projection registration framework. Our method introduces two key innovations: (1) a novel patch-to-pixel differentiable attention matching mechanism that enables fine-grained cross-modal correspondence; and (2) a multi-scale CNN architecture jointly modeling multi-view 2D projections of point clouds and image features, significantly enhancing robustness under small overlap conditions. The framework requires no calibration boards or manual feature engineering, ensuring strong generalizability. Evaluated on KITTI, it achieves a registration accuracy of 99.2% with real-time inference capability. On nuScenes, it surpasses state-of-the-art methods in both accuracy and speed.
To address the misalignment of cross-platform 3D models caused by the absence of geospatial metadata in ground- and airborne imagery, this paper proposes a fully metadata-free georegistration method. Our approach leverages only publicly available satellite imagery and a digital surface model (DSM), integrating neural radiance fields (NeRF) modeling with differentiable rendering to achieve robust geolocalization of airborne imagery via multi-view geometric optimization, and subsequently enables precise alignment of ground-level imagery to the airborne 3D model. We introduce, for the first time, a novel “satellite + DSM + NeRF” collaborative registration paradigm, supporting unified georegistration of non-overlapping and disconnected 3D models. Evaluated across multiple real-world scenes, our method achieves an average geolocation error of under 5 meters—without requiring any raw sensor metadata such as GPS or IMU readings.
Point cloud registration suffers from poor robustness in geometrically ambiguous scenes (e.g., symmetric structures, planar regions). To address this, this paper proposes a cross-modal registration method that jointly leverages RGB radiometric information and 3D geometric features. The core contribution is a novel two-stage cross-modal feature fusion mechanism: first, point clouds are enhanced via pixel-level projection onto RGB images; second, coarse matching is performed by jointly reasoning over image patches and superpoints, with explicit modeling of geometric ambiguities. The method integrates joint point cloud–image encoding, superpoint graph construction, and a coarse-to-fine matching strategy. It achieves state-of-the-art performance on 3DMatch, 3DLoMatch, IndoorLRS, and ScanNet++, with registration recall rates of 95.9% on 3DMatch and 81.6% on 3DLoMatch—substantially outperforming geometry-only approaches.
Calibrating spatiotemporal parameters across asynchronous multi-LiDAR, multi-camera, and IMU systems remains challenging due to hardware-trigger dependency, requirement of calibration targets or overlapping fields-of-view, and sensitivity to scene structure. Method: This paper proposes a target-free, overlap-agnostic continuous-time joint autocalibration framework. It introduces the first targetless continuous-time bundle adjustment formulation, enables cross-modal feature matching between LiDAR intensity images and camera images, and jointly optimizes 6-DoF extrinsics and sensor time delays by integrating Structure-from-Motion, adaptive voxel mapping, and nonlinear optimization (Ceres). Contribution/Results: The method eliminates reliance on hardware synchronization and structured environments, supporting arbitrary numbers of heterogeneous sensors. Evaluated on real-world structured scenes, it achieves sub-0.8-pixel image reprojection error and centimeter-level point-cloud registration accuracy. Calibration results satisfy motion consistency and geometric constraints without drift or cumulative error.
This work addresses the system-level inconsistency in extrinsic calibration of multi-camera–LiDAR systems caused by independent estimation for each camera. To resolve this, the authors propose a two-stage joint calibration framework: first, the CMRNext network is employed to obtain initial extrinsics and 2D–3D correspondences for each camera–LiDAR pair; subsequently, a multi-frame bundle adjustment jointly optimizes all extrinsics by integrating reprojection errors with single-camera priors and inter-camera relative pose constraints, yielding globally consistent estimates. This approach uniquely combines learning-based pairwise initialization with explicit multi-camera geometric constraints. Evaluated on KITTI, it achieves a translation error of 0.89 cm and a rotation error of 0.038°, and on the Walkley dataset, it reduces translation error from 108.6 cm to 3.1 cm, significantly enhancing cross-domain robustness and calibration accuracy.
This work addresses the challenges of poor sensor compatibility and low robustness to large initialization errors in LiDAR–camera extrinsic calibration under target-free scenarios. We propose a calibration method that requires neither calibration targets nor initial pose estimates. Given only a single RGB image and one LiDAR scan line (i.e., two 3D points), our approach leverages GlueStick to establish automatic 2D–3D point–line feature correspondences. An adaptive weighting scheme—based on geometric distribution—and an optimal feature selection strategy jointly guide a nonlinear optimization to estimate extrinsics. Evaluated across diverse LiDAR–camera configurations (e.g., Velodyne, Ouster, Livox with RGB cameras), our method achieves superior accuracy and robustness over state-of-the-art approaches, converging even with translation and rotation initialization errors up to ±1 m and ±30°. The implementation is open-sourced, significantly enhancing practical deployment efficiency.
This work addresses the boundary blurring and distortion in LiDAR–camera extrinsic calibration caused by laser beam footprint effects and mixed-intensity returns. To this end, the authors propose a joint calibration method that integrates boundary response modeling. By co-observing visual fiducials on a printable planar calibration board and LiDAR-visible circular reflective boundaries, the approach iteratively refines 3D LiDAR edge feature points. It further incorporates an intensity- and geometry-constrained refinement strategy and a confidence-weighted reprojection optimization framework. Notably, this is the first method to embed explicit modeling of LiDAR boundary response characteristics into the extrinsic calibration pipeline. Evaluated on real-world data, it achieves sub-pixel reprojection accuracy and millimeter-level feature consistency, substantially improving downstream visual–LiDAR odometry performance.