CalibBEV: LiDAR-Camera Calibration via BEV Alignment

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenging problem of cross-modal calibration between LiDAR and cameras by proposing the first bird’s-eye-view (BEV)-based alignment framework. The method unifies multimodal data into a shared BEV space and employs a two-stage optimization strategy: it first implicitly regresses coarse calibration parameters and then explicitly aligns cross-modal features, enhanced by a CLIP-inspired contrastive loss to enforce semantic consistency. By integrating domain-specific BEV feature extraction with contrastive learning constraints, the approach significantly outperforms existing methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, it reduces relative rotation error by 51% and 68%, and translation error by 80% and 91%, respectively.
📝 Abstract
We present CalibBEV, a novel Bird's Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross-modal calibration. CalibBEV extracts sensor-wise BEV features from each modality using domain-specific architectures and estimates the calibration matrix through a two-step alignment process. First, we perform an implicit alignment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce semantic consistency between BEV representations across modalities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the initial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to-pixel matching methods, achieving state-of-the-art calibration accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.
Problem

Research questions and friction points this paper is trying to address.

LiDAR-camera calibration
BEV alignment
cross-modal calibration
sensor fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

BEV alignment
LiDAR-camera calibration
cross-modal representation
contrastive learning
sensor fusion
🔎 Similar Papers
2024-03-18European Conference on Computer VisionCitations: 8
F
Filippo D'Addeo
University of Bologna, Department of Industrial Engineering, Italy
L
Lorenzo Cipelli
University of Parma, Department of Engineering and Architecture, Italy
Adriano Cardace
Adriano Cardace
Computer Vision Research Scientist at Stanford, Enigma Project
Computer VisionDeep Learning
E
Emanuele Ghelfi
VisLab srl, an Ambarella Inc. company, Italy
A
Andrea Zinelli
VisLab srl, an Ambarella Inc. company, Italy
Massimo Bertozzi
Massimo Bertozzi
University of Parma
deep learningintelligent transportation systemsmachine vision