🤖 AI Summary
Existing cross-modal contrastive distillation methods overemphasize modality-shared features while neglecting modality-specific information, leading to incomplete 3D representations. To address this, we propose CMCR—a novel framework that systematically identifies, for the first time, the theoretical limitations of contrastive distillation in 3D representation learning. CMCR innovatively integrates geometrically enhanced masked image modeling, voxel occupancy estimation, and a unified multimodal vector-quantized codebook to jointly model both shared and modality-specific features. The unified codebook enables cross-modal semantic alignment, while geometric priors embedded via occupancy modeling enforce structural 3D constraints. Extensive experiments on multiple downstream 3D perception tasks—including 3D object detection and semantic segmentation—demonstrate that CMCR consistently outperforms state-of-the-art image-LiDAR contrastive distillation approaches, substantiating significant improvements in representation completeness and generalization capability.
📝 Abstract
Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the pre-training process, which leads to suboptimal representations. In this paper, we theoretically analyze the limitations of current contrastive methods for 3D representation learning and propose a new framework, namely CMCR, to address these shortcomings. Our approach improves upon traditional methods by better integrating both modality-shared and modality-specific features. Specifically, we introduce masked image modeling and occupancy estimation tasks to guide the network in learning more comprehensive modality-specific features. Furthermore, we propose a novel multi-modal unified codebook that learns an embedding space shared across different modalities. Besides, we introduce geometry-enhanced masked image modeling to further boost 3D representation learning. Extensive experiments demonstrate that our method mitigates the challenges faced by traditional approaches and consistently outperforms existing image-to-LiDAR contrastive distillation methods in downstream tasks. Code will be available at https://github.com/Eaphan/CMCR.