🤖 AI Summary
To address the degradation of 3D detection performance across vehicle platforms with heterogeneous sensor configurations in autonomous driving, this paper proposes a multi-sensor-aware domain adaptation method. Our approach introduces, for the first time, a synergistic fine-tuning strategy combining downstream task adaptation and partial-layer fine-tuning, leveraging co-located, multi-sensor point cloud pairs for adaptation. Unlike conventional joint training, we selectively fine-tune only the detection head and critical backbone layers, thereby mitigating negative transfer induced by point cloud distribution shifts. Extensive experiments across diverse real-world sensor configurations demonstrate an average +4.2% improvement in BEV mAP. The method exhibits strong generalization and scalability, offering an efficient, lightweight domain adaptation solution for deploying 3D detectors across heterogeneous platforms.
📝 Abstract
Recent advances in autonomous driving have underscored the importance of accurate 3D object detection, with LiDAR playing a central role due to its robustness under diverse visibility conditions. However, different vehicle platforms often deploy distinct sensor configurations, causing performance degradation when models trained on one configuration are applied to another because of shifts in the point cloud distribution. Prior work on multi-dataset training and domain adaptation for 3D object detection has largely addressed environmental domain gaps and density variation within a single LiDAR; in contrast, the domain gap for different sensor configurations remains largely unexplored. In this work, we address domain adaptation across different sensor configurations in 3D object detection. We propose two techniques: Downstream Fine-tuning (dataset-specific fine-tuning after multi-dataset training) and Partial Layer Fine-tuning (updating only a subset of layers to improve cross-configuration generalization). Using paired datasets collected in the same geographic region with multiple sensor configurations, we show that joint training with Downstream Fine-tuning and Partial Layer Fine-tuning consistently outperforms naive joint training for each configuration. Our findings provide a practical and scalable solution for adapting 3D object detection models to the diverse vehicle platforms.