🤖 AI Summary
To address the performance degradation of 3D object detection in roadside intelligent infrastructure under aerial-view settings caused by domain shift, this paper proposes a vehicle-infrastructure cooperative 2.5D detection framework. Methodologically, it eliminates redundant height modeling and instead introduces a parallelogram-based parameterization for vehicle ground projections—encoding position, size, and orientation—to better align with top-down geometric priors. The framework jointly leverages real-world and synthetically generated data for training and enables collaborative perception via V2X communication. Experiments demonstrate substantial improvements in detection accuracy and cross-view generalization, particularly under unseen viewpoints and adverse weather conditions, confirming strong environmental robustness. The model weights and inference code are publicly released.
📝 Abstract
On-board sensors of autonomous vehicles can be obstructed, occluded, or limited by restricted fields of view, complicating downstream driving decisions. Intelligent roadside infrastructure perception systems, installed at elevated vantage points, can provide wide, unobstructed intersection coverage, supplying a complementary information stream to autonomous vehicles via vehicle-to-everything (V2X) communication. However, conventional 3D object-detection algorithms struggle to generalize under the domain shift introduced by top-down perspectives and steep camera angles. We introduce a 2.5D object detection framework, tailored specifically for infrastructure roadside-mounted cameras. Unlike conventional 2D or 3D object detection, we employ a prediction approach to detect ground planes of vehicles as parallelograms in the image frame. The parallelogram preserves the planar position, size, and orientation of objects while omitting their height, which is unnecessary for most downstream applications. For training, a mix of real-world and synthetically generated scenes is leveraged. We evaluate generalizability on a held-out camera viewpoint and in adverse-weather scenarios absent from the training set. Our results show high detection accuracy, strong cross-viewpoint generalization, and robustness to diverse lighting and weather conditions. Model weights and inference code are provided at: https://gitlab.kit.edu/kit/aifb/ATKS/public/digit4taf/2.5d-object-detection