CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving
Camera-only autonomous driving is inherently limited by occlusion and monocular depth uncertainty. This work proposes a Bayesian collaborative perception framework that leverages a VGGT feedforward network to generate uncertainty-aware 3D Gaussian representations, effectively modeling geometric ambiguity. To facilitate efficient multi-agent observation sharing over C-V2X communication, the method introduces dynamic object primitives (DOPs) requiring only 35 bytes each, thereby resolving depth ambiguity without relying on LiDAR. Extensive evaluations demonstrate that the proposed approach achieves performance improvements of 11.48% and 10.62% on the OPV2V+ and DAIR-V2X-C datasets, respectively, significantly outperforming existing vision-only methods.