CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Camera-only autonomous driving is inherently limited by occlusion and monocular depth uncertainty. This work proposes a Bayesian collaborative perception framework that leverages a VGGT feedforward network to generate uncertainty-aware 3D Gaussian representations, effectively modeling geometric ambiguity. To facilitate efficient multi-agent observation sharing over C-V2X communication, the method introduces dynamic object primitives (DOPs) requiring only 35 bytes each, thereby resolving depth ambiguity without relying on LiDAR. Extensive evaluations demonstrate that the proposed approach achieves performance improvements of 11.48% and 10.62% on the OPV2V+ and DAIR-V2X-C datasets, respectively, significantly outperforming existing vision-only methods.
📝 Abstract
Autonomous vehicles often suffer from limited perception due to occlusions, blind spots, limited sensor range, and the complex nature of surrounding environments. Multi-agent collaborative perception (CP) addresses these challenges by allowing vehicles to share sensory information and reconstruct the scene cooperatively. However, camera-only perception remains fundamentally limited by the uncertainty of distance-dependent monocular depth estimation. We propose CoCam4D, a Bayesian framework for collaborative perception that explicitly models geometric uncertainty. It uses a VGGT-based feedforward network to generate 3D Gaussian scene representations with associated uncertainty estimates, enabling multiple vehicles or agents to efficiently combine their observations. By sharing compact Gaussian primitives, reliable observations from one agent can reduce the depth uncertainty of another without requiring LiDAR sensors. To support real-world deployment, we introduce Dynamic Object Primitives (DOPs), a compact 35-byte representation designed for efficient C-V2X communication. Extensive experiments show that our proposed method consistently outperforms recent vision-only methods, achieving improvements of 11.48% on OPV2V+ and 10.62% on DAIR-V2X-C, demonstrating the potential of geometrically grounded collaborative perception for LiDAR-free autonomous driving.
Problem

Research questions and friction points this paper is trying to address.

collaborative perception
camera-only
autonomous driving
monocular depth uncertainty
multi-agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Collaborative Perception
3D Gaussian Representation
Bayesian Framework
Geometric Uncertainty
Dynamic Object Primitives
🔎 Similar Papers
No similar papers found.