Sparse2comm: Towards Robust Cooperative 3D Object Detection

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical challenge in collaborative perception where bandwidth constraints, coupled with packet loss, latency, and spatial misalignment, severely degrade detection performance. To tackle these issues, this work proposes a robust 3D object detection framework based on sparse communication. It introduces a novel sparse-to-dense semantic reconstruction mechanism that uniformly mitigates the effects of limited bandwidth and packet loss. Furthermore, the framework integrates motion flow prediction for latency alignment and incorporates self-supervised spatial offset estimation to enable self-calibrated fusion, effectively suppressing spatiotemporal perturbations. Extensive experiments demonstrate that under mixed degradation scenarios across multiple datasets, the proposed method achieves up to a 20% improvement in Average Precision (AP) over baselines, successfully balancing detection accuracy with system robustness.
📝 Abstract
Cooperative perception improves autonomous driving by sharing complementary observations among vehicles and roadside infrastructure for 3D object detection. However, practical deployment is constrained by limited bandwidth and unreliable cooperation, where packet loss, transmission delay, and spatial misalignment jointly degrade the cooperative feature stream. Existing methods often reduce communication cost or compensate for one degradation type, leaving coupled disturbances insufficiently addressed. To address this problem, we propose Sparse2comm, a bandwidth-efficient and robust cooperative 3D object detection framework that treats unreliable cooperation as progressive restoration over degraded cooperative features. Sparse Feature Encoding first encodes communication as randomly mask-sampled foreground features transmitted by collaborating agents, from which the ego vehicle reconstructs dense semantic representations. This sparse-to-dense mechanism learns to infer missing object-centric content from sparse observations, enabling ultra-low-bandwidth communication and packet-loss recovery within the same representation. On the semantically restored features, Latency-Aware Alignment predicts motion flow to compensate delayed messages, and Self-Calibrating Fusion estimates residual spatial offsets in a self-supervised manner before adaptive cross-agent fusion. Sparse2comm therefore restores semantic completeness, temporal consistency, and spatial alignment in an ordered pipeline. Extensive experiments on DAIR-V2X, OpenV2V, and V2V4Real show that Sparse2comm maintains competitive clean accuracy and consistently improves robustness under individual and mixed real-world degradations. Compared with the selective feature communication baseline Where2comm, Sparse2comm improves mixed-setting AP@0.5/AP@0.7 by +20.15/+11.79, +12.66/+11.07, and +15.36/+12.61 on the three datasets, respectively.
Problem

Research questions and friction points this paper is trying to address.

Cooperative 3D Object Detection
Bandwidth Constraint
Packet Loss
Transmission Delay
Spatial Misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cooperative 3D Object Detection
Sparse Feature Encoding
Robustness
Bandwidth Efficiency
Self-Calibrating Fusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lei Yang
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore
Boqi Li
Boqi Li
Postdoc Research Fellow, Civil and Environmental Engineering, University of Michigan
connected mobility systems
C
Chunmian Lin
School of Transportation Science and Engineering, Beihang University, Beijing, China
L
Li Wang
School of Mechanical Engineering, Beijing Institute of Technology, Beijing, China
Ziying Song
Ziying Song
Beijing Jiaotong University
Object DetectionComputer VisionDeep Learning
Shaoqing Xu
Shaoqing Xu
University of Macau, BUAA, Xiaomi EV
3D Computer Vision3D GenerationVision and Language ModelEnd2EndWorld Model
Heye Huang
Heye Huang
University of Wisconsin–Madison
Autonomous SystemsMulti-AgentsRisk AssessmentInteractive Decision-MakingHuman-Centered AI
H
Haibao Yu
University of Hong Kong, China
C
Chen Lv
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore