๐ค AI Summary
This study addresses feature misalignment caused by transmission latency and the neglect of spatial quality variations during fusion in asynchronous cooperative perception. To tackle these challenges, we propose an ego-vehicle-referenced framework integrating predictive alignment with reliability-aware fusion. Specifically, current ego-vehicle features guide trajectory field prediction to achieve precise spatiotemporal alignment. Subsequently, dual-stream features are adaptively reweighted based on trajectory discrepancies and refinement magnitudes, enabling reliability-aware fusion. Extensive experiments on the V2V4Real and DAIR-V2X-Seq datasets demonstrate that the proposed method outperforms TraF-Align by 1.6 and 2.9 percentage points in AP@0.5 and AP@0.7, respectively, significantly improving 3D object detection performance in asynchronous scenarios.
๐ Abstract
Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating occlusion. Under asynchronous communication, however, cooperative features arrive with temporal delay. Existing prediction-based methods compensate for these features mainly from the transmitting agent's own history, leaving residual misalignment with the ego agent's current observation; subsequent fusion also often overlooks spatial variations in alignment quality. We propose EgoRefine, an ego-referenced predictive alignment and reliability-aware fusion framework for asynchronous collaborative perception. Its Ego-referenced Predictive Alignment module uses the current ego feature to guide cooperative trajectory-field prediction and refines the sampling offsets along an ego-referenced trajectory direction. Its Trajectory-conditioned Reliability-aware Fusion module treats the trajectory discrepancy between the ego and cooperative streams and the directional refinement magnitude as alignment cues, using them to condition the relation between aligned features and adaptively reweight the two streams before convolutional fusion. Experiments on V2V4Real and DAIR-V2X-Seq show that EgoRefine outperforms TraF-Align by 1.6 and 2.9 points on average in AP@0.5 and AP@0.7, respectively. The source code will be made publicly available at https://github.com/godk0509/EgoRefine.