🤖 AI Summary
This work addresses the challenge of pedestrian appearance distortion caused by drastic viewpoint changes in multi-view, multi-platform cooperative perception scenarios—such as those involving drones and ground cameras—and proposes the FUSION framework to achieve viewpoint-robust cross-view identity association and temporally consistent trajectory tracking. The core innovations include a Multi-cue Adaptive Combination (MAC) module that integrates viewpoint-invariant cues with appearance features, and an Online Multi-view Feature Synchronization (OMFS) strategy for cross-frame feature aggregation. The authors also introduce RealMvMoAT, the largest and most diverse benchmark dataset to date, featuring complex platform motion and heterogeneous viewpoints. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on RealMvMoAT and six public datasets, confirming its effectiveness and robustness in complex, dynamic multi-view environments.
📝 Abstract
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.