Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of pedestrian appearance distortion caused by drastic viewpoint changes in multi-view, multi-platform cooperative perception scenarios—such as those involving drones and ground cameras—and proposes the FUSION framework to achieve viewpoint-robust cross-view identity association and temporally consistent trajectory tracking. The core innovations include a Multi-cue Adaptive Combination (MAC) module that integrates viewpoint-invariant cues with appearance features, and an Online Multi-view Feature Synchronization (OMFS) strategy for cross-frame feature aggregation. The authors also introduce RealMvMoAT, the largest and most diverse benchmark dataset to date, featuring complex platform motion and heterogeneous viewpoints. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on RealMvMoAT and six public datasets, confirming its effectiveness and robustness in complex, dynamic multi-view environments.
📝 Abstract
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.
Problem

Research questions and friction points this paper is trying to address.

Multi-view Multi-object Association and Tracking
viewpoint variation
pedestrian tracking
cross-view association
mobile aerial-ground platforms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-view Multi-object Association and Tracking
Viewpoint-invariant Feature Fusion
Online Feature Synchronization
Large-scale Benchmark
Mobile Aerial-Ground Platforms
🔎 Similar Papers
R
Ruiqi Wu
School of Computer Science, Northwestern Polytechnical University, Xi’an, China; Ningbo Institute, Northwestern Polytechnical University, Ningbo 315000, China; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi’an, China
Bingliang Jiao
Bingliang Jiao
City University of Hong Kong, Postdoc
GeneralizationRetrievalAgent AI
Ruize Han
Ruize Han
SUAT
Computer VisionMultimedia AnalysisVideo UnderstandingActive Vision
H
Hangzheng Yu
School of Computer Science, Northwestern Polytechnical University, Xi’an, China; Ningbo Institute, Northwestern Polytechnical University, Ningbo 315000, China; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi’an, China
X
Xunkai Jiang
School of Computer Science, Northwestern Polytechnical University, Xi’an, China; Ningbo Institute, Northwestern Polytechnical University, Ningbo 315000, China; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi’an, China
S
Shining Wang
School of Computer Science, Northwestern Polytechnical University, Xi’an, China; Ningbo Institute, Northwestern Polytechnical University, Ningbo 315000, China; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi’an, China
Y
Yuanqi Hu
School of Computer Science, Northwestern Polytechnical University, Xi’an, China; Ningbo Institute, Northwestern Polytechnical University, Ningbo 315000, China; National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, Xi’an, China
Wenxuan Wang
Wenxuan Wang
Northwestern Polytechnical University
Computer ScienceArtifical Intelligence
Peng Wang
Peng Wang
School of Computer Science, Northwestern Polytechnical University, China
Computer VisionMachine LearningArtificial Intelligence