🤖 AI Summary
This work addresses the scalability limitations of centralized fusion approaches in 3D multi-object tracking under overlapping multi-camera fields of view. The authors propose MV3DT, the first distributed, real-time multi-view 3D tracking framework that operates without a central node or scene-specific training. MV3DT enables peer-to-peer collaboration among camera nodes through a lightweight, modular pipeline comprising monocular 3D perception, distributed data association, and message-passing-based cooperative fusion, thereby supporting identity propagation and occlusion recovery. Evaluated on WILDTRACK, the method achieves 94.3% IDF1 and 93.3% MOTA, scales to 100 cameras at 30 FPS, maintains inter-camera latency below 10 ms, and incurs only 2.2% communication overhead.
📝 Abstract
Multi-camera tracking with overlapping fields of view typically relies on centralized fusion, which creates computational bottlenecks that prevent deployment at scale. We present MV3DT, a fully distributed framework for real-time multi-view 3D tracking that achieves accurate identity propagation and occlusion recovery through peer-to-peer coordination, eliminating the need for central aggregation. Each camera node executes a lightweight modular pipeline comprising monocular 3D perception, distributed multi-view association, and collaborative fusion via lightweight messaging. MV3DT achieves 94.3% IDF1 and 93.3% MOTA on WILDTRACK, competitive with state-of-the-art centralized methods, while demonstrating superior scalability by sustaining 30 FPS on 100 cameras with less than 10 ms inter-camera latency and only 2.2% communication overhead. MV3DT operates in a zero-shot regime given camera calibrations, requiring no scene-specific learning and making it directly deployable in new environments. These results establish MV3DT as a practical solution for real-time multi-view tracking in large-scale overlapping camera networks.