CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking

πŸ“… 2026-06-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses critical challenges in panoramic multi-object tracking, where the periodic nature of equirectangular projection violates the planar motion assumption, causing conventional IoU-based association to fail at the 0Β°/360Β° seam. Additionally, wide fields of view lead to dense object distributions, drastic scale variations, and unstable depth estimates, severely degrading online association performance. To overcome these issues, the authors propose CylindTrack, a novel framework that elevates depth modeling from the frame level to the trajectory level. It introduces spherical spatio-temporal consistency learning to enhance depth representation and designs a topology-aware cylindrical motion model to enable seam-consistent motion prediction and data association. This approach significantly improves identity preservation and trajectory continuity in complex panoramic scenes, effectively mitigating seam-induced association failures and depth instability.
πŸ“ Abstract
Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360Β° field of view reduces blind spots and keeps surrounding targets observable for longer durations. However, panoramic MOT is not a straightforward extension of perspective MOT. In equirectangular panoramic videos, the horizontal image domain is periodic rather than Euclidean, which breaks planar motion assumptions and makes IoU-based association unreliable near the 0Β°/360Β° seam. Meanwhile, large-FoV scenes often contain more objects, stronger scale variation, and more frequent interactions, making online association particularly sensitive to unstable frame-wise depth cues. To address these issues, we propose CylindTrack, a depth-aware cylindrical tracking-by-detection framework for panoramic MOT. CylindTrack first introduces Depth-Temporal Trajectory Modeling (DTM), which promotes instance depth from an isolated frame-wise cue to a temporally filtered trajectory-level state. To improve the reliability of depth observations, we further develop Spherical Spatio-Temporal Consistency Learning (SSTC), which combines a Temporal Mixer and Spherical Geometry-aware Attention to enhance temporal coherence and panoramic geometric alignment in depth-aware representations. Finally, we design a Topology-Aware Cylindrical Motion Model (TCMM) that lifts horizontal motion into a continuous angular state space and performs seam-consistent motion prediction and association in the periodic panoramic domain. By jointly modeling trajectory-level depth consistency and panoramic topology, CylindTrack improves identity preservation and trajectory continuity in challenging panoramic scenes. The source code will be released at https://github.com/warriordby/CylindTrack.
Problem

Research questions and friction points this paper is trying to address.

Panoramic Multi-Object Tracking
Periodic Domain
Depth Cues
Motion Modeling
Equirectangular Projection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-Temporal Trajectory Modeling
Spherical Spatio-Temporal Consistency
Topology-Aware Cylindrical Motion Model
Panoramic Multi-Object Tracking
Equirectangular Video
B
Buyin Deng
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China
K
Kai Luo
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China
L
Lingxin Huang
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China
X
Xinqi Liu
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China
F
Fei Cheng
School of Advanced Technology, Xi’an Jiaotong-Liverpool University, China
Hang Zheng
Hang Zheng
Zhejiang University
array signal processingDOA estimationbeamformingtensor signal processingmachine learning
L
Liming Yin
Suzhou VSDeep Intelligent Technology Co., Ltd., China
Kailun Yang
Kailun Yang
Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous DrivingRobotics