Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of static appearance descriptors in traditional multi-object tracking, which lack temporal modeling and often lead to identity switches. The authors reformulate appearance modeling as a recursive estimation problem, maintaining and continuously updating an independent appearance state for each target to predict future representations by accumulating observations. The proposed object-centric recursive appearance state estimation mechanism, integrated within a tracking-by-detection framework and paired with a tailored learning strategy, enables efficient online appearance refinement and generalizes well to unseen categories. Evaluated on KITTI, Waymo, and MOT17, the method significantly reduces ID switches; when incorporated into RobMOT, it achieves a state-of-the-art MOTA of 92.27% on KITTI while running at 90.57 Hz.
📝 Abstract
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. However, these descriptors are frame-independent, limiting their robustness as visual cues. Since such descriptors are often obtained from computationally intensive pretrained backbones, real-time MOT systems frequently abandon appearance cues altogether and rely solely on motion prediction and geometric association. In this work, we introduce Polycepta, an object-centric appearance state estimation framework that reformulates appearance modeling as a recursive estimation problem rather than a frame-wise matching task. Polycepta constructs and continuously updates an independent appearance state for each tracked object, enabling future appearance representations to be estimated from accumulated observations. Polycepta is encouraged to learn the appearance-state construction of object-specific representations rather than memorize them through a proposed learning strategy, enabling appearance estimation for unseen classes. A key property of Polycepta is that the quality of appearance estimation improves as object states evolve during inference. While conventional appearance descriptors remain static or degrade over time, Polycepta progressively refines appearance estimates as additional observations are accumulated. Extensive experiments on KITTI, the Waymo Open Dataset, and MOT17 demonstrate consistent reductions in identity switches and improvements in tracking performance when integrated into the tracking-by-detection pipelines. Polycepta operates at 90.57 Hz and delivers state-of-the-art performance on the KITTI benchmark when integrated into the RobMOT framework, achieving a MOTA of 92.27\%.
Problem

Research questions and friction points this paper is trying to address.

multi-object tracking
appearance modeling
tracking-by-detection
identity switches
real-time MOT
Innovation

Methods, ideas, or system contributions that make the work stand out.

object-centric
appearance estimation
recursive state estimation
multi-object tracking
tracking-by-detection
🔎 Similar Papers
No similar papers found.