🤖 AI Summary
This study addresses the challenges of non-rigid deformation, dense occlusion, and distractors in generic multi-object tracking (GMOT) by proposing a unified architecture. Methodologically, it integrates exemplar-conditioned detection with an instance propagation head to achieve precise pixel-level segmentation and tracking. A novel training-free energy minimization algorithm is introduced to resolve conflicts among overlapping proposals, alongside a hierarchical memory management protocol designed to enhance long-term association capabilities. Experimental results demonstrate that the proposed method establishes new state-of-the-art performance on GMOT benchmarks and video counting tasks, achieving accuracy comparable to that of specialized, expert-level MOT models.
📝 Abstract
General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.