VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of large-scale, multi-class, and high-quality annotations in existing benchmarks that hinders generalizable multi-animal tracking (MAT) research. We construct an ultra-large-scale MAT benchmark encompassing 337 animal species and millions of frames, alongside an iterative expert-reviewed annotation pipeline and a Seen/Unseen dual-protocol evaluation mechanism. Methodologically, we propose a lightweight association module integrating adaptive IoU with normalized center similarity to optimize identity matching under low-overlap conditions, and introduce a training-free CDA post-processing algorithm. Experiments demonstrate that our approach establishes a state-of-the-art HOTA baseline of 66.37%, while CDA further improves TrackTrack performance by 1.58 percentage points, significantly enhancing cross-category generalization capability.
📝 Abstract
Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large-scale video data, and extensive within-video multi-instance association. To address this gap, we introduce VastMAT, which has four key characteristics: (1) Large scale. It comprises 2,947 videos with 1,002,562 annotated frames, totaling 27.85 hours. (2) Broad category coverage. These videos cover 337 animal categories with diverse morphologies and motion patterns. (3) Extensive instance annotations. It provides 3,663,248 bounding boxes and 22,883 identity trajectories---to our knowledge, the largest numbers of both among dedicated MAT benchmarks. (4) High-quality annotations. To ensure reliability, annotations undergo iterative expert review and correction, and quality is assessed through an independent reannotation audit. To systematically assess tracking performance and cross-category generalization, we establish Seen-category and category-disjoint Unseen-category protocols, and evaluate eight representative MOT methods under both protocols. Under these protocols, the highest baseline HOTA scores are 66.37\% and 52.90\%, respectively, highlighting the challenge of tracking unseen animals. To address the low-overlap association challenge revealed by our analysis, we propose Center-Distance-Augmented Association (CDA), a lightweight module that adaptively combines IoU with center similarity normalized by the boxes' own scales. Without additional training, CDA improves TrackTrack's HOTA by 1.58 and 1.31 percentage points under the two protocols, respectively. To facilitate further MAT research, we will publicly release our benchmark and code.
Problem

Research questions and friction points this paper is trying to address.

Multi-animal tracking
Benchmark
Cross-category generalization
Low-overlap association
Multi-object tracking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Animal Tracking
Benchmark
Cross-Category Generalization
Center-Distance-Augmented Association
Data Association
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhizhen Li
Zhizhen Li
NCSU CS student
Z
Zan Wang
University of North Texas
H
Huidong Peng
Wuhan University
B
Bohan Tan
The Hong Kong University of Science and Technology
S
Shimin Shan
Dalian University of Technology
Yu Liu
Yu Liu
Dalian University of Technology
computer visionmultimodal learning
L
Liang Peng
Wuhan University