TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of severe occlusion, inter-individual similarity, and absent identity annotations in multi-view 3D fish tracking by proposing a geometry-driven self-supervised framework. Methodologically, pseudo-association labels are generated via triangulation and reprojection consistency to enable unsupervised training. Contrastive learning is introduced to disambiguate co-visible individuals, while a temporal prediction network maintains identity continuity. The core architecture integrates a geometric encoder with a global association Transformer. Experimental results demonstrate that the proposed method achieves a MOTA of 95.8%, significantly outperforming baselines, and attains 81.1% on a zebrafish dataset. Furthermore, the framework successfully generalizes to bird tracking tasks, underscoring its robustness and broad applicability across species.
📝 Abstract
Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 87.7% for the strongest baseline to 95.8%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking.
Problem

Research questions and friction points this paper is trying to address.

3D multi-object tracking
schooling fish
occlusion
identity annotation
multi-view video
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised 3D tracking
Multi-view geometry
Global association transformer
Contrastive learning
Cross-view correspondence
💼 Related Jobs
No related jobs found.
P
Patt Phurtivilai
The University of Hong Kong, Hong Kong, China
Z
Zhiyang Dou
The University of Hong Kong, Hong Kong, China
Y
Yifan Wu
The University of Hong Kong, Hong Kong, China
K
Kinfung Chu
Centre for Transformative Garment Production
Y
Yuan Liu
Hong Kong University of Science and Technology, Hong Kong, China
L
Lei Yang
The University of Hong Kong, Hong Kong, China
Wenping Wang
Wenping Wang
Texas A&M University
Computer GraphicsGeometric Computing
Taku Komura
Taku Komura
The University of Hong Kong
Character AnimationComputer GraphicsRobotics