Promptable Animal Pose Tracking Across Species

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Animal pose estimation faces significant challenges due to large morphological variations across species and the scarcity of annotated data, which hinders the simultaneous achievement of high generalization and precision. This work proposes a dual-path approach leveraging vision foundation models: one path injects structural priors through a supervised keypoint prompt encoder, while the other enables unsupervised cross-species feature matching to support efficient, user-defined keypoint tracking in videos. By uniquely integrating vision foundation models with a promptable keypoint mechanism, the method achieves high-accuracy and highly generalizable cross-species pose tracking with minimal annotations. It attains state-of-the-art performance on the APTv2 and TigDog benchmarks without any additional training, offering a practical tool for wildlife behavior analysis and conservation.
📝 Abstract
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data. Existing approaches either optimise generic keypoint localisation from annotated datasets (such as APTv2) with poor generalisation, or track custom keypoints using visual tracking, at the cost of performance. In this paper, we demonstrate that vision foundation models trained on large datasets can be used effectively to track animal pose with limited labelled data. We propose two models, one unsupervised and the other supervised, to track user-selected keypoints in videos. The supervised approach delivers superior tracking accuracy by employing a keypoint prompt encoder to explicitly inject structural priors from a reference frame into feature matching. In parallel, the unsupervised route provides strong cross-species robustness by leveraging diverse foundation-model features for training-free correspondence matching. Extensive evaluation on challenging animal video benchmarks APTv2 and TigDog demonstrates that our framework achieves strong performance while maintaining an effective balance between accuracy and generalisation, offering a practical solution for real-world animal behaviour analysis and conservation applications.
Problem

Research questions and friction points this paper is trying to address.

animal pose estimation
cross-species generalization
limited annotated data
keypoint tracking
wildlife monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision foundation models
keypoint prompt encoder
cross-species pose tracking
unsupervised correspondence
animal behavior analysis
🔎 Similar Papers
No similar papers found.