🤖 AI Summary
This study addresses the limited generalizability and low geometric accuracy in monocular 4D animal reconstruction caused by morphological diversity and scarce supervision. We propose a training-free, cross-species 4D reconstruction framework that innovatively decouples pose and shape estimation. Specifically, our method integrates the SMAL+ parametric model with DINO-based self-supervised correspondence matching and leverages generative 3D priors to refine geometric details, while recovering camera poses to ensure global motion consistency. Furthermore, we introduce PAW4D, a new benchmark dataset for evaluation. Experimental results demonstrate that, without requiring species-specific training, the proposed framework achieves high-fidelity, globally consistent 4D reconstructions of diverse quadrupeds from in-the-wild videos.
📝 Abstract
Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.