🤖 AI Summary
This study addresses the limitation of the Ego-Exo4D dataset, which provides only sparse 3D pose annotations insufficient for dense human motion reconstruction. To overcome this challenge, we propose a comprehensive pipeline that reconstructs dense human meshes from first-person and multi-view third-person videos. Our approach integrates synchronized multi-view video processing, 4D human motion reconstruction algorithms, and large-scale data generation techniques to establish a high-quality reconstruction framework. As our primary contribution, we release the Ego-Exo4D-HM dataset alongside its accompanying codebase, bridging the critical gap in dense geometric reconstruction for this setting. This work introduces the first large-scale 4D human mesh benchmark specifically designed for skill learning and embodied AI, substantially enhancing data availability and facilitating future research in these domains.
📝 Abstract
Ego-Exo4D is a large-scale dataset providing synchronized egocentric and multi-view exocentric video, a rich resource for skill learning and assessment, procedural activity understanding, and embodied AI. However, the dataset ships with only sparse 3D human pose annotations, and reconstructing dense human motion from its multi-view captures is nontrivial. To this end, we present Ego-Exo4D-HM, a large-scale dataset of 4D human motion reconstructions for Ego-Exo4D's captures, and release the accompanying reconstruction pipeline. The code, dataset, and documentation can be found at https://abhiram824.github.io/egoexo4d_human_meshes.