🤖 AI Summary
This work addresses the challenge of robustly reconstructing 3D human motion in real-world multi-person interactions, a task at which existing head-worn motion capture methods typically falter due to their reliance on either egocentric or exocentric viewpoints alone. The paper proposes a lightweight, distributed framework that, for the first time, unifies both perspectives within a single system, enabling global 3D motion reconstruction using only off-the-shelf smart glasses worn by multiple participants. By fusing head (and wrist) tracking signals with context-aware image features derived from DINOv3, the method substantially enhances robustness against occlusion and sensor noise. Experiments on two in-the-wild datasets demonstrate that the approach consistently achieves accurate and stable multi-person motion reconstruction even in complex, unconstrained environments.
📝 Abstract
Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer's surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.