🤖 AI Summary
This study addresses the error accumulation, pipeline complexity, and high computational overhead inherent in existing cascaded methods for recovering world-space hand motion from egocentric videos. To overcome these limitations, this work proposes an end-to-end streaming feedforward framework that jointly estimates MANO parameters, camera trajectories, and hand positions. By integrating persistent spatiotemporal memory with hand-centric features to unify camera motion and local geometry, and employing a two-stage progressive training strategy to mitigate drift, the approach achieves real-time inference without requiring an independent SLAM module, leveraging pretraining on 5,000 hours of multi-source data. Evaluated on the ARCTIC benchmark, the proposed method reduces PA-p error by 21.4% and attains 11.19 FPS—more than twice the speed of HaMoR—while demonstrating significant generalization capability in unconstrained in-the-wild scenarios.
📝 Abstract
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhead. To address these limitations, we present InfiniHand, an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video. InfiniHand integrates persistent spatiotemporal memory with hand-centered visual features, explicitly coupling camera motion with local hand geometry within a unified architecture. We train InfiniHand in two progressive stages by first learning robust camera-space hand priors and then extending to streaming world-space reconstruction. To support this process, we aggregate a pretraining corpus of approximately 5,000 hours of egocentric data across multiple public datasets. Extensive evaluations demonstrate that InfiniHand outperforms state-of-the-art baselines on in-domain benchmarks, achieving a 21.4% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift. Furthermore, InfiniHand generalizes robustly to in-the-wild videos and operates at 11.19 FPS, delivering more than twice the throughput of HaWoR.