🤖 AI Summary
This study addresses the loss of hand-object interaction details and object consistency degradation caused by drastic viewpoint changes in third-person to first-person video translation. To this end, we propose Exo2EgoHOI, a novel framework that introduces the first unified 4D HOI prior integrating geometry, hand rendering, and relational fields, alongside a dual-branch residual adapter designed to inject structural and relational cues. Furthermore, a factorized gated cross-attention mechanism is developed to synergistically enhance object-centric anchoring and multi-scale feature fusion. Experimental results on the ARCTIC-HOI dataset demonstrate that our framework improves object mIoU by 32.3% while reducing MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively. These gains significantly outperform existing baselines, establishing Exo2EgoHOI as a high-fidelity, scalable data source for embodied intelligence.
📝 Abstract
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.