Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of hand-object interaction details and object consistency degradation caused by drastic viewpoint changes in third-person to first-person video translation. To this end, we propose Exo2EgoHOI, a novel framework that introduces the first unified 4D HOI prior integrating geometry, hand rendering, and relational fields, alongside a dual-branch residual adapter designed to inject structural and relational cues. Furthermore, a factorized gated cross-attention mechanism is developed to synergistically enhance object-centric anchoring and multi-scale feature fusion. Experimental results on the ARCTIC-HOI dataset demonstrate that our framework improves object mIoU by 32.3% while reducing MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively. These gains significantly outperform existing baselines, establishing Exo2EgoHOI as a high-fidelity, scalable data source for embodied intelligence.
📝 Abstract
Egocentric videos of human manipulation provide valuable visual experience for embodied intelligence, yet collecting such data at scale is costly. Exocentric-to-egocentric video generation offers a scalable alternative by transforming abundant third-person manipulation videos into first-person observations. However, existing methods often struggle to faithfully preserve demonstrated hand-object interactions (HOI) across large viewpoint changes due to insufficient fine-grained interaction guidance and weak object-centric anchoring. We present Exo2EgoHOI, an HOI-aware video generative framework for interaction-preserving exocentric-to-egocentric translation. To preserve fine-grained HOI, we introduce a unified 4D HOI prior that combines scene geometry, articulated hand renderings, and dense hand-object relation fields, together with a dual-branch residual adapter for injecting structural and relational cues into the video generation backbone. To preserve object consistency, we introduce Decomposed Gated Cross-Attention, which separately encodes object and background references and adaptively integrates global semantic and local appearance features as object-centric anchors. Experiments on ARCTIC-HOI and Ego-Exo4D demonstrate substantial improvements in object consistency and HOI preservation while maintaining competitive visual fidelity. In particular, on ARCTIC-HOI, Exo2EgoHOI improves object mIoU by 32.3% and reduces MPJPE and PA-MPJPE by 34.7% and 50.0%, respectively, relative to the respective best baseline results. Project page: https://rcl-robotics.github.io/Exo2EgoHOI/.
Problem

Research questions and friction points this paper is trying to address.

Exocentric-to-Egocentric Video Generation
Hand-Object Interaction
Viewpoint Change
Object Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Exocentric-to-Egocentric Video Generation
Hand-Object Interaction
4D HOI Prior
Dual-Branch Residual Adapter
Decomposed Gated Cross-Attention