Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of accurate hand motion prediction from egocentric views, which is hindered by limited field of view, high dynamics, and scarcity of robotic data. To this end, we propose a vision–language guided cross-view 3D hand pose prediction framework that leverages globally stable third-person (exo) demonstration videos to supervise first-person (ego) hand pose estimation—a novel paradigm in this domain. Our approach introduces a dual-level exo reconstruction module (DERM) and a global-to-local modulation mechanism (GLMM), which jointly fuse visual, linguistic, and pose cues through attention and adaptive modulation to achieve cross-view, cross-modal alignment between human and robot actions. Extensive experiments demonstrate significant performance gains over state-of-the-art methods on AssemblyHands, Ego-Exo4D, and our newly curated EgoMe-pose benchmark, while also showing effective human-to-robot transfer capabilities on the CALVIN platform.
📝 Abstract
Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulation. However, previous studies either rely mainly on under-informed visual inputs to predict coarse human motions or follow the VRM/VLA paradigm, which suffers from insufficient robot data and the gap between human and robot embodiments. We observe that 3D hand pose naturally serves as a unified representation to bridge human-robot actions. Hence, we investigate an under-explored Vision-Language guided Egocentric 3D Hand Pose Forecasting (VL-EHPF) task, which aims to predict future Ego 3D hand poses from visual observations, a language instruction, and pose states. To overcome the limited field-of-view and highly dynamic motions in the Ego view, we propose a framework dubbed Exo2EgoPose, which innovatively leverages holistic and stable exocentric (Exo) demonstrations as guidance to compensate for partial and dynamic Ego-view cues. Specifically, we introduce a Dual-level Exocentric Reconstruction Module (DERM), which incorporates the paired Exo videos as supervision to reconstruct their video-level and chunked frame-level representations, thereby modeling spatial contexts and temporal dynamics. Then, the Global-to-Local Modulation Module (GLMM) utilizes the reconstructed hierarchical Exo representations for progressive feature refinement via attention mechanisms and adaptive modulation, enabling comprehensive Exo guidance for accurate Ego hand pose forecasting. Extensive experiments on \textit{AssemblyHands}, \textit{Ego-Exo4D}, and our newly constructed \textit{EgoMe-pose} benchmarks show the superiority of our method, which outperforms state-of-the-art methods by a large margin. Moreover, it demonstrates an effective human-to-robot transfer capability and yields improvements on the \textit{CALVIN} dataset. Code will be released.
Problem

Research questions and friction points this paper is trying to address.

Egocentric 3D Hand Pose Forecasting
Vision-Language Guidance
Exocentric Demonstrations
Human-Robot Action Transfer
Fine-grained Action Prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Exocentric-to-Egocentric
3D Hand Pose Forecasting
Vision-Language Guidance
Dual-level Reconstruction
Human-Robot Transfer
🔎 Similar Papers
No similar papers found.