EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of reconstructing hands in world coordinates from egocentric videos, where camera motion and occlusions hinder accuracy, multi-model inference remains inefficient, and interaction attributes are typically absent. To overcome these limitations, this work proposes a unified end-to-end framework that jointly estimates metric-scale world-space hand geometry and pointwise interaction attributes, including visibility, contact, and distance. Structured hand prompts are leveraged to transfer geometric priors, while an explicit representation is designed to decouple geometry from attributes. Furthermore, a shared-parameter multi-rate architecture is adopted to enhance computational efficiency. Extensive evaluations on multiple benchmarks demonstrate that the proposed method reduces MPJPE by up to 43.2%, predicts interaction attributes within a single forward pass, and achieves approximately a sixfold increase in throughput.
📝 Abstract
Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.
Problem

Research questions and friction points this paper is trying to address.

egocentric video
hand reconstruction
world coordinates
interaction attributes
throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Hand Reconstruction
End-to-End Model
Point-Wise Interaction Attributes
Structured Hand Prompts
Shared-Parameter Multi-Rate Design
🔎 Similar Papers