🤖 AI Summary
This study addresses the low execution success rates encountered when transferring hand-object interactions reconstructed from monocular videos to dexterous robotic hands, a limitation primarily caused by insufficient physical consistency. To overcome this, we propose OmniHOI, a staged pipeline that progressively integrates visual, contact-geometric, and dynamic evidence to optimize representations, effectively converting RGB videos into high-fidelity robot trajectories while preventing error propagation without requiring task-specific reinforcement learning. The core contribution lies in unifying visual reconstruction, geometric retargeting, and physics-in-the-loop dynamic refinement. Extensive evaluations demonstrate that our approach achieves success rates of 39–89% across diverse dexterous hands and a 53% video-to-robot transfer success rate, significantly outperforming existing methods. Furthermore, its practical efficacy is validated on a real-world bimanual robotic system.
📝 Abstract
Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods. On 60 monocular video clips, we achieve 53% success, compared with 28% for the best prior video-to-robot pipeline. Its trajectories also execute on a real bimanual robot across diverse tasks.