Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of poor generalization in motion imitation and inefficient exploration in reinforcement learning when acquiring dexterous manipulation skills from single-segment human videos. To overcome these limitations, we propose a real-sim-real framework that abstracts video demonstrations into scene graphs. These scene graphs serve as generative constraints in place of strict pose matching, thereby guiding reinforcement learning sampling and constructing dense stage-wise rewards that balance diverse reset states with efficient exploration. This approach effectively resolves the bottlenecks of generalization and exploration. Experimental results demonstrate that our method surpasses baseline approaches across five tasks, achieving a 71% performance improvement in unseen scenarios and enabling zero-shot sim-to-real deployment for multi-fingered dexterous hands.
📝 Abstract
While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions. Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video. Alternatively, discovering a policy via reinforcement learning (RL) allows for broad generalization, but without prior guidance, it struggles with high-dimensional exploration in complex, multi-stage tasks. To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video. Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability. The graphs serve as generative constraints for sampling diverse reset states and provide dense rewards for each stage. Because the graphs constrain relations rather than exact poses, these reset states cover object poses and grasps beyond the video, while initializing each stage from them with dense rewards keeps exploration short and guided. Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand. Across five tool-use and manipulation tasks, Dex-One2Many exceeds baselines by 6.5% in seen configurations, while its robust generalization widens this gap to 71% in unseen scenarios.
Problem

Research questions and friction points this paper is trying to address.

Dexterous Manipulation
Single Human Demonstration
Generalization
Reinforcement Learning
Exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dexterous Manipulation
Reinforcement Learning
Scene Graphs
Sim-to-Real Transfer
Single Demonstration
🔎 Similar Papers
No similar papers found.