🤖 AI Summary
This study addresses the exploration bottleneck in reinforcement learning for manipulating multi-degree-of-freedom objects by proposing a policy training framework guided by demonstration-informed reset distributions. Rather than directly imitating actions, the method leverages human hand-object interaction demonstrations to construct reset distributions through kinematic retargeting and state filtering, thereby driving efficient exploration under a task-agnostic generic reward. The proposed approach successfully trains universal manipulation policies across three heterogeneous embodiments, achieving zero-shot cross-embodiment transfer and zero-shot sim-to-real deployment. By decoupling exploration guidance from action-level imitation, this work provides a highly generalizable solution for complex dexterous manipulation tasks.
📝 Abstract
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.