X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the exploration bottleneck in reinforcement learning for manipulating multi-degree-of-freedom objects by proposing a policy training framework guided by demonstration-informed reset distributions. Rather than directly imitating actions, the method leverages human hand-object interaction demonstrations to construct reset distributions through kinematic retargeting and state filtering, thereby driving efficient exploration under a task-agnostic generic reward. The proposed approach successfully trains universal manipulation policies across three heterogeneous embodiments, achieving zero-shot cross-embodiment transfer and zero-shot sim-to-real deployment. By decoupling exploration guidance from action-level imitation, this work provides a highly generalizable solution for complex dexterous manipulation tasks.
📝 Abstract
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Exploration Problem
Dexterous Manipulation
Cross-Embodiment
Generalist Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Embodiment Resets
Object-Centric Reinforcement Learning
Kinematic Retargeting
Human Hand-Object Demonstrations
Sim-to-Real Transfer
🔎 Similar Papers
No similar papers found.