DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent conflict between exploration noise and precise control in reinforcement learning for dexterous manipulation by proposing an explicit exploration scale scheduling strategy. By modeling exploration intensity as an annealing function of training steps, the method enables a smooth broad-to-narrow transition that decouples exploration from control, alongside a terminal task success-based evaluation criterion. Experiments integrating PPO, GRPO, and flow-parameterized FPO algorithms are conducted on the YCB dataset using a RealMan robotic arm and an Inspire dexterous hand. Results demonstrate that the proposed approach significantly improves simulation success rates. In real-world experiments, FPO increases grasping success from 25% to 85%, thoroughly validating the effectiveness of the exploration scheduling mechanism.
📝 Abstract
Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized PPO variant (FPO). Across five YCB objects and three training seeds, mean deterministic Target success rises from 49.4% to 68.1% (FPO), 14.1% to 45.4% (GRPO), and 32.0% to 35.7% (PPO). On a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials over three objects raise mean Target success from 25.0% to 85.0% (FPO), 10.0% to 63.3% (GRPO), and 8.3% to 43.3% (PPO), with one trained model per object-method condition. PPO component screening favors noise control over the tested optimizer contraction; the selected PPO schedule yields higher mean Target success than linear decay with the same endpoints on three tested objects. Training return, deterministic Target success, and tolerance to execution noise dissociate; schedules should therefore be judged by terminal task success under the intended execution conditions, per task and policy-optimization setting. Code: https://github.com/AIGeeksGroup/DexPolicy. Website: https://aigeeksgroup.github.io/DexPolicy.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
reinforcement learning
exploration noise
trajectory-guided control
action noise interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dexterous Manipulation
Scheduled Exploration
Reinforcement Learning
Flow-Parameterized Policy
Trajectory Guidance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoyu Wang
School of Computer Science, Peking University
S
Siyuan Qian
School of Computer Science, Peking University
Y
Yanjun Li
School of Computer Science, Peking University
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
Y
Yandong Guo
AI2 Robotics
Boxin Shi
Boxin Shi
Peking University
Computer VisionComputational Photography
Hao Tang
Hao Tang
Peking University
computer vision