Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges in reinforcement learning for dexterous grasping, where coordinating interception with impact mitigation is difficult and teacher demonstrations are often imperfect. We propose an outcome-sensitive motion search method for constructing demonstration manifolds. By introducing the concept of an outcome-sensitive window, our approach integrates local geodesic search to optimize successful trajectories and rectify failure cases. Furthermore, a calibrated action error model is employed to validate action quality, thereby overcoming demonstration data bottlenecks. Experimental results demonstrate that the proposed method effectively mitigates the limitations of teacher policies, enabling the resulting policy to surpass privileged reinforcement learning teachers in both grasping success rate and impact mitigation performance.
📝 Abstract
Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
Problem

Research questions and friction points this paper is trying to address.

dexterous catching
impact mitigation
reinforcement learning
imitation learning
outcome sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Outcome-Sensitive Motion Search
Impact-Aware Dexterous Catching
Imitation Learning
Interventional Outcome Sensitivity
Reinforcement Learning
💼 Related Jobs
No related jobs found.
G
Guorui Pei
College of Robotics Science and Engineering, Taiyuan University of Technology, Taiyuan, China
Jinsong Wu
Jinsong Wu
University of Chile, Chile
green technologiesdata-driven sustainabilitysustainable engineeringbig dataInternet of things
S
Songyuan Su
Southern University of Science and Technology, Shenzhen, China
Jiaming Qi
Jiaming Qi
Northeast Forestry University
RobotShape deformationModel-free adaptive control
S
Sichao Liu
Department of Production Engineering, KTH Royal Institute of Technology, Stockholm, Sweden
David Navarro-Alarcon
David Navarro-Alarcon
The Hong Kong Polytechnic University
Robotics
B
Bin Liu
Kaiyang Laboratory, Chery Automobile Co., Ltd., Wuhu, Anhui, China
P
Peng Zhou
School of Advanced Engineering, Great Bay University, Dongguan, Guangdong, China