🤖 AI Summary
This study addresses the limitation of visual rewards in capturing local physical interactions during contact-rich tasks, which leads to inefficient and brittle reinforcement learning. To overcome this, we propose a tactile reward learning framework that leverages tactile deformation maps to regress task progress, extracting generalizable physical feedback signals from both successful and failed demonstrations. By integrating visuotactile multimodal rewards for policy optimization, our approach transcends the constraints of purely visual observations. Experiments demonstrate that the proposed framework significantly improves sample efficiency and generalization across scene layouts and object instances, achieving strong performance in both simulated and real-world tasks. Notably, success rates increase from 34% to 56% for nut threading and from 37% to 97% for cube grasping.
📝 Abstract
Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at https://embodiedai-ntu.github.io/tarl.