🤖 AI Summary
This study addresses the limited generalization and high demonstration costs of imitation learning, as well as the policy retraining requirements in conventional human-in-the-loop reinforcement learning. We propose a residual correction framework built upon a frozen imitation policy. By superimposing residual actions onto a pretrained policy, our approach integrates dual signals—direct supervision from human interventions and reward shaping—while introducing a zero-initialization mechanism to ensure stability during online learning. Requiring only 20 demonstrations and 10 minutes of training, the proposed method outperforms state-of-the-art baselines and pure imitation policies trained with five times more data across five contact-rich dexterous manipulation tasks, substantially reducing reliance on initial demonstration data.
📝 Abstract
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.