🤖 AI Summary
This study addresses two critical challenges: low-quality demonstrations by novice teachers and robots’ limited skill acquisition and generalization capabilities. We propose a novel human-robot collaborative teaching framework that synergistically integrates Machine Teaching (MT) with Reinforcement Learning from Demonstrations (RLfD). To our knowledge, this is the first work to apply MT systematically to enhance human teaching proficiency—leveraging quantifiable feedback to refine demonstration quality—enabling task mastery with only eight high-fidelity demonstrations and effective zero-shot transfer to unseen skills. The framework comprises three core components: an MT-driven teacher capability enhancement module, a lightweight RLfD policy learning module, and a multi-dimensional human-robot co-evaluation framework. Experimental results demonstrate an 89% improvement in robot performance on trained tasks and a 70% gain in generalization to novel tasks, significantly overcoming the bottlenecks of few-shot skill learning and cross-task transfer.
📝 Abstract
Learning from demonstration (LfD) is a technique that allows expert teachers to teach task-oriented skills to robotic systems. However, the most effective way of guiding novice teachers to approach expert-level demonstrations quantitatively for specific teaching tasks remains an open question. To this end, this paper investigates the use of machine teaching (MT) to guide novice teachers to improve their teaching skills based on reinforcement learning from demonstration (RLfD). The paper reports an experiment in which novices receive MT-derived guidance to train their ability to teach a given motor skill with only 8 demonstrations and generalise this to previously unseen ones. Results indicate that the MT-guidance not only enhances robot learning performance by 89% on the training skill but also causes a 70% improvement in robot learning performance on skills not seen by subjects during training. These findings highlight the effectiveness of MT-guidance in upskilling human teaching behaviours, ultimately improving demonstration quality in RLfD.