🤖 AI Summary
This study addresses the bottlenecks of low data collection efficiency and the difficulty of focusing on failure states in robotic imitation learning under fixed budgets. To overcome these challenges, this work proposes a performance-guided data collection strategy that treats the initial state distribution as a decision variable to enable targeted sampling of failure cases. Furthermore, it integrates interactive imitation learning with Deep Q-Learning (IDQL) to train a value function, combining value-based action selection with human-robot collaborative supervision to optimize deployment. Experimental results demonstrate that the proposed approach significantly outperforms uniform sampling baselines across both real-world and simulated tasks, improving final success rates by 10 to 34 percentage points. Notably, with operator intervention, the task completion rate reaches 98%, thereby achieving highly efficient data utilization for robotic policy learning.
📝 Abstract
Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in a small subset of initial states, so uniform collection spends much of the operator's time on states the policy already handles. Mulligan makes the initial-state distribution a decision, starting each round's episodes at observed failures and untried states. To further improve data efficiency, we augment interactive imitation learning with a value function trained on all data, including failures that imitation discards. Across three real-world tasks evaluated on 2,550 held-out, blinded episodes and two simulated tasks, Mulligan outperforms uniform initial-state sampling at matched collection budgets, and combined with value-based action selection, HiL-IDQL+Mulligan, improves final real-task success by 10-34 percentage points. With operator interventions, the human-robot team completes 98% of collection episodes, remaining productive while the policy learns. Videos, code, and data are available at https://mulligan.page/.