🤖 AI Summary
This study addresses the heavy reliance of humanoid robot skill learning on successful demonstration data by proposing the TRACC framework, which for the first time converts failure videos into effective supervisory signals. The method extracts valid motion prefixes from failure videos for imitation and, based on inferred task goals, integrates imitation learning with reinforcement learning to complete the remaining actions guided by task-completion rewards. Experiments conducted on six real-world failure tasks from the Oops! dataset validate the effectiveness of this approach. The results demonstrate that skill learning can be achieved without successful trajectories, establishing failure videos as a viable source of supervision for humanoid robot policy optimization.
📝 Abstract
Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.