🤖 AI Summary
This study addresses the challenge in cross-embodiment imitation learning where human demonstration actions are frequently infeasible due to robot physical constraints, leading to degraded policy performance. To overcome this, we propose the EF-GAIfO framework, which integrates generative adversarial imitation learning with reinforcement learning. By leveraging the robot's own interaction experience, this approach dynamically evaluates state trajectory feasibility and expands the feasible region. Notably, the mechanism operates without requiring explicit dynamics models or large-scale exploration data, adapting progressively through policy iterations. We validate the effectiveness of our method on simulated locomotion and quadrupedal grasping tasks. Experimental results demonstrate that the proposed framework significantly enhances imitation learning performance in cross-embodiment scenarios, offering a practical solution for transferring human demonstrations to morphologically distinct robotic agents.
📝 Abstract
With the increasing use of robot-free demonstration interfaces that provide state trajectories without action labels, imitation from observation has become a promising approach for learning robot behaviors from human demonstrations. However, due to differences in embodiment and dynamics between humans and robots, demonstrated human motions may not be feasible for the robot, potentially degrading policy performance. In this study, we propose Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), which estimates the feasibility of state-only demonstrations from the robot's own experience rather than relying on explicit dynamics models or large prior exploration datasets. A key feature of EF-GAIfO is that the notion of feasibility evolves with policy learning: as the policy improves and the robot experiences a broader range of state transitions, the feasible region is progressively expanded, allowing additional demonstrations to be incorporated into learning. This enables feasibility-aware imitation that adapts to the current stage of policy learning, rather than relying on a pre-designed feasibility criterion. We validate the effectiveness of EF-GAIfO on a locomotion task in simulation and on a real quadruped robot performing a object-reaching-and-grasping task.