🤖 AI Summary
This study addresses the issue of subtask behavioral confusion in existing methods that rely on a single description to guide multi-stage robotic manipulation. To overcome this limitation, we propose an instruction-decomposition-based, multi-granularity language-guided imitation learning framework. Specifically, our method leverages natural language processing techniques to decompose global instructions into fine-grained subtask instructions and introduces a novel multi-granularity conditional policy modeling mechanism, enabling the model to receive precise semantic guidance across different execution stages. Experimental results demonstrate that the proposed framework significantly enhances behavioral disambiguation between stages and improves learning efficiency in multi-task imitation learning scenarios. Consequently, it effectively boosts robot performance on complex, multi-stage manipulation tasks.
📝 Abstract
Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed method decomposes an overall task description into more fine-grained, concrete subtask-level language instructions, thereby enhancing learning efficiency and improving performance. We evaluate the proposed method in the setting of multi-task imitation learning and validate its effectiveness.