🤖 AI Summary
This study addresses the data scarcity and generalization bottlenecks encountered by small models in robotic manipulation by proposing an LLM-SLM orchestration framework based on task decomposition and skill composition. The core methodology leverages large language models to drive data synthesis for training small language models, designs a skill-aware context-free grammar (CFG) to constrain the action space, and introduces a progressive orchestration strategy to ensure reliable task planning and natural language instruction parsing. Experimental evaluations conducted on both unmanned aerial vehicles and ground vehicles demonstrate that the proposed approach significantly outperforms existing baselines, substantially enhancing the model's zero-shot generalization capability to unseen tasks.
📝 Abstract
Small language models (SLMs) have been increasingly adopted for onboard robot operation because they enable intelligent decision-making. However, existing approaches are mainly distillation-oriented and rely on enumerating representative task-solution pairs. This makes dataset construction difficult and limits generalization to diverse robot tasks whose possible forms grow rapidly. This paper proposes Skill-SLM, a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem. Given a natural language task instruction, Skill-SLM decomposes the task into subtasks, selects appropriate skills from the skill library, and orchestrates the selected skills into executable robot operations. First, to support the skill-driven workflow, we propose a novel robot operational skill aware context-free grammar (CFG) to extract the skills required to accomplish tasks and build the skill library accordingly. Then, we configure LLM teachers to induce and synthesize training datasets for the SLMs, enabling SLMs to decompose tasks and orchestrate skills reliably. Additionally, we employ a progressive skill orchestration strategy to improve the reliability of skill implementation and overall robot operation. Experiments on UAV operation tasks indicate that Skill-SLM substantially outperforms distillation-oriented baselines, especially on unseen tasks that require generalization of capabilities. Additional experiments on ground vehicle tasks further demonstrate that Skill-SLM can be applied to different robot platforms.