🤖 AI Summary
This study addresses the scarcity and synthesis difficulty of training data for skill invocation in large language model (LLM) agents by proposing an automated data synthesis and training framework. The framework introduces a novel method for automatically constructing offline skill environments with controllable difficulty levels and built-in execution verifiers. By integrating web crawling with a builder-reviewer pipeline, it enables large-scale synthesis of high-quality trajectories, followed by agent training via supervised fine-tuning (SFT) and automated verification mechanisms. Experimentally, the approach constructs 6.8k environments and 19k trajectories, improving the skill reading rate of the Qwen3.5-9B model from 28% to 96%. Notably, the resulting trained model surpasses the performance of an untrained 397B-parameter baseline.
📝 Abstract
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training