🤖 AI Summary
This study addresses the limitations of existing large language model (LLM) skills, which typically rely on expert authoring, exhibit poor adaptability, and are difficult to optimize automatically for emerging tasks due to scarce training data. To overcome these challenges, this work proposes the first unsupervised framework that automatically generates and optimizes external LLM skills using only natural language instructions, without requiring manually constructed training sets. The framework achieves iterative skill refinement through task specification derivation, dataset synthesis, and reflection-based closed-loop editing. Experimental results demonstrate that the proposed approach consistently outperforms direct prompting baselines across four domains, including question answering and reading comprehension, yielding an average performance improvement of 10.8% for both open-source and frontier models.
📝 Abstract
Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.