🤖 AI Summary
This work addresses the limitations of existing agent skill evolution methods that rely on execution trajectories, which often overlook implicit task requirements and tightly couple revision with skill behaviors. We propose a novel skill evolution framework grounded in a fixed capability space, which decomposes recurring task demands into discrete capabilities and precisely identifies the weakest unresolved capability via task outcome mapping for targeted optimization, thereby decoupling task requirements from skill behaviors. By integrating large language models, capability decomposition, iterative revision, and priority evidence matching strategies, our approach achieves state-of-the-art accuracy across four benchmarks. It outperforms the strongest baseline by an average of 5.7 points while reducing evolution token consumption by 24%, demonstrating the effectiveness of capability-driven skill evolution.
📝 Abstract
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24\% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.