🤖 AI Summary
Current approaches to skill evolution in large language model agents are often limited to local updates, neglecting inter-skill relationships and consequently suffering from overfitting and poor generalization. This work proposes a Global Skill Evolution (GSE) framework that models skill dependencies through a Skill Relationship Graph (SRG), jointly optimizing skill compatibility and generalization. GSE further incorporates a clustering-driven skill integration mechanism and a replay-based validation strategy to enable the continual evolution of reusable, encoded skills. Evaluated on test generation and false positive filtering tasks, GSE achieves up to 34.1% higher precision and 180.0% higher recall compared to existing methods, with an industrial deployment yielding a 61.4% improvement in F1-score.
📝 Abstract
Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evolution as a sequence of local updates, overlooking relationships among skills and often producing overfitted skill updates that fail to generalize across tasks. We propose GSE, a globalized skill evolution framework that jointly optimizes skill compatibility and skill generalization. To preserve consistency across the skill bank, GSE maintains a Skill Relation Graph (SRG) that explicitly models and co-evolves inter-skill relationships. To improve generalization, GSE performs cluster-based skill consolidation to abstract reusable capabilities from local updates and employs replay-driven verification to prevent overfitting and behavioral regressions. We evaluate GSE on two representative software engineering tasks: bug-revealing test generation and false-positive bug report filtering. Across two state-of-the-art coding agents, OpenHands and mini-SWE-agent, GSE consistently achieves the best precision, recall, and F1-score. Compared with existing evolution techniques, GSE improves precision and recall by 6.1%~34.1% and 31.8%~180.0% for test generation, and by 15.4%~96.4% and 13.1%~19.8% for false-positive filtering. Deployment on an internal industrial agent further yields a 61.4% improvement in F1-score, demonstrating the effectiveness and generalizability of GSE for evolving effective skills.