Score
Designs and implements algorithms and systems that autonomously identify, create, and organize reusable skills or options from ongoing experience or data; this includes methods to detect coordination primitives, decide when to expand or partition the skill set as new tasks arrive, and consolidate or prune the library. Engineers competence in managing interference and transfer between tasks by structuring, merging, or isolating skills to support continual learning and reuse.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
This work addresses catastrophic forgetting and feature drift in embodied continual learning caused by closed-loop control. To tackle these challenges, the authors propose the Skill-Compositional Experts (SCE) framework, which leverages Compositional Skill Grounding (CSG) to construct a reusable and explicitly structured skill library from task demonstrations. SCE further introduces a dual-mechanism architecture—comprising Execution Experts and Transfer Experts (DETE)—to enable efficient compositional learning of new tasks and smooth transitions between skills. Evaluated on the LIBERO benchmark and real-world robotic platforms, SCE significantly improves task retention and overall performance. Ablation studies and feature drift analyses confirm its effectiveness, marking the first approach to achieve structured, composable embodied continual learning.
This work addresses the critical bottleneck in cross-system multi-agent collaboration caused by the lack of portability and self-optimization capabilities in existing coordination protocols. To overcome this, the authors propose Swarm Skills—a portable semantic specification for multi-agent systems that encapsulates collaborative workflows as first-class assets, embedding roles, procedural logic, execution boundaries, and self-evolution semantics. A multidimensional scoring mechanism based on validity, utilization, and novelty drives the autonomous distillation and updating of strategies. This approach achieves, for the first time, fully automated self-evolution of collaboration policies without human intervention and enables zero-adapter cross-framework portability through progressive disclosure. Experimental results demonstrate the feasibility of deploying Swarm Skills across diverse systems and its capacity for continuous autonomous optimization.
This study addresses the current lack of systematic understanding of reusable agent skills in software engineering, particularly regarding their coverage across the software lifecycle. For the first time, it adopts an activity-oriented perspective and conducts a large-scale empirical investigation to collect, categorize, and model software engineering skills from public skill repositories. The work systematically characterizes the types of encapsulated activities, their evolutionary patterns, and evaluation mechanisms. Findings reveal that engineering activities with high contextual dependency are progressively being transformed into reusable skills. Building on these insights, the paper outlines promising future directions, including skill recommendation, structured organization of skill repositories, and enhanced encapsulation strategies for high-context skills, thereby providing both theoretical foundations and practical guidance for agent-driven software engineering.
This study addresses the lack of systematic understanding regarding the creation, reuse, customization, and maintenance of AI agent skills as reusable software artifacts. Treating AI skills as engineered software artifacts for the first time, the authors conduct an empirical investigation based on over 40,000 skill instances drawn from public registries and GitHub repositories, employing a mixed-methods approach that combines large-scale data mining, LLM-driven SWEBOK-based knowledge classification, topic modeling, and qualitative coding. The findings reveal that 53% of reused skills remain unmodified, with reuse predominantly involving one-time copying; customization primarily serves to adapt skills to local environments; and evolution tends to follow an incremental addition pattern while preserving highly stable behavioral contracts. The study identifies six content categories of skills and six modification themes, establishing an empirical foundation for the engineering-oriented management of AI skills.
This study addresses the unclear adaptation patterns and impacts of large language model (LLM) agent skills when reused in downstream applications. Through an empirical analysis of 1,126 adaptation instances from six prominent skill repositories, the work systematically characterizes LLM skill adaptation behaviors and constructs a taxonomy comprising 46 patterns grouped into 13 families. The research uncovers critical phenomena including a “reuse paradox,” strong cross-component dependencies, and the introduction of security-sensitive content in nearly one-fifth of adaptations. It further identifies prevalent challenges such as logic rewriting, fixing discoverability issues, and cross-tool or cross-language translation. These findings offer new empirical insights and foundational support for improving skill design, standardizing interfaces, and enabling automated adaptation of LLM-based agents.
Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns shift over deployment. Recent work seeks to automate skill curation, but it largely evaluates against automated baselines and treats human maintenance as an unmeasured bottleneck. We study that missing process directly. We mine the full commit histories of five public AI-skill repositories, a purposive sample of AI-tooling organizations, covering 873 commits, 143 skill files, and 254 substantive post-creation edits from October 2025 to June 2026. We code each edit with pre-registered governance, operation, and trigger-evidence codebooks. Three findings emerge. First, every substantive edit is authored or merged through a named human account, while 62% carry an AI co-author trailer, with large repository-level variation. Second, these edits are genuine curation: an audited sample shows that most change skill content, and the coded operations are dominated by additions and corrections. Third, a pre-registered rule-likeness axis fails its reliability gate; reliably coding rule-likeness from commit artifacts remains an open measurement problem. We release the corpus, codebooks, mining scripts, and a replay protocol for automated skill curators. For self-evolving agents, public skill maintenance currently looks less like an autonomous pipeline than a human-governed, AI-assisted loop that future curators must measure against and operate within.
This work addresses the lack of a unified engineering framework for modeling and managing AI agent skills as independent software artifacts. It introduces the Skillware abstraction, formally defining Skill Artifacts and Skillware Units to represent skills as persistent behavioral components with distinct identities, continuous lifecycles, and explicit execution relationships. Grounded in three core tenets—behavioral primacy, identity independence, and execution relationality—the approach integrates ontology modeling, task specification, metadata management, and version control to establish a reusable and composable skill engineering framework. Empirical validation through analysis of 138,133 deduplicated SKILL.md files and 20,556 repositories demonstrates the method’s effectiveness in enhancing encapsulation, identity resolution, execution compatibility, and evolutionary maintainability.
This study addresses the challenge of filtering erroneous and non-transferable knowledge from skill libraries of self-evolving agents by proposing a Proposer-Builder-Verifier architecture to validate skill reusability in unseen tasks. The method dynamically synthesizes test scenarios through conditional constraint generation, overcoming the limitations of traditional static evaluation. Furthermore, it employs an automated execution comparison mechanism to quantitatively assess skill utility and efficiency, enabling reliable retention or rejection decisions. Experimental results demonstrate that the proposed framework significantly improves downstream task performance and execution efficiency on the ALFWorld and WebShop benchmarks, while also confirming the accuracy of its skill reusability assessment.
研究通过在线技能进化框架,将交互轨迹和评估反馈转化为持久可复用的程序库,以解决计算机使用代理在图形界面中执行复杂任务时经验无法积累的问题。
This work addresses the inadequacy of traditional engineering education in meeting the urgent demand for a new generation of engineers equipped for the era of generative and autonomous intelligence. It proposes the ACCEL educational framework, which introduces a novel competency model for “autonomous intelligence engineers.” The framework systematically cultivates core capabilities—including intent articulation, multi-agent orchestration, output evaluation, and ethical judgment—through three integrated pathways: curriculum restructuring, collaborative mechanisms, and lifelong learning. Grounded in principal–agent theory, research on trust in automation, and international AI competency standards, ACCEL incorporates a commission–verification instructional cycle, governance-oriented ethics education, and an innovative assessment system. It identifies critical risks such as automation bias and skill atrophy, and advances engineering education beyond incremental reform by shifting its focus from artifact-centric outcomes to cultivating judgment over autonomous sociotechnical systems.