Score
Designs and builds reusable skill artifacts and supporting infrastructure — skill libraries, templates, frameworks, indexing and standardization schemes, authoring tools, and integration/management systems — that encode and version discrete skills. Creates and analyzes skill development plans, programs, assessments, templates and governance needed to operationalize, integrate, measure, and manage skills for skills-based execution.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
Current large language model (LLM) agents face challenges in real-world deployment, including inefficiency, error-proneness, and poor maintainability, largely due to their reliance on on-the-fly reasoning and low-level tool invocation. This work introduces, for the first time, a skill-centric agent architecture that formalizes a comprehensive skill lifecycle framework encompassing representation, acquisition, retrieval, and evolution. It positions skills as a complementary mechanism bridging high-level reasoning and operational execution. By integrating key techniques—such as skill representation learning, automated acquisition, semantic retrieval, and continual evolution—and synergizing them with tool use, memory mechanisms, and contextual constraints, the proposed framework establishes a reusable and composable skill system. The paper further surveys representative approaches, open-source resources, and application scenarios, offering both theoretical foundations and practical guidance to enhance the scalability, robustness, and maintainability of intelligent agent systems.
Current agent capabilities lack large-scale infrastructure for production, governance, and evolution akin to Wikipedia and GitHub. This work proposes SkillWiki, the first dynamic knowledge infrastructure that enables symbiotic co-evolution of knowledge, skills, and execution experiences. By integrating knowledge extraction, skill assetization, provenance tracking, and an execution-driven feedback loop, SkillWiki establishes an end-to-end framework for the full lifecycle management of skills, supporting reusable organization, verifiable provenance, and continuous evolution. The system has been fully validated through a pipeline spanning knowledge ingestion, skill generation, and execution-driven refinement. Both the codebase and the system are publicly open-sourced.
This study addresses the unclear adaptation patterns and impacts of large language model (LLM) agent skills when reused in downstream applications. Through an empirical analysis of 1,126 adaptation instances from six prominent skill repositories, the work systematically characterizes LLM skill adaptation behaviors and constructs a taxonomy comprising 46 patterns grouped into 13 families. The research uncovers critical phenomena including a “reuse paradox,” strong cross-component dependencies, and the introduction of security-sensitive content in nearly one-fifth of adaptations. It further identifies prevalent challenges such as logic rewriting, fixing discoverability issues, and cross-tool or cross-language translation. These findings offer new empirical insights and foundational support for improving skill design, standardizing interfaces, and enabling automated adaptation of LLM-based agents.
Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns shift over deployment. Recent work seeks to automate skill curation, but it largely evaluates against automated baselines and treats human maintenance as an unmeasured bottleneck. We study that missing process directly. We mine the full commit histories of five public AI-skill repositories, a purposive sample of AI-tooling organizations, covering 873 commits, 143 skill files, and 254 substantive post-creation edits from October 2025 to June 2026. We code each edit with pre-registered governance, operation, and trigger-evidence codebooks. Three findings emerge. First, every substantive edit is authored or merged through a named human account, while 62% carry an AI co-author trailer, with large repository-level variation. Second, these edits are genuine curation: an audited sample shows that most change skill content, and the coded operations are dominated by additions and corrections. Third, a pre-registered rule-likeness axis fails its reliability gate; reliably coding rule-likeness from commit artifacts remains an open measurement problem. We release the corpus, codebooks, mining scripts, and a replay protocol for automated skill curators. For self-evolving agents, public skill maintenance currently looks less like an autonomous pipeline than a human-governed, AI-assisted loop that future curators must measure against and operate within.
This work addresses the inefficiency in generating and sharing new capabilities for AI agents, stemming from a lack of reusable skills during runtime. To overcome this, we propose a demand-driven, agent-centric skill production platform that introduces a novel “demand-first” paradigm for skill generation. The platform natively integrates full lifecycle skill management into Git workflows, enabling collaborative development, review, and version control among humans, scripts, and external agents within a unified state space. Leveraging mechanisms such as scoped push URLs, range-based commit ingestion, workflow state reading, and event tracing—combined with hosted repositories and registries—the system ensures auditability, recoverability, and multi-interface access (Web/REST/MCP) to skills. Empirical validation demonstrates end-to-end execution of an OS detection skill, conversion of Docker research bundles into reusable skills, and versioned submission of high-quality skill artifacts.
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
This study addresses the lack of systematic understanding regarding the creation, reuse, customization, and maintenance of AI agent skills as reusable software artifacts. Treating AI skills as engineered software artifacts for the first time, the authors conduct an empirical investigation based on over 40,000 skill instances drawn from public registries and GitHub repositories, employing a mixed-methods approach that combines large-scale data mining, LLM-driven SWEBOK-based knowledge classification, topic modeling, and qualitative coding. The findings reveal that 53% of reused skills remain unmodified, with reuse predominantly involving one-time copying; customization primarily serves to adapt skills to local environments; and evolution tends to follow an incremental addition pattern while preserving highly stable behavioral contracts. The study identifies six content categories of skills and six modification themes, establishing an empirical foundation for the engineering-oriented management of AI skills.
Current LLM agents lack explicit dependency management for their skills, leading to redundant dependencies, environmental inconsistencies, and security risks. This work proposes the Agent Skill Supply Chain (ASSC) framework, which introduces a systematic model of hybrid dependencies among skills, packages, and services. Inspired by Software Bill of Materials (SBOM), we design SkillDepAnalyzer—a tool that automatically extracts dependency evidence from natural language descriptions to construct skill dependency graphs. Experiments on the SKILL-DEP benchmark demonstrate that our approach significantly outperforms both LLM-based baselines and conventional SBOM tools. An analysis of 1.43 million skills reveals four distinct dependency patterns and enables effective identification and reporting of malicious skills, offering a novel paradigm for secure governance in skill ecosystems.
This study addresses the current lack of systematic understanding of reusable agent skills in software engineering, particularly regarding their coverage across the software lifecycle. For the first time, it adopts an activity-oriented perspective and conducts a large-scale empirical investigation to collect, categorize, and model software engineering skills from public skill repositories. The work systematically characterizes the types of encapsulated activities, their evolutionary patterns, and evaluation mechanisms. Findings reveal that engineering activities with high contextual dependency are progressively being transformed into reusable skills. Building on these insights, the paper outlines promising future directions, including skill recommendation, structured organization of skill repositories, and enhanced encapsulation strategies for high-context skills, thereby providing both theoretical foundations and practical guidance for agent-driven software engineering.
This study addresses the disconnect between agent skill relevance and task utility, along with the lack of mechanistic explanations, through an empirical investigation across 87 tasks. Methodologically, we define a downstream utility metric based on pass-rate differentials and employ both large language models and human review to analyze skill content, execution trajectories, and artifacts. Furthermore, we propose a re-ranking strategy supporting essential operations alongside a DAG-based dependency organization method. Results reveal that in 36.78% of tasks, identical skills exhibit opposite utility under different configurations. The proposed re-ranking strategy improves preferred pass rates by 4.35 to 5.80 percentage points. Additionally, this work distills 17 practical guidelines for effective skill authoring.
This study addresses the gap between the claimed reusability and actual cross-scenario transferability of publicly available agent skills (SKILL.md), which often originate from single-task contexts. Building upon official specifications, we establish a two-tier defect taxonomy and conduct the first large-scale empirical analysis of 138,133 skill files, revealing that reuse barriers stem predominantly from routine packaging flaws rather than sophisticated attacks. We propose a quality-assured generation pipeline integrating specification-aware prompting, lightweight linting, automated repair, and security gating. Our experiments show that 91.8% of skills contain at least one defect; skills with valid routing metadata achieve significantly higher retrieval success rates; specification-aware skills exhibit fewer defects overall, whereas AI-generated skills suffer from more severe security and portability issues.