Score
Design and implement pipelines and data structures that create structured skill nodes from collections of skill examples, encoding node features and metadata and aggregating related skills into compact, queryable representations; link these nodes into skill graphs and reconcile or ensemble outputs from multiple models during node generation.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.
为解决大规模技能库中有效检索难题,SE-GoS通过执行轨迹自进化图结构,无需训练即可优化技能关系与描述,提高任务奖励并减少输入标记。
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
This study addresses the challenge of organizing vast volumes of unstructured, multilingual professional skill statements by proposing a hybrid knowledge graph construction approach that integrates top-down and bottom-up strategies. Anchoring large language models to the Wikidata multilingual knowledge graph and incorporating an agent-based reflection mechanism, the method dynamically aligns known entities, creates nodes for emerging skills, and establishes relationships through a five-stage pipeline. Leveraging techniques such as multilingual normalization, active curation, and deduplication, the system yields a scalable, interpretable, and self-repairing knowledge graph. The resulting resource encompasses a comprehensive skill ontology and structured taxonomy across five European languages, effectively managing noisy textual inputs and continuously recovering unmapped concepts.
This work addresses the performance degradation of large language model agents as their skill repertoire expands, a problem rooted in the absence of effective skill orchestration mechanisms. To resolve this, the authors propose GraSP, an architecture that introduces a compilation layer between skill retrieval and execution. GraSP organizes a flat set of skills into a typed directed acyclic graph enriched with precondition-effect dependency edges. By incorporating node-level validation and five locally bounded repair operators, the framework reduces replanning complexity from O(N) to O(d^h). GraSP is the first method to produce executable skill graphs, achieving consistent improvements over baselines such as ReAct across four benchmarks—including ALFWorld—with up to a 19-point increase in task reward and a 41% reduction in interaction steps, while maintaining robustness under skill overload and degraded skill quality conditions.
This work addresses the challenge that large language models struggle to effectively follow textual skill instructions in long-context scenarios, which hinders practical skill deployment. To overcome this limitation, the authors propose ParametricSkills, a novel framework that, for the first time, dynamically converts free-form textual skills into LoRA adapter parameters at test time, enabling context-independent skill invocation and establishing a new paradigm for test-time continual learning. Leveraging a large-scale skill repository and invocation trajectories derived from OpenCode, a hypernetwork is trained to map textual descriptions to parameterized skills. Evaluated across six software engineering subtasks, the method outperforms in-context learning by an average of 6.44 points under DeepSeek-V4-Flash assessment, with significant improvements in both BERTScore and F1 metrics.
This study addresses the unclear adaptation patterns and impacts of large language model (LLM) agent skills when reused in downstream applications. Through an empirical analysis of 1,126 adaptation instances from six prominent skill repositories, the work systematically characterizes LLM skill adaptation behaviors and constructs a taxonomy comprising 46 patterns grouped into 13 families. The research uncovers critical phenomena including a “reuse paradox,” strong cross-component dependencies, and the introduction of security-sensitive content in nearly one-fifth of adaptations. It further identifies prevalent challenges such as logic rewriting, fixing discoverability issues, and cross-tool or cross-language translation. These findings offer new empirical insights and foundational support for improving skill design, standardizing interfaces, and enabling automated adaptation of LLM-based agents.
This work addresses the challenge that large language model agents struggle to effectively retrieve and compose executable skill sequences for complex tasks. To this end, the authors propose a unified three-layer graph structure that jointly models the semantic hierarchy of user queries, query-skill alignment, and inter-skill dependency constraints. By integrating graph traversal and semantic propagation algorithms with large language models, the approach enables end-to-end automatic discovery of executable skill paths. This is the first method to formalize skill composition as a unified graph-based reasoning framework. It achieves state-of-the-art performance with success rates of 53.17% on SkillsBench and 91.43% on ALFWorld, significantly outperforming existing approaches while demonstrating consistent gains across different backbone language models.
This work addresses the challenge of efficiently reusing fine-grained, executable, and contract-consistent skill units from large agent skill repositories under limited context budgets. To this end, we propose SkillZip, a framework that performs node-level compression of skills through an execution-aware graph structure, abstracting redundant contract-valid subgraphs into invertible, ported macros. These macros preserve boundary signatures and dependency closures while supporting on-demand expansion. SkillZip is the first approach to achieve contract-preserving, dependency-closed, verifiable, and scalable skill graph compression, effectively bridging the unit mismatch among skill retrieval, compression, and execution. Experiments demonstrate that SkillZip improves performance by an average of 12.2 points across multiple agent benchmarks, achieves a 3.46× compression ratio, retains 99.2% of dependencies, attains 98.7% verifiability coverage, and enables stable retrieval in skill libraries with over 100,000 entries.
为解决技能优化中流程指导不明确及搜索空间过大问题,提出将技能表示为图结构,并使用进化优化框架GraphSkillEvo进行优化。