Score
Design and build algorithms and pipelines that extract, distill, and organize reusable, parameterized procedural skills from expert demonstrations, recordings, or execution traces, producing step-level segments, hierarchical skills, dependency graphs, and skill taxonomies or SBOM-like inventories. Analyze and transform expert macros into generalized skills, induce skill-aligned clustering and hierarchies, and encode operation ordering and dependencies to support retrieval, prompt instantiation, and compositional reuse.
Current large language model (LLM) agents face challenges in real-world deployment, including inefficiency, error-proneness, and poor maintainability, largely due to their reliance on on-the-fly reasoning and low-level tool invocation. This work introduces, for the first time, a skill-centric agent architecture that formalizes a comprehensive skill lifecycle framework encompassing representation, acquisition, retrieval, and evolution. It positions skills as a complementary mechanism bridging high-level reasoning and operational execution. By integrating key techniques—such as skill representation learning, automated acquisition, semantic retrieval, and continual evolution—and synergizing them with tool use, memory mechanisms, and contextual constraints, the proposed framework establishes a reusable and composable skill system. The paper further surveys representative approaches, open-source resources, and application scenarios, offering both theoretical foundations and practical guidance to enhance the scalability, robustness, and maintainability of intelligent agent systems.
This work addresses the challenge that agents struggle to efficiently reuse successful experiences when repeatedly performing similar tasks, often resulting in redundant reasoning and excessive interaction rounds. To overcome this limitation, the paper introduces a novel framework that formalizes procedural skills as parameterized finite state machine (PFSM) subgraphs and automatically extracts, verifies, and reuses structured skills through distillation and compilation of successful execution trajectories. Evaluated on the ALFWorld and WebArena benchmarks, the proposed method significantly improves task success rates while reducing the number of required interactions, demonstrating effectiveness across language models of varying scales.
Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.
This work addresses the limited reusability and governability of procedural capabilities in large language model (LLM) agents when performing long-horizon tasks, a challenge exacerbated by existing tool-calling mechanisms that struggle to support cross-task generalization. The paper introduces, for the first time, seven system-level design patterns for skills alongside an orthogonal “representation × scope” taxonomy, and constructs a comprehensive skill lifecycle framework encompassing metadata encapsulation, executable code, self-evolving libraries, marketplace-based distribution, and trust-tiered execution. Through benchmark evaluations and case studies, the authors demonstrate that structured skill representations significantly improve task success rates, while also uncovering critical risks: performance degradation in self-generated skills and severe security vulnerabilities within the skill supply chain.
This work addresses the limited procedural expertise of large language models in autonomous workflows, which hinders their practical deployment. To overcome this, we propose the first large-scale framework for automatic extraction of procedural knowledge tailored to multi-agent repositories. By analyzing open-source agent projects from platforms like GitHub, our approach combines repository structure parsing with dense retrieval to identify high-value skills—such as visualization and pedagogy—and leverages the Manim engine to generate instructional content, uniformly formatted as standardized SKILL.md files. The framework enables skill expansion without model retraining and incorporates safety governance alongside a multidimensional evaluation mechanism. Experimental results demonstrate that the generated instructional materials achieve a 40% improvement in knowledge transfer efficiency while matching the quality of human-authored tutorials.
Existing large language model agents struggle to effectively decompose complex tasks, retrieve appropriate skills, and generate executable multi-skill composition plans. This work proposes SkillWeaver, a framework comprising a three-stage pipeline—task decomposition, skill retrieval, and dependency-aware DAG planning—and introduces the first iterative Skill-Aware Decomposition (SAD) mechanism, which leverages a retrieval feedback loop to enhance alignment between subtasks and the skill repository. The study also constructs CompSkillBench, the first benchmark dedicated to compositional skills. Experimental results demonstrate that a single SAD iteration improves decomposition accuracy from 51.0% to 67.7% while reducing context consumption by over 99%, and achieves a 35.6% relative planning gain on unseen skill categories.
This work addresses the semantic misalignment between retrieved general-purpose skills and the current task, environment, or other skills during execution—where skills are semantically relevant but suffer from execution-level mismatches. To resolve this, the paper proposes SkillAligner, a framework that treats retrieved skills as tunable drafts at inference time without requiring additional training. SkillAligner performs a one-shot joint adaptation to simultaneously customize skills for the target task, align their interfaces, and coordinate multi-skill interactions by resolving dependencies, eliminating conflicts, and removing redundancies, thereby producing a compact and unified execution plan. Experiments demonstrate that SkillAligner significantly improves task success rates across diverse agent benchmarks and model backbones, effectively mitigates performance degradation caused by skill integration, and reduces inference overhead.
研究探讨了通过激活空间中的向量表示来操纵大型语言模型的程序技能,发现这些技能可以作为方向被激活和组合,以实现更高级别的功能和个人化优化。
本文提出SkillAlchemy框架,通过对比证据识别隐含需求并编译成技能包,解决开放环境下从有限信息创建可靠智能体技能的问题。
This work addresses the challenge of efficiently reusing fine-grained, executable, and contract-consistent skill units from large agent skill repositories under limited context budgets. To this end, we propose SkillZip, a framework that performs node-level compression of skills through an execution-aware graph structure, abstracting redundant contract-valid subgraphs into invertible, ported macros. These macros preserve boundary signatures and dependency closures while supporting on-demand expansion. SkillZip is the first approach to achieve contract-preserving, dependency-closed, verifiable, and scalable skill graph compression, effectively bridging the unit mismatch among skill retrieval, compression, and execution. Experiments demonstrate that SkillZip improves performance by an average of 12.2 points across multiple agent benchmarks, achieves a 3.46× compression ratio, retains 99.2% of dependencies, attains 98.7% verifiability coverage, and enables stable retrieval in skill libraries with over 100,000 entries.
This study addresses the disconnect between agent skill relevance and task utility, along with the lack of mechanistic explanations, through an empirical investigation across 87 tasks. Methodologically, we define a downstream utility metric based on pass-rate differentials and employ both large language models and human review to analyze skill content, execution trajectories, and artifacts. Furthermore, we propose a re-ranking strategy supporting essential operations alongside a DAG-based dependency organization method. Results reveal that in 36.78% of tasks, identical skills exhibit opposite utility under different configurations. The proposed re-ranking strategy improves preferred pass rates by 4.35 to 5.80 percentage points. Additionally, this work distills 17 practical guidelines for effective skill authoring.