SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the imprecise evaluation in existing agent skill package self-evolution benchmarks, which fail to distinguish among documentation repair, script repair, and behavior preservation. To overcome this limitation, we construct the first fine-grained skill evolution benchmark comprising 350 tasks that separately evaluates these three capabilities. Furthermore, this work proposes an Abstract Syntax Tree (AST)-guided revision method that leverages static call graph constraints to restrict the editing scope, thereby enabling coordinated updates and precise repairs across both documentation and code. Experimental results demonstrate that the proposed approach achieves an absolute improvement of over 20% in repair success rate compared to a pure Markdown baseline, while significantly enhancing consistency across multiple execution runs.
📝 Abstract
Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey of over 35,000 GitHub-hosted Skill roots, we select 100 packages and construct 150 repair tasks. Each task pairs a package containing injected script faults with a maintenance request and executable checks of the required behavior. A complementary controlled track contains 200 tasks from 50 packages, each evaluated under the same maintenance request in four states: clean, documentation faults, script faults, and faults in both. Across four LLMs, methods that edit both documentation and scripts can repair script faults but do not consistently outperform Markdown-only revision on documentation repair or preservation. We therefore introduce AST-Guided Skill Revision, which uses abstract syntax trees and calling relationships to link maintenance requirements to relevant code locations. It restricts script edits to these locations and updates the documentation to match the revised scripts. Averaged across models, this revision stage yields absolute gains in repair success of 21.9% for Raw Package and 27.7% for CoEvoSkills on faulty packages. Absolute gains in the proportion of tasks solved in all three runs reach 20.8% and 31.5%, respectively, indicating more consistent repair success across repeated runs.
Problem

Research questions and friction points this paper is trying to address.

Executable Agent Skills
Skill Self-Evolution
Benchmarking
Script Repair
Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

SkillScriptBench
AST-Guided Skill Revision
Executable Agent Skills
Self-Evolution
Abstract Syntax Tree
🔎 Similar Papers
No similar papers found.