Agent Skill Evolution: How Revisions Affect Coding Agents

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear impact of skill file revisions on the behavior and performance of LLM-based coding agents. It presents the first systematic quantification of skill evolution patterns by employing empirical analysis, automated rule checking, multi-model comparisons, and blind evaluations to assess the effects and costs of rule modifications in both single-turn and sandbox environments. The findings reveal that introducing new skill rules improves compliance rates by 0.41, increases agent action rates by 0.23, and enhances final accuracy by 0.10. Furthermore, this work quantifies the associated token overhead incurred by these modifications. By providing rigorous empirical evidence on how skill rules shape agent capabilities and computational costs, this research offers a principled foundation for optimizing the design and deployment of LLM-powered coding agents.
πŸ“ Abstract
Agent Skills, the SKILL.md files that tell an LLM coding agent how a project works, are revised like code, yet what a revision does to the agent is unknown. From 2,608 first/last revision pairs of 3,159 Skills, we characterize how Skills evolve and how they change together with the configuration of the agent's harness. We then focus on rule changes, revisions that add or remove a rule we can check automatically, such as"run allium check". We measure their effect on 21 models in single answers and on four agents in a sandbox, and their cost on 20 of these models and the four agents. Most revisions (55%) change a rule or procedure, and commits that revise a Skill change harness files such as CLAUDE.md more often than other commits of the same size. Across 16 open-weight models, an added rule raises compliance in a single answer by +0.41 on average. Across the four agents, the rate at which the agent takes the required action rises by +0.23 on average (+0.16 to +0.36), and for the three agents that blind judges assessed, final correctness rises by +0.10 on average (+0.06 to +0.14). The gain comes mainly from rules that name a command or path the old Skill did not mention. Real tools load a Skill's body only when the agent decides it needs it. In that setting the four agents keep about half of the action gain on average (51%), and the three open models about 38%. A revision adds 18-19% input tokens to a single answer and no detectable cost to an agent episode, while loading a Skill's body raises the tokens of an episode by 50% on average.
Problem

Research questions and friction points this paper is trying to address.

Coding Agents
Agent Skills
Skill Revision
LLM Compliance
Rule Changes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Skill Evolution
Coding Agents
Rule Compliance
Large Language Models
Prompt Engineering