🤖 AI Summary
This study addresses the limitation of existing skill evolution methods for LLM agents, where behavioral evidence loss and global verification mechanisms frequently lead to the erroneous discarding of locally effective modifications. To overcome this, we propose an evidence-driven skill optimization framework featuring a novel replayable evidence card mechanism that structurally encapsulates execution observations, enabling explicit binding between skill edits and contextual evidence. Through targeted replay for local verification, the framework progressively retains supportive edits to continuously refine the skill library. By integrating LLM agents, continual learning, and experience replay techniques, this work establishes a cohesive approach to agent skill improvement. Extensive experiments conducted across three interactive benchmarks and six LLM backbones comprehensively validate the effectiveness of the proposed method.
📝 Abstract
Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justified and when it should become persistent guidance. However, existing experience-driven methods can lose the behavioral evidence and task contexts supporting edits. Moreover, a global validation outcome provides an incomplete judgment of its constituent changes: locally supported corrections may be discarded with a rejected revision, while evidence may require further experience to inform useful updates. To this end, we introduce EVISKILL, an evidence-driven framework that organizes execution observations into Replayable Evidence Cards and synthesizes edits with explicit links to their supporting contexts. Targeted replay verifies these edits through re-execution and provides feedback for correction. Across epochs, EVISKILL preserves evidence and provisionally retains supported edits for further refinement, while global validation governs their incorporation into the final skill. Experiments on three interactive benchmarks across six LLM backbones demonstrate the effectiveness of this approach.