Score
Designs and analyzes algorithms, systems, and update mechanisms that iteratively evolve and optimize meta-skills — i.e., skills whose purpose is to acquire, compose, orchestrate, or improve other skills — including processes for recursive self-improvement and closed-loop meta-skill optimization. Work in this area also builds and evaluates two-timescale or hierarchical evolution dynamics and non-parametric update mechanisms that change meta-skills over time rather than only changing base-model weights.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
This work investigates how agents can achieve controllable, continual evolution and adaptation with minimal human intervention. We model self-improving agents as operational scaffolds comprising a foundation model integrated with prompts, memory, tools, and control logic, and formalize self-improvement as self-triggered updates to either model parameters or scaffold components. Based on the update objectives and driving signals, we propose the first systematic taxonomy that unifies existing approaches, clarifies application scenarios and evaluation metrics, identifies key open challenges, and establishes a dynamically maintained repository of technical advances in the field.
Current large language models (LLMs) lack dynamic parameter updating capabilities, hindering adaptation to novel tasks, evolving knowledge, and real-time interaction in open-world environments—thus impeding their progression toward autonomous agents. To address this, we propose the “self-evolving agent” paradigm, introducing the first unified framework for dynamic adaptation along three dimensions: *what* evolves (evolutionary content), *when* it evolves (timing), and *how* it evolves (mechanism). We establish the first systematic taxonomy covering component-level evolution, adaptation stages, and driving mechanisms. Our approach integrates continual learning, test-time adaptation, reinforcement learning with reward shaping, natural-language feedback, and multi-agent coordination to enable real-time model evolution. We further design a unified evaluation framework, specifying application domains—including code generation, education, and healthcare—and critical challenges such as safety, scalability, and robustness. This work lays a theoretical foundation and provides a concrete technical pathway toward artificial superintelligence.
Current large language model (LLM) agents lack recursive self-improvement, as they optimize only task-specific skills without evolving their own improvement mechanisms. This work proposes a dual-timescale framework enabling agents to concurrently maintain and co-evolve both task skills and meta-skills through a unified process. For the first time, this approach achieves self-evolution of the improvement mechanism within a single framework, without requiring auxiliary models or external objectives. The meta-skill module comprises five components—Analyzer, Retriever, Allocator, Proposer, and Evolver—built upon a shared, frozen backbone, facilitating rapid optimization of task skills in a fast loop and gradual evolution of meta-skills in a slow loop. The method outperforms the strongest baselines by 23.54, 16.09, and 1.92 percentage points on the OfficeQA, SealQA, and ALFWorld benchmarks, respectively.
This work addresses the limitation of existing large language model post-training methods in modeling self-evolutionary meta-skills—such as introspection driven by environmental feedback—hindering autonomous, continual improvement in open-ended scenarios. The authors propose a novel paradigm that integrates synthetic evolutionary trajectory data, verifiable reward-based reinforcement learning grounded in code execution feedback, and inference-time evolutionary search. For the first time, test case execution outcomes are leveraged as supervision signals to systematically train models for self-reflection and cross-domain generalization without explicit annotations. The approach achieves substantial performance gains across seven programming benchmarks, improving absolute accuracy by 10.01% on in-distribution tasks and by 24.12% on out-of-distribution tasks, with a remarkable 46.9% relative improvement on out-of-domain open-ended algorithmic optimization problems.
Current agent systems struggle to efficiently evolve skills at test time, often relying on hard-coded strategies or costly updates to large model parameters, and lack the capability to continuously optimize the skill evolution mechanism itself. This work proposes HiSME—a lightweight hierarchical skill meta-evolution framework—that, for the first time, directly optimizes the skill evolution mechanism during testing. By extracting meta-skills from task execution trajectories, HiSME employs a hierarchical architecture to jointly refine both skills and their evolution policies, enabling cross-scenario adaptation without updating large model parameters. Experiments demonstrate that HiSME significantly enhances skill repertoire quality across multiple agent benchmarks and generates diverse, transferable meta-skills for various downstream tasks, effectively supporting continual experiential learning.
This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.
Existing approaches to agent skill evolution typically assume a fixed set of tools and evaluate skills in isolation, struggling to handle tool-level failures and inter-skill interactions. This work proposes SkillSmith, the first framework that co-evolves skills and tools in a unified proposal space, jointly optimizing both through an ecological utility model inspired by Lotka-Volterra dynamics. SkillSmith incorporates an anti-pattern logging system to avoid conflicts and repeated failures, supports bundled operations such as encapsulation, composition, and decomposition of skill-tool pairs, and leverages execution trace analysis to enhance coordination. Evaluated across three benchmarks—including WildClawBench—and five scales of Qwen3.5 models, SkillSmith significantly outperforms strong baselines, with particularly pronounced gains in high-complexity tasks requiring multi-skill collaboration.
This study investigates whether performance gains from self-evolved skills of AI agents on training tasks generalize to unseen test tasks. By evaluating five self-evolution methods across six benchmarks, we find that only 5 out of 21 skills fully retain their training benefits, revealing significant overfitting limitations in existing approaches. To address this, we propose the General Skill Optimization (GSO) framework, which preserves general writing guidelines while dynamically generating task-specific skills for each new task. Evaluations using an LLM-as-a-judge paradigm demonstrate that GSO achieves state-of-the-art performance across all benchmarks, validating the superiority of the per-task skill generation strategy.
This study addresses the challenges of overfitting, redundant instruction accumulation, and behavioral degradation in agent skill evolution by proposing the first general regularization framework tailored for external skill evolution. Drawing inspiration from neural network training paradigms, the framework employs skill dropout and complexity-aware local regularization to prevent unnecessary expansion of the skill library. Furthermore, it introduces a causal counterexample verification mechanism that precisely detects update-level regressions often overlooked by structural metrics. Experimental results across multiple benchmarks demonstrate that the proposed approach effectively constrains skill library size while maintaining competitive performance, significantly enhancing downstream transferability and late-stage evolutionary outcomes.
针对技能自进化中的稳定性和效率问题,提出SkillAdam框架,通过优化记忆和波动驱动的编辑预算来稳定更新方向并自适应控制更新幅度。
Existing skill evolution methods for LLM agents force failed trajectories to match fixed successful paths, overlooking valid progress within failure prefixes. To address this limitation, this work proposes SkillPivot, a framework that introduces the first deviation-point-guided mechanism. By precisely detecting the transition point from a valid prefix to an erroneous suffix, SkillPivot guides a teacher model to generate alternative actions for local updates, thereby preserving verified effective strategies and avoiding inefficient global reflection. Evaluated on benchmarks such as ToolQA, the proposed method surpasses existing baselines, significantly enhancing the performance of multiple models while generating compact and transferable skill updates.
This study addresses the stagnation of skill evolution in formal theorem proving caused by the absence of successful trajectories and the neglect of dynamically evolving reference knowledge. To overcome these limitations, we propose a mutation-enhanced skill self-evolution framework. By leveraging large language model agents and Monte Carlo Tree Search sampling, our method integrates progressive evolution with a verification feedback-driven, concept-oriented mutation mechanism. This enables the joint evolution of high-level strategies and mathematical concepts, effectively breaking the update bottleneck when no successful samples are available. Experimental results demonstrate that the proposed framework achieves a 100% success rate on MiniF2F and 90.6% on PutnamBench, while solving four problems each from the IMO and USAMO competitions, significantly outperforming existing baselines.