Score
Designs and evaluates representations of skills intended for transfer across tasks, agents, or environments; builds encoding schemes, latent spaces, or modular controllers that capture reusable, compositional, and transferable aspects of behavior, and analyzes how these representations affect transfer efficiency, generalization, and adaptation.
This study systematically investigates the management of dynamically evolving skill repositories in large language model agents. Based on a comprehensive review of 124 publications from 2023 to 2026, it introduces the first integrated framework that treats skill repositories as evolvable artifacts, comprising a six-dimensional skill taxonomy, an eight-stage lifecycle architecture, and a ten-operator provenance vocabulary. The work uncovers the critical roles of skill admission and repair mechanisms, demonstrates how validator quality influences reinforcement learning efficacy, and identifies performance bottlenecks of flat retrieval strategies under scaling conditions. Building on these insights, the paper proposes standardized evaluation criteria for dynamic skill repositories and outlines key open challenges in the field.
Existing research lacks a systematic understanding of the full lifecycle of model-generated skills—spanning experience generation, skill extraction, and skill consumption—making it difficult to evaluate their effectiveness and applicability. This work proposes the first utility-based evaluation framework to systematically analyze key factors influencing skill extraction and consumption across five task domains. Through multi-model comparisons, cross-consumer transfer tests, and analyses of experience composition, we find that extracted skills are on average beneficial but exhibit significant negative transfer, with utility independent of model scale. Furthermore, we introduce a meta-skill guidance strategy that substantially improves cross-domain skill quality and mitigates negative transfer, revealing a notable inconsistency between extractor and consumer performance.
研究探讨了通过激活空间中的向量表示来操纵大型语言模型的程序技能,发现这些技能可以作为方向被激活和组合,以实现更高级别的功能和个人化优化。
研究通过对比任务级与子任务级技能诱导及文本与代码格式,解决LLM代理技能转移不可靠问题,提出技能效用评分以预测任务成功。
This work addresses the prevalent issue of capability degradation and catastrophic forgetting in large language models following task-specific fine-tuning. To mitigate this, the authors propose Activation-difference-guided Channel Targeting (ACT), a method that identifies a sparse, decoupled, and stable subset of model channels where task-specific capabilities are concentrated. By selectively transferring only these critical channel parameters, the approach enables efficient capability fusion and recovery of forgotten skills. Experimental results demonstrate that ACT effectively preserves original competencies while restoring lost abilities across multilingual mathematical and scientific reasoning tasks, and successfully consolidates multiple specialized models into a single, versatile model without significant performance trade-offs.
Existing training and evaluation frameworks lack controllable shared latent structures, making it difficult to systematically analyze how agents leverage cross-task experience to improve decision-making. This work proposes LatentGym—the first benchmark suite grounded in real, controllable latent variables—that decouples exploration (acquiring latent knowledge) from exploitation (applying learned knowledge), thereby enabling fine-grained assessment of cross-task adaptation mechanisms. Experiments demonstrate that the platform can uncover the root causes of large language models’ failures in cross-task generalization, validate the efficacy of post-training on task sequences, and elucidate how design choices such as inter-task feedback critically shape learning dynamics and generalization performance.
This study addresses the challenge of transferring unsupervised skill discovery to unseen environment layouts, which is hindered by the absence of expert data and limited generalization. To overcome this, the authors propose an action-aware representation objective that is functionally equivariant while preserving temporal structure. Grounded in bisimulation theory, they introduce a novel skill discovery paradigm conditioned on directly executable relevant states, achieving cross-layout generalization by constraining skill behaviors to remain invariant over specific subsets of state features. By integrating deep reinforcement learning with unsupervised representation learning, this work demonstrates that the acquired skills can be efficiently transferred across diverse environmental layouts for various downstream tasks, exhibiting superior out-of-distribution generalization performance.
This study addresses the issue of erroneous policy merging caused by semantic similarity during experience integration in LLM agents. To mitigate this, we propose an Online Skill Evolution framework that abstracts experiences into a hierarchical, reusable skill library through behavior-verified scope expansion, cross-instance replay, and mechanism validation. This approach effectively prevents semantic misguidance while ensuring behavioral consistency. Experimental results demonstrate that the framework consistently enhances agent performance across multiple benchmarks. Furthermore, the acquired skills exhibit transferability across different model scales and families, achieving continuous learning and generalization without compromising behavioral fidelity.
为了解决现有递归自我改进方法的局限性,提出了一种模块化、可泛化的框架ModularRSI,通过独立进化五个功能模块并整合来改善代理执行机制。
This study addresses the unclear activation mechanisms and failure boundaries of skills in LLM agents by employing controlled experiments and trajectory analysis to construct a taxonomy comprising three categories and twelve patterns. The research identifies "procedural anchoring" as the dominant mechanism for stable execution, accounting for 65.7% of cases, and quantifies the degradation of retrieval precision as pool size increases. Results demonstrate that skills outperform workflow memory by 6.06 points. Moving beyond traditional aggregate evaluation paradigms, this work systematically elucidates the underlying mechanisms and failure conditions of skill utilization. Consequently, it provides both theoretical foundations and practical guidance for designing reliable self-evolving agents, offering critical insights into optimizing agent architectures through rigorous mechanistic analysis rather than mere performance benchmarking.
This study addresses the challenge of filtering erroneous and non-transferable knowledge from skill libraries of self-evolving agents by proposing a Proposer-Builder-Verifier architecture to validate skill reusability in unseen tasks. The method dynamically synthesizes test scenarios through conditional constraint generation, overcoming the limitations of traditional static evaluation. Furthermore, it employs an automated execution comparison mechanism to quantitatively assess skill utility and efficiency, enabling reliable retention or rejection decisions. Experimental results demonstrate that the proposed framework significantly improves downstream task performance and execution efficiency on the ALFWorld and WebShop benchmarks, while also confirming the accuracy of its skill reusability assessment.