SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of overfitting, redundant instruction accumulation, and behavioral degradation in agent skill evolution by proposing the first general regularization framework tailored for external skill evolution. Drawing inspiration from neural network training paradigms, the framework employs skill dropout and complexity-aware local regularization to prevent unnecessary expansion of the skill library. Furthermore, it introduces a causal counterexample verification mechanism that precisely detects update-level regressions often overlooked by structural metrics. Experimental results across multiple benchmarks demonstrate that the proposed approach effectively constrains skill library size while maintaining competitive performance, significantly enhancing downstream transferability and late-stage evolutionary outcomes.
📝 Abstract
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neural-network training. SkillEvoReg combines training-time skill dropout, which perturbs update generation, and complexity-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation (CCV), which provides targeted behavioral validation of candidate-specific regressions. We instantiate the framework across heterogeneous skill-evolution systems while retaining each system's native skill evolver and task evaluator. Across SkillOpt, SkillEvolBench, and ContinualSkillBench, SkillEvoReg consistently controls skill-state growth while preserving competitive downstream capability, improves several transfer and later-stage evolution outcomes, and identifies update-level regressions that structural metrics alone cannot reveal. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters.
Problem

Research questions and friction points this paper is trying to address.

skill evolution
overfitting
language-model agents
reusable skills
catastrophic forgetting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Skill Evolution Overfitting
Regularization Framework
Skill Dropout
Causal Counterexample Validation
Language-Model Agents
🔎 Similar Papers
No similar papers found.