🤖 AI Summary
This study addresses the optimization bias in skill evolution for large language models caused by sparse single-trajectory signals and local noise. To overcome these limitations, we propose a skill evolution framework based on group contrastive analysis and a dual-memory architecture. By generating multiple trajectories and performing inter-group comparisons, our method precisely identifies behavioral discrepancies. Furthermore, a dual-memory system continuously accumulates historical evidence to distill experiences into reusable skills, achieving a paradigm shift from isolated path dependence to collective pattern recognition. Experimental results demonstrate that the proposed framework consistently outperforms baseline methods across six benchmarks, yielding improvements of up to 5.69 points across five model categories.
📝 Abstract
Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for analysis, making the optimization direction susceptible to noisy evidence. To address these, we propose SkillCome, a Skill-evolution method based on group Contrast optimization with dual memory. For each question, SkillCome generates trajectories and performs group contrast analysis to precisely identify key behavioral divergences between successful and failed trajectories, offering reliable optimization signals. The dual memory system further accumulates evidence from historical steps to track patterns shared across different groups, leading to more generalized optimization directions. Together, SkillCome builds a systematic optimization process that transforms experience from observed successful trajectories into reusable skills. Extensive experiments on six benchmarks spanning question answering, reasoning, and agentic tasks demonstrate the effectiveness of our method. SkillCome consistently outperforms baselines across five models of varying families and scales, with gains up to +5.69 points.