A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing skill evolution methods for LLM agents force failed trajectories to match fixed successful paths, overlooking valid progress within failure prefixes. To address this limitation, this work proposes SkillPivot, a framework that introduces the first deviation-point-guided mechanism. By precisely detecting the transition point from a valid prefix to an erroneous suffix, SkillPivot guides a teacher model to generate alternative actions for local updates, thereby preserving verified effective strategies and avoiding inefficient global reflection. Evaluated on benchmarks such as ToolQA, the proposed method surpasses existing baselines, significantly enhancing the performance of multiple models while generating compact and transferable skill updates.
📝 Abstract
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
skill self-evolution
tool-use tasks
trajectory deviation
failed trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Skill Self-Evolution
Deviation-Point Guidance
LLM Agents
Trajectory Contrast
Localized Skill Update
🔎 Similar Papers
No similar papers found.