StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalizability and rigid strategies of fixed revision operators in large language model (LLM) skill optimization by proposing StraTune, a novel framework that introduces an adaptive operator selection mechanism conditioned on optimization states. Specifically, StraTune leverages a frozen optimizer LLM to dynamically select revision operators based on execution feedback, combined with hierarchical evaluation for candidate skill filtering, thereby achieving state-aware self-evolution without parameter updates. Experimental results demonstrate that this framework consistently outperforms baseline methods across four benchmarks under two distinct LLM configurations. Furthermore, skills acquired by smaller models can be effectively transferred to more capable ones.
📝 Abstract
Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision forms applied under it. However, we observe that no single revision operator consistently performs best across tasks, and repeatedly applying an unsuitable operator can limit further improvement. We propose StraTune (strategy-guided skill tuning), which lets a frozen optimizer LLM choose the revision operator at every round from the optimization state, which is defined as the current execution feedback together with the recorded outcomes of earlier strategies and forms. Candidate skills from every revision operator pass one candidate evaluation, which screens for gains and regressions on a small sample set and validates them on a larger one, and every outcome is written back to the optimization state for later choices. Across four benchmarks and two LLM settings, StraTune outperforms all five baselines in most settings. Ablations attribute the gains to the adaptive choice of the revision operator, since fixed, random, scheduled, and bandit strategy choices all score lower, and skills learned with a small target LLM also improve a stronger one. Code and learned skills are available at https://github.com/seai-lab/StraTune.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Skill Revision
Self-Evolving Skills
Revision Operators
Execution Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Revision Operator
Self-Evolving Skills
Strategy-Guided Tuning
Candidate Evaluation
Optimization State