🤖 AI Summary
This study addresses the efficiency bottleneck arising from separately training challengers in self-evolving language models by proposing Direct Evolutionary Optimization (DEO). DEO introduces a novel self-evolution paradigm that eliminates explicit challenger training, replacing parameter updates with solver-guided task sampling. Specifically, a frozen LLM generates mutated tasks, and an exponentially tilted distribution defined via a KL-regularized objective is employed for sampling. Efficient optimization is achieved through an approximate Metropolis rule combined with theoretical gradient dominance condition analysis. Experiments demonstrate that DEO attains reasoning performance comparable to R-Zero while reducing training time by over 50%, significantly outperforming random-walk-free baselines. These results validate both the feasibility and efficiency of removing challenger training in self-evolving frameworks.
📝 Abstract
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.