Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the efficiency bottleneck arising from separately training challengers in self-evolving language models by proposing Direct Evolutionary Optimization (DEO). DEO introduces a novel self-evolution paradigm that eliminates explicit challenger training, replacing parameter updates with solver-guided task sampling. Specifically, a frozen LLM generates mutated tasks, and an exponentially tilted distribution defined via a KL-regularized objective is employed for sampling. Efficient optimization is achieved through an approximate Metropolis rule combined with theoretical gradient dominance condition analysis. Experiments demonstrate that DEO attains reasoning performance comparable to R-Zero while reducing training time by over 50%, significantly outperforming random-walk-free baselines. These results validate both the feasibility and efficiency of removing challenger training in self-evolving frameworks.
📝 Abstract
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}volving \textbf{O}ptimization (DEO), which replaces challenger parameter updates with solver-guided task sampling. The KL-regularized challenger objective defines an exponential tilt of a fixed base task distribution. DEO uses this distribution as a sampling target: a frozen LLM generates and mutates tasks, the solver scores them, and an approximate Metropolis selection rule refines the training pool. Only the solver is trained. Theoretically, for an idealized variant that samples exactly from the tilted distribution, and under regularity, local gradient-dominance, and initialization conditions, we show that DEO learns distributionally robust reasoning ability. In experiments, DEO achieves reasoning performance competitive with R-Zero while using over $50\%$ less wall-clock training time, and improves reasoning accuracy over a no-walk ablation. Replacing the task generator with a frozen API-only LLM further improves the local solver, illustrating a capability enabled by removing challenger training.
Problem

Research questions and friction points this paper is trying to address.

self-evolving language models
challenger training
task generation
reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Optimization
Challenger-Free Training
Metropolis Selection
Task Sampling
Distributionally Robust Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yuyang Deng
Yuyang Deng
Columbia University
OptimizationMachine Learning TheoryDistributed Machine LearningDeep Learning
Y
Yu Wang
Accenture, Center for Advanced AI
J
Jiayun Wang
Georgia Institute of Technology