🤖 AI Summary
This study addresses the limitations of Evolution Strategies (ES) in fine-tuning large language models, specifically their inability to dynamically perceive problem difficulty and the efficiency degradation caused by feedback compression. To overcome these challenges, this work proposes the DART-ES framework, which constructs a dynamic difficulty state by estimating local solvability without requiring auxiliary difficulty models or backpropagation. This mechanism jointly guides continuous difficulty reweighting and targeted replay of rare samples, thereby optimizing perturbation evaluation and data allocation. Evaluated on tasks such as mathematical reasoning, DART-ES significantly outperforms standard ES and achieves performance comparable to GRPO while substantially reducing runtime and memory overhead. These results establish DART-ES as a promising new paradigm for efficient full-parameter fine-tuning.
📝 Abstract
Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem changes with model capability. To address this limitation, we propose Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies (DART-ES). DART-ES estimates the local solvability of each problem from its pass rate across the perturbation population and aggregates historical observations to construct a dynamic difficulty state. This shared state jointly guides continuous difficulty reweighting and rare solvable sample replay, thereby improving perturbation direction evaluation and training data allocation without introducing an additional difficulty model or backpropagation. Extensive experiments show that DART-ES achieves good fine-tuning performance. DART-ES outperforms ES on all five base models and improves the average accuracy from 72.07\% to 73.53\%, exceeding the 73.26\% achieved by GRPO on GSM8K. Across five challenging mathematical reasoning benchmarks, DART-ES achieves an average accuracy of 49.20\%, compared with 48.34\% for ES and remains competitive with strong 7B models trained with RL. Further experiments show consistent gains in instruction tuning, code generation and the Countdown task with a 14B model, demonstrating strong generalization across tasks and scalability to larger models. Beyond performance gains, DART-ES also shows clear advantages in system efficiency. It reduces runtime per step by 15.2\%--50.2\% and peak memory usage per GPU by 21.1\%--51.1\% compared with GRPO. Despite performing full parameter fine-tuning, DART-ES also requires less runtime and GPU memory than GRPO+LoRA.