Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

📅 2026-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of directly applying evolutionary strategies (ES) to billion-scale language model inference, where high-dimensional random perturbations become nearly orthogonal to effective update directions. To overcome this, the authors propose Hyper-ES, which first constructs a low-dimensional adaptation subspace via limited gradient fine-tuning and then employs CMA-ES within this subspace to optimize layer-wise DARE-TIES fusion coefficients, thereby restricting the search to meaningful directional combinations. By integrating gradient-informed subspace construction with gradient-free evolutionary optimization, Hyper-ES consistently outperforms GRPO-LoRA by approximately 1% across Qwen2.5-Instruct and DeepSeek-R1-Distill models on six mathematical reasoning benchmarks, while reducing space-intensive gradient updates by 10%.
📝 Abstract
Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper-ES first performs a small number of inexpensive gradient-based fine-tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper-ES then applies CMA-ES to optimize layer-wise DARE-TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full-model perturbations. We evaluate Hyper-ES on three Qwen2.5-Instruct and DeepSeek-R1-Distill backbones across six mathematical reasoning datasets. Results show that Hyper-ES consistently outperforms GRPO-LoRA by 1% while requiring 10% fewer space-consuming gradient updates. Code at https://github.com/kuangrepi/Hyper-ES.
Problem

Research questions and friction points this paper is trying to address.

Evolution Strategy
Large Language Model
High-dimensional Optimization
Parameter Space
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evolution Strategy
Subspace Optimization
DARE-TIES Merging
Gradient-Free Fine-Tuning
Large Language Models
Y
Yu Gu
School of Intelligence Science and Technology, Nanjing University, China
Zhi Zheng
Zhi Zheng
National University of Singapore, NUS; Southern University of Science and Technology, SUSTech
Large Language ModelMachine LearningNeural Combinatorial Optimization
Y
Yunpeng Ba
School of Automation and Intelligent Manufacturing, Southern University of Science and Technology, China
X
Xialiang Tong
Noah’s Ark Lab, Huawei Technologies Ltd., China
M
Mingxuan Yuan
Noah’s Ark Lab, Huawei Technologies Ltd., China
Z
Zhenkun Wang
School of Automation and Intelligent Manufacturing, Southern University of Science and Technology, China