🤖 AI Summary
This work addresses the high memory cost of gradient-based post-training for large language model agents in resource-constrained environments, where long interaction trajectories render backpropagation through time impractical, while traditional evolutionary strategies, though memory-efficient, suffer from slow convergence. To overcome these limitations, the authors propose Cooperative Parameter Subspace Evolution Strategy (CoPES), which introduces cooperative coevolution into large language model post-training for the first time. CoPES decomposes the full parameter space into low-dimensional subspaces and optimizes them collaboratively using evolution strategies instead of backpropagation. Experiments on tool-augmented tasks with Qwen3.5-4B show that CoPES achieves 92% of the validation accuracy gain obtained by GRPO under the same GPU-hour budget while using less than one-eighth of its memory footprint—significantly outperforming standard evolutionary strategies, which recover only 67%. CoPES also surpasses baseline methods across five mathematical reasoning and question-answering benchmarks.
📝 Abstract
Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES translate into prohibitively long training times. To address this, we introduce Cooperative Parameter-subspace Evolution Strategy (CoPES), a cooperative coevolutionary method that decomposes the full parameter space into lower-dimensional subspaces and searches over them cooperatively to improve optimization efficiency. We post-train a Qwen3.5-4B tool-using agent for the math task and evaluate it on five benchmarks of varying difficulty. Under the GPU-hour budget of full-parameter GRPO's best validation checkpoint, CoPES recovers 92% of GRPO's validation-accuracy gain, versus 67% for standard ES, while its theoretical GPU memory requirement is less than one-eighth that of full-parameter GRPO. It consistently outperforms standard ES and LoRA-based GRPO on all evaluated pass@k metrics across the five benchmarks. Additional experiments further show the advantage of CoPES on the question-answering task. These results demonstrate an improved trade-off between memory requirements and training time for agentic LLM post-training under resource constraints. The code is open-sourced in https://github.com/MetaronWang/CoPES