RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poor scalability and high evaluation costs associated with the joint search of reinforcement learning algorithm components by proposing a large language model (LLM)-driven self-evolutionary framework. Through LLM-guided program evolution, the framework advances from single-component editing to multi-component co-evolution. By integrating progressive training, probabilistic evaluation, and repeated verification mechanisms, it effectively balances search breadth with evaluation fidelity. Experimental results demonstrate that the proposed method improves returns by 32%–84% on benchmarks such as SAC while significantly reducing computational overhead. Furthermore, the framework discovers novel algorithms with enhanced interpretability.
📝 Abstract
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Algorithm Self-Evolution
Large Language Models
Joint Search
Evaluation Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
LLM-driven Evolution
Progressive Co-Evolution
Algorithm Discovery
Probabilistic Evaluation
Haoran Li
Haoran Li
Institute of Automation,Chinese Academy of Sciences
Artificial IntelligenceRoboticsReinforcement LearningEmbodied Intelligence
Z
Zengle Ge
Baidu
X
Xiaomin Yuan
Baidu
Y
Yui Lo
Baidu
S
Songlin Zhou
Tsinghua University
J
Jiahua Ying
Baidu
Haoxin Li
Haoxin Li
Nanyang Technological University
Computer VisionVision and Language
Q
Qianhui Liu
Institute of Automation, Chinese Academy of Sciences
Yuanhang Liu
Yuanhang Liu
Assistant Professor in Biomedical Informatics at Mayo Clinic
Bioinformaticsgenomicssingle cell and spatial omicsmachine learning
J
Jiaqun Liu
Peking University
G
Guokai Chen
Tsinghua University
M
Mingju Chen
Shanghai University
R
Ruinan Wang
University of Bristol
A
Annan Li
Baidu
J
Jianmin Wu
Baidu
Dawei Yin
Dawei Yin
Senior Director, Head of Search Science at Baidu
Machine LearningWeb MiningData Mining
Dou Shen
Dou Shen
Baidu Inc
Data MiningMachine LearningOnline Advertising