Code Evolution for Control: Synthesizing Policies via LLM-Driven Evolutionary Search

πŸ“… 2026-01-11
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Traditional reinforcement learning approaches suffer from low sample efficiency, difficulty in reward design, and lack of interpretability, while handcrafted policies rely heavily on expert knowledge and exhibit limited generalization. This work proposes to frame control policy synthesis as a code evolution problem, introducing EvoToolkitβ€”a novel framework that integrates the programming priors of large language models (LLMs) with evolutionary search to enable training-free, automatic policy generation. By combining LLM-driven code mutation, task-specific fitness evaluation, and evolutionary selection, the method produces compact, human-readable, and executable control policies. These policies achieve strong performance across multiple tasks and inherently support direct inspection, manual modification, and formal verification.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Evolutionary LearningHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
πŸ“ Abstract
Designing effective control policies for autonomous systems remains a fundamental challenge, traditionally addressed through reinforcement learning or manual engineering. While reinforcement learning has achieved remarkable success, it often suffers from high sample complexity, reward shaping difficulties, and produces opaque neural network policies that are hard to interpret or verify. Manual design, on the other hand, requires substantial domain expertise and struggles to scale across diverse tasks. In this work, we demonstrate that LLM-driven evolutionary search can effectively synthesize interpretable control policies in the form of executable code. By treating policy synthesis as a code evolution problem, we harness the LLM's prior knowledge of programming patterns and control heuristics while employing evolutionary search to explore the solution space systematically. We implement our approach using EvoToolkit, a framework that seamlessly integrates LLM-driven evolution with customizable fitness evaluation. Our method iteratively evolves populations of candidate policy programs, evaluating them against task-specific objectives and selecting superior individuals for reproduction. This process yields compact, human-readable control policies that can be directly inspected, modified, and formally verified. This work highlights the potential of combining foundation models with evolutionary computation for synthesizing trustworthy control policies in autonomous systems. Code is available at https://github.com/pgg3/EvoControl.
Problem

Research questions and friction points this paper is trying to address.

control policies
autonomous systems
interpretability
reinforcement learning
code synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-driven evolution
code synthesis
interpretable control policies
evolutionary search
executable policy
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
P
Ping Guo
Shenzhen Yuyi Tech. Co. Ltd., Shenzhen, China
C
Chao Li
City University of Hong Kong, Hong Kong
Y
Yinglan Feng
Shenzhen University, Shenzhen, China
Chaoning Zhang
Chaoning Zhang
Professor at UESTC (η”΅ε­η§‘ζŠ€ε€§ε­¦, China)
Computer VisionLLM and VLMGenAI and AIGC Detection