AGAR: a reinforcement learning substrate for LLM program evolution

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key limitations in LLM-based program evolution, including reliance on handcrafted constants and the absence of credit assignment and diversity control. To this end, it formally casts program evolution as a Markov Decision Process (MDP) for the first time, utilizing module prefixes as the action space to construct a reinforcement learning framework that decouples the controller from the estimator. This formulation reveals the implicit zero-discount property of existing algorithms, enables mechanism-level auditing, and facilitates automated algorithm discovery without gradient-based training. Experimental results demonstrate that the proposed framework significantly outperforms baseline methods across 19 tasks, with particularly notable improvements in competitive programming, thereby validating its effectiveness and transferability.
📝 Abstract
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, rather than the program it emits. Credit assignment, value estimation, adaptive exploration, and experience memory can then attach to distinct components. AGAR (Algorithm Generation As RL) provides the resulting substrate: any estimator can be replaced or switched off without changing the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends, and three seeds under one harness, AGAR improves on the stronger of two published baselines on most tasks, with gains concentrated in the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice, but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.
Problem

Research questions and friction points this paper is trying to address.

program evolution
reinforcement learning
algorithm discovery
credit assignment
Markov decision process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Markov Decision Process
Program Evolution
Credit Assignment
Large Language Models