🤖 AI Summary
This study addresses the excessive real-time control latency inherent in existing large language model (LLM)-based control schemes that rely on online inference and planning. To overcome this limitation, this work proposes a structure-parameter decoupling strategy that leverages LLMs to offline generate controller code structures, subsequently fitting parameters via derivative-free optimization to synthesize directly executable Python control policies. By eliminating LLM involvement during the decision-making phase, the approach achieves ultra-low-latency real-time control. Evaluated on reinforcement learning benchmarks such as Atari, the proposed method outperforms conventional planning-based approaches and demonstrates significantly faster action selection than PPO. Furthermore, it exhibits notable advantages in sample efficiency and cross-task transferability, offering a practical framework for deploying LLM-synthesized controllers in latency-sensitive environments.
📝 Abstract
Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.