🤖 AI Summary
This work addresses the limitations of large language models (LLMs) in optimizing actions over long-horizon sequential decision-making tasks and the inability of reinforcement learning (RL) to perform high-level abstraction and task decomposition. To bridge this gap, the paper proposes a hierarchical hybrid architecture that, for the first time, effectively integrates the semantic reasoning capabilities of LLMs with the precise control mechanisms of RL. In this framework, the LLM generates subgoals and structured plans, while the RL agent optimizes low-level action policies through environmental interaction. Evaluated across multiple sequential decision tasks, the approach significantly improves sample efficiency, task success rates, and trajectory coherence, outperforming both pure RL and pure LLM baselines, thereby demonstrating the efficacy of synergistically combining high-level planning with low-level execution.
📝 Abstract
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise action optimization and environment interaction. Reinforcement Learning (RL), while effective for sequential control, often lacks the high-level abstraction and task decomposition abilities needed for complex scenarios. This paper introduces an LLM-Augmented Reinforcement Learning Agent that integrates LLM-driven planning with RL-based action optimization. The proposed architecture leverages the LLM to generate subgoals, structured plans, and contextual guidance, while the RL agent refines low-level actions through interaction with the environment. Experiments on sequential decision tasks demonstrate improved sample efficiency, higher success rates, and more coherent action trajectories compared to RL-only and LLM-only baselines. This hybrid paradigm highlights a promising direction for building more capable autonomous systems.