🤖 AI Summary
The rapid evolution of deep reinforcement learning (DRL) and sequential decision-making has fragmented research across paradigms and emerging modalities, hindering systematic understanding and cross-paradigm integration. Method: This work presents a comprehensive survey of state-of-the-art DRL, organizing advances around four core paradigms—value-based methods, policy gradients, model-based prediction, and multi-agent RL—and uniquely unifies them with cross-modal frontiers including large language models (LLMs) and reasoning-augmented agents. Through comparative analysis, paradigm mapping, and technical taxonomy, it constructs a structured knowledge graph for general-purpose intelligent agents. Contribution/Results: We propose a full-stack unified analytical framework for DRL; introduce the first taxonomy integrating classical RL paradigms with LLM-driven agent architectures; and distill scalable methodological guidelines alongside a curated list of key open challenges—providing an authoritative reference for both theoretical advancement and applications in embodied intelligence and decision-focused foundation models.
📝 Abstract
This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based method, policy-gradient methods, model-based methods, and various other topics (e.g., multi-agent RL, RL+LLMs, and RL+inference).