Score
Designs and trains reinforcement-learning (including deep RL) agents or policies that select among available devices, producing decision rules that map observed environment and device states to selection actions. Builds and evaluates systems that optimize provisioning, reliability, performance, and constraint satisfaction (for example security or resource limits) under dynamic, partially observable conditions.
This work addresses the low task completion and win rates of reinforcement learning (RL) agents in text-based games. We propose a novel end-to-end architecture that jointly integrates deep language understanding with policy gradient optimization. Methodologically, we employ pretrained language models (e.g., BERT or T5) to parse game text and implicitly construct a differentiable world model, which is co-optimized with an enhanced policy gradient algorithm—specifically Proximal Policy Optimization (PPO)—to enable efficient mapping from textual observations to action policies. Our key contributions are: (i) the first incorporation of a differentiable world modeling mechanism directly into the policy network, improving long-horizon reasoning and state consistency; and (ii) the integration of multi-task pretraining and curriculum learning to enhance generalization. Evaluated on standard benchmarks including Zork and TextWorld, our approach significantly outperforms existing RL baselines, achieving average improvements of 23.6% in task completion rate and 31.2% in win rate, thereby demonstrating both effectiveness and cross-game transferability.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
The rapid evolution of deep reinforcement learning (DRL) and sequential decision-making has fragmented research across paradigms and emerging modalities, hindering systematic understanding and cross-paradigm integration. Method: This work presents a comprehensive survey of state-of-the-art DRL, organizing advances around four core paradigms—value-based methods, policy gradients, model-based prediction, and multi-agent RL—and uniquely unifies them with cross-modal frontiers including large language models (LLMs) and reasoning-augmented agents. Through comparative analysis, paradigm mapping, and technical taxonomy, it constructs a structured knowledge graph for general-purpose intelligent agents. Contribution/Results: We propose a full-stack unified analytical framework for DRL; introduce the first taxonomy integrating classical RL paradigms with LLM-driven agent architectures; and distill scalable methodological guidelines alongside a curated list of key open challenges—providing an authoritative reference for both theoretical advancement and applications in embodied intelligence and decision-focused foundation models.
Deep reinforcement learning (DRL) suffers from low data efficiency and poor generalization across tasks. Meta-reinforcement learning (Meta-RL) addresses this by treating algorithm design itself as a learning problem, enabling agents to rapidly adapt to new tasks with few samples drawn from a task distribution. This paper introduces the first unified classification framework for Meta-RL, characterized along two orthogonal dimensions: (i) whether the task distribution is explicitly modeled, and (ii) whether the per-task learning budget is constrained. We systematically survey problem formulations, core paradigms, and representative algorithms—including MAML-based, RNN-based, contextual, and Bayesian approaches. Our contributions are threefold: (i) clarifying the field’s evolutionary trajectory and identifying key open challenges; (ii) constructing the first structured pedagogical guide; and (iii) advancing Meta-RL toward becoming a standard tool in the DRL toolkit—now widely adopted for both education and practical onboarding.
Reinforcement learning (RL) faces dual challenges of low accessibility and poor generalization: existing environments require manual implementation using low-level frameworks (e.g., CUDA, JAX), imposing high engineering barriers that hinder adoption by non-specialist teams; moreover, the absence of a unified, formalizable environment representation impedes agent transfer across tasks. This paper introduces, for the first time, the “linguistic environment modeling” paradigm—a framework that formalizes RL environments via domain-specific languages (DSLs) and natural language, integrated with semantic parsing and describability-aware modeling. By elevating environment specification from code-level implementation to high-level semantic abstraction, our approach drastically lowers the entry barrier for RL application. It enables small teams to efficiently construct, reuse, and transfer environments, while enhancing zero-shot generalization of agents over describable environment families. This work opens a new pathway toward democratizing and generalizing RL.
The application of reinforcement learning (RL) in software engineering lacks a systematic, comprehensive survey. Method: We conduct the first panoramic systematic literature review (SLR), covering 22 top-tier conferences and synthesizing 115 deep RL studies. Using multidimensional analysis—spanning algorithm types, datasets, model architectures, and evaluation methodologies—we develop a taxonomy organized along four software engineering activities: design, development, quality assurance, and maintenance. Contribution/Results: Our analysis reveals prevailing practices, shared limitations—including data scarcity, inconsistent evaluation protocols, and poor reproducibility—and emerging evolutionary trends. We propose actionable recommendations to address these challenges and publicly release all research artifacts, including the curated literature corpus, coding scheme, and analysis tools. This open resource establishes an extensible benchmarking framework and identifies concrete directions for future research in RL for software engineering.
Traditional deep reinforcement learning (DRL) assumes perfect action execution, neglecting system dynamics, hardware constraints, and actuation delays—leading to substantial performance degradation upon real-world deployment. To address this, we propose a control-aware DRL framework featuring a novel two-stage action mechanism: first generating a desired action, then producing a compensatory control signal via a learnable controller; execution errors and their dynamic compensation are explicitly modeled during training. Leveraging control-theoretic principles, we refactor five open-source robotic simulation environments to incorporate realistic uncertainties, enabling precise error modeling and online correction. Experiments demonstrate significant improvements in policy robustness and task success rates across diverse perturbation scenarios. Our approach provides a scalable, engineering-friendly solution for reliable decision-making in autonomous systems operating under non-ideal actuation conditions.
This paper addresses the persistent gap between theoretical advances in reinforcement learning (RL) and their practical deployment in robotics and control systems. To bridge this divide, we propose a structured taxonomy tailored to real-world robotic applications, grounded in the Markov decision process (MDP) framework and systematically incorporating mainstream deep RL algorithms—including DDPG, TD3, PPO, and SAC—across canonical domains such as motion control, dexterous manipulation, and multi-agent coordination. The taxonomy explicitly integrates training paradigms and deployment maturity metrics. Crucially, we identify recurring design patterns and evolutionary trends in high-dimensional continuous control tasks, thereby unifying theoretical insights with engineering constraints. Our framework advances reproducibility, transferability, and robustness in RL deployment on physical robots, offering both a methodological foundation and actionable guidelines for practitioners. (149 words)
Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.