Score
Designs, implements, and evaluates agents, policies, and training pipelines that optimize sequential-decision objectives using reinforcement learning methods (e.g., policy-gradient, Q-learning, distributional, model-free, online, and sample-efficient algorithms). This work includes formulating reward signals (including validator-guided or validator-rewarded shaping), integrating RL into control or system workflows, running step- and episode-level experiments, and analyzing algorithmic performance and sample efficiency.
Reinforcement learning (RL) faces critical challenges in real-world deployment, including poor scalability, low sample efficiency, insufficient training stability, and suboptimal exploration-exploitation trade-offs. Method: This work systematically surveys over 120 RL algorithms and introduces the first multidimensional unified evaluation framework integrating theoretical properties with engineering requirements—spanning tabular methods to deep RL approaches (e.g., DQN, PPO, A3C). It proposes a deployment-oriented algorithm selection guideline, specifying adaptation strategies for seven representative application scenarios. Contribution/Results: Through extensive empirical evaluation and case studies, the framework quantifies algorithmic performance across scalability, convergence rate, stability, and sample efficiency. The study bridges the gap between academic survey and industrial adoption, and its outputs have been adopted as core reference material in university AI curricula and as the de facto standard introductory resource for RL practitioners in industry.
This work investigates the statistical and algorithmic foundations of reinforcement learning (RL) under sample scarcity, aiming to improve sample and computational efficiency. Motivated by real-world constraints—such as expensive data acquisition and high-stakes decision-making—it systematically analyzes major RL paradigms: simulator-based, online, offline, robust, and human-feedback-driven RL, all modeled as Markov decision processes. A unified theoretical framework is developed to characterize the sample complexity and convergence rates of model-based, value-based, and policy-optimization methods. Innovatively, the study establishes a non-asymptotic, algorithm-dependent analysis framework tightly coupled with information-theoretic lower bounds. This yields provably efficient algorithms with sharp, instance-dependent guarantees across diverse settings. The results provide rigorous theoretical foundations and practical design principles for low-sample, robust decision-making systems in safety-critical domains—including healthcare and robotics—where data efficiency and reliability are paramount.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
Potential-based reward shaping (PBRS) introduces bias in finite-horizon settings, degrading sample efficiency. Method: This paper presents the first systematic analysis of PBRS bias under finite horizons and proposes an automatic potential function construction method integrated with state abstraction: interpretable state abstractions yield theoretically grounded potential functions that intrinsically suppress bias at its source. The approach replaces CNNs with a lightweight fully connected architecture. Results: Evaluated on navigation tasks and three ALE games, it matches CNN-based baselines in performance while reducing model complexity by ~60% in parameter count and improving sample efficiency by 2.1–3.4×. The method achieves synergistic optimization of theoretical interpretability and empirical sample efficiency.
This work addresses the online synthesis of control policies satisfying Linear Temporal Logic (LTL) specifications for safety-critical systems operating under unknown Markov Decision Processes (MDPs). Existing approaches provide only asymptotic performance guarantees and lack instantaneous performance assurances during learning. To overcome this limitation, we propose the first online no-regret reinforcement learning algorithm applicable to arbitrary LTL specifications. Our method reformulates LTL synthesis as a reach-avoid graph game and introduces a dedicated probabilistic graph structure learning module, integrated with MDP modeling, LTL automaton construction, and hierarchical control synthesis. We theoretically prove that the algorithm achieves zero cumulative regret within a finite number of steps, delivering rigorous, verifiable finite-time performance guarantees for any LTL specification over finite-state/finite-action MDPs—thereby breaking the reliance on asymptotic convergence inherent in prior methods.
Current reinforcement learning (RL) post-training of large language models (LLMs) is overly focused on policy gradient methods such as PPO and GRPO, largely neglecting the broader RL algorithmic landscape. This work proposes a modular analytical framework centered on three core dimensions—MDP formulation, exploration strategies, and learning mechanisms—and systematically maps classical RL techniques—including value functions, off-policy learning, bootstrapped credit assignment, intrinsic motivation, tree search, and curriculum learning—onto the LLM training context for the first time. The study reveals a predominant reliance in existing approaches on actor-only, Monte Carlo–style policy optimization and explicitly identifies underexplored yet promising directions, thereby offering a clear roadmap for future algorithmic innovation in LLM alignment and training.
This work addresses the challenge of efficiently learning Pareto-optimal policies in multi-objective reinforcement learning under non-Markovian environments. It introduces Reward Machines (RMs) into the multi-objective reinforcement learning framework for the first time, integrating them with Pareto Q-learning (PQL). By leveraging the automaton-based decomposition of reward structures offered by RMs, the method maintains a set of vector-valued Q-function estimates in the cross-product MDP to approximate the Pareto front. This approach significantly improves sample efficiency, overcoming the limitation of conventional QRM methods that cannot handle multi-objective optimization, and enables the synthesis of Pareto-optimal policies inaccessible to standard QRM. Experimental results demonstrate that the proposed method converges faster and achieves superior performance compared to a naive PQL baseline directly applied to the cross-product MDP.
This work addresses the challenge of efficiently solving reinforcement learning tasks subject to complex temporal logic constraints by proposing a novel framework that integrates Signal Temporal Logic (STL) into reward machines. The approach leverages STL specifications to generate events and construct structured rewards, dynamically guiding the agent’s policy learning through an online STL monitoring algorithm to satisfy formal specifications. As the first study to combine STL with reward machines, it achieves compact reward representations and efficient training for intricate tasks. Empirical evaluations in Minigrid, Cart-Pole, and Highway environments demonstrate the method’s effectiveness and strong generalization capabilities on non-trivial tasks requiring precise temporal reasoning.
This work aims to bridge the gap between reinforcement learning and dynamic programming in terms of objective formulation, modeling assumptions, and optimization criteria. By developing a derandomized reinforcement learning framework, it establishes both theoretical and empirical connections to value iteration and Dijkstra’s algorithm, thereby unifying cost-minimization and reward-maximization paradigms. The core contributions include identifying the equivalence conditions between these two objective formulations, demonstrating the equivalence between single-episode termination tasks and infinite-horizon learning settings, and proposing an optimization objective centered on true cost. The approach is validated in both deterministic and stochastic environments, with precise conditions under which discounting leads to objective misalignment clearly characterized. Performance comparisons are enabled through planning-oriented evaluation metrics.
This work addresses the high computational cost of high-fidelity simulation models, which hinders efficient training and retraining of reinforcement learning agents in dynamic environments. To overcome this limitation, the paper proposes a learnable surrogate modeling framework tailored for dynamic settings, which approximates the input–output mapping of high-fidelity simulations to substantially reduce training overhead when system dynamics, parameters, or reward structures change. By integrating discrete-event simulation, reinforcement learning, and data-driven surrogate modeling techniques, the framework enables rapid adaptation of policies to environmental shifts. Empirical evaluation in stochastic service systems demonstrates significant acceleration in both initial training and retraining processes, thereby enhancing the adaptability of reinforcement learning policies to evolving conditions.