reinforcement learning

Designs, implements, and evaluates agents, policies, and training pipelines that optimize sequential-decision objectives using reinforcement learning methods (e.g., policy-gradient, Q-learning, distributional, model-free, online, and sample-efficient algorithms). This work includes formulating reward signals (including validator-guided or validator-rewarded shaping), integrating RL into control or system workflows, running step- and episode-level experiments, and analyzing algorithmic performance and sample efficiency.

reinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-3.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$254K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Statistical and Algorithmic Foundations of Reinforcement Learning

Jul 18, 2025
YC
Yuejie Chi
🏛️ Yale University | University of Pennsylvania

This work investigates the statistical and algorithmic foundations of reinforcement learning (RL) under sample scarcity, aiming to improve sample and computational efficiency. Motivated by real-world constraints—such as expensive data acquisition and high-stakes decision-making—it systematically analyzes major RL paradigms: simulator-based, online, offline, robust, and human-feedback-driven RL, all modeled as Markov decision processes. A unified theoretical framework is developed to characterize the sample complexity and convergence rates of model-based, value-based, and policy-optimization methods. Innovatively, the study establishes a non-asymptotic, algorithm-dependent analysis framework tightly coupled with information-theoretic lower bounds. This yields provably efficient algorithms with sharp, instance-dependent guarantees across diverse settings. The results provide rigorous theoretical foundations and practical design principles for low-sample, robust decision-making systems in safety-critical domains—including healthcare and robotics—where data efficiency and reliability are paramount.

Addressing computational challenges in complex nonconvex RL modelsEnhancing sample efficiency in data-scarce RL scenariosEstablishing theoretical bounds for RL algorithms' performance

Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Jul 22, 2024
SI
Sinan Ibrahim
🏛️ Skolkovo Institute of Science and Technology | Innopolis University

Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)

Complex Real-World ProblemsReinforcement LearningReward Mechanism

This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.

function approximationMarkov decision processesmathematical foundations

Potential-based reward shaping (PBRS) introduces bias in finite-horizon settings, degrading sample efficiency. Method: This paper presents the first systematic analysis of PBRS bias under finite horizons and proposes an automatic potential function construction method integrated with state abstraction: interpretable state abstractions yield theoretically grounded potential functions that intrinsically suppress bias at its source. The approach replaces CNNs with a lightweight fully connected architecture. Results: Evaluated on navigation tasks and three ALE games, it matches CNN-based baselines in performance while reducing model complexity by ~60% in parameter count and improving sample efficiency by 2.1–3.4×. The method achieves synergistic optimization of theoretical interpretability and empirical sample efficiency.

Addressing bias from finite horizons in PBRSChoosing optimal potential function for PBRS in RLImproving sample efficiency using abstractions in RL

Regret-Free Reinforcement Learning for LTL Specifications

Nov 18, 2024
RM
R. Majumdar
🏛️ MPI-SWS

This work addresses the online synthesis of control policies satisfying Linear Temporal Logic (LTL) specifications for safety-critical systems operating under unknown Markov Decision Processes (MDPs). Existing approaches provide only asymptotic performance guarantees and lack instantaneous performance assurances during learning. To overcome this limitation, we propose the first online no-regret reinforcement learning algorithm applicable to arbitrary LTL specifications. Our method reformulates LTL synthesis as a reach-avoid graph game and introduces a dedicated probabilistic graph structure learning module, integrated with MDP modeling, LTL automaton construction, and hierarchical control synthesis. We theoretically prove that the algorithm achieves zero cumulative regret within a finite number of steps, delivering rigorous, verifiable finite-time performance guarantees for any LTL specification over finite-state/finite-action MDPs—thereby breaking the reliance on asymptotic convergence inherent in prior methods.

Learn regret-free control for LTL specifications with unknown dynamicsProvide finite-time performance bounds for LTL controller synthesisReduce LTL synthesis to reach-avoid problems using MDPs

Latest Papers

What's happening recently
View more

Current reinforcement learning (RL) post-training of large language models (LLMs) is overly focused on policy gradient methods such as PPO and GRPO, largely neglecting the broader RL algorithmic landscape. This work proposes a modular analytical framework centered on three core dimensions—MDP formulation, exploration strategies, and learning mechanisms—and systematically maps classical RL techniques—including value functions, off-policy learning, bootstrapped credit assignment, intrinsic motivation, tree search, and curriculum learning—onto the LLM training context for the first time. The study reveals a predominant reliance in existing approaches on actor-only, Monte Carlo–style policy optimization and explicitly identifies underexplored yet promising directions, thereby offering a clear roadmap for future algorithmic innovation in LLM alignment and training.

Credit AssignmentExplorationLarge Language Models

This work addresses the challenge of efficiently learning Pareto-optimal policies in multi-objective reinforcement learning under non-Markovian environments. It introduces Reward Machines (RMs) into the multi-objective reinforcement learning framework for the first time, integrating them with Pareto Q-learning (PQL). By leveraging the automaton-based decomposition of reward structures offered by RMs, the method maintains a set of vector-valued Q-function estimates in the cross-product MDP to approximate the Pareto front. This approach significantly improves sample efficiency, overcoming the limitation of conventional QRM methods that cannot handle multi-objective optimization, and enables the synthesis of Pareto-optimal policies inaccessible to standard QRM. Experimental results demonstrate that the proposed method converges faster and achieves superior performance compared to a naive PQL baseline directly applied to the cross-product MDP.

multi-objective reinforcement learningnon-Markovian rewardsPareto optimality

This work addresses the challenge of efficiently solving reinforcement learning tasks subject to complex temporal logic constraints by proposing a novel framework that integrates Signal Temporal Logic (STL) into reward machines. The approach leverages STL specifications to generate events and construct structured rewards, dynamically guiding the agent’s policy learning through an online STL monitoring algorithm to satisfy formal specifications. As the first study to combine STL with reward machines, it achieves compact reward representations and efficient training for intricate tasks. Empirical evaluations in Minigrid, Cart-Pole, and Highway environments demonstrate the method’s effectiveness and strong generalization capabilities on non-trivial tasks requiring precise temporal reasoning.

Complex TasksReinforcement LearningReward Machines

This work aims to bridge the gap between reinforcement learning and dynamic programming in terms of objective formulation, modeling assumptions, and optimization criteria. By developing a derandomized reinforcement learning framework, it establishes both theoretical and empirical connections to value iteration and Dijkstra’s algorithm, thereby unifying cost-minimization and reward-maximization paradigms. The core contributions include identifying the equivalence conditions between these two objective formulations, demonstrating the equivalence between single-episode termination tasks and infinite-horizon learning settings, and proposing an optimization objective centered on true cost. The approach is validated in both deterministic and stochastic environments, with precise conditions under which discounting leads to objective misalignment clearly characterized. Performance comparisons are enabled through planning-oriented evaluation metrics.

Cost MinimizationDynamic ProgrammingOptimal Planning

This work addresses the high computational cost of high-fidelity simulation models, which hinders efficient training and retraining of reinforcement learning agents in dynamic environments. To overcome this limitation, the paper proposes a learnable surrogate modeling framework tailored for dynamic settings, which approximates the input–output mapping of high-fidelity simulations to substantially reduce training overhead when system dynamics, parameters, or reward structures change. By integrating discrete-event simulation, reinforcement learning, and data-driven surrogate modeling techniques, the framework enables rapid adaptation of policies to environmental shifts. Empirical evaluation in stochastic service systems demonstrates significant acceleration in both initial training and retraining processes, thereby enhancing the adaptability of reinforcement learning policies to evolving conditions.

Discrete-Event SimulationReinforcement LearningSimulation Surrogate Models

Hot Scholars

AK

Aviral Kumar

Carnegie Mellon University
AIReinforcement Learning
MH

Marco Hutter

Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl
WX

Wayne Xin Zhao

Professor, Renmin University of China
Recommender SystemNatural Language ProcessingLarge Language Model
JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
HM

Haitao Mi

Principal Researcher, Tencent US
Large Language Models