simulation-guided reinforcement learning

Designs, implements, and evaluates reinforcement learning systems that use a simulator integrated directly into the training or control loop so that simulator-generated observations, rewards, and constraint signals guide policy learning and decision-making. Builds simulation-in-the-loop training pipelines, defines simulator-based reward/constraint functions, and analyzes policy behavior to improve operational objectives (e.g., efficiency) while preserving correctness and safe transfer to execution outside the simulator.

simulation-guidedreinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Simulation Pipeline to Facilitate Real-World Robotic Reinforcement Learning Applications

Feb 21, 2025
JS
Jefferson Silveira
🏛️ Queen's University | Ingenuity Labs Research Institute

To address safety risks, high training costs, and the sim-to-real gap in deploying reinforcement learning (RL) on physical robots, this paper proposes a four-stage progressive RL training framework: system identification → core simulation training → high-fidelity simulation → real-robot deployment. The framework integrates domain randomization, policy distillation, and online fine-tuning, and is implemented using PyTorch, MuJoCo, and ROS2, specifically optimized for the Boston Dynamics Spot platform. Its key innovation lies in enabling cross-fidelity policy transfer and iterative refinement, substantially improving sim-to-real generalization. Evaluated on a robotic inspection task, the approach achieves high-precision control of position and orientation, with >92% success rate in real-world deployment, 60% reduction in training cost, and 3.5× faster convergence compared to baseline methods.

High costs and safety risks in physical robot trainingIterative policy improvement for real-world deploymentSimulation-to-reality gap in robotic reinforcement learning

This work addresses the high computational cost of high-fidelity simulation models, which hinders efficient training and retraining of reinforcement learning agents in dynamic environments. To overcome this limitation, the paper proposes a learnable surrogate modeling framework tailored for dynamic settings, which approximates the input–output mapping of high-fidelity simulations to substantially reduce training overhead when system dynamics, parameters, or reward structures change. By integrating discrete-event simulation, reinforcement learning, and data-driven surrogate modeling techniques, the framework enables rapid adaptation of policies to environmental shifts. Empirical evaluation in stochastic service systems demonstrates significant acceleration in both initial training and retraining processes, thereby enhancing the adaptability of reinforcement learning policies to evolving conditions.

Discrete-Event SimulationReinforcement LearningSimulation Surrogate Models

This study addresses a prevalent conflation in reinforcement learning research between two distinct objectives involving simulators: solving the simulator as an end in itself versus treating it as a proxy for a real-world deployment environment. The former seeks high returns within the simulated domain, while the latter aims to transfer learned policies to the physical world. These goals entail fundamentally different algorithmic constraints, methodological requirements, and evaluation criteria. Through conceptual clarification, illustrative case studies, and controlled experiments, this work exposes the pitfalls arising from conflating these roles, delineates their respective appropriate use cases, and calls upon the research community to align experimental design, evaluation metrics, and algorithm development with the intended purpose of simulator usage.

deploymentevaluation metricsreinforcement learning

Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Jul 22, 2024
SI
Sinan Ibrahim
🏛️ Skolkovo Institute of Science and Technology | Innopolis University

Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)

Complex Real-World ProblemsReinforcement LearningReward Mechanism

When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning

Jun 27, 2022
HN
Haoyi Niu
🏛️ Tsinghua University | Indian Institute of Technology, Bombay | Shanghai Jiaotong University | Beijing National Research Center for Information Science and Technology | Shanghai AI Laboratory

To address policy transfer failure caused by insufficient coverage in offline reinforcement learning (RL) data and dynamics discrepancies between simulation and reality, this paper proposes a hybrid RL framework that synergistically integrates limited real-world offline data with online exploration in an imperfect simulator. Its core innovation is a novel dynamics-aware Q-function adaptive penalty mechanism, which dynamically estimates and suppresses interference from high-bias simulated transitions—using either model residuals or contrastive encoding—to improve policy robustness. The framework jointly incorporates BCQ-style offline policy constraints, SAC-style online policy optimization, and explicit modeling of dynamics mismatch. Evaluated across multiple simulated and real-robot tasks, the method significantly outperforms pure offline, pure online, and cross-domain baselines. Theoretically, it yields a tighter upper bound on policy error, and empirically achieves a 2.3× improvement in sample efficiency.

Addressing sim-to-real gaps and insufficient offline data coverageCombining limited real data with imperfect simulators for reinforcement learningDeveloping dynamics-aware hybrid offline-online RL framework H2O

Latest Papers

What's happening recently
View more

This work investigates the impact of dynamics and reward model errors on policy performance in imagination-based trajectory learning. By extending MDP error analysis to settings with learned reward models, the authors reformulate policy optimization under reward noise as a one-dimensional problem, targeting representations with low Lipschitz constants. Integrating theoretical error bounds, REINFORCE gradient estimation, and sample complexity analysis, they derive the optimal allocation ratio between dynamics and reward samples under a fixed sampling budget. The analysis theoretically establishes that zero-mean reward noise introduces no bias and that its variance decays with the number of imagined trajectories, yielding clear practical guidelines for sample-efficient policy training.

dynamics model errorimagined rolloutsmodel-based reinforcement learning

This work addresses the limitations of conventional sim2real transfer, where excessive alignment between simulation and reality often restricts policy learning, hampers exploration, and leads to simulator overfitting. To overcome these issues, the authors propose a novel “sim2sim2real” paradigm that leverages only robotic kinematics as a constraint, thereby preserving real-world feasibility while substantially enhancing policy flexibility and exploratory capacity. By integrating a lightweight transfer framework with purposefully designed simulation environments, the approach effectively mitigates simulator locking and significantly improves both generalization performance and training efficiency when deploying policies on real hardware.

policy explorationpolicy learningreal-world constraints

This work addresses the longstanding methodological, objective, and cultural divide between reinforcement learning and control theory by proposing a novel paradigm that integrates adaptive control with actor-critic reinforcement learning. The resulting framework enables data-driven optimization of controllers by unifying dynamic programming and online learning mechanisms, thereby reconciling modeling and optimization perspectives from both fields within classical motion control tasks. Theoretical analysis elucidates fundamental differences between the two approaches, while empirical results demonstrate the efficacy of the integrated strategy. This synthesis offers a solution for controlling systems with unknown dynamics that simultaneously guarantees stability and retains strong learning capabilities, fostering interoperability and synergistic development across disciplinary boundaries.

Actor-Critic AlgorithmsAdaptive ControlControl Theory

This study addresses the sim-to-real gap arising from calibration data confusion and drift in pretrained simulators by investigating how to effectively integrate cheap but biased simulations with expensive yet unbiased real-world experiments in sequential decision-making. The authors extend the simulation lemma to decompose policy value error, revealing that under passive learning, the reachability gap is irreducible. To overcome this limitation, they propose Fisher-SEP, an active experimentation strategy based on the Fisher information matrix, which minimizes the predictive variance of the target policy’s value through Bayesian posterior inference. Empirical validation on two real-world domains—vending machine supply chains and mobile HIV testing—demonstrates that early real-world trials yield substantial long-term benefits in the former, while only active exploration effectively covers low-monitoring regions in the latter.

experimental designpolicy evaluationsequential decision making

Hot Scholars

FS

Fan Shi

Assistant Professor in National University of Singapore
Robotics
DR

Denis Rakitin

HSE University
Bayesian methodsDeep learningProbabilistic modeling
JZ

Jiayu Zhou

University of Michigan
Machine LearningAI + Health Informatics
RA

Raed Al Kontar

Associate Professor at University of Michigan
Data sciencedistributed learningpersonalizationuncertainty quantification
MS

Martin Smit

PhD in Artificial Intelligence, University of Amsterdam
multi-agent systemsreinforcement learning