Score
Designs, implements, and evaluates reinforcement learning systems that use a simulator integrated directly into the training or control loop so that simulator-generated observations, rewards, and constraint signals guide policy learning and decision-making. Builds simulation-in-the-loop training pipelines, defines simulator-based reward/constraint functions, and analyzes policy behavior to improve operational objectives (e.g., efficiency) while preserving correctness and safe transfer to execution outside the simulator.
To address safety risks, high training costs, and the sim-to-real gap in deploying reinforcement learning (RL) on physical robots, this paper proposes a four-stage progressive RL training framework: system identification → core simulation training → high-fidelity simulation → real-robot deployment. The framework integrates domain randomization, policy distillation, and online fine-tuning, and is implemented using PyTorch, MuJoCo, and ROS2, specifically optimized for the Boston Dynamics Spot platform. Its key innovation lies in enabling cross-fidelity policy transfer and iterative refinement, substantially improving sim-to-real generalization. Evaluated on a robotic inspection task, the approach achieves high-precision control of position and orientation, with >92% success rate in real-world deployment, 60% reduction in training cost, and 3.5× faster convergence compared to baseline methods.
This work addresses the high computational cost of high-fidelity simulation models, which hinders efficient training and retraining of reinforcement learning agents in dynamic environments. To overcome this limitation, the paper proposes a learnable surrogate modeling framework tailored for dynamic settings, which approximates the input–output mapping of high-fidelity simulations to substantially reduce training overhead when system dynamics, parameters, or reward structures change. By integrating discrete-event simulation, reinforcement learning, and data-driven surrogate modeling techniques, the framework enables rapid adaptation of policies to environmental shifts. Empirical evaluation in stochastic service systems demonstrates significant acceleration in both initial training and retraining processes, thereby enhancing the adaptability of reinforcement learning policies to evolving conditions.
This study addresses a prevalent conflation in reinforcement learning research between two distinct objectives involving simulators: solving the simulator as an end in itself versus treating it as a proxy for a real-world deployment environment. The former seeks high returns within the simulated domain, while the latter aims to transfer learned policies to the physical world. These goals entail fundamentally different algorithmic constraints, methodological requirements, and evaluation criteria. Through conceptual clarification, illustrative case studies, and controlled experiments, this work exposes the pitfalls arising from conflating these roles, delineates their respective appropriate use cases, and calls upon the research community to align experimental design, evaluation metrics, and algorithm development with the intended purpose of simulator usage.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
To address policy transfer failure caused by insufficient coverage in offline reinforcement learning (RL) data and dynamics discrepancies between simulation and reality, this paper proposes a hybrid RL framework that synergistically integrates limited real-world offline data with online exploration in an imperfect simulator. Its core innovation is a novel dynamics-aware Q-function adaptive penalty mechanism, which dynamically estimates and suppresses interference from high-bias simulated transitions—using either model residuals or contrastive encoding—to improve policy robustness. The framework jointly incorporates BCQ-style offline policy constraints, SAC-style online policy optimization, and explicit modeling of dynamics mismatch. Evaluated across multiple simulated and real-robot tasks, the method significantly outperforms pure offline, pure online, and cross-domain baselines. Theoretically, it yields a tighter upper bound on policy error, and empirically achieves a 2.3× improvement in sample efficiency.
This work investigates the impact of dynamics and reward model errors on policy performance in imagination-based trajectory learning. By extending MDP error analysis to settings with learned reward models, the authors reformulate policy optimization under reward noise as a one-dimensional problem, targeting representations with low Lipschitz constants. Integrating theoretical error bounds, REINFORCE gradient estimation, and sample complexity analysis, they derive the optimal allocation ratio between dynamics and reward samples under a fixed sampling budget. The analysis theoretically establishes that zero-mean reward noise introduces no bias and that its variance decays with the number of imagined trajectories, yielding clear practical guidelines for sample-efficient policy training.
This work addresses the limitations of conventional sim2real transfer, where excessive alignment between simulation and reality often restricts policy learning, hampers exploration, and leads to simulator overfitting. To overcome these issues, the authors propose a novel “sim2sim2real” paradigm that leverages only robotic kinematics as a constraint, thereby preserving real-world feasibility while substantially enhancing policy flexibility and exploratory capacity. By integrating a lightweight transfer framework with purposefully designed simulation environments, the approach effectively mitigates simulator locking and significantly improves both generalization performance and training efficiency when deploying policies on real hardware.
本文探讨了如何利用强化学习解决运筹学中的动态决策问题,通过与传统方法结合来提高解决方案的质量和效率,并为未来研究提供了路线图。
This work addresses the longstanding methodological, objective, and cultural divide between reinforcement learning and control theory by proposing a novel paradigm that integrates adaptive control with actor-critic reinforcement learning. The resulting framework enables data-driven optimization of controllers by unifying dynamic programming and online learning mechanisms, thereby reconciling modeling and optimization perspectives from both fields within classical motion control tasks. Theoretical analysis elucidates fundamental differences between the two approaches, while empirical results demonstrate the efficacy of the integrated strategy. This synthesis offers a solution for controlling systems with unknown dynamics that simultaneously guarantees stability and retains strong learning capabilities, fostering interoperability and synergistic development across disciplinary boundaries.
This study addresses the sim-to-real gap arising from calibration data confusion and drift in pretrained simulators by investigating how to effectively integrate cheap but biased simulations with expensive yet unbiased real-world experiments in sequential decision-making. The authors extend the simulation lemma to decompose policy value error, revealing that under passive learning, the reachability gap is irreducible. To overcome this limitation, they propose Fisher-SEP, an active experimentation strategy based on the Fisher information matrix, which minimizes the predictive variance of the target policy’s value through Bayesian posterior inference. Empirical validation on two real-world domains—vending machine supply chains and mobile HIV testing—demonstrates that early real-world trials yield substantial long-term benefits in the former, while only active exploration effectively covers low-monitoring regions in the latter.