Score
Designs and trains control policies end-to-end in simulation and develops methods to transfer those policies to real-world systems with little or no real-world fine-tuning. This involves building simulators and training pipelines, reducing the simulation-to-reality gap (e.g., via domain randomization or adaptation), and evaluating real-world success and data-efficiency to minimize costly real data collection.
Deep reinforcement learning (DRL) policies trained in simulation often suffer from performance degradation and safety risks when deployed in the real world—commonly termed the sim-to-real gap. Method: This paper proposes the first unified taxonomy for sim-to-real transfer, grounded in the four core components of Markov Decision Processes (MDPs): states, actions, transitions, and rewards. It systematically organizes existing techniques—including domain adaptation, representation alignment, simulation modeling, and foundation-model–guided policy transfer (e.g., LLMs and multimodal models)—along this MDP-centric axis. Contribution/Results: We introduce the first openly maintained knowledge graph and reproducible benchmark platform for sim-to-real research, featuring over 100 algorithms and standardized evaluation protocols. Our analysis identifies six fundamental open challenges and uncovers novel pathways for leveraging foundation models to bridge the sim-to-real gap, thereby providing both theoretical foundations and practical paradigms for robust cross-domain policy transfer.
In reinforcement learning, policies trained in simulation often suffer significant performance degradation when deployed in the real world—termed the Sim2Real gap. Existing approaches optimize simulators using proxy metrics (e.g., simulation fidelity or variability), which exhibit weak correlation with actual real-world performance. To address this, we propose a bilevel reinforcement learning framework that directly optimizes for real-world performance: the inner loop trains the policy in simulation, while the outer loop jointly adapts simulator parameters and the reward function based on real-world feedback. This eliminates reliance on imperfect proxies and enables adaptive calibration of both the dynamics model and reward structure. Theoretical analysis establishes convergence guarantees under mild assumptions, and extensive experiments across robotic control benchmarks demonstrate that our method substantially narrows the Sim2Real gap and significantly improves policy generalization to physical environments.
This work addresses the challenge of robust sim-to-real transfer for reinforcement learning policies that converge in simulation but fail to generalize reliably on physical robots. We propose a post-convergence robust transfer paradigm that replaces heuristic, simulation-performance-based policy selection with a theoretically grounded optimization framework. Specifically, we formulate policy selection as a convex quadratically constrained linear program, optimizing for worst-case real-world performance—thereby providing provable robustness guarantees. Our method integrates convex optimization, worst-case performance modeling, and empirical policy evaluation, eliminating ad hoc “cherry-picking.” Evaluated on legged robot locomotion control tasks, the approach significantly improves deployment success rates in the real world. Experiments demonstrate consistent superiority over conventional selection strategies that prioritize policies with the highest simulated reward.
Addressing the sim-to-real transfer challenge in deep reinforcement learning (DRL) for bipedal robots, this paper systematically analyzes simulation discrepancies arising from dynamics modeling, contact dynamics, state estimation, and numerical solvers. We propose a dual-track协同 framework integrating “model-centric calibration” and “policy robustification.” Specifically, we develop a simulation error diagnostic framework, a physics-based simulation calibration mechanism, domain randomization combined with online adaptive training, and integrate robust control with high-fidelity contact modeling. These components jointly enhance policy generalizability and robustness in real-world deployment. Experimental results demonstrate that our approach enables stable locomotion of bipedal robots on unseen complex terrains, reducing the sim-to-real performance gap by over 40%. The method provides a systematic, reusable solution for practical sim-to-real deployment of DRL-based locomotion controllers.
To address insufficient policy generalization in sim-to-real transfer due to dynamical discrepancies, this paper proposes a context-aware reinforcement learning framework. The method enables adaptive control in unknown real-world environments by online estimating physical dynamics parameters—such as friction, mass, and inertia—and feeding them as conditional inputs to the policy network. It integrates domain randomization, state inference, and conditional policy networks, incorporating a learnable dynamic context encoding module during training. Evaluated on standard control benchmarks (CartPole, Reacher) and a real robotic pushing task, the approach significantly outperforms context-agnostic baselines, achieving an average 32.7% improvement in task success rate under unseen dynamical configurations, while maintaining real-time inference capability. The core contribution lies in explicitly modeling implicit dynamics via a lightweight, online context estimation mechanism—and empirically demonstrating its critical role in enhancing cross-domain robustness.
This study addresses the sim-to-real gap arising from calibration data confusion and drift in pretrained simulators by investigating how to effectively integrate cheap but biased simulations with expensive yet unbiased real-world experiments in sequential decision-making. The authors extend the simulation lemma to decompose policy value error, revealing that under passive learning, the reachability gap is irreducible. To overcome this limitation, they propose Fisher-SEP, an active experimentation strategy based on the Fisher information matrix, which minimizes the predictive variance of the target policy’s value through Bayesian posterior inference. Empirical validation on two real-world domains—vending machine supply chains and mobile HIV testing—demonstrates that early real-world trials yield substantial long-term benefits in the former, while only active exploration effectively covers low-monitoring regions in the latter.
This study addresses the optimal allocation of limited real-world measurement time between system identification and domain randomization to enhance sim-to-real transfer performance in robot learning. Through controlled simulation-to-simulation experiments on a pendulum system, the work presents the first quantitative analysis of the trade-off between parameter identification accuracy and the breadth of domain randomization under a fixed real-data budget. The results demonstrate that, under identifiable dynamics, even a small amount of real data used for precise system identification significantly reduces the reality gap. In contrast, broad domain randomization—even when encompassing the true system parameters—fails to match the effectiveness of accurate parameter estimation. These findings reveal a key strategy for efficiently leveraging scarce real-world data and establish a new paradigm for improving sim-to-real transfer.
This work addresses the challenge of sim-to-real policy transfer failures caused by unobservable dynamics—such as abrupt contacts—by introducing an inverse dynamics extraction mechanism that recovers implicit dynamical information from real-world transition data. The approach formulates dynamics transfer between simulation and reality as an unpaired domain translation task, preserving domain-specific styles while enabling effective cross-domain adaptation. By integrating physics-based simulation with real robot data, it overcomes the limitations of conventional methods that rely solely on observed history to infer latent variables. Experimental validation across humanoid, quadrupedal, and robotic arm platforms demonstrates substantially improved dynamics modeling accuracy, particularly in scenarios where observation history is insufficient or misleading. Real-world trials on the Go2 quadruped further confirm a marked enhancement in policy transfer performance.
This work addresses the performance degradation of reinforcement learning policies when transferring from simulation to real-world robotic systems, particularly in contact-rich tasks such as cutting unknown materials, where domain discrepancies and scarce real-world data—especially the absence of ground-truth reward signals—pose significant challenges. To overcome this, the authors propose an end-to-end, example-driven approach that reinterprets neural style transfer as temporal trajectory stylization. By integrating variational autoencoders with self-supervised representation learning, the method leverages unpaired and unlabeled real-world data to generate physically plausible, weakly aligned trajectories, enabling policy adaptation without requiring real reward signals. Experiments demonstrate that the approach substantially outperforms baselines such as CycleGAN and conditional variational autoencoders across diverse geometries and materials, achieving high task success rates and behavioral stability with only minimal real-world data.
This work addresses the limitations of conventional sim2real transfer, where excessive alignment between simulation and reality often restricts policy learning, hampers exploration, and leads to simulator overfitting. To overcome these issues, the authors propose a novel “sim2sim2real” paradigm that leverages only robotic kinematics as a constraint, thereby preserving real-world feasibility while substantially enhancing policy flexibility and exploratory capacity. By integrating a lightweight transfer framework with purposefully designed simulation environments, the approach effectively mitigates simulator locking and significantly improves both generalization performance and training efficiency when deploying policies on real hardware.