Score
Designs, builds, and evaluates training pipelines, domain-adaptation algorithms, and simulator-to-sensor models that enable models and control policies trained on synthetic or simulated data to operate reliably on real sensor-equipped systems; this includes pretrain-and-finetune approaches, sensor-aware simulation, and visuomotor policy transfer. It also covers methods and metrics for reducing simulation-to-reality domain shift, minimizing real-data annotation needs, and deploying or zero‑shot testing learned models and controllers in real environments.
Deep reinforcement learning (DRL) policies trained in simulation often suffer from performance degradation and safety risks when deployed in the real world—commonly termed the sim-to-real gap. Method: This paper proposes the first unified taxonomy for sim-to-real transfer, grounded in the four core components of Markov Decision Processes (MDPs): states, actions, transitions, and rewards. It systematically organizes existing techniques—including domain adaptation, representation alignment, simulation modeling, and foundation-model–guided policy transfer (e.g., LLMs and multimodal models)—along this MDP-centric axis. Contribution/Results: We introduce the first openly maintained knowledge graph and reproducible benchmark platform for sim-to-real research, featuring over 100 algorithms and standardized evaluation protocols. Our analysis identifies six fundamental open challenges and uncovers novel pathways for leveraging foundation models to bridge the sim-to-real gap, thereby providing both theoretical foundations and practical paradigms for robust cross-domain policy transfer.
This work addresses the performance degradation of locomotion policies in quadrupedal robots during sim-to-real transfer caused by dynamics mismatches. To bridge this gap, the authors propose a simulator adaptation method based on proprioceptive distribution matching, which aligns the distributions of observations and actions between simulation and real hardware without requiring temporal alignment, external sensing, motion capture, or precise initial conditions. The approach jointly optimizes simulator dynamics parameters through parameter identification, an action delta model, and a residual actuator model to achieve efficient adaptation. Experiments on the Go2 robot demonstrate that with less than five minutes of real-world data, the method substantially reduces trajectory drift and significantly enhances policy performance—even enabling challenging bipedal walking tasks.
This work addresses the sim-to-real transfer failure commonly encountered in reinforcement learning due to mismatches between idealized actuator models used in simulation and the nonlinear, hardware-dependent motor dynamics of real robots. To bridge this gap, the authors propose “actuator reality shaping,” a method that deploys a two-degree-of-freedom feedforward–feedback controller on physical hardware to shape the closed-loop actuator response to closely match an ideal second-order reference model assumed in simulation. Notably, this approach requires no system identification or learned actuator models and enables zero-shot policy deployment through a standardized actuator interface. Experiments across diverse platforms—including single-joint servos, a 7-DoF manipulator, wheeled-legged robots, and humanoids—demonstrate substantial reductions in tracking error and successful zero-shot transfer across multiple tasks and systems, thereby shifting the paradigm from increasing simulation fidelity to unifying real-world actuator behavior to conform to simulation assumptions.
This work addresses the performance degradation commonly observed in sim-to-real transfer of reinforcement learning for robotic navigation, which stems from domain discrepancies and a lack of systematic analysis linking training strategies to deployment outcomes. The authors propose an end-to-end training and deployment pipeline that decouples key influencing factors, introducing perturbation-aware fine-tuning and a Transformer-based temporal reasoning policy to significantly enhance zero-shot transfer robustness and control smoothness. By integrating perturbation modeling, post-training fine-tuning, and system-level domain gap analysis, the method outperforms existing learning-based baselines in both static and dynamic environments, matching the performance of optimization-based planners in static scenes and achieving successful zero-shot deployment across multiple real-world robotic platforms.
To address key bottlenecks in simulation-to-reality (Sim2Real) transfer—including poor adaptability in dynamic, complex environments, lengthy deployment cycles, and strong reliance on high-fidelity simulation—this paper proposes a novel “abstraction-to-reality” paradigm. Rather than narrowing the sim-to-real gap, our approach extracts semantic commonalities across diverse low-fidelity simulation sources to construct a generalized autonomous stack with cross-simulation semantic alignment. The method integrates semantic abstraction representation, multi-simulation joint training, online self-adaptation, lightweight dynamic optimization, and cross-domain meta-policy distillation. Experiments demonstrate substantial improvements in real-time adaptation to unseen dynamic environments, reduce deployment time by several orders of magnitude, and validate strong generalization across heterogeneous robot platforms and varying simulation fidelity levels.
This work addresses the sim-to-real performance gap in visual navigation, where policies trained in simulation underperform when deployed on real-world platforms. We propose a novel approach integrating pretrained visual representations, end-to-end deep reinforcement learning, and a lightweight real-time inference architecture, leveraging online learning from simulated data to enhance cross-domain generalization. Key contributions include: (i) freezing a pretrained image encoder to extract robust visual features that mitigate appearance discrepancies between simulation and reality; and (ii) retaining the online adaptation mechanism of the simulated policy—enabling it to surpass purely real-data-trained baselines. On wheeled robots, our method achieves a 31% higher success rate than real-data training and outperforms current state-of-the-art methods by 50%. Furthermore, it successfully transfers to quadrotor drones without architectural modification, empirically validating its cross-platform generalizability.
This work addresses the failure of policy transfer from simulation to reality caused by discrepancies in visual rendering and physical dynamics. To bridge this sim-to-real gap, the authors propose a cross-domain dual imitation learning approach that leverages a shared history encoder to jointly model domain-invariant features across visual and dynamical domains in observation space. By mapping observation–action sequences that yield equivalent long-term behaviors to nearby latent states, the method enables zero-shot sim-to-real policy transfer without requiring separate adaptation modules. Evaluated on visually guided navigation, contact-rich manipulation, and visual servoing tasks, the approach substantially outperforms existing domain adaptation and co-training baselines, achieving the first end-to-end, highly effective zero-shot sim-to-real transfer.
Addressing the sim-to-real transfer challenge in deep reinforcement learning (DRL) for bipedal robots, this paper systematically analyzes simulation discrepancies arising from dynamics modeling, contact dynamics, state estimation, and numerical solvers. We propose a dual-track协同 framework integrating “model-centric calibration” and “policy robustification.” Specifically, we develop a simulation error diagnostic framework, a physics-based simulation calibration mechanism, domain randomization combined with online adaptive training, and integrate robust control with high-fidelity contact modeling. These components jointly enhance policy generalizability and robustness in real-world deployment. Experimental results demonstrate that our approach enables stable locomotion of bipedal robots on unseen complex terrains, reducing the sim-to-real performance gap by over 40%. The method provides a systematic, reusable solution for practical sim-to-real deployment of DRL-based locomotion controllers.
This work addresses the challenge of sim-to-real policy transfer failures caused by unobservable dynamics—such as abrupt contacts—by introducing an inverse dynamics extraction mechanism that recovers implicit dynamical information from real-world transition data. The approach formulates dynamics transfer between simulation and reality as an unpaired domain translation task, preserving domain-specific styles while enabling effective cross-domain adaptation. By integrating physics-based simulation with real robot data, it overcomes the limitations of conventional methods that rely solely on observed history to infer latent variables. Experimental validation across humanoid, quadrupedal, and robotic arm platforms demonstrates substantially improved dynamics modeling accuracy, particularly in scenarios where observation history is insufficient or misleading. Real-world trials on the Go2 quadruped further confirm a marked enhancement in policy transfer performance.
Autonomous driving in the real world faces significant challenges, including data scarcity, stringent safety constraints, and limited generalization across diverse environments. This work presents a systematic review of synthetic data and virtual simulation techniques applied to perception, planning, and validation in autonomous systems. It proposes an integrated three-dimensional framework that combines synthetic data generation, digital twin–based validation, and domain adaptation, further enhanced by vision–language models to improve simulation fidelity and semantic generalization. By establishing a comprehensive taxonomy of current methodologies, the study identifies critical research directions—such as safety verification, cooperative autonomy, and simulation-driven policy learning—to advance the development of scalable, safe, and generalizable autonomous driving systems.
Existing autonomous driving datasets lack sufficient diversity, coordination, and cross-domain support, limiting their utility for training multi-agent, multi-sensor systems. To address this gap, this work proposes a modular data generation pipeline built upon the AVstack framework and the CARLA simulator, capable of efficiently producing terabyte-scale, ground-truth-annotated multimodal data. The pipeline encompasses perspectives from ground vehicles, aerial platforms, and infrastructure sensors, and supports flexible single- or multi-agent configurations under controllable, complex scenarios. This approach represents the first scalable, cross-domain collaborative data generation methodology for autonomous driving, substantially enhancing the customization, training efficacy, and practical applicability of perception and sensor fusion models in cooperative autonomous systems.