sim-to-real benchmarking

Design and implement experimental frameworks that run identical tasks and metrics in both simulated and real environments to quantify how policies, models, or systems transfer from simulation to reality. Build unified task suites, metrics, and analysis pipelines to measure generalization, robustness, long-horizon behavior, precision, and other performance gaps between simulation and real-world execution.

sim-to-realbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In reinforcement learning, policies trained in simulation often suffer significant performance degradation when deployed in the real world—termed the Sim2Real gap. Existing approaches optimize simulators using proxy metrics (e.g., simulation fidelity or variability), which exhibit weak correlation with actual real-world performance. To address this, we propose a bilevel reinforcement learning framework that directly optimizes for real-world performance: the inner loop trains the policy in simulation, while the outer loop jointly adapts simulator parameters and the reward function based on real-world feedback. This eliminates reliance on imperfect proxies and enables adaptive calibration of both the dynamics model and reward structure. Theoretical analysis establishes convergence guarantees under mild assumptions, and extensive experiments across robotic control benchmarks demonstrate that our method substantially narrows the Sim2Real gap and significantly improves policy generalization to physical environments.

Addressing simulator accuracy limitations for real-world deploymentClosing Sim2Real performance gap in policy transferDeveloping bi-level RL framework for direct simulator adaptation

MultiCoSim: A Python-based Multi-Fidelity Co-Simulation Framework

Jun 12, 2025
QT
Quinn Thibeault
🏛️ Arizona State University

Existing CPS co-simulation tools suffer from limited portability, modularity, and automation. To address these limitations, this paper proposes a Python-based programmable co-simulation framework. The framework enables declarative orchestration and runtime dynamic substitution of multi-fidelity heterogeneous components—breaking away from conventional static configuration paradigms. It adopts a componentized architecture, supports distributed communication via ZeroMQ and ROS, and provides standardized adaptation interfaces for third-party platforms (e.g., PX4), thereby enabling cross-platform, reconfigurable co-simulation. Its core innovation is the first-ever declarative component orchestration mechanism, which significantly enhances simulation system reusability, reproducibility, and development efficiency. The framework is validated through co-simulation of unmanned aerial vehicles and autonomous controllers, demonstrating its flexibility and practicality. This work establishes a novel paradigm for CPS benchmark construction and automated evaluation.

Addresses rigid configurations and lack of automation in existing toolsEnables multi-fidelity co-simulation for complex cyber-physical systemsSupports flexible composition and integration of heterogeneous components

Accelerating scientific discovery with the common task framework

Nov 06, 2025
JK
J. Kutz
🏛️ University of Washington | Google DeepMind | Harvard University | Meta Platforms, Inc. | Los Alamos National Laboratory | Flatiron Institute | Google Research | Amazon Research | Columbia University | University of California at Berkeley | International Computer Science Institute | Lawrence Berkeley National Laboratory | Freie Universität Berlin | Microsoft Research | University of Amsterdam | New York University | University of Hawai‘i at Mānoa

The scientific and engineering communities lack unified, reproducible benchmarks for evaluating AI/ML methods in dynamical systems modeling. Method: This paper introduces the Common Task Framework (CTF), a general-purpose framework targeting multiple scientific objectives—including prediction, state reconstruction, generalization, and control—under realistic constraints of limited data and noisy measurements. CTF establishes standardized datasets, objective evaluation metrics, and an open benchmarking platform. Contribution/Results: CTF enables the first cross-disciplinary, physics-constrained comparison of system identification and machine learning algorithms, facilitating rapid iterative development and integration. Experimental results demonstrate that CTF significantly improves model development efficiency and deployment reliability, thereby addressing a critical gap in AI evaluation frameworks oriented toward scientific discovery.

Addressing limited data scenarios and noisy measurements in modelingCreating standardized challenge datasets for algorithm evaluation across disciplinesEstablishing comparative metrics for diverse scientific objectives in ML/AI

GRS: Generating Robotic Simulation Tasks from Real-World Images

Oct 20, 2024
AZ
Alex Zook
🏛️ NVIDIA | Stanford

Bridging the domain gap between real-world RGB-D images and robot simulation environments remains challenging for digital twin task generation. Method: This paper proposes a simulation-task alignment framework leveraging vision-language models (VLMs) and an iterative routing mechanism to generate executable simulation tasks end-to-end from single-frame RGB-D input. The method integrates SAM2 for precise object segmentation, VLM-driven semantic understanding, dynamic matching against a simulation asset library, and automated generation of self-validating test suites—forming a closed-loop “perceive–match–generate–verify” optimization pipeline. Contribution/Results: It achieves the first high-fidelity geometric-semantic alignment between real-scene objects and simulation assets while ensuring physical feasibility and executability within physics engines. Evaluated on multiple real-world benchmarks, the approach significantly improves object correspondence accuracy (+23.6%), task success rate (+31.4%), and cross-scene generalization.

Align simulation tasks using vision-language modelsConvert real-world images to robotic simulation tasksGenerate digital twin simulations from RGB-D observations

Coupled Requirements-Driven Testing of CPS: From Simulation to Reality

Mar 24, 2024
AA
Ankit Agrawal
🏛️ St. Louis University | University of Innsbruck

Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.

Cyber-Physical Systems TestingSafety Analysis IntegrationStandardization

Latest Papers

What's happening recently
View more

Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.

behavioral evaluationcross-benchmark generalizationlong-horizon agents

Current evaluations of vision-language-action (VLA) policies suffer from a lack of reliable correlation between simulation performance and real-world outcomes, limiting their utility in guiding practical deployment. This work presents the first systematic quantification of sim-to-real alignment across multiple simulation platforms, analyzing consistency in policy ranking, performance correlation, and failure modes under perturbations across diverse tasks. Through extensive multi-platform simulations, policy fine-tuning analyses, post-training data scaling studies, and robustness evaluations, the study identifies key characteristics of high-fidelity simulators and proposes design principles to enhance simulation utility. These findings substantially improve the reliability and practical guidance value of simulation in VLA policy development.

policy correlationroboticssim-to-real

This work addresses the challenge in scientific computing where models struggle to transfer knowledge from individual tasks to broader, reusable capabilities. The authors propose SciConsolidate, a novel framework that, for the first time, extracts procedural knowledge from runtime success and failure trajectories and employs a development-validation gating mechanism to filter effective knowledge. To bridge the gap between abstract knowledge and executable code, the approach integrates failure-driven synthesis of unanswerable queries with strong-model-guided concretization supervision. Additionally, a matched teacher-branch architecture is introduced to significantly enhance the performance of smaller models. Experimental results demonstrate that Qwen3.5-9B achieves gains of 6.25 and 3.89 points over program-free supervised fine-tuning—and improvements of 11.25 and 5.62 points over the base model—on main tasks and sub-steps, respectively, validating the efficacy of the proposed methodology.

abstraction-execution gapexperience consolidationprocedural knowledge

Hot Scholars

YD

Yilun Du

Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
DF

Dieter Fox

University of Washington and AI2
RoboticsArtificial IntelligenceComputer Vision
YL

Yunzhu Li

Columbia University
RoboticsComputer VisionMachine Learning
TM

Takamitsu Matsubara

Nara Institute of Science and Technology
Robot LearningMachine LearningReinforcement LearningRobotics