Score
Design and implement experimental frameworks that run identical tasks and metrics in both simulated and real environments to quantify how policies, models, or systems transfer from simulation to reality. Build unified task suites, metrics, and analysis pipelines to measure generalization, robustness, long-horizon behavior, precision, and other performance gaps between simulation and real-world execution.
In reinforcement learning, policies trained in simulation often suffer significant performance degradation when deployed in the real world—termed the Sim2Real gap. Existing approaches optimize simulators using proxy metrics (e.g., simulation fidelity or variability), which exhibit weak correlation with actual real-world performance. To address this, we propose a bilevel reinforcement learning framework that directly optimizes for real-world performance: the inner loop trains the policy in simulation, while the outer loop jointly adapts simulator parameters and the reward function based on real-world feedback. This eliminates reliance on imperfect proxies and enables adaptive calibration of both the dynamics model and reward structure. Theoretical analysis establishes convergence guarantees under mild assumptions, and extensive experiments across robotic control benchmarks demonstrate that our method substantially narrows the Sim2Real gap and significantly improves policy generalization to physical environments.
Existing CPS co-simulation tools suffer from limited portability, modularity, and automation. To address these limitations, this paper proposes a Python-based programmable co-simulation framework. The framework enables declarative orchestration and runtime dynamic substitution of multi-fidelity heterogeneous components—breaking away from conventional static configuration paradigms. It adopts a componentized architecture, supports distributed communication via ZeroMQ and ROS, and provides standardized adaptation interfaces for third-party platforms (e.g., PX4), thereby enabling cross-platform, reconfigurable co-simulation. Its core innovation is the first-ever declarative component orchestration mechanism, which significantly enhances simulation system reusability, reproducibility, and development efficiency. The framework is validated through co-simulation of unmanned aerial vehicles and autonomous controllers, demonstrating its flexibility and practicality. This work establishes a novel paradigm for CPS benchmark construction and automated evaluation.
The scientific and engineering communities lack unified, reproducible benchmarks for evaluating AI/ML methods in dynamical systems modeling. Method: This paper introduces the Common Task Framework (CTF), a general-purpose framework targeting multiple scientific objectives—including prediction, state reconstruction, generalization, and control—under realistic constraints of limited data and noisy measurements. CTF establishes standardized datasets, objective evaluation metrics, and an open benchmarking platform. Contribution/Results: CTF enables the first cross-disciplinary, physics-constrained comparison of system identification and machine learning algorithms, facilitating rapid iterative development and integration. Experimental results demonstrate that CTF significantly improves model development efficiency and deployment reliability, thereby addressing a critical gap in AI evaluation frameworks oriented toward scientific discovery.
Bridging the domain gap between real-world RGB-D images and robot simulation environments remains challenging for digital twin task generation. Method: This paper proposes a simulation-task alignment framework leveraging vision-language models (VLMs) and an iterative routing mechanism to generate executable simulation tasks end-to-end from single-frame RGB-D input. The method integrates SAM2 for precise object segmentation, VLM-driven semantic understanding, dynamic matching against a simulation asset library, and automated generation of self-validating test suites—forming a closed-loop “perceive–match–generate–verify” optimization pipeline. Contribution/Results: It achieves the first high-fidelity geometric-semantic alignment between real-scene objects and simulation assets while ensuring physical feasibility and executability within physics engines. Evaluated on multiple real-world benchmarks, the approach significantly improves object correspondence accuracy (+23.6%), task success rate (+31.4%), and cross-scene generalization.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.
Current evaluations of vision-language-action (VLA) policies suffer from a lack of reliable correlation between simulation performance and real-world outcomes, limiting their utility in guiding practical deployment. This work presents the first systematic quantification of sim-to-real alignment across multiple simulation platforms, analyzing consistency in policy ranking, performance correlation, and failure modes under perturbations across diverse tasks. Through extensive multi-platform simulations, policy fine-tuning analyses, post-training data scaling studies, and robustness evaluations, the study identifies key characteristics of high-fidelity simulators and proposes design principles to enhance simulation utility. These findings substantially improve the reliability and practical guidance value of simulation in VLA policy development.
This work addresses the challenge in scientific computing where models struggle to transfer knowledge from individual tasks to broader, reusable capabilities. The authors propose SciConsolidate, a novel framework that, for the first time, extracts procedural knowledge from runtime success and failure trajectories and employs a development-validation gating mechanism to filter effective knowledge. To bridge the gap between abstract knowledge and executable code, the approach integrates failure-driven synthesis of unanswerable queries with strong-model-guided concretization supervision. Additionally, a matched teacher-branch architecture is introduced to significantly enhance the performance of smaller models. Experimental results demonstrate that Qwen3.5-9B achieves gains of 6.25 and 3.89 points over program-free supervised fine-tuning—and improvements of 11.25 and 5.62 points over the base model—on main tasks and sub-steps, respectively, validating the efficacy of the proposed methodology.