Score
Design and build systems that capture data produced by simulators—including in-engine capture of rendered frames, physics and telemetry, and simulated sensor streams—serialize and annotate that data, and store it for downstream use. Implement synchronization, timestamping, efficient batching and format conversion, metadata recording, and reliable transfer/ingestion interfaces so simulation outputs become usable datasets for training, validation, or analysis pipelines.
Subsymbolic AI struggles with training under few-shot and low-quality data, while existing virtual simulation approaches lack systematic, standards-aligned frameworks. Method: We conduct a systematic literature review covering 22 state-of-the-art works and propose, for the first time, a unified reference framework for digital twin–driven AI simulation—deeply integrating digital twins with AI agents to establish a closed-loop, cyber-physical data orchestration mechanism. We further achieve systematic alignment of this framework with the ISO 23247 international standard for digital twins. Contribution/Results: We distill key technological evolution trends, identify five core challenges and several open research directions, and deliver a reusable architectural guideline and reference framework. This work provides a standardized, methodology-driven foundation for high-fidelity AI simulation, advancing both theoretical rigor and practical deployability in industrial AI applications.
Large-scale, heterogeneous metadata in scientific simulations impede result reproducibility and cross-team sharing. To address this, we propose a hardware- and software-agnostic, user-definable two-stage metadata governance framework: (1) non-intrusive acquisition of raw metadata, and (2) on-demand, dynamic structuring. Our key contribution is the first lightweight, general-purpose metadata governance paradigm that decouples acquisition from structuring, enabling zero-code integration into existing HPC simulation workflows. Implemented via the Python-based tool Archivist, the framework supports dynamic schema mapping, declarative configuration, and HPC-adapted interfaces. Evaluated in neuroscience and hydrology simulation use cases, it significantly improves metadata completeness, queryability, and cross-team sharing efficiency—thereby strengthening reproducible and sustainable numerical experimentation.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
To address the “simulation-to-reality gap”—the difficulty of reproducing simulation-identified failure scenarios in real-world autonomous driving—this paper proposes a verification method based on formal scenario modeling and time-series matching. The method formally translates abstract scenario programs written in the Scenic probabilistic programming language into computable temporal matching rules, enabling precise retrieval of failure-relevant patterns from large-scale real-world sensor data. A key contribution is the design of an efficient, linearly scalable query algorithm that supports real-time pattern matching over long temporal sequences. Experimental evaluation demonstrates that the approach achieves higher recall accuracy for critical failure scenarios than state-of-the-art commercial vision-language models, while accelerating query throughput by several orders of magnitude. This significantly improves both the efficiency and trustworthiness of transferring simulation-discovered failures to real-vehicle validation.
In injection molding, acquiring high-fidelity real-world data is time-consuming and costly, severely limiting the generalizability of machine learning models. To address this, we propose an LSTM-based modeling framework that synergistically integrates synthetic and real data. A high-fidelity simulation model of the production process generates physically consistent synthetic data, while a tunable data injection strategy enhances dataset diversity without compromising physical plausibility. This approach alleviates reliance on large-scale labeled real data, significantly improving model robustness and prediction accuracy under complex operating conditions. Experimental results demonstrate that judicious incorporation of synthetic data boosts model performance by 12.7%, while concurrently reducing annotation effort, equipment wear, and material waste. Our work establishes a reusable and scalable data augmentation paradigm for data-scarce industrial applications.
为了解决动作条件视频模型数据获取难题,本文提出基于虚幻引擎的两阶段合成数据生成管道,生成大规模动作条件多视角视频。
为解决工业能源数据采集问题,提出STREAM框架,通过目标驱动和不确定性评估方法确保数据满足能源性能评估需求。
This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.
Existing autonomous driving datasets lack sufficient diversity, coordination, and cross-domain support, limiting their utility for training multi-agent, multi-sensor systems. To address this gap, this work proposes a modular data generation pipeline built upon the AVstack framework and the CARLA simulator, capable of efficiently producing terabyte-scale, ground-truth-annotated multimodal data. The pipeline encompasses perspectives from ground vehicles, aerial platforms, and infrastructure sensors, and supports flexible single- or multi-agent configurations under controllable, complex scenarios. This approach represents the first scalable, cross-domain collaborative data generation methodology for autonomous driving, substantially enhancing the customization, training efficacy, and practical applicability of perception and sensor fusion models in cooperative autonomous systems.
This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.