simulator dataset capture and ingestion

Design and build systems that capture data produced by simulators—including in-engine capture of rendered frames, physics and telemetry, and simulated sensor streams—serialize and annotate that data, and store it for downstream use. Implement synchronization, timestamping, efficient batching and format conversion, metadata recording, and reliable transfer/ingestion interfaces so simulation outputs become usable datasets for training, validation, or analysis pipelines.

simulatordatasetcaptureand

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Metadata practices for simulation workflows

Aug 30, 2024
JV
Jose Villamar
🏛️ Jülrich Research Centre | RWTH Aachen University | Helmholtz-Centre for Environmental Research | University of Sussex

Large-scale, heterogeneous metadata in scientific simulations impede result reproducibility and cross-team sharing. To address this, we propose a hardware- and software-agnostic, user-definable two-stage metadata governance framework: (1) non-intrusive acquisition of raw metadata, and (2) on-demand, dynamic structuring. Our key contribution is the first lightweight, general-purpose metadata governance paradigm that decouples acquisition from structuring, enabling zero-code integration into existing HPC simulation workflows. Implemented via the Python-based tool Archivist, the framework supports dynamic schema mapping, declarative configuration, and HPC-adapted interfaces. Evaluated in neuroscience and hydrology simulation use cases, it significantly improves metadata completeness, queryability, and cross-team sharing efficiency—thereby strengthening reproducible and sustainable numerical experimentation.

Ensuring replicability and data sharing in simulation researchHandling heterogeneous metadata from complex computational modelsTracking and organizing metadata for simulation workflows

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Querying Labeled Time Series Data with Scenario Programs

Nov 13, 2025
EK
Edward Kim
🏛️ University of California, Berkeley | Korea University | Chalmers University of Technology | University of Gothenburg | University of California, Santa Cruz

To address the “simulation-to-reality gap”—the difficulty of reproducing simulation-identified failure scenarios in real-world autonomous driving—this paper proposes a verification method based on formal scenario modeling and time-series matching. The method formally translates abstract scenario programs written in the Scenic probabilistic programming language into computable temporal matching rules, enabling precise retrieval of failure-relevant patterns from large-scale real-world sensor data. A key contribution is the design of an efficient, linearly scalable query algorithm that supports real-time pattern matching over long temporal sequences. Experimental evaluation demonstrates that the approach achieves higher recall accuracy for critical failure scenarios than state-of-the-art commercial vision-language models, while accelerating query throughput by several orders of magnitude. This significantly improves both the efficiency and trustworthiness of transferring simulation-discovered failures to real-vehicle validation.

Bridging the sim-to-real gap in autonomous vehicle failure scenario validationDeveloping efficient algorithms to query labeled time series data matching abstract scenariosIdentifying real-world occurrences of simulated failure scenarios in sensor data

Advancements in Synthetic Data Extraction for Industrial Injection Molding

Nov 11, 2025
GR
Georg Rottenwalter
🏛️ Rosenheim Technical University of Applied Sciences

In injection molding, acquiring high-fidelity real-world data is time-consuming and costly, severely limiting the generalizability of machine learning models. To address this, we propose an LSTM-based modeling framework that synergistically integrates synthetic and real data. A high-fidelity simulation model of the production process generates physically consistent synthetic data, while a tunable data injection strategy enhances dataset diversity without compromising physical plausibility. This approach alleviates reliance on large-scale labeled real data, significantly improving model robustness and prediction accuracy under complex operating conditions. Experimental results demonstrate that judicious incorporation of synthetic data boosts model performance by 12.7%, while concurrently reducing annotation effort, equipment wear, and material waste. Our work establishes a reusable and scalable data augmentation paradigm for data-scarce industrial applications.

Balancing synthetic and real data to improve model robustnessOptimizing injection molding processes using machine learning with synthetic dataSynthetic data addresses costly industrial data acquisition challenges

Latest Papers

What's happening recently
View more

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

Existing autonomous driving datasets lack sufficient diversity, coordination, and cross-domain support, limiting their utility for training multi-agent, multi-sensor systems. To address this gap, this work proposes a modular data generation pipeline built upon the AVstack framework and the CARLA simulator, capable of efficiently producing terabyte-scale, ground-truth-annotated multimodal data. The pipeline encompasses perspectives from ground vehicles, aerial platforms, and infrastructure sensors, and supports flexible single- or multi-agent configurations under controllable, complex scenarios. This approach represents the first scalable, cross-domain collaborative data generation methodology for autonomous driving, substantially enhancing the customization, training efficacy, and practical applicability of perception and sensor fusion models in cooperative autonomous systems.

autonomous systemslarge-scale learningmulti-agent

This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.

Data CurationSim-to-Real TransferSynthetic Data

Hot Scholars

XL

Xiaomeng Li

Assistant Professor, The Hong Kong University of Science and Technology
Medical Image AnalysisAI in HealthcareDeep Learning
JO

Jim O'Connor

Connecticut College
Game AIRoboticsArtificial IntelligenceEvolutionary Computation
DG

Derin Gezgin

Undergraduate Student Researcher, Connecticut College
Computer VisionEvolutionary RoboticsArtificial Intelligence for Games
JO

Jacob O. Wobbrock

Professor, University of Washington
Human-Computer InteractionInteraction TechniquesResearch MethodsMobile Computing
XL

Xueying Liu

St. Jude Children’s Research Hospital
Machine learningComputational biology