benchmark and validate simulation results

Designs and implements numerical simulations and the experimentation pipelines that generate simulation data and test cases. Builds and applies benchmarks, verification and validation procedures, and numerical precision/error analyses to evaluate, compare, and ensure correctness, stability, and reproducibility of simulation results.

benchmarkandvalidatesimulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.

Automating benchmarking of diverse scientific software alternativesManaging expanding parameter spaces in benchmark setupsStreamlining re-evaluation when adding new metrics or cases

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Towards Experiment Execution in Support of Community Benchmark Workflows for HPC

Jul 29, 2025
GV
Gregor von Laszewski
🏛️ University of Virginia | Oak Ridge National Laboratory | Hewlett Packard Enterprise Canada | University of Florida | Cummins | San Diego Supercomputer Center | University of California, San Diego

To address low reusability of HPC benchmarks, poor cross-platform portability, and inefficient resource validation, this paper proposes the “benchmark carpentry” paradigm—a lightweight, reusable experimental execution framework. Methodologically, it integrates Cloudmesh’s experiment executor with HPE SmartSim, incorporating standardized workflow templates, AI/ML–simulation coupling mechanisms, and a unified experimental management interface. Its key contribution is the first application of craftsmanship principles to benchmarking process design, enabling automated, cross-domain and cross-architecture benchmark deployment and capability assessment. Evaluated on representative scientific computing workloads—including cloud masking analysis, seismic forecasting, and CFD surrogate modeling—the framework achieves ≥92% workflow reproducibility and reduces average deployment time by 68%, significantly improving resource configuration efficiency. It establishes a scalable, community-driven paradigm for HPC capability validation.

Creating adaptable workflow templates for scientific applicationsDemonstrating HPC compute capability with limited benchmarksImproving experiment management tools for broader workflow adaptability

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

Latest Papers

What's happening recently
View more

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

This work addresses the challenge of directly verifying numerical algorithms in large-scale scientific simulation frameworks, where system complexity and insufficient unit test coverage hinder validation. Focusing on Cloud-in-Cell (CIC) algorithm development for Flash-X, we propose a "model-checking-in-the-loop" workflow. By abstracting infrastructure details to extract interfaces, we construct small, self-contained C models and embed the CIVL model checker with symbolic execution into the development cycle as design guardrails. This approach successfully verifies physical properties such as mass conservation alongside memory safety, and uncovers concurrency defects overlooked by randomized testing. Consequently, it significantly enhances algorithmic reliability, offering an efficient new paradigm for the formal verification of complex scientific computing systems.

concurrency defectsmodel checkingnumerical algorithm verification

This study addresses the limitation of existing generative model benchmarks, which focus solely on code compilation or structural similarity and fail to verify engineering compliance. We construct a text-to-executable Simulink model generation benchmark spanning ten domains and propose an executable system archive-based method for aligning with engineering specifications. Furthermore, we design a hierarchical automated evaluation mechanism alongside a six-dimensional dynamic response metric system, enabling end-to-end assessment from deliverability and executability to engineering qualification. Experimental results demonstrate that the best-performing agent achieves a score of only 42.86, confirming a significant disconnect between structural similarity and engineering performance. These findings reveal critical capability bottlenecks in current large language models regarding the generation of qualified systems that satisfy multidimensional engineering requirements.

benchmarking agentsengineering requirementsmodel evaluation

Traditional computer architecture simulation suffers from poor scalability, limited reproducibility, and excessive customization due to its reliance on implicit scripts and directory conventions. This work proposes the first end-to-end explicit and service-oriented simulation framework, which models hardware topologies declaratively via graph representations, automatically generates executable simulation code, and employs a stateless runner for automated task scheduling and structured result management. The approach eliminates the need for manual simulation programming and enables systematic exploration through automatic expansion of configuration–benchmark matrices. Evaluated across 96 GPU workloads, the framework achieves a median kernel time error of only 0.18% compared to hand-tuned MGPUSim configurations—covering 95.8% of all configurations—with a negligible per-simulation overhead of just 1.6 seconds.

computer architecture simulationreproducibilityscalability

This study addresses the absence of benchmark datasets for Modelica, which has hindered empirical research on model evolution. We propose ModBench, an automated pipeline that establishes a novel paradigm for generating model snapshot benchmarks directly from code repositories by mining Git history, filtering commits, extracting simulatable classes, and normalizing representations. Applying this approach to the Modelica Standard Library, we constructed a comprehensive dataset comprising 85,562 class snapshots spanning all versions since v3, complete with API access and traceability links. This work fills a critical data gap in the domain, providing essential infrastructure to support research in model evolution analysis, compiler testing, and automated program repair.

benchmark datasetscyber-physical systemsmodel evolution

Hot Scholars

MS

Mikael Skoglund

KTH Royal Institute of Technology
Information TheoryCommunicationsSignal Processing
BH

Barbara Hammer

Professor, Bielefeld University
machine learningdata miningneural networksbioinformatics
NA

Nail Akar

Professor of Electrical and Electronics Eng. Dept., Bilkent University
Computer networksperformance evaluationqueuing theorystochastic models
TW

Tobias Weinzierl

Durham University
Scientific ComputingParallel AlgorithmsHigh Performance Computing
AA

Anima Anandkumar

California Institute of Technology and NVIDIA
Machine Learning and Artificial Intelligence