benchmark and validate simulation results

Designs and implements numerical simulations and the experimentation pipelines that generate simulation data and test cases. Builds and applies benchmarks, verification and validation procedures, and numerical precision/error analyses to evaluate, compare, and ensure correctness, stability, and reproducibility of simulation results.

benchmarkandvalidatesimulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Scientific software selection frequently suffers from non-reproducible benchmarks due to multi-library, multi-metric evaluation and dynamic evolution—such as the introduction of new algorithms or modifications to test cases and evaluation criteria. This paper addresses numerical integration over arbitrary 2D/3D domains with implicit or parameterized boundaries (cut-cell quadrature), proposing the first automated benchmarking framework that systematically integrates CI/CD engineering practices into scientific computing workflows. The framework unifies GitHub Actions, Docker, Python-based scheduling, Jupyter-based report generation, and semantically versioned result archiving. It supports automated configuration, execution, visualization, and historical result comparison. It achieves >90% automation for benchmark tasks and regression detection; reduces integration time for new libraries or algorithms by 70%; and enables precise attribution of performance deviations to specific code commits. The framework significantly enhances reliability, reproducibility, and evolutionary adaptability in scientific software evaluation.

Automating benchmarking of diverse scientific software alternativesManaging expanding parameter spaces in benchmark setupsStreamlining re-evaluation when adding new metrics or cases

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Towards Experiment Execution in Support of Community Benchmark Workflows for HPC

Jul 29, 2025
GV
Gregor von Laszewski
🏛️ University of Virginia | Oak Ridge National Laboratory | Hewlett Packard Enterprise Canada | University of Florida | Cummins | San Diego Supercomputer Center | University of California, San Diego

To address low reusability of HPC benchmarks, poor cross-platform portability, and inefficient resource validation, this paper proposes the “benchmark carpentry” paradigm—a lightweight, reusable experimental execution framework. Methodologically, it integrates Cloudmesh’s experiment executor with HPE SmartSim, incorporating standardized workflow templates, AI/ML–simulation coupling mechanisms, and a unified experimental management interface. Its key contribution is the first application of craftsmanship principles to benchmarking process design, enabling automated, cross-domain and cross-architecture benchmark deployment and capability assessment. Evaluated on representative scientific computing workloads—including cloud masking analysis, seismic forecasting, and CFD surrogate modeling—the framework achieves ≥92% workflow reproducibility and reduces average deployment time by 68%, significantly improving resource configuration efficiency. It establishes a scalable, community-driven paradigm for HPC capability validation.

Creating adaptable workflow templates for scientific applicationsDemonstrating HPC compute capability with limited benchmarksImproving experiment management tools for broader workflow adaptability

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

Latest Papers

What's happening recently
View more

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

Traditional computer architecture simulation suffers from poor scalability, limited reproducibility, and excessive customization due to its reliance on implicit scripts and directory conventions. This work proposes the first end-to-end explicit and service-oriented simulation framework, which models hardware topologies declaratively via graph representations, automatically generates executable simulation code, and employs a stateless runner for automated task scheduling and structured result management. The approach eliminates the need for manual simulation programming and enables systematic exploration through automatic expansion of configuration–benchmark matrices. Evaluated across 96 GPU workloads, the framework achieves a median kernel time error of only 0.18% compared to hand-tuned MGPUSim configurations—covering 95.8% of all configurations—with a negligible per-simulation overhead of just 1.6 seconds.

computer architecture simulationreproducibilityscalability

This work addresses the challenge of accurately and physically consistently capturing sharp gradients such as shock waves in hypersonic flow fields, which are poorly resolved by conventional reduced-order models or neural surrogates. To this end, the authors propose an end-to-end fully GPU-accelerated workflow that leverages the differentiable high-fidelity solver JAX-Fluids for efficient data generation and introduces a residual-driven, physics-aware refinement mechanism. The resulting neural surrogate is trained using only mesh coordinates and input parameters, significantly reducing residuals of the governing equations while improving the physical fidelity of predicted shock locations and strengths. The method demonstrates strong generalization and reliability even under out-of-distribution operating conditions.

hypersonic flowsneural emulatorsphysical consistency

Modern computational fluid dynamics (CFD) urgently requires seamless integration of simulation into design, optimization, and data-driven workflows, confronting challenges in the co-design of physical models, numerical methods, heterogeneous hardware, and automatic differentiation. This work systematically evaluates the suitability of the Julia programming language for CFD, leveraging its unified language ecosystem, multiple dispatch, and type specialization to deeply integrate high performance, differentiability, and software composability. Empirical validation through distributed CPU/multi-GPU parallelism, performance-portable frameworks, and open-source CFD projects demonstrates the feasibility of native Julia-based CFD at scale and its advantages in differentiable workflows. Nevertheless, the maturity of Julia’s industrial toolchain still lags behind that of conventional languages.

automatic differentiationcomposabilitycomputational fluid dynamics

Hot Scholars

BH

Barbara Hammer

Professor, Bielefeld University
machine learningdata miningneural networksbioinformatics
MS

Mikael Skoglund

KTH Royal Institute of Technology
Information TheoryCommunicationsSignal Processing
NA

Nail Akar

Professor of Electrical and Electronics Eng. Dept., Bilkent University
Computer networksperformance evaluationqueuing theorystochastic models
TW

Tobias Weinzierl

Durham University
Scientific ComputingParallel AlgorithmsHigh Performance Computing
AA

Anima Anandkumar

California Institute of Technology and NVIDIA
Machine Learning and Artificial Intelligence