TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D turbulent flow datasets suffer from limited diversity and reproducibility, and most prior research remains confined to 2D settings, hindering rigorous evaluation of models’ ability to learn genuine physical dynamics. To address this gap, this work introduces TIDE—the first high-resolution (256³) benchmark dataset of incompressible 3D turbulence from direct numerical simulations (DNS), encompassing 15 distinct configurations and multiple independent ensembles. TIDE includes full pressure fields and equation-level validation, and proposes five standardized tasks to assess prediction accuracy, generalization, and physical fidelity. It establishes the first multi-ensemble, multi-parameter-controlled 3D turbulence benchmark, featuring physics-consistency metrics and controlled generalization splits. Empirical evaluation reveals that current learning-based models only marginally outperform persistence baselines, exhibit errors roughly twice those of spectral solvers, and show pronounced deficiencies in capturing small-scale dynamics and transitioning from forced to decaying turbulence regimes.
📝 Abstract
Turbulence is a central testbed for machine learning on physical dynamics because its governing laws are known exactly. However, most existing studies remain in 2D, while 3D turbulence has fundamentally different physics and is far more costly to simulate. Existing 3D resources also typically provide only one realization per configuration, making it difficult to distinguish learning the dynamics from fitting the statistics of a single flow. In this paper, we introduce TIDE (Turbulent Incompressible DNS Ensembles), a 256^3 DNS corpus and benchmark for 3D incompressible turbulence, with 15 configurations on eight controlled axes, independent ensembles, pressure fields, and equation-level verification. The benchmark includes five tasks, standardized learned baselines, controlled generalization splits, and physical-fidelity metrics alongside pointwise error. Across the main forecasting configurations, current learned models barely outperform persistence and still make about twice the error of a spectral solver given the true equations. Moreover, lower pointwise error can coincide with severely distorted small-scale dynamics, showing that accuracy alone does not ensure physical fidelity. Generalization results further show that most regime shifts reflect limited training coverage, whereas forced-to-decay transfer exposes a missing conditioning variable: operators trained under forcing continue to predict driven evolution when the external drive is removed. Closing these accuracy, fidelity, and conditioning gaps is the central open problem made measurable by TIDE.
Problem

Research questions and friction points this paper is trying to address.

3D turbulence
scientific machine learning
physical fidelity
generalization
dynamical modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D turbulence
scientific machine learning
DNS ensembles
physical fidelity
generalization benchmark
🔎 Similar Papers
No similar papers found.