🤖 AI Summary
This work addresses the pervasive distribution shift problem in machine learning–enhanced hybrid simulation. We first establish a formal mathematical modeling and theoretical analysis framework, revealing the root causes of distribution shift and its error-amplification mechanism over long-term simulation. To mitigate shift propagation, we propose the Tangent Space Regularized Estimator (TSRE), which explicitly enforces consistency of the underlying manifold’s tangent space during surrogate model training. We provide rigorous theoretical guarantees showing that TSRE significantly tightens the long-horizon simulation error bound. Extensive experiments on strongly nonlinear reaction–diffusion systems and high-Reynolds-number Navier–Stokes simulations demonstrate that TSRE reduces average prediction error by 42% beyond 100 time steps compared to baseline methods, with especially pronounced gains under severe distribution shift. This work delivers the first theoretically grounded, distributionally robust solution for ML-enhanced simulation.
📝 Abstract
We study the problem of distribution shift generally arising in machine-learning augmented hybrid simulation, where parts of simulation algorithms are replaced by data-driven surrogates. We first establish a mathematical framework to understand the structure of machine-learning augmented hybrid simulation problems, and the cause and effect of the associated distribution shift. We show correlations between distribution shift and simulation error both numerically and theoretically. Then, we propose a simple methodology based on tangent-space regularized estimator to control the distribution shift, thereby improving the long-term accuracy of the simulation results. In the linear dynamics case, we provide a thorough theoretical analysis to quantify the effectiveness of the proposed method. Moreover, we conduct several numerical experiments, including simulating a partially known reaction-diffusion equation and solving Navier-Stokes equations using the projection method with a data-driven pressure solver. In all cases, we observe marked improvements in simulation accuracy under the proposed method, especially for systems with high degrees of distribution shift, such as those with relatively strong non-linear reaction mechanisms, or flows at large Reynolds numbers.