pde modeling of training dynamics

Designs and analyzes continuum PDE models that approximate the discrete-time and stochastic dynamics of training algorithms by deriving coarse-grained or homogenized evolution equations (for example viscous Hamilton–Jacobi or Burgers-type PDEs) which capture multiscale, anisotropic, and contact-like effects. Applies homogenization and local-entropy coarse-graining techniques to reduce degrees of freedom, formalize SDE-to-PDE limits, derive effective continuous-time dynamics, detect phenomena such as shock formation, and expose algorithmic equivalences for analysis or faster simulation.

pdemodelingoftraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.67
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional numerical methods—prohibitive computational cost in high-dimensional and geometrically complex settings—and the lack of theoretical guarantees and poor generalization in current machine learning approaches for solving partial differential equations (PDEs). To bridge this gap, the authors propose a hybrid PDE-solving paradigm that integrates deductive numerical schemes with inductive learning models. They establish a unified evaluation framework encompassing six core computational challenges and, from an epistemological perspective, formally distinguish between the two methodological classes. By introducing a structure inheritance mechanism and an error budget decomposition, they clarify the conditions under which theoretical guarantees propagate through the hybrid system. Leveraging physics-informed neural networks, differentiable programming, foundation models, and quantum algorithms, the study constructs a multi-paradigm collaborative framework and articulates responsible criteria for method selection, revealing three forms of complementarity that enable scalable, theoretically grounded simulation of high-dimensional complex systems.

Computational ChallengesHybrid MethodsMachine Learning

SPDEBench: An Extensive Benchmark for Learning Regular and Singular Stochastic PDEs

May 24, 2025
ZL
Zheyan Li
🏛️ Shanghai Jiaotong University | University College London | Fujian Normal University | Chinese Academy of Sciences

Existing machine learning research on stochastic partial differential equations (SPDEs) suffers from the absence of a unified benchmark, particularly neglecting noise-induced sampling errors and renormalization requirements for singular SPDEs—leading to biased model evaluation. Method: We introduce the first comprehensive, standardized benchmark covering both regular and singular SPDEs—including Φ⁴, Navier–Stokes, and KdV equations—featuring a Wick-renormalization-aware data generation pipeline and a novel noise-sensitivity evaluation paradigm. The open-source, reproducible, and extensible platform integrates white-noise-driven solvers, FNO, NSPDE, DLR-Net, and other state-of-the-art models, all validated against high-accuracy numerical reference solutions. Contribution/Results: Experiments reveal that omitting proper numerical discretization severely distorts predictions for singular SPDEs; multiple tasks achieve new state-of-the-art performance; the codebase and standardized datasets have been widely adopted by the community.

Impact of noise sampling and renormalization on ML modelsLack of extensive unified datasets for SPDE learningNeed for high-quality test data to evaluate ML models accurately

Reinforcement Learning Closures for Underresolved Partial Differential Equations using Synthetic Data

May 16, 2025
LH
Lothar Heimbach
🏛️ ETH Zurich | Harvard University | Argonne National Laboratory

In partial differential equation (PDE) coarse-grid simulations under computational resource constraints, unresolved spatiotemporal interactions degrade accuracy, while conventional closure models suffer from physical inconsistency and strong data dependency. Method: This work proposes a novel closure modeling paradigm integrating manufactured solutions with proximal policy optimization (PPO)—a deep reinforcement learning (DRL) algorithm—marking the first application of RL to PDE closure modeling. High-fidelity synthetic data are generated via manufactured solutions, and physics-informed constraints ensure training stability. The model inherently generalizes across homogeneity classes (non-homogeneous → homogeneous). Results: Evaluated on 1D/2D Burgers equations and the 2D advection equation, the method achieves significant accuracy improvements in coarse-grid simulations using only minimal training data—demonstrating robustness in sparse-observation regimes.

Developing closure models for underresolved PDEs using synthetic dataGeneralizing closure models from inhomogeneous to homogeneous PDEsImproving coarse-grained PDE accuracy with reinforcement learning closures

On the Benefits of Memory for Modeling Time-Dependent PDEs

Sep 03, 2024
RB
Ricardo Buitrago Ruiz
🏛️ Carnegie Mellon University | Cartesia AI

This work investigates the necessity and benefits of incorporating historical state memory when modeling time-dependent partial differential equations (PDEs), challenging the conventional Markovian assumption. Leveraging the Mori–Zwanzig formalism, we provide the first rigorous theoretical proof that explicit memory modeling yields fundamental representational gains for linear PDEs. Building on this insight, we propose the Memory Neural Operator (MemNO), a novel architecture that synergistically integrates the S4 state-space model—capable of capturing long-range temporal dependencies—with the Fourier Neural Operator—designed to represent spatial nonlinearities. Evaluated on challenging benchmarks—including low-resolution data, noisy observations, and high-frequency-dominated PDEs (e.g., low-viscosity fluid dynamics)—MemNO achieves up to a 6× reduction in test error compared to state-of-the-art baselines. The method significantly enhances generalization accuracy and robustness under distributional shifts and data scarcity.

Introduces Memory Neural Operator for better PDE performanceInvestigates benefits of memory for time-dependent PDE modelingProves memory-based solutions outperform Markovian ones theoretically

Latest Papers

What's happening recently
View more

This work addresses the challenge posed by parameter symmetries in deep neural network training, which induce redundancy and impede analytical understanding of training dynamics and phase transitions. For the first time, it integrates shock wave theory, differential geometry, and Lie group symmetry reduction into the analysis of deep learning dynamics. By performing symmetry reduction on quotient manifolds and applying local entropy-based coarse-graining, the study establishes a rigorous mathematical connection between stochastic gradient descent dynamics and viscous Hamilton–Jacobi or Burgers-type partial differential equations. The proposed observables defined on the quotient space effectively diagnose training phase transitions and demonstrate universality across diverse architectures—including multilayer perceptrons, convolutional networks, Transformers, and mean-field networks—thereby offering a principled framework for monitoring and controlling the training process.

deep learning diagnosticsparameter symmetriesquotient observables

Nonlocal partial differential equations are challenging to model and predict due to the complexity of their nonlocal operators. This work proposes a flow map learning framework that directly learns the finite-time evolution operator from solution data, bypassing the need for explicit modeling or approximation of the nonlocal operator. The approach is compatible with both spectral and grid-based representations and integrates evolution operator learning with spectral and finite difference methodologies. Demonstrated on one- and two-dimensional fractional diffusion and wave equations, the method achieves accurate and stable long-term dynamical predictions using only short-time observational windows, substantially enhancing the capability to model unknown nonlocal systems.

data-driven modelingflow map learningnonlocal operators

本文提出了一种基于新型多尺度核框架函数逼近技术的多尺度算子学习方法,用于求解多尺度偏微分方程,并在文献中的难题上展示了其优越性。

multiscaleoperator learningpartial differential equations

This work addresses the significant degradation in solution trajectories and spatial fidelity caused by insufficient resolution in coarse-grid partial differential equation (PDE) solvers. To overcome this limitation, the authors propose RECAST, a novel framework that uniquely integrates recurrent error correction with super-resolution reconstruction into the PDE solving pipeline. Specifically, a recurrent neural network models and corrects temporal errors on the coarse grid, and high-resolution solutions are subsequently reconstructed from the corrected historical states. Evaluated across six classes of one-dimensional PDE systems, RECAST achieves high-fidelity long-horizon rollouts, demonstrating strong generalization to unseen initial conditions and parameters. Compared to uncorrected coarse-grid solvers, it reduces time-averaged relative errors by 50%–92% and outperforms existing methods over 5,000-step predictions.

coarse-grid PDE solverscomputational efficiencynumerical accuracy

This work addresses the physical inaccuracies introduced by coarse-grid simulations due to their inability to resolve small-scale turbulence. The authors propose a decoupled dual-network architecture based on structure-preserving neural networks and entropy variables to learn a subgrid-scale parameterization for the Burgers equation. In this framework, the subgrid flux is decomposed into a conservative flux potential network and an eddy-viscosity network. The approach rigorously preserves conservation properties while significantly enhancing model robustness and generalization in extrapolation scenarios. Numerical experiments demonstrate that the model faithfully reproduces the energy spectrum, spatiotemporal correlation functions, and dynamical characteristics of the fully resolved system, maintaining high fidelity even when applied beyond the range of training parameters.

Burgers' equationcoarse simulationpartial differential equations

Hot Scholars

JE

Jan Eliáš

Brno University of Technology, Faculty of Civil Engineering, Institute of Structural Mechanics
mechanicsengineering
TJ

Ting-Ju Wei

National Taiwan University
Computational MechanicsArtificial intelligenceMolecular dynamics
CG

Christophe Geuzaine

Full Professor, University of Liège
Computational ElectromagneticsScientific ComputingFinite ElementsMesh Generation
SS

Sebastian Schöps

Technische Universität Darmstadt
Computational ElectromagneticsMultiphysicsComputer Aided DesignHigh-Performance Computing
KA

Karl A. Kalina

Institute of Solid Mechanics, TU Dresden
Computational MechanicsArtificial Neural NetworksMultiscale ProblemsCoupled Problems