Score
Designs and analyzes continuum PDE models that approximate the discrete-time and stochastic dynamics of training algorithms by deriving coarse-grained or homogenized evolution equations (for example viscous Hamilton–Jacobi or Burgers-type PDEs) which capture multiscale, anisotropic, and contact-like effects. Applies homogenization and local-entropy coarse-graining techniques to reduce degrees of freedom, formalize SDE-to-PDE limits, derive effective continuous-time dynamics, detect phenomena such as shock formation, and expose algorithmic equivalences for analysis or faster simulation.
This work addresses the limitations of traditional numerical methods—prohibitive computational cost in high-dimensional and geometrically complex settings—and the lack of theoretical guarantees and poor generalization in current machine learning approaches for solving partial differential equations (PDEs). To bridge this gap, the authors propose a hybrid PDE-solving paradigm that integrates deductive numerical schemes with inductive learning models. They establish a unified evaluation framework encompassing six core computational challenges and, from an epistemological perspective, formally distinguish between the two methodological classes. By introducing a structure inheritance mechanism and an error budget decomposition, they clarify the conditions under which theoretical guarantees propagate through the hybrid system. Leveraging physics-informed neural networks, differentiable programming, foundation models, and quantum algorithms, the study constructs a multi-paradigm collaborative framework and articulates responsible criteria for method selection, revealing three forms of complementarity that enable scalable, theoretically grounded simulation of high-dimensional complex systems.
Existing machine learning research on stochastic partial differential equations (SPDEs) suffers from the absence of a unified benchmark, particularly neglecting noise-induced sampling errors and renormalization requirements for singular SPDEs—leading to biased model evaluation. Method: We introduce the first comprehensive, standardized benchmark covering both regular and singular SPDEs—including Φ⁴, Navier–Stokes, and KdV equations—featuring a Wick-renormalization-aware data generation pipeline and a novel noise-sensitivity evaluation paradigm. The open-source, reproducible, and extensible platform integrates white-noise-driven solvers, FNO, NSPDE, DLR-Net, and other state-of-the-art models, all validated against high-accuracy numerical reference solutions. Contribution/Results: Experiments reveal that omitting proper numerical discretization severely distorts predictions for singular SPDEs; multiple tasks achieve new state-of-the-art performance; the codebase and standardized datasets have been widely adopted by the community.
In partial differential equation (PDE) coarse-grid simulations under computational resource constraints, unresolved spatiotemporal interactions degrade accuracy, while conventional closure models suffer from physical inconsistency and strong data dependency. Method: This work proposes a novel closure modeling paradigm integrating manufactured solutions with proximal policy optimization (PPO)—a deep reinforcement learning (DRL) algorithm—marking the first application of RL to PDE closure modeling. High-fidelity synthetic data are generated via manufactured solutions, and physics-informed constraints ensure training stability. The model inherently generalizes across homogeneity classes (non-homogeneous → homogeneous). Results: Evaluated on 1D/2D Burgers equations and the 2D advection equation, the method achieves significant accuracy improvements in coarse-grid simulations using only minimal training data—demonstrating robustness in sparse-observation regimes.
本文提出一种时空神经算子,用于解决多尺度PDE系统的粗粒度动力学预测问题,通过傅里叶卷积和因果核操作实现空间和时间混合。
This work investigates the necessity and benefits of incorporating historical state memory when modeling time-dependent partial differential equations (PDEs), challenging the conventional Markovian assumption. Leveraging the Mori–Zwanzig formalism, we provide the first rigorous theoretical proof that explicit memory modeling yields fundamental representational gains for linear PDEs. Building on this insight, we propose the Memory Neural Operator (MemNO), a novel architecture that synergistically integrates the S4 state-space model—capable of capturing long-range temporal dependencies—with the Fourier Neural Operator—designed to represent spatial nonlinearities. Evaluated on challenging benchmarks—including low-resolution data, noisy observations, and high-frequency-dominated PDEs (e.g., low-viscosity fluid dynamics)—MemNO achieves up to a 6× reduction in test error compared to state-of-the-art baselines. The method significantly enhances generalization accuracy and robustness under distributional shifts and data scarcity.
This work addresses the challenge posed by parameter symmetries in deep neural network training, which induce redundancy and impede analytical understanding of training dynamics and phase transitions. For the first time, it integrates shock wave theory, differential geometry, and Lie group symmetry reduction into the analysis of deep learning dynamics. By performing symmetry reduction on quotient manifolds and applying local entropy-based coarse-graining, the study establishes a rigorous mathematical connection between stochastic gradient descent dynamics and viscous Hamilton–Jacobi or Burgers-type partial differential equations. The proposed observables defined on the quotient space effectively diagnose training phase transitions and demonstrate universality across diverse architectures—including multilayer perceptrons, convolutional networks, Transformers, and mean-field networks—thereby offering a principled framework for monitoring and controlling the training process.
Nonlocal partial differential equations are challenging to model and predict due to the complexity of their nonlocal operators. This work proposes a flow map learning framework that directly learns the finite-time evolution operator from solution data, bypassing the need for explicit modeling or approximation of the nonlocal operator. The approach is compatible with both spectral and grid-based representations and integrates evolution operator learning with spectral and finite difference methodologies. Demonstrated on one- and two-dimensional fractional diffusion and wave equations, the method achieves accurate and stable long-term dynamical predictions using only short-time observational windows, substantially enhancing the capability to model unknown nonlocal systems.
本文提出了一种基于新型多尺度核框架函数逼近技术的多尺度算子学习方法,用于求解多尺度偏微分方程,并在文献中的难题上展示了其优越性。
This work addresses the significant degradation in solution trajectories and spatial fidelity caused by insufficient resolution in coarse-grid partial differential equation (PDE) solvers. To overcome this limitation, the authors propose RECAST, a novel framework that uniquely integrates recurrent error correction with super-resolution reconstruction into the PDE solving pipeline. Specifically, a recurrent neural network models and corrects temporal errors on the coarse grid, and high-resolution solutions are subsequently reconstructed from the corrected historical states. Evaluated across six classes of one-dimensional PDE systems, RECAST achieves high-fidelity long-horizon rollouts, demonstrating strong generalization to unseen initial conditions and parameters. Compared to uncorrected coarse-grid solvers, it reduces time-averaged relative errors by 50%–92% and outperforms existing methods over 5,000-step predictions.
This work addresses the physical inaccuracies introduced by coarse-grid simulations due to their inability to resolve small-scale turbulence. The authors propose a decoupled dual-network architecture based on structure-preserving neural networks and entropy variables to learn a subgrid-scale parameterization for the Burgers equation. In this framework, the subgrid flux is decomposed into a conservative flux potential network and an eddy-viscosity network. The approach rigorously preserves conservation properties while significantly enhancing model robustness and generalization in extrapolation scenarios. Numerical experiments demonstrate that the model faithfully reproduces the energy spectrum, spatiotemporal correlation functions, and dynamical characteristics of the fully resolved system, maintaining high fidelity even when applied beyond the range of training parameters.