🤖 AI Summary
Existing PDE foundation models operate exclusively on numerical or symbolic modalities, limiting their ability to handle real-world multimodal inputs—such as natural-language problem descriptions—and produce interpretable, explanatory outputs.
Method: We propose the first multimodal foundation model for PDE solving and scientific text generation, uniquely integrating natural language deeply into the PDE modeling framework. It jointly encodes parametric ODEs/PDEs, initial/boundary conditions, and textual descriptions of physical processes via a Transformer architecture, enabling end-to-end numerical prediction and attribution-aware text generation.
Contribution/Results: The model achieves mean relative errors of <3.3% (in-distribution) and <7.8% (out-of-distribution), 100% accuracy in generating physically faithful textual descriptions, and robust time extrapolation. It significantly enhances applicability and interpretability of PDE models in scenarios with incomplete information or no closed-form analytical solutions.
📝 Abstract
Neural networks are one tool for approximating non-linear differential equations used in scientific computing tasks such as surrogate modeling, real-time predictions, and optimal control. PDE foundation models utilize neural networks to train approximations to multiple differential equations simultaneously and are thus a general purpose solver that can be adapted to downstream tasks. Current PDE foundation models focus on either learning general solution operators and/or the governing system of equations, and thus only handle numerical or symbolic modalities. However, real-world applications may require more flexible data modalities, e.g. text analysis or descriptive outputs. To address this gap, we propose a novel multimodal deep learning approach that leverages a transformer-based architecture to approximate solution operators for a wide variety of ODEs and PDEs. Our method integrates numerical inputs, such as equation parameters and initial conditions, with text descriptions of physical processes or system dynamics. This enables our model to handle settings where symbolic representations may be incomplete or unavailable. In addition to providing accurate numerical predictions, our approach generates interpretable scientific text descriptions, offering deeper insights into the underlying dynamics and solution properties. The numerical experiments show that our model provides accurate solutions for in-distribution data (with average relative error less than 3.3%) and out-of-distribution data (average relative error less than 7.8%) together with precise text descriptions (with correct descriptions generated 100% of times). In certain tests, the model is also shown to be capable of extrapolating solutions in time.