🤖 AI Summary
This study addresses the overreliance of AI-based Earth science models on statistical metrics, which often neglects physical reliability and applicability under non-stationary climate conditions. To overcome these limitations, this work proposes a systematic governance framework that rigorously distinguishes predictive skill from physical reliability. Methodologically, it integrates foundation model fine-tuning, mechanistic interpretability analysis, and behavioral testing to establish five core priorities for physics-informed evaluation spanning data construction to output validation. As a primary contribution, this paper formulates safety assessment guidelines for the field over the next decade and advocates for the development of open datasets and shared standards, thereby advancing the establishment of standardized benchmarking systems for AI-driven Earth science models.
📝 Abstract
AI foundation models pretrained on weather and climate data are increasingly fine-tuned to Earth science tasks well beyond weather forecasting. Their development and adoption are outpacing the scientific community's ability to evaluate them. These models are judged almost entirely by benchmark skill metrics, which measure how closely a forecast reproduces a reference product but not whether a model represents the physical processes governing the system it predicts. Forecast skill and physical reliability are therefore distinct properties. The distinction is most consequential under the nonstationary conditions of a changing climate for which these models were never trained. We identify five priorities for the physical evaluation of AI models in Earth science from task-specific emulators to foundation models, spanning training data, fine-tuning, behavioral testing, mechanistic interpretability, and output validation. We recommend three activities for the coming decade: 1) open AI-ready evaluation datasets, 2) a shared reporting standard for physics-based evaluation, and 3) a dedicated research program on the safety of these models.