🤖 AI Summary
This study addresses the limitation of conventional evaluations, which focus solely on prediction error and physical consistency yet fail to reveal whether models genuinely reproduce underlying dynamics. We propose a dynamical fidelity evaluation framework that translates domain-expert-defined physical representations, such as planetary waves and the Northern Annular Mode (NAM), into systematic tests. By establishing mappings via reference trajectories, this approach captures structural deviations overlooked by standard metrics without requiring prior assumptions. Empirical evaluations using ERA5 data and large-scale weather foundation models, including Pangu-Weather, demonstrate that the proposed framework effectively identifies distinctive structural deviations in model predictions. These results validate its diagnostic capability for assessing learned dynamics.
📝 Abstract
Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.