🤖 AI Summary
This study addresses the limitation that existing research predominantly relies on reanalysis data, making it difficult to determine whether AI weather models systematically underestimate extreme events. To overcome this, we present the first evaluation based on real-world station observations across Europe, defining extreme intervals using ERA5 and employing the Integrated Forecasting System (IFS) as a benchmark. We conduct a multidimensional statistical verification of twelve physical and AI models regarding extreme wind and temperature performance. Our findings reveal that observed error patterns stem from conditional biases rather than inherent deficiencies in AI architectures. Furthermore, we demonstrate that AI models do not exhibit a uniform degradation in tail skill; certain models outperform their physical counterparts under extreme conditions, with overall performance contingent upon the specific meteorological variable and experimental configuration.
📝 Abstract
AI weather models are often reported to underestimate extremes, but most evidence concerns deterministic regression models verified against reanalysis. We evaluate twelve physical and AI forecast models against ECMWF IFS using ten months of European station observations. The evaluation covers 10 m wind, 2 m temperature, solar radiation, and precipitation within regimes defined from a fixed ERA5 1991-2020 climatology. We find no uniform AI-specific deficit in the tails. Several AI models remain more accurate than IFS under extreme conditions, while others deteriorate markedly; comparable variation occurs among physical models. Every model nevertheless exhibits a common conditional-error pattern, overpredicting low observations and underpredicting high observations. Attenuation of extreme values therefore does not imply a uniform loss of relative skill: tail performance depends on the model, variable, and evaluation setting rather than on whether the forecast is produced by AI or physical numerical modelling.