🤖 AI Summary
This study addresses the lack of a unified benchmark with real observations for generative meteorological data assimilation models by constructing the first standardized, controlled evaluation benchmark using data from 11,849 NOAA stations across the United States. By fixing architectures and operators, it systematically compares diffusion models, flow matching, and various inference strategies against traditional 3D-Var. Results demonstrate that full gradient guidance significantly outperforms stop-gradient strategies, and deep generative priors substantially surpass Gaussian priors. The proposed method reduces RMSE by 35.7% relative to the baseline, exceeding the 33.3% improvement achieved by 3D-Var, with advantages becoming more pronounced under sparse observation conditions. These findings provide reliable empirical evidence for design choices in generative meteorological data assimilation.
📝 Abstract
Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.