🤖 AI Summary
This study addresses the challenge of generating joint probabilistic weather forecasts with inter-variable dependencies using only station observations to assess compound meteorological risks. To this end, we propose CLARA, a lightweight, CPU-friendly architecture comprising approximately 28,000 parameters. By leveraging a calibrated advection-routing attention mechanism, CLARA directly learns joint Gaussian predictive distributions for five surface variables from station data without requiring numerical weather predictions or reanalysis products. Furthermore, we develop a consistent covariance scale estimator and demonstrate that neglecting inter-variable correlations significantly degrades negative log-likelihood performance. Experiments across ten global regions reveal that CLARA reduces energy scores by 4.9%–65% relative to baselines, outperforming the persistence baseline in all 60 comparisons and a comparable-scale model in 57 instances.
📝 Abstract
Assessing compound weather risks requires forecasts representing dependence between variables. CLARA (Calibrated Advection-Routing Attention) learns joint Gaussian predictive distributions of five surface variables from station observations alone, without numerical weather prediction or reanalysis; the approximately 28,000-parameter model supports CPU training and prediction. Across six multi-year folds on 96 stations, its lead-mean energy score is 4.9% lower than that of a learned comparator with matched temporal inputs (4.7% with a similar parameter count) and 11-65% lower than those of statistical baselines. Holding marginal variances fixed, removing learned correlations worsens joint negative log-likelihood by 1.0-2.8 nats per station. A covariance-scale estimator, proved consistent under stated assumptions, improves short-lead calibration but over-corrects at long leads. Synthetic interventions show an attention-bias coefficient alone does not measure forecast influence. Retrained in ten regions on six continents, CLARA outperforms persistence in all 60 multi-year region-lead comparisons and a similarly sized learned model in 57 of 60.