M$^2$Weather: A Benchmark for Joint Multi-Station and Multi-Variable Weather Forecasting
This study addresses the limitation of existing meteorological forecasting research, which models station-level spatial dependencies and variable-level physical couplings in isolation without a unified benchmark to evaluate their joint contributions. To bridge this gap, we construct a multi-scale, high-quality meteorological benchmark and introduce the first standardized evaluation framework that jointly assesses multi-station and multi-variable modeling. Furthermore, we design lightweight, plug-and-play adapters that efficiently recover missing spatial and inter-variable correlations without requiring model retraining. Extensive experiments across 16 representative models demonstrate that joint modeling significantly reduces prediction errors, confirming that complementary information across stations and variables constitutes a critical resource for enhancing forecasting accuracy.