🤖 AI Summary
This study addresses the limited accuracy of precipitation nowcasting under strong spatiotemporal variability by exploring the potential of standard diffusion architectures. We construct a forecasting framework based on the standard Diffusion Transformer (DiT), innovatively introducing a dynamics-aware temporally consistent noise prior, and optimize it via end-to-end reinforcement learning coupled with a timestep-aware reward mechanism. Our findings demonstrate that a standard DiT can effectively accomplish this task without requiring complex architectural customizations. Extensive experiments on the SEVIR and MRMS benchmarks indicate that the proposed method achieves state-of-the-art performance in both visual perceptual quality and meteorological skill scores.
📝 Abstract
Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.