๐ค AI Summary
This study addresses the challenge of reconstructing cloud-contaminated and noisy NDVI time series in remote sensing by proposing the GloSSR framework. To circumvent the need for paired training data, GloSSR employs a self-supervised strategy that synthetically degrades clean observations using realistic cloud masks, thereby generating training pairs. The method introduces an end-to-end spatiotemporal network integrating bidirectional Transformers and ConvLSTM modules, enhanced with channelโtemporal attention mechanisms and spatiotemporal prior constraints to jointly model long-term temporal dependencies and short-range spatiotemporal correlations. Experiments demonstrate that GloSSR significantly outperforms existing approaches on MODIS data, accurately capturing vegetation dynamics and critical phenological states, while also exhibiting strong transferability and suitability for large-scale environmental monitoring when applied to AVHRR data.
๐ Abstract
Accurate and efficient reconstruction of cloud-contaminated and noise-corrupted NDVI time series remains a challenge in remote sensing. Deep learning provides a promising solution for modeling complex spatiotemporal dependencies; however, its application is often limited by the difficulty of obtaining paired clear-sky and degraded NDVI data for identical spatiotemporal locations. To address this issue, we propose GloSSR, a Global-scale Self-supervised Spatiotemporal framework for NDVI Reconstruction. The framework constructs supervisory signals by artificially degrading relatively clean NDVI observations with realistic cloud contamination patterns, producing self-supervised training pairs that closely mimic real-world degradation. It further introduces an end-to-end spatiotemporal learning network that jointly captures long-range temporal dependencies and short-term spatiotemporal correlation through a bidirectional Transformer with a ConvLSTM architecture. A temporal-channel attention-based reconstruction module is incorporated to enhance informative features, while a spatiotemporal prior constraint is designed to preserve both fine-scale structures and long-term phenological trends during optimization. Extensive evaluations on MODIS NDVI data demonstrate the effectiveness of the proposed framework across both artificial and real-world scenarios. In artificial degraded-pixel reconstruction experiments, GloSSR consistently outperforms the comparison methods. Time-series analyses based on real observations further demonstrate that the proposed framework can accurately characterize vegetation dynamics and capture the key phenological states. Long-term vegetation trend analysis and the transferability analysis to AVHRR data validate the scalability of the framework and illustrate its broad applicability for large-scale environmental monitoring.