Score
Designs and implements algorithms and processing pipelines that detect and label clouds in optical remote‑sensing imagery, producing per‑acquisition cloud masks (binary or probabilistic) and associated quality/confidence layers. Builds procedures to interpolate or gap‑fill masked areas for mosaicking and multi‑temporal analysis and to manage cloud masks across large, multi‑acquisition image collections.
Clouds are a common phenomenon that distorts optical satellite imagery, which poses a challenge for remote sensing. However, in the literature cloudless analysis is often performed where cloudy images are excluded from machine learning datasets and methods. Such an approach cannot be applied to time sensitive applications, e.g., during natural disasters. A possible solution is to apply cloud removal as a preprocessing step to ensure that cloudfree solutions are not failing under such conditions. But cloud removal methods are still actively researched and suffer from drawbacks, such as generated visual artifacts. Therefore, it is desirable to develop cloud robust methods that are less affected by cloudy weather. Cloud robust methods can be achieved by combining optical data with radar, a modality unaffected by clouds. While many datasets for machine learning combine optical and radar data, most researchers exclude cloudy images. We identify this exclusion from machine learning training and evaluation as a limitation that reduces applicability to cloudy scenarios. To investigate this, we assembled a dataset, named CloudyBigEarthNet (CBEN), of paired optical and radar images with cloud occlusion for training and evaluation. Using average precision (AP) as the evaluation metric, we show that state-of-the-art methods trained on combined clear-sky optical and radar imagery suffer performance drops of 23-33 percentage points when evaluated on cloudy images. We then adapt these methods to cloudy optical data during training, achieving relative improvement of 17.2-28.7 percentage points on cloudy test cases compared with the original approaches. Code and dataset are publicly available at: https://github.com/mstricker13/CBEN
This study addresses the challenge of accurately detecting thin clouds, fragmented cloud cover, and fine boundary details in remote sensing imagery. To this end, the authors propose an uncertainty-guided two-stage cloud detection method featuring a dual-scale network architecture that integrates CNN and Mamba components. In the first stage, the model produces an initial segmentation mask while estimating pixel-level uncertainty; in the second stage, it refines predictions specifically within low-confidence regions. An embedded uncertainty estimation module guides the optimization across both stages, enabling effective modeling of multi-scale structures and boundary details while maintaining linear computational complexity. Experimental results demonstrate that the proposed approach significantly outperforms existing methods on the GF1-WHU and LEVIR-CS datasets, achieving high accuracy, computational efficiency, and interpretability throughout the detection process.
To address the low accuracy and poor interpretability of infrared remote sensing–based cloud detection, this paper proposes a cloud classification method integrating physical constraints with data-driven learning. Leveraging IASI far-infrared radiance data, we develop a physics-guided support vector machine (CISVM) framework. Brightness temperature features are extracted and dimensionally reduced via joint principal component analysis and cloud-sensitive channel selection, thereby enhancing both model generalizability and physical interpretability. Evaluated on an independent test set, the method achieves 88.30% classification accuracy and exhibits strong agreement with MODIS cloud masks—minor discrepancies in polar regions are attributable to sensor-specific characteristics. These results demonstrate the method’s robustness and operational applicability. This work establishes a novel paradigm for high-accuracy, physically interpretable, automated cloud identification tailored to the far-infrared spectral band.
Cloud occlusion in remote sensing imagery causes critical loss of underlying surface information, severely hindering downstream applications. To address this, we propose DC4CR—the first prompt-driven multimodal diffusion framework for cloud removal, capable of selectively removing both thin and thick clouds without requiring pre-generated cloud masks. Methodologically, DC4CR innovatively integrates prompt-based control mechanisms, Low-Rank Adaptation (LoRA), subject-driven generation, and grouped learning strategies, substantially enhancing few-shot generalization capability and inference efficiency; its modular architecture enables plug-and-play deployment. Evaluated on the RICE and CUHK-CR benchmarks, DC4CR achieves state-of-the-art performance, demonstrating superior reconstruction accuracy and robustness under complex cloud conditions compared to existing methods, thereby exhibiting strong practical potential.
To address the challenge of real-time cloud and cloud shadow detection in on-board preprocessing of hyperspectral satellite imagery, this paper proposes an ultra-lightweight CNN model specifically designed for spaceborne AI systems. The model contains only 597 trainable parameters and integrates feature dimensionality reduction with structural compression to achieve over 93% classification accuracy while drastically reducing memory footprint and computational overhead. Compared to conventional XGBoost/LightGBM and standard CNNs, it enables millisecond-level inference on both CPU and GPU platforms and compresses model size to the kilobyte scale. Experimental results demonstrate its efficiency and robustness under resource-constrained onboard conditions. This work is the first to validate the feasibility of ultra-lightweight deep learning models for hyperspectral cloud detection, establishing a novel paradigm for intelligent on-board preprocessing of remote sensing data.
This study addresses the challenge of cloud occlusion in optical remote sensing imagery during flood events, which severely hinders accurate reconstruction of inundation dynamics using conventional cloud removal methods. To overcome this limitation, the work proposes a novel cloud removal framework based on denoising diffusion probabilistic models, introducing for the first time the Masked Diffusion Transformer to flood-related remote sensing tasks. By leveraging self-attention mechanisms and masked token modeling, the method explicitly reconstructs multispectral information beneath cloud cover. Evaluated on Sentinel-2B data, the approach effectively preserves both spatial continuity and spectral consistency of water bodies, significantly outperforming existing techniques across standard image quality metrics as well as hydrology-specific indicators such as water detection indices.
This work addresses the high cost of pixel-level annotation and heavy reliance on labeled data in cloud detection by proposing CloudMatch, a novel framework that integrates cross-scene and within-scene mixed augmentation to generate semantically consistent yet structurally diverse augmented views. By enforcing consistency learning between a weakly augmented view and two strongly augmented views, and incorporating semi-supervised strategies such as pseudo-labeling, CloudMatch effectively captures the structural diversity and contextual variability of clouds. Evaluated on multiple remote sensing datasets, CloudMatch significantly outperforms existing methods, demonstrating superior accuracy and generalization in semi-supervised cloud detection through efficient utilization of unlabeled data.
This study addresses the challenge of balancing spectral fidelity and perceptual quality in Sentinel-2 satellite image super-resolution by introducing, for the first time, a flow-matching model for large-scale 4× remote sensing image enhancement. The model is trained on co-located same-day Sentinel-2 and NAIP image pairs and, without requiring retraining, enables flexible trade-offs between perceptual quality and pixel-level accuracy during inference simply by adjusting the number of sampling steps. It also supports efficient generation of terapixel-scale high-resolution products. Experimental results demonstrate that with a single sampling step, the method achieves superior pixel accuracy compared to diffusion models and Real-ESRGAN. Furthermore, the generated 2.5-meter-resolution CONUS imagery attains an overall accuracy of 89.11% in land cover classification over the Chesapeake Bay watershed.
Frequent cloud cover in tropical regions severely limits the availability of optical remote sensing imagery, while deep learning models often lose critical spatial and spectral details during downsampling. Method: This paper proposes (1) a lightweight Normalized Difference Index (NDI) injection technique integrated at the decoder’s end to preserve key spatial features; (2) a physically constrained, realistic cloud synthesis and injection framework to systematically evaluate model robustness under cloud occlusion; and (3) a multimodal fusion strategy combining Sentinel-1 SAR and Sentinel-2 optical data. Results: On the DFC2020 dataset, NDI injection improves mIoU by 1.99% and 2.78% for U-Net and DeepLabV3, respectively, under cloud-free conditions. Under cloudy conditions, radar–optical fusion significantly outperforms optical-only input, demonstrating the effectiveness and generalizability of the proposed approach under complex meteorological conditions.