🤖 AI Summary
Existing global crop-type reference datasets are often hindered by heterogeneous sources, labeling noise, and regional biases, which limit the performance of crop mapping models. This work proposes a local-aware anomaly detection framework—Embedding-based Anomaly detection (EBA)—that leverages embeddings from pretrained geospatial foundation models for automated data cleaning. By measuring local neighborhood similarity and incorporating a confidence-weighting mechanism, EBA assigns anomaly scores to samples within the same region and class, adapting to regional variations without requiring handcrafted rules. Experiments demonstrate the method’s effectiveness on both synthetic and real-world data, achieving an AUROC of 0.84. After conservative filtering using EBA, the WorldCereal crop-type classification model shows significantly improved accuracy across five major macro-regions.
📝 Abstract
High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCereal aggregate labels from heterogeneous sources such as parcel registers, national databases, field surveys, and map-derived products, each with their own biases, coverage gaps and unknown label noise. Simple global rules are inadequate, since crop phenology and observation conditions vary strongly across regions and seasons.
In this study, we focus on a single, operationally relevant question: whether embeddings produced through geospatial foundation models are a viable basis for cleaning the reference data. We propose a practical, locality-aware, embedding-based anomaly (EBA) detection framework that operates on the embeddings of a pretrained Earth-observation encoder. We score each labelled sample against other samples of the same crop in the same area using a pretrained embedding, flag the ones that stand out, and test whether removing or down-weighting them before training yields a better model.
We establish that the flagged points are genuinely mislabelled or misplaced in two independent ways: against synthetic ground truth, the detector concentrates injected label errors 2.5-5x above chance in its flagged set (detection AUROC up to 0.84); and on real data, a model-independent test shows that removing or confidence-weighting the flagged held-out points raises measured accuracy in trained models, for both crop type and land cover. Acting on the flags then improves the WorldCereal crop-type model across five macro-regions, evaluated on a fixed held-out split under three views. We find conservative cleaning helps while over-cleaning hurts.
The EBA detector approach is designed to be reproducible and extensible, and can serve as a template for cleaning large, noisy Earth observation reference datasets beyond crop mapping.