🤖 AI Summary
This study addresses the limitations of conventional self-supervised pretraining in Earth observation foundation models, which often employ physically unconstrained random masking and thus fail to meet the trustworthiness requirements of high-stakes applications such as public health decision-making. To overcome this, we propose SpecTM—a spectrally targeted masking strategy that integrates biophysical priors into the masking process, enabling band-specific reconstruction guided by bio-optical constraints for the first time. SpecTM is trained and evaluated within a multi-task self-supervised framework on NASA PACE hyperspectral data, jointly optimizing spectral reconstruction, bio-optical index inference, and 8-day temporal forecasting. On the task of predicting microcystin concentrations in Lake Erie, our method achieves R² scores of 0.695 for current-week and 0.620 for 8-day-ahead predictions—improving over the strongest baseline by 34% and 99%, respectively—and demonstrates a 2.2× gain in label efficiency.
📝 Abstract
Foundation models are now increasingly being developed for Earth observation (EO), yet they often rely on stochastic masking that do not explicitly enforce physics constraints; a critical trustworthiness limitation, in particular for predictive models that guide public health decisions. In this work, we propose SpecTM (Spectral Targeted Masking), a physics-informed masking design that encourages the reconstruction of targeted bands from cross-spectral context during pretraining. To achieve this, we developed an adaptable multi-task (band reconstruction, bio-optical index inference, and 8-day-ahead temporal prediction) self-supervised learning (SSL) framework that encodes spectrally intrinsic representations via joint optimization, and evaluated it on a downstream microcystin concentration regression model using NASA PACE hyperspectral imagery over Lake Erie. SpecTM achieves R^2 = 0.695 (current week) and R^2 = 0.620 (8-day-ahead) predictions surpassing all baseline models by (+34% (0.51 Ridge) and +99% (SVR 0.31)) respectively. Our ablation experiments show targeted masking improves predictions by +0.037 R^2 over random masking. Furthermore, it outperforms strong baselines with 2.2x superior label efficiency under extreme scarcity. SpecTM enables physics-informed representation learning across EO domains and improves the interpretability of foundation models.