Score
Designs and implements models and training procedures that randomly mask portions of spectral measurements and learn to reconstruct the missing wavelengths or predict masked spectral values, producing contextualized spectral representations. Practitioners build masking schedules, reconstruction losses, and evaluation pipelines to analyze and improve spectral reconstruction and downstream tasks such as material identification or spectral unmixing.
To address the scarcity of labeled data for hyperspectral image (HSI) analysis, which severely limits Transformer-based model performance, this paper proposes Spatial-Frequency Masked Imaging Modeling (SFMIM), a self-supervised pretraining framework. SFMIM is the first method to explicitly couple spatial locality and spectral frequency-domain structure in HSI pretraining: it applies random masking to non-overlapping spectral patches in the spatial domain, while simultaneously masking spectral Fourier coefficients in the frequency domain via FFT, followed by inverse FFT reconstruction—enabling joint spatial-spectral modeling. Built upon a Transformer encoder, SFMIM performs end-to-end optimization of dual-domain reconstruction objectives. Evaluated on three public HSI classification benchmarks, it achieves state-of-the-art (SOTA) performance. Moreover, fine-tuning converges significantly faster and attains high accuracy with only a small number of labeled samples, demonstrating strong transferability and data efficiency.
This work identifies fundamental limitations of RGB-to-hyperspectral reconstruction: data bias, model overfitting, and strong dependence on camera optical characteristics—leading to poor robustness against noise, JPEG compression, and metamerism. To address these issues, we propose the first optics-aware analytical framework, integrating chameleon-style metameric data augmentation with physics-driven lens aberration modeling to explicitly enhance RGB images’ capacity to encode spectrally ambiguous information. Through spectral modeling, aberration simulation, and mapping network analysis, we systematically validate the framework’s efficacy: it significantly improves reconstruction accuracy and stability, achieving robust performance even in challenging metameric scenarios. Our approach establishes a new paradigm for hyperspectral reconstruction research by unifying physical interpretability with generalizability—bridging the gap between data-driven learning and optical physics.
This work addresses the weak cross-model and cross-resolution generalization in AI-generated image detection. We propose a self-supervised detection framework grounded in spectral priors. Methodologically, we introduce— for the first time—spectral reconstruction similarity metrics and spectral contextual attention, integrated with masked spectral learning and frequency-domain reconstruction pretraining, enabling robust perception of subtle spectral anomalies in images of arbitrary resolution. Our core contribution lies in leveraging the invariance and discriminability of natural image spectral distributions, thereby eliminating reliance on specific generative models. Evaluated on 13 unseen generative models, our method achieves a 5.5% AUC improvement over state-of-the-art methods. Moreover, it demonstrates strong robustness against common online perturbations, including JPEG compression and spatial rescaling.
This paper addresses the pervasive data incompleteness in remote sensing imagery—caused by cloud cover, occlusion, and sensor failures—by systematically surveying mask image modeling (MIM) for self-supervised pretraining in this domain. It establishes the first comprehensive taxonomy of remote sensing MIM methodologies, clarifying masking strategies (pixel-, patch-, and feature-level), architectural evolution (Transformer- vs. CNN-based), multi-source data fusion techniques, and downstream task adaptation (e.g., cloud removal, super-resolution). Synthesizing over 100 state-of-the-art works, it identifies key performance bottlenecks and proposes a unified evaluation protocol. The study further outlines three critical future directions: scalable pretraining, cross-modal alignment, and physics-informed modeling. Collectively, this work formalizes the first systematic research framework for MIM in remote sensing, bridging methodological rigor with practical applicability.
This work addresses the challenge of limited labeled samples in vibrational spectroscopy, which constrains both the discriminative power and interpretability of deep learning models. To overcome this, we propose Task-enhanced Augmentation Network (TeaNet), a novel approach that generates in-domain augmented samples through random spectral masking and reconstruction, trained end-to-end with the classifier in a task-driven manner. This reconstruction objective explicitly encourages the model to learn informative wavenumber features, thereby improving classification accuracy under few-shot conditions while enhancing model interpretability. Experimental results demonstrate that TeaNet consistently outperforms conventional CNNs on both synthetic and real-world datasets, achieving up to a 17% accuracy gain in the most challenging synthetic scenario and more accurately identifying diagnostically relevant wavenumbers.
This work addresses the challenge of hyperspectral image restoration, which is hindered by data scarcity, sensor-specific characteristics, and the high dimensionality of spectral information, making it difficult to learn robust priors. The authors propose a lightweight transfer framework that projects hyperspectral data into a low-dimensional subspace, leverages a frozen pre-trained RGB denoiser for noise removal, and reconstructs the hyperspectral cube through a lightweight adapter coupled with constrained linear aggregation. This approach is the first to efficiently transfer large-scale RGB image priors to hyperspectral restoration tasks, achieving plug-and-play performance with minimal training. It consistently outperforms specialized hyperspectral methods across multiple datasets in denoising, deblurring, and super-resolution, demonstrating the remarkable transferability of RGB-based priors.
This work addresses the challenge of effectively learning spatial-spectral features from multispectral remote sensing imagery, where complex backgrounds, ambiguous targets, and lack of semantic guidance hinder existing Masked Autoencoders. To overcome this, we propose SIGMAE, a novel approach that incorporates domain-specific spectral indices as prior knowledge to design a Semantic Saliency-guided Dynamic Token Masking (SSDTM) strategy. SSDTM adaptively selects and prioritizes the reconstruction of information-rich regions, progressively increasing task difficulty through curriculum learning. Evaluated on five remote sensing datasets, SIGMAE substantially outperforms current geospatial foundation models, enabling high-quality image reconstruction even at mask ratios up to 90% and significantly improving complex target recognition performance under limited labeled data conditions.
This study addresses the limitations of conventional self-supervised pretraining in Earth observation foundation models, which often employ physically unconstrained random masking and thus fail to meet the trustworthiness requirements of high-stakes applications such as public health decision-making. To overcome this, we propose SpecTM—a spectrally targeted masking strategy that integrates biophysical priors into the masking process, enabling band-specific reconstruction guided by bio-optical constraints for the first time. SpecTM is trained and evaluated within a multi-task self-supervised framework on NASA PACE hyperspectral data, jointly optimizing spectral reconstruction, bio-optical index inference, and 8-day temporal forecasting. On the task of predicting microcystin concentrations in Lake Erie, our method achieves R² scores of 0.695 for current-week and 0.620 for 8-day-ahead predictions—improving over the strongest baseline by 34% and 99%, respectively—and demonstrates a 2.2× gain in label efficiency.
This work investigates the generalization performance and spectral structure of matrix-valued predictors formed by aggregating multiple masks in masked self-supervised learning under high-dimensional settings. Leveraging random matrix theory within an asymptotic framework where sample size and dimension grow proportionally, the study establishes the first high-dimensional theoretical analysis for masked self-supervised learning. The core contributions include deriving an explicit expression for the generalization error, characterizing the spectral properties of the aggregated predictor, revealing a BBP-type phase transition under spiked covariance models, and identifying the precise threshold conditions under which latent signals can be effectively recovered. The analysis further provides theoretical evidence that this approach outperforms classical PCA in certain structured scenarios.