🤖 AI Summary
To address the representation learning challenges posed by the high-dimensional spatial-spectral coupling in hyperspectral imagery, this paper proposes a Transformer-based dual-masked self-supervised pretraining framework. Methodologically, it introduces a novel joint 50% masking strategy that simultaneously masks spatial patches and spectral bands, enforcing cross-dimensional complementary representation learning; incorporates wavelength-aware learnable harmonic Fourier positional encoding; and jointly optimizes pixel-wise fidelity and spectral shape preservation via MSE and Spectral Angle Mapper (SAM) losses. Extensive unsupervised pretraining is conducted on Hyperion and EnMAP datasets, followed by fine-tuning on Indian Pines, achieving state-of-the-art classification accuracy. The model comprises 180 million parameters and outputs 768-dimensional embeddings. Ablation studies validate the efficacy of both the dual-masking mechanism and the wavelength-aware positional encoding for hyperspectral representation learning.
📝 Abstract
Hyperspectral imagery provides rich spectral detail but poses unique challenges because of its high dimensionality in both spatial and spectral domains. We propose extit{HyperspectralMAE}, a Transformer-based foundation model for hyperspectral data that employs a extit{dual masking} strategy: during pre-training we randomly occlude 50% of spatial patches and 50% of spectral bands. This forces the model to learn representations capable of reconstructing missing information across both dimensions. To encode spectral order, we introduce learnable harmonic Fourier positional embeddings based on wavelength. The reconstruction objective combines mean-squared error (MSE) with the spectral angle mapper (SAM) to balance pixel-level accuracy and spectral-shape fidelity. The resulting model contains about $1.8 imes10^{8}$ parameters and produces 768-dimensional embeddings, giving it sufficient capacity for transfer learning. We pre-trained HyperspectralMAE on two large hyperspectral corpora -- NASA EO-1 Hyperion ($sim$1,600 scenes, $sim$$3 imes10^{11}$ pixel spectra) and DLR EnMAP Level-0 ($sim$1,300 scenes, $sim$$3 imes10^{11}$ pixel spectra) -- and fine-tuned it for land-cover classification on the Indian Pines benchmark. HyperspectralMAE achieves state-of-the-art transfer-learning accuracy on Indian Pines, confirming that masked dual-dimensional pre-training yields robust spectral-spatial representations. These results demonstrate that dual masking and wavelength-aware embeddings advance hyperspectral image reconstruction and downstream analysis.