🤖 AI Summary
This study addresses the challenges of detail loss and dataset discrepancies in satellite image time series crop segmentation by proposing PAtteRNS, a hybrid Transformer-convolutional model. Its core innovation lies in the first parallel dimension self-attention architecture tailored for Sentinel-2 multispectral data, which independently decouples attention computation across temporal, spectral, and spatial dimensions, achieving fully factorized attention while significantly reducing computational complexity. Experimental results demonstrate that PAtteRNS surpasses existing state-of-the-art methods on the PASTIS and MTLCC datasets, exhibiting particularly notable advantages in parcel boundary delineation quality. Furthermore, this work reveals the potential impact of dataset grouping deficiencies on model performance.
📝 Abstract
The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.