🤖 AI Summary
This study investigates whether attention mechanisms outperform convolutional networks in classifying texture-dominant, imbalanced ice sheet imagery. We evaluate AlexNet, ConvNeXt, Swin Transformer, and CoAtNet on a Greenland Ice Sheet dataset using a multi-metric framework to assess the alignment between architectural inductive biases and classification reliability. Results demonstrate that increased model complexity does not necessarily yield performance gains; the lightweight AlexNet achieves the highest overall accuracy, purely convolutional architectures prove more reliable for texture-based classes, while hybrid models effectively improve recall for rare categories. These findings underscore the necessity of aligning architectural inductive biases with the intrinsic properties of ice sheet data, providing empirical guidance for architecture selection in polar remote sensing classification tasks.
📝 Abstract
Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice Sheet image dataset, previously shown to favor convolutional neural networks (CNNs), to examine whether modern attention-based and hybrid architectures improve class-wise reliability. We conduct a controlled comparison between a classical CNN (AlexNet), a modern CNN (ConvNeXt-Tiny), a pure attention-based model (Swin-Tiny), and a hybrid convolution-attention model (CoAtNet-0) under identical training and evaluation protocols. Results show that AlexNet achieves the highest accuracy and the strongest balanced performance as measured by macro-averaged F1, while ConvNeXt-Tiny exhibits the highest macro-averaged AUC, indicating strong class separability but less consistent final decision quality. Class-wise analysis reveals that hybrid architectures improve recall for rare and structurally distinct surface classes, whereas convolutional models remain more reliable for texture-dominated categories. These findings highlight the importance of aligning architectural inductive bias with cryospheric data characteristics and suggest that increased model complexity does not necessarily translate to improved reliability for ice-sheet surface classification.