🤖 AI Summary
This study addresses the challenges of code fragmentation, incompatible tensor formats, and absent evaluation standards in deep learning for hyperspectral imaging by introducing a unified, modular framework. The framework integrates 55 models spanning six architectural paradigms—including CNNs, Vision Transformers, and Mamba—across 24 benchmark scenes, enabling automatic tensor adaptation and standardized evaluation. Furthermore, it pioneers a spatially disjoint partitioning strategy incorporating Chebyshev guard bands to eliminate data leakage. Extensive experiments comprising 1,320 trials demonstrate that scene complexity, rather than architectural paradigm, predominantly governs model performance, and that lightweight models with fewer parameters can achieve accuracy comparable to larger counterparts (56.7%–96.4%). The source code is publicly available.
📝 Abstract
Hyperspectral remote sensing has advanced across diverse deep learning paradigms, including spectral spatial CNNs, Vision Transformers, Mamba, graph neural networks, Kolmogorov Arnold networks, and self supervised masked autoencoding. Yet progress remains hindered by fragmented repositories, incompatible tensor conventions, and non standardized evaluation. Hyperspectral Image Models addresses these challenges through a modular framework unifying 55 representative models across six paradigms with a common registry, automatic 4D/5D tensor adaptation, and standardized constructors. It integrates 24 benchmark scenes from Airborne, Spaceborne, UAV, and Mars CRISM sensors, with caching, label remapping, PCA, explicit band selection or raw spectra, optional spatial max pooling, and arbitrary PxP patch extraction. To prevent inflated accuracy from overlapping windows, it supports class balanced random partitioning and spatially disjoint regional blocking with Chebyshev guard bands that eliminate train test pixel overlap. Experiments use a single config.yaml with deterministic seeds and complete provenance, generating LaTeX benchmark tables and classification maps. Across 1,320 model scene evaluations and 6,600 seeded runs, scene difficulty dominates architecture, with mean accuracy ranging from 96.40% on Botswana to 56.70% on Houston 2018, versus a 15 point spread across paradigm means. No paradigm universally dominates, while sub 1 M parameter models can match architectures two orders of magnitude larger. Code is publicly available at https://github.com/Tanishq251/Hyperspectral-Image-Models.