🤖 AI Summary
This work addresses the longstanding challenge in multispectral and hyperspectral image fusion—balancing spatial detail enhancement with spectral fidelity—where existing methods struggle with cross-scale interaction and joint spatial-spectral modeling. The authors propose CoFusion, a novel framework featuring a three-level multiscale pyramid architecture. Each level employs a dual-branch design: SpaCAM captures multiscale contextual information through a spatial coordinate-aware mixing mechanism, while SpeCAM enhances spectral representation by integrating frequency-domain decomposition with coordinate attention. A spatial-spectral cross-fusion module (SSCFM) further enables dynamic cross-modal alignment and complementary feature integration. CoFusion is the first to jointly model cross-scale and cross-modal dependencies, achieving state-of-the-art performance across multiple benchmark datasets with superior spatial reconstruction quality and spectral fidelity.
📝 Abstract
Multispectral and Hyperspectral Image Fusion (MHIF) aims to reconstruct high-resolution images by integrating low-resolution hyperspectral images (LRHSI) and high-resolution multispectral images (HRMSI). However, existing methods face limitations in modeling cross-scale interactions and spatial-spectral collaboration, making it difficult to achieve an optimal trade-off between spatial detail enhancement and spectral fidelity. To address this challenge, we propose CoFusion: a unified spatial-spectral collaborative fusion framework that explicitly models cross-scale and cross-modal dependencies. Specifically, a Multi-Scale Generator (MSG) is designed to construct a three-level pyramidal architecture, enabling the effective integration of global semantics and local details. Within each scale, a dual-branch strategy is employed: the Spatial Coordinate-Aware Mixing module (SpaCAM) is utilized to capture multi-scale spatial contexts, while the Spectral Coordinate-Aware Mixing module (SpeCAM) enhances spectral representations through frequency decomposition and coordinate mixing. Furthermore, we introduce the Spatial-Spectral Cross-Fusion Module (SSCFM) to perform dynamic cross-modal alignment and complementary feature fusion. Extensive experiments on multiple benchmark datasets demonstrate that CoFusion consistently outperforms state-of-the-art methods, achieving superior performance in both spatial reconstruction and spectral consistency.