Efficient Image Super-Resolution with Multi-Scale Spatial Adaptive Attention Networks

📅 2026-02-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing image super-resolution methods struggle to balance reconstruction quality and model complexity. To address this challenge, this work proposes a lightweight Multi-scale Spatially Adaptive Attention Network (MSAAN), whose core component is the Multi-scale Spatially Adaptive Attention (MSAA) module. The MSAA module integrates Global Feature Modulation (GFM) and Multi-scale Feature Aggregation (MFA), complemented by a Local Enhancement Block (LEB) and a Feature Interaction Gated Feed-Forward (FIGFF) module to efficiently model both local details and long-range dependencies. Extensive experiments demonstrate that MSAAN and its lightweight variant achieve state-of-the-art or competitive performance in terms of PSNR and SSIM across multiple benchmark datasets, while significantly reducing model parameters and computational overhead.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Large Multimodal Models (LMMs)Multiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Large pretrained models with web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
This paper introduces a lightweight image super-resolution (SR) network, termed the Multi-scale Spatial Adaptive Attention Network (MSAAN), to address the common dilemma between high reconstruction fidelity and low model complexity in existing SR methods. The core of our approach is a novel Multi-scale Spatial Adaptive Attention Module (MSAA), designed to jointly model fine-grained local details and long-range contextual dependencies. The MSAA comprises two synergistic components: a Global Feature Modulation Module (GFM) that learns coherent texture structures through differential feature extraction, and a Multi-scale Feature Aggregation Module (MFA) that adaptively fuses features from local to global scales using pyramidal processing. To further enhance the network's capability, we propose a Local Enhancement Block (LEB) to strengthen local geometric perception and a Feature Interactive Gated Feed-Forward Module (FIGFF) to improve nonlinear representation while reducing channel redundancy. Extensive experiments on standard benchmarks (Set5, Set14, B100, Urban100, Manga109) across $\times2$, $\times3$, and $\times4$ scaling factors demonstrate that both our lightweight (MSAAN-light) and standard (MSAAN) versions achieve superior or competitive performance in terms of PSNR and SSIM, while maintaining significantly lower parameters and computational costs than state-of-the-art methods. Ablation studies validate the contribution of each component, and visual results show that MSAAN reconstructs sharper edges and more realistic textures.
Problem

Research questions and friction points this paper is trying to address.

Image Super-Resolution
Model Complexity
Reconstruction Fidelity
Lightweight Network
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-scale Spatial Adaptive Attention
Lightweight Super-Resolution
Global Feature Modulation
Feature Aggregation
Gated Feed-Forward
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sushi Rao
Zhejiang Gongshang University
J
Jingwei Li
Zhejiang Gongshang University