🤖 AI Summary
To address limitations of conventional CNN/GAN-based super-resolution (SR) models—including restricted receptive fields, insufficient global modeling, and difficulty in recovering high-frequency details—this paper presents a systematic survey and advancement of Transformer-based SR methods. We propose the first comprehensive taxonomy of Transformer-SR evolution and introduce a multi-scale global-local collaborative architectural paradigm, uncovering the critical role of cross-scale attention in texture reconstruction. Our approach integrates ViT, Swin Transformer, cross-attention mechanisms, and CNN-Transformer hybrid designs, augmented with degradation modeling, frequency-domain supervision, and interpretability analysis. On standard benchmarks (Set5, Set14, Urban100), our method achieves significant PSNR/SSIM gains over state-of-the-art methods—up to +0.82 dB—and we publicly release both a technical roadmap and representative failure cases. Key contributions include: (i) a novel architectural paradigm; (ii) mechanistic insights into cross-scale attention; and (iii) rigorous, large-scale empirical validation.
📝 Abstract
Image Super-Resolution (SR) aims to recover a high-resolution image from its low-resolution counterpart, which has been affected by a specific degradation process. This is achieved by enhancing detail and visual quality. Recent advancements in transformer-based methods have remolded image super-resolution by enabling high-quality reconstructions surpassing previous deep-learning approaches like CNN and GAN-based. This effectively addresses the limitations of previous methods, such as limited receptive fields, poor global context capture, and challenges in high-frequency detail recovery. Additionally, the paper reviews recent trends and advancements in transformer-based SR models, exploring various innovative techniques and architectures that combine transformers with traditional networks to balance global and local contexts. These neoteric methods are critically analyzed, revealing promising yet unexplored gaps and potential directions for future research. Several visualizations of models and techniques are included to foster a holistic understanding of recent trends. This work seeks to offer a structured roadmap for researchers at the forefront of deep learning, specifically exploring the impact of transformers on super-resolution techniques.