Attention-Guided Multi-scale Interaction Network for Face Super-Resolution

📅 2024-09-01
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient and weakly complementary multi-scale feature fusion in face super-resolution (FSR), this paper proposes a hybrid CNN-Transformer architecture. Our method introduces three key innovations: (1) a Local-Global Feature Interaction (LGFI) module that enables cross-scale and cross-stage collaborative modeling of local details and global structure; (2) a Selective Kernel Attention Fusion (SKAF) module that adaptively weights and fuses multi-scale features via kernel-wise attention; and (3) Residual Deep Feature Extraction (RDFE), which enhances deep representation learning through hierarchical residual connections. Evaluated on standard benchmarks—including CelebA and FFHQ—our approach achieves state-of-the-art (SOTA) performance, significantly improving high-frequency texture recovery and visual fidelity. Moreover, it reduces computational overhead and accelerates inference speed without compromising accuracy.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalization
📝 Abstract
Recently, CNN and Transformer hybrid networks demonstrated excellent performance in face super-resolution (FSR) tasks. Since numerous features at different scales in hybrid networks, how to fuse these multi-scale features and promote their complementarity is crucial for enhancing FSR. However, existing hybrid network-based FSR methods ignore this, only simply combining the Transformer and CNN. To address this issue, we propose an attention-guided Multi-scale interaction network (AMINet), which contains local and global feature interactions and encoder-decoder phase feature interactions. Specifically, we propose a Local and Global Feature Interaction Module (LGFI) to promote fusions of global features and different receptive fields' local features extracted by our Residual Depth Feature Extraction Module (RDFE). Additionally, we propose a Selective Kernel Attention Fusion Module (SKAF) to adaptively select fusions of different features within LGFI and encoder-decoder phases. Our above design allows the free flow of multi-scale features from within modules and between encoder and decoder, which can promote the complementarity of different scale features to enhance FSR. Comprehensive experiments confirm that our method consistently performs well with less computational consumption and faster inference.
Problem

Research questions and friction points this paper is trying to address.

Fuse multi-scale features in hybrid networks for face super-resolution
Enhance complementarity of global and local features in FSR
Adaptively select feature fusions across different network phases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention-guided multi-scale interaction network
Local and global feature interaction module
Selective kernel attention fusion module
🔎 Similar Papers
No similar papers found.
Nanjing University of Posts and Telecommunications | Soochow University | Beijing University of Posts and Telecommunications | Southeast University | Nanjing University of Science and Technology | National Tsing Hua University
X
Xujie Wan
Institute of Advanced Technology, Nanjing University of Posts and Telecommunications, Key Laboratory of Artificial Intelligence, Ministry of Education, Provincial Key Laboratory for Computer Information Processing Technology, Soochow University
W
Wenjie Li
Pattern Recognition and Intelligent System Laboratory, School of Artificial Intelligence, Beijing University of Posts and Telecommunications
Guangwei Gao
Guangwei Gao
Professor of PCALab@NJUST, IEEE/CCF/CSIG/CAAI/CAA Senior Member
Pattern RecognitionImage UnderstandingMachine Learning
Huimin Lu
Huimin Lu
National University of Defense Technology
Robot VisionMulti-robot CoordinationRobot SoccerRobot Rescue
J
Jian Yang
School of Computer Science and Technology, Nanjing University of Science and Technology
C
Chia-Wen Lin
Department of Electrical Engineering, National Tsing Hua University