Residual Prior-driven Frequency-aware Network for Image Fusion

📅 2025-07-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Image fusion faces two key challenges: high computational cost in modeling long-range spatial dependencies and difficulty in capturing complementary features due to the absence of ground-truth labels. To address these, we propose a residual-prior-guided frequency-domain fusion framework. Our method introduces a residual prior module to explicitly model cross-modal discrepancies and employs a dual-branch frequency-domain convolutional network enhanced by a bidirectional cross-promotion mechanism and a frequency-domain contrastive loss, enabling synergistic global–local feature learning. Additionally, an auxiliary decoder and an adaptive-weighted saliency-structure loss are incorporated to improve fine-detail preservation and target representation. Extensive experiments on multi-task benchmarks demonstrate that our approach significantly enhances fused image quality and consistently improves performance on downstream high-level vision tasks—including object detection and semantic segmentation—outperforming state-of-the-art methods.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multimodal Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalization
📝 Abstract
Image fusion aims to integrate complementary information across modalities to generate high-quality fused images, thereby enhancing the performance of high-level vision tasks. While global spatial modeling mechanisms show promising results, constructing long-range feature dependencies in the spatial domain incurs substantial computational costs. Additionally, the absence of ground-truth exacerbates the difficulty of capturing complementary features effectively. To tackle these challenges, we propose a Residual Prior-driven Frequency-aware Network, termed as RPFNet. Specifically, RPFNet employs a dual-branch feature extraction framework: the Residual Prior Module (RPM) extracts modality-specific difference information from residual maps, thereby providing complementary priors for fusion; the Frequency Domain Fusion Module (FDFM) achieves efficient global feature modeling and integration through frequency-domain convolution. Additionally, the Cross Promotion Module (CPM) enhances the synergistic perception of local details and global structures through bidirectional feature interaction. During training, we incorporate an auxiliary decoder and saliency structure loss to strengthen the model's sensitivity to modality-specific differences. Furthermore, a combination of adaptive weight-based frequency contrastive loss and SSIM loss effectively constrains the solution space, facilitating the joint capture of local details and global features while ensuring the retention of complementary information. Extensive experiments validate the fusion performance of RPFNet, which effectively integrates discriminative features, enhances texture details and salient objects, and can effectively facilitate the deployment of the high-level vision task.
Problem

Research questions and friction points this paper is trying to address.

Integrate complementary information across modalities for high-quality fused images
Reduce computational costs of long-range feature dependencies in spatial domain
Effectively capture complementary features without ground-truth data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Prior Module extracts modality-specific differences
Frequency Domain Fusion enables efficient global modeling
Cross Promotion Module enhances local-global feature synergy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guan Zheng
the School of Information Science and Engineering, Yunnan University, Kunming, Yunnan, China
X
Xue Wang
the School of Information Science and Engineering, Yunnan University, Kunming, Yunnan, China
W
Wenhua Qian
the School of Information Science and Engineering, Yunnan University, Kunming, Yunnan, China
P
Peng Liu
the School of Information Science and Engineering, Yunnan University, Kunming, Yunnan, China
Runzhuo Ma
Runzhuo Ma
Department of Urology, NewYork-Presbyterian Hospital, Cornell Medical Center
Urology