Rethinking Normalization Strategies and Convolutional Kernels for Multimodal Image Fusion

📅 2024-11-15
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing multimodal image fusion (MMIF) methods overlook architectural constraints—such as normalization schemes and convolutional kernel design—on feature representation, particularly where batch normalization attenuates sparse, salient features. Addressing fundamental disparities between natural and medical image fusion, this work proposes: (1) a hybrid normalization strategy combining instance and group normalization to enhance both feature independence and intrinsic correlation; (2) large-kernel convolutions to expand receptive fields; and (3) a multi-path adaptive fusion module for cross-scale feature recalibration. The resulting end-to-end network achieves state-of-the-art performance across diverse fusion benchmarks. Moreover, it significantly improves downstream task accuracy and fine-detail preservation in medical diagnosis and remote sensing analysis, demonstrating superior generalizability and fidelity.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Multimodal image fusion (MMIF) aims to integrate information from different modalities to obtain a comprehensive image, aiding downstream tasks. However, existing methods tend to prioritize natural image fusion and focus on information complementary and network training strategies. They ignore the essential distinction between natural and medical image fusion and the influence of underlying components. This paper dissects the significant differences between the two tasks regarding fusion goals, statistical properties, and data distribution. Based on this, we rethink the suitability of the normalization strategy and convolutional kernels for end-to-end MMIF.Specifically, this paper proposes a mixture of instance normalization and group normalization to preserve sample independence and reinforce intrinsic feature correlation.This strategy promotes the potential of enriching feature maps, thus boosting fusion performance. To this end, we further introduce the large kernel convolution, effectively expanding receptive fields and enhancing the preservation of image detail. Moreover, the proposed multipath adaptive fusion module recalibrates the decoder input with features of various scales and receptive fields, ensuring the transmission of crucial information. Extensive experiments demonstrate that our method exhibits state-of-the-art performance in multiple fusion tasks and significantly improves downstream applications. The code is available at https://github.com/HeDan-11/LKC-FUNet.
Problem

Research questions and friction points this paper is trying to address.

Reevaluates UNet's normalization and convolution for multimodal image fusion
Proposes hybrid normalization to preserve sparse features and enhance correlations
Introduces adaptive fusion module for dynamic multi-scale feature calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid instance and group normalization for feature preservation
Large kernel convolution to enhance receptive field usage
Multi-path adaptive fusion module for dynamic feature calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.