MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion

📅 2025-09-16
📈 Citations: 0
Influential: 0
📄 PDF

career value

199K/year
🤖 AI Summary
Existing infrared–visible image fusion methods predominantly rely on low-level visual features or unstructured textual descriptions, limiting their capacity to model high-level semantics and spatial relationships. To address this, we propose MSGFusion—a novel framework that introduces structured scene graphs (explicitly encoding entities, attributes, and spatial relations) as multimodal semantic guidance for the first time. MSGFusion jointly learns semantically aligned graph representations across vision and language modalities, and achieves synergistic optimization of high-level semantics and low-level details via scene graph representation learning, hierarchical feature aggregation, and graph-driven fusion. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly enhances fused images in terms of detail preservation, structural fidelity, and semantic consistency. Moreover, it exhibits superior generalization performance on downstream tasks, including low-light object detection and semantic segmentation.

Technology Category

Application Category

📝 Abstract
Infrared and visible image fusion has garnered considerable attention owing to the strong complementarity of these two modalities in complex, harsh environments. While deep learning-based fusion methods have made remarkable advances in feature extraction, alignment, fusion, and reconstruction, they still depend largely on low-level visual cues, such as texture and contrast, and struggle to capture the high-level semantic information embedded in images. Recent attempts to incorporate text as a source of semantic guidance have relied on unstructured descriptions that neither explicitly model entities, attributes, and relationships nor provide spatial localization, thereby limiting fine-grained fusion performance. To overcome these challenges, we introduce MSGFusion, a multimodal scene graph-guided fusion framework for infrared and visible imagery. By deeply coupling structured scene graphs derived from text and vision, MSGFusion explicitly represents entities, attributes, and spatial relations, and then synchronously refines high-level semantics and low-level details through successive modules for scene graph representation, hierarchical aggregation, and graph-driven fusion. Extensive experiments on multiple public benchmarks show that MSGFusion significantly outperforms state-of-the-art approaches, particularly in detail preservation and structural clarity, and delivers superior semantic consistency and generalizability in downstream tasks such as low-light object detection, semantic segmentation, and medical image fusion.
Problem

Research questions and friction points this paper is trying to address.

Overcoming reliance on low-level visual cues in image fusion
Incorporating structured semantic guidance beyond unstructured text descriptions
Enhancing fine-grained fusion performance with explicit entity-relationship modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scene graph-guided fusion framework
Structured scene graph coupling
Hierarchical aggregation and graph-driven fusion