Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between spatial-domain supervision and downstream objectives in multimodal image fusion caused by the absence of ground-truth references. To overcome this, it proposes a relation-constrained supervision paradigm that shifts supervision from the spatial domain to a learned relational space. Methodologically, shared, dominant, and coordination losses are defined using features from pretrained models. A learnable feature adapter is introduced to align heterogeneous DINO and CLIP representations, while a self-supervised contrastive ranking objective optimizes this supervision space. Training employs frozen pretrained backbones with an alternating optimization strategy. Experiments demonstrate that the proposed method consistently improves fusion performance across various mainstream backbone architectures, establishing a supervision mechanism better aligned with downstream tasks.
📝 Abstract
Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: github.com/GMY628/RCS-Fusion.
Problem

Research questions and friction points this paper is trying to address.

Multi-modal image fusion
Supervision misalignment
Ground-truth absence
Modality bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Modal Image Fusion
Relation-Constrained Supervision
Feature Adapter
Self-Supervised Contrastive Ranking
Alternating Optimization
Z
Zeyu Wang
College of Computer Science and Engineering, Dalian Minzu University
M
Mingyu Ge
College of Computer Science and Engineering, Dalian Minzu University
H
Haiyu Song
College of Computer Science and Engineering, Dalian Minzu University
Haoran Duan
Haoran Duan
Tsinghua/Newcastle/Durham University
Multimodal AIGenerative AI