🤖 AI Summary
This study addresses the misalignment between spatial-domain supervision and downstream objectives in multimodal image fusion caused by the absence of ground-truth references. To overcome this, it proposes a relation-constrained supervision paradigm that shifts supervision from the spatial domain to a learned relational space. Methodologically, shared, dominant, and coordination losses are defined using features from pretrained models. A learnable feature adapter is introduced to align heterogeneous DINO and CLIP representations, while a self-supervised contrastive ranking objective optimizes this supervision space. Training employs frozen pretrained backbones with an alternating optimization strategy. Experiments demonstrate that the proposed method consistently improves fusion performance across various mainstream backbone architectures, establishing a supervision mechanism better aligned with downstream tasks.
📝 Abstract
Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Rather than relying solely on direct source approximation, we further leverage frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Code: github.com/GMY628/RCS-Fusion.