Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion
This study addresses the misalignment between spatial-domain supervision and downstream objectives in multimodal image fusion caused by the absence of ground-truth references. To overcome this, it proposes a relation-constrained supervision paradigm that shifts supervision from the spatial domain to a learned relational space. Methodologically, shared, dominant, and coordination losses are defined using features from pretrained models. A learnable feature adapter is introduced to align heterogeneous DINO and CLIP representations, while a self-supervised contrastive ranking objective optimizes this supervision space. Training employs frozen pretrained backbones with an alternating optimization strategy. Experiments demonstrate that the proposed method consistently improves fusion performance across various mainstream backbone architectures, establishing a supervision mechanism better aligned with downstream tasks.