🤖 AI Summary
This study addresses the limitations of existing feature decomposition methods in multimodal image fusion, which lack explicit supervision and rely on potentially conflicting pixel-level metrics. We propose a signal-level self-supervised feature decomposition paradigm whose core innovation lies in reformulating two-dimensional image supervision as a one-dimensional integral optimization problem. By replacing ambiguous 2D supervision with well-defined integral constraints, this approach fundamentally resolves the challenge of loss conflicts. Building upon this formulation, we design a two-stage self-supervised learning framework that integrates signal-level integral-driven decomposition, structure-preserving reconstruction, and a distinctive feature fusion strategy. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple representative multimodal fusion tasks. The source code has been made publicly available.
📝 Abstract
Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art (SOTA) performance. Code: github.com/Wangjiayu0512/SIDFusion.