Institution profile

Dalian Minzu University

Academic institutionasia · cn
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

Sep 30, 2026

This study addresses the misalignment between spatial-domain supervision and downstream objectives in multimodal image fusion caused by the absence of ground-truth references. To overcome this, it proposes a relation-constrained supervision paradigm that shifts supervision from the spatial domain to a learned relational space. Methodologically, shared, dominant, and coordination losses are defined using features from pretrained models. A learnable feature adapter is introduced to align heterogeneous DINO and CLIP representations, while a self-supervised contrastive ranking objective optimizes this supervision space. Training employs frozen pretrained backbones with an alternating optimization strategy. Experiments demonstrate that the proposed method consistently improves fusion performance across various mainstream backbone architectures, establishing a supervision mechanism better aligned with downstream tasks.

0 citationsRead paper

When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion

Sep 30, 2026

This study addresses the limitations of existing feature decomposition methods in multimodal image fusion, which lack explicit supervision and rely on potentially conflicting pixel-level metrics. We propose a signal-level self-supervised feature decomposition paradigm whose core innovation lies in reformulating two-dimensional image supervision as a one-dimensional integral optimization problem. By replacing ambiguous 2D supervision with well-defined integral constraints, this approach fundamentally resolves the challenge of loss conflicts. Building upon this formulation, we design a two-stage self-supervised learning framework that integrates signal-level integral-driven decomposition, structure-preserving reconstruction, and a distinctive feature fusion strategy. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple representative multimodal fusion tasks. The source code has been made publicly available.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Spatial-Domain Supervision: A Relation Constrained Space for Multi-Modal Image Fusion

Sep 30, 2026

This study addresses the misalignment between spatial-domain supervision and downstream objectives in multimodal image fusion caused by the absence of ground-truth references. To overcome this, it proposes a relation-constrained supervision paradigm that shifts supervision from the spatial domain to a learned relational space. Methodologically, shared, dominant, and coordination losses are defined using features from pretrained models. A learnable feature adapter is introduced to align heterogeneous DINO and CLIP representations, while a self-supervised contrastive ranking objective optimizes this supervision space. Training employs frozen pretrained backbones with an alternating optimization strategy. Experiments demonstrate that the proposed method consistently improves fusion performance across various mainstream backbone architectures, establishing a supervision mechanism better aligned with downstream tasks.

0 citationsRead paper

When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion

Sep 30, 2026

This study addresses the limitations of existing feature decomposition methods in multimodal image fusion, which lack explicit supervision and rely on potentially conflicting pixel-level metrics. We propose a signal-level self-supervised feature decomposition paradigm whose core innovation lies in reformulating two-dimensional image supervision as a one-dimensional integral optimization problem. By replacing ambiguous 2D supervision with well-defined integral constraints, this approach fundamentally resolves the challenge of loss conflicts. Building upon this formulation, we design a two-stage self-supervised learning framework that integrates signal-level integral-driven decomposition, structure-preserving reconstruction, and a distinctive feature fusion strategy. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across multiple representative multimodal fusion tasks. The source code has been made publicly available.

0 citationsRead paper