AngularFuse: A Closer Look at Angle-based Perception for Spatial-Sensitive Multi-Modality Image Fusion

📅 2025-10-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing unsupervised visible-infrared image fusion methods suffer from handcrafted loss functions: reference images often lack fine details and exhibit uneven brightness, while gradient losses model only magnitude—ignoring directional information—leading to spatial structural distortion. To address these issues, we propose AngularFuse, the first angle-aware fusion framework that jointly constrains both gradient magnitude and direction in the gradient domain. We design a cross-modal complementary masking module and integrate Laplacian edge enhancement with adaptive histogram equalization to collaboratively generate high-quality reference images. Extensive experiments on MSRS, RoadScene, and M3FD datasets demonstrate that AngularFuse significantly outperforms state-of-the-art methods, especially under low-light conditions and complex backgrounds. It achieves superior detail fidelity and edge accuracy, thereby enhancing robustness and applicability for downstream vision tasks.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multi-instance/Multi-view Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
Visible-infrared image fusion is crucial in key applications such as autonomous driving and nighttime surveillance. Its main goal is to integrate multimodal information to produce enhanced images that are better suited for downstream tasks. Although deep learning based fusion methods have made significant progress, mainstream unsupervised approaches still face serious challenges in practical applications. Existing methods mostly rely on manually designed loss functions to guide the fusion process. However, these loss functions have obvious limitations. On one hand, the reference images constructed by existing methods often lack details and have uneven brightness. On the other hand, the widely used gradient losses focus only on gradient magnitude. To address these challenges, this paper proposes an angle-based perception framework for spatial-sensitive image fusion (AngularFuse). At first, we design a cross-modal complementary mask module to force the network to learn complementary information between modalities. Then, a fine-grained reference image synthesis strategy is introduced. By combining Laplacian edge enhancement with adaptive histogram equalization, reference images with richer details and more balanced brightness are generated. Last but not least, we introduce an angle-aware loss, which for the first time constrains both gradient magnitude and direction simultaneously in the gradient domain. AngularFuse ensures that the fused images preserve both texture intensity and correct edge orientation. Comprehensive experiments on the MSRS, RoadScene, and M3FD public datasets show that AngularFuse outperforms existing mainstream methods with clear margin. Visual comparisons further confirm that our method produces sharper and more detailed results in challenging scenes, demonstrating superior fusion capability.
Problem

Research questions and friction points this paper is trying to address.

Addresses limitations in visible-infrared image fusion methods
Proposes angle-based perception for spatial-sensitive multimodal fusion
Enhances texture intensity and edge orientation in fused images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Angle-based perception framework for image fusion
Cross-modal complementary mask module for learning
Angle-aware loss constrains gradient magnitude and direction
💼 Related Jobs
No related jobs found.
X
Xiaopeng Liu
School of Information Engineering, Guangdong University of Technology, Guangzhou, 510006, China
Y
Yupei Lin
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, 510006, China
S
Sen Zhang
TikTok, ByteDance Inc, Sydney, NSW 2000, Australia
X
Xiao Wang
School of Computer Science, Anhui University, Hefei, 230000, China
Y
Yukai Shi
School of Information Engineering, Guangdong University of Technology, Guangzhou, 510006, China
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis