Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection

📅 2025-08-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address detail loss, blurred boundaries, and insufficient multimodal fusion in salient object detection under complex scenes, this paper proposes a Dynamic Uncertainty Propagation and Multimodal Collaborative Reasoning Network. Methodologically: (1) a sparse graph based on spatial-semantic distance is constructed to enable inter-layer uncertainty propagation, coupled with channel-adaptive interaction for enhanced detail perception; (2) a learnable modality-gating mechanism is designed for multimodal collaborative fusion, balancing cross-modal semantic complementarity and interference suppression; (3) the framework integrates dynamic uncertainty-aware graph convolution, multimodal attention-weighted fusion, multi-scale BCE/IoU joint loss, cross-scale consistency constraints, and uncertainty-guided supervision. Extensive experiments demonstrate significant improvements over state-of-the-art methods across multiple benchmarks—particularly in localizing small objects, preserving boundary sharpness, and enhancing robustness under occlusion and weak-texture conditions.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
In view of the problems that existing salient object detection (SOD) methods are prone to losing details, blurring edges, and insufficient fusion of single-modal information in complex scenes, this paper proposes a dynamic uncertainty propagation and multimodal collaborative reasoning network (DUP-MCRNet). Firstly, a dynamic uncertainty graph convolution module (DUGC) is designed to propagate uncertainty between layers through a sparse graph constructed based on spatial semantic distance, and combined with channel adaptive interaction, it effectively improves the detection accuracy of small structures and edge regions. Secondly, a multimodal collaborative fusion strategy (MCF) is proposed, which uses learnable modality gating weights to weightedly fuse the attention maps of RGB, depth, and edge features. It can dynamically adjust the importance of each modality according to different scenes, effectively suppress redundant or interfering information, and strengthen the semantic complementarity and consistency between cross-modalities, thereby improving the ability to identify salient regions under occlusion, weak texture or background interference. Finally, the detection performance at the pixel level and region level is optimized through multi-scale BCE and IoU loss, cross-scale consistency constraints, and uncertainty-guided supervision mechanisms. Extensive experiments show that DUP-MCRNet outperforms various SOD methods on most common benchmark datasets, especially in terms of edge clarity and robustness to complex backgrounds. Our code is publicly available at https://github.com/YukiBear426/DUP-MCRNet.
Problem

Research questions and friction points this paper is trying to address.

Improves detection of small structures and edge regions in complex scenes
Dynamically fuses RGB, depth, and edge features using adaptive weighting
Enhances robustness against occlusion, weak textures, and background interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-based uncertainty propagation for edge accuracy
Multimodal fusion with adaptive gating weights
Multi-scale loss and uncertainty-guided supervision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuqi Xiong
Guangdong Key Laboratory of Intelligent Information Processing, College of Electronics and Information Engineering, Shenzhen University, Shenzhen, China
Wuzhen Shi
Wuzhen Shi
Shenzhen University
Image/Video Compression and EnhancementAffective ComputingAIGC
Yang Wen
Yang Wen
Guangdong Key Laboratory of Intelligent Information Processing, College of Electronics and Information Engineering, Shenzhen University, Shenzhen, China
R
Ruhan Liu
Furong Laboratory, Central South University, Changsha, China