Cross-modal State Space Modeling for Real-time RGB-thermal Wild Scene Semantic Segmentation

📅 2025-06-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of real-time RGB-thermal multimodal semantic segmentation under resource-constrained野外 settings, this paper proposes CM-SSM—the first efficient architecture to introduce linear-complexity State Space Models (SSMs) into multimodal segmentation. Its core innovation lies in a cross-modal 2D selective scanning mechanism coupled with a state-space correlation module, enabling synergistic fusion of global contextual modeling and local feature representation. Integrated with a lightweight convolutional backbone and a dual-stream feature interaction mechanism, CM-SSM achieves significant reductions in both parameter count and FLOPs while preserving linear computational complexity. On the CART dataset, CM-SSM establishes new state-of-the-art performance; on PST900, it demonstrates strong cross-dataset generalization. Overall, CM-SSM provides a deployable, efficient paradigm for edge-device multimodal perception.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
The integration of RGB and thermal data can significantly improve semantic segmentation performance in wild environments for field robots. Nevertheless, multi-source data processing (e.g. Transformer-based approaches) imposes significant computational overhead, presenting challenges for resource-constrained systems. To resolve this critical limitation, we introduced CM-SSM, an efficient RGB-thermal semantic segmentation architecture leveraging a cross-modal state space modeling (SSM) approach. Our framework comprises two key components. First, we introduced a cross-modal 2D-selective-scan (CM-SS2D) module to establish SSM between RGB and thermal modalities, which constructs cross-modal visual sequences and derives hidden state representations of one modality from the other. Second, we developed a cross-modal state space association (CM-SSA) module that effectively integrates global associations from CM-SS2D with local spatial features extracted through convolutional operations. In contrast with Transformer-based approaches, CM-SSM achieves linear computational complexity with respect to image resolution. Experimental results show that CM-SSM achieves state-of-the-art performance on the CART dataset with fewer parameters and lower computational cost. Further experiments on the PST900 dataset demonstrate its generalizability. Codes are available at https://github.com/xiaodonguo/CMSSM.
Problem

Research questions and friction points this paper is trying to address.

Improve RGB-thermal semantic segmentation in wild environments
Reduce computational overhead in multi-source data processing
Achieve linear computational complexity with image resolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-modal state space modeling for RGB-thermal segmentation
CM-SS2D module enables cross-modal hidden state derivation
Linear complexity CM-SSM outperforms Transformer-based methods
🔎 Similar Papers
2024-04-05IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 56
💼 Related Jobs
No related jobs found.
X
Xiaodong Guo
School of Automation, Beijing Institute of Technology, Beijing 100081, China
Z
Zi'ang Lin
School of Automation, Beijing Institute of Technology, Beijing 100081, China
L
Luwen Hu
School of Automation, Beijing Institute of Technology, Beijing 100081, China
Zhihong Deng
Zhihong Deng
Faculty of Engineering and Information Technology, University of Technology Syndey
reinforcement learningrecommender systemscausal inference
T
Tong Liu
School of Automation, Beijing Institute of Technology, Beijing 100081, China
W
Wujie Zhou
School of Information and Electronic Engineering, Zhejiang University of Science and Technology, Hangzhou 310023, China, and also with the School of Computer Science and Engineering, Nanyang Technological University, Singapore 308232