TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion

๐Ÿ“… 2025-09-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing RGB-T semantic segmentation models suffer from inadequate thermal feature extraction, inefficient cross-modal fusion, and encoder redundancy, hindering real-time perception for autonomous driving in complex scenarios. This paper proposes a lightweight unified multimodal architecture: (1) a streamlined thermal branch to enhance thermalโ€“infrared representation learning; (2) an adaptive cosine similarity mechanism for efficient local cross-modal feature fusion; and (3) joint optimization via large-scale RGB-to-pseudo-thermal pretraining and a lightweight encoder. The method achieves state-of-the-art performance on the FMB, PST900, and CART benchmarks, reducing model parameters by 32% and FLOPs by 41% compared to prior approaches, while attaining 27 FPS real-time inference on the Jetson Orin NX platform.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multimodal Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
๐Ÿ“ Abstract
RGB-thermal (RGB-T) semantic segmentation improves the environmental perception of autonomous platforms in challenging conditions. Prevailing models employ encoders pre-trained on RGB images to extract features from both RGB and infrared inputs, and design additional modules to achieve cross-modal feature fusion. This results in limited thermal feature extraction and suboptimal cross-modal fusion, while the redundant encoders further compromises the model's real-time efficiency. To address the above issues, we propose TUNI, with an RGB-T encoder consisting of multiple stacked blocks that simultaneously perform multi-modal feature extraction and cross-modal fusion. By leveraging large-scale pre-training with RGB and pseudo-thermal data, the RGB-T encoder learns to integrate feature extraction and fusion in a unified manner. By slimming down the thermal branch, the encoder achieves a more compact architecture. Moreover, we introduce an RGB-T local module to strengthen the encoder's capacity for cross-modal local feature fusion. The RGB-T local module employs adaptive cosine similarity to selectively emphasize salient consistent and distinct local features across RGB-T modalities. Experimental results show that TUNI achieves competitive performance with state-of-the-art models on FMB, PST900 and CART, with fewer parameters and lower computational cost. Meanwhile, it achieves an inference speed of 27 FPS on a Jetson Orin NX, demonstrating its real-time capability in deployment. Codes are available at https://github.com/xiaodonguo/TUNI.
Problem

Research questions and friction points this paper is trying to address.

Improves RGB-thermal semantic segmentation in challenging conditions
Addresses limited thermal feature extraction and suboptimal fusion
Enhances real-time efficiency while maintaining competitive performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified RGB-T encoder for simultaneous extraction and fusion
Large-scale pre-training with RGB and pseudo-thermal data
Adaptive cosine similarity for cross-modal local feature fusion
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
X
Xiaodong Guo
School of Automation, Beijing Institute of Technology, Beijing 100081, China
T
Tong Liu
School of Automation, Beijing Institute of Technology, Beijing 100081, China
Yike Li
Yike Li
School of Automation, Beijing Institute of Technology, Beijing 100081, China
Z
Zi'ang Lin
School of Automation, Beijing Institute of Technology, Beijing 100081, China
Zhihong Deng
Zhihong Deng
Faculty of Engineering and Information Technology, University of Technology Syndey
reinforcement learningrecommender systemscausal inference