DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of high-level semantic constraints in RGB-D segmentation by proposing a Transformer-based tri-modal (RGB-depth-text) segmentation framework. The method introduces language priors as strong semantic regularization and designs a text-guided semantic alignment module that explicitly aligns multi-modal features with semantic prototypes across multiple levels of both the encoder and decoder, thereby significantly enhancing feature discriminability. Through deep cross-modal fusion, this approach effectively compensates for the insufficient semantic representations inherent in conventional methods. Extensive experiments demonstrate consistent performance improvements across multiple benchmark datasets, while maintaining architectural simplicity and computational efficiency.
📝 Abstract
RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel-level supervision, lacking direct high-level semantic constraints. This raises a central question-can external knowledge such as language priors inject stronger semantic discriminability into mainstream RGB-D segmentation models. We present DTFormer, a novel tri-modal (RGB-D-Text) semantic segmentation framework. At its core is Text-guided Semantic Alignment Module (TSAM) that first encodes textual cues into a set of semantic prototypes and then explicitly aligns multi-modal RGB-D features with these prototypes at multiple encoder and decoder layers. This design imposes strong semantic regularization on representation learning, guiding the network toward more discriminative features. Extensive experiments on multiple benchmarks show that DTFormer delivers consistent gains while remaining simple and efficient. Our results demonstrate that explicit semantic alignment offers an effective and practical route to improving RGB-D semantic segmentation. The code will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

RGB-D semantic segmentation
pixel-level supervision
semantic constraints
language priors
semantic discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

RGB-D Semantic Segmentation
Tri-modal Fusion
Text-Guided Alignment
Semantic Prototypes
TSAM
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziang Wei
Ziang Wei
University Laval
ThermographyMachine LearningArtificial Intelligence
Y
Yinlong Liu
City University of Macau, China
Y
Yan Xia
University of Science and Technology of China, China
Alois Knoll
Alois Knoll
Technische Universität München
RoboticsAISensor Data FusionAutonomous DrivingCyber Physical Systems
H
Hu Cao
School of Automation, Southeast University, Nanjing, China