A Contrastive Fewshot RGBD Traversability Segmentation Framework for Indoor Robotic Navigation

📅 2026-03-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of reliably detecting fine-grained obstacles, such as chair legs, in indoor robot navigation using vision-only models, which often suffer from limited robustness and require extensive labeled data. To overcome these limitations, the authors propose a few-shot traversable area segmentation framework that fuses RGB images with sparse 1D LiDAR depth measurements. The approach introduces negative prototype contrastive learning to mitigate overfitting under few-shot conditions and incorporates a two-stage attention module to effectively align RGB and sparse depth features. Evaluated on a newly curated indoor RGB-D traversability dataset, the method achieves up to a 9% improvement in mIoU under both 1-shot and 5-shot settings, significantly outperforming existing few-shot and RGB-D segmentation approaches.

Technology Category

Intelligent Robots: Multimodal Perception & Sensor FusionComputer Vision: SegmentationNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Indoor traversability segmentation aims to identify safe, navigable free space for autonomous agents, which is critical for robotic navigation. Pure vision-based models often fail to detect thin obstacles, such as chair legs, which can pose serious safety risks. We propose a multi-modal segmentation framework that leverages RGB images and sparse 1D laser depth information to capture geometric interactions and improve the detection of challenging obstacles. To reduce the reliance on large labeled datasets, we adopt the few-shot segmentation (FSS) paradigm, enabling the model to generalize from limited annotated examples. Traditional FSS methods focus solely on positive prototypes, often leading to overfitting to the support set and poor generalization. To address this, we introduce a negative contrastive learning (NCL) branch that leverages negative prototypes (obstacles) to refine free-space predictions. Additionally, we design a two-stage attention depth module to align 1D depth vectors with RGB images both horizontally and vertically. Extensive experiments on our custom-collected indoor RGB-D traversability dataset demonstrate that our method outperforms state-of-the-art FSS and RGB-D segmentation baselines, achieving up to 9\% higher mIoU under both 1-shot and 5-shot settings. These results highlight the effectiveness of leveraging negative prototypes and sparse depth for robust and efficient traversability segmentation.
Problem

Research questions and friction points this paper is trying to address.

traversability segmentation
few-shot learning
RGB-D
indoor navigation
obstacle detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

few-shot segmentation
negative contrastive learning
RGB-D fusion
traversability segmentation
sparse depth alignment
💼 Related Jobs
No related jobs found.
Qiyuan An
Qiyuan An
University of Texas at Arlington
Machine Learning
T
Tuan Dang
Cognitive Robotics Lab, Department of Electrical Engineering and Computer Science, University of Arkansas, Fayetteville, AR, USA
F
Fillia Makedon
Department of Computer Science and Engineering, University of Texas at Arlington, 1225 West Mitchell, Arlington, TX 76019, USA