RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing RGB-D semantic segmentation benchmarks, which suffer from insufficient category diversity, limited scale, and annotation noise that severely constrain model generalization. To overcome these issues, we construct a large-scale, high-quality RGB-D benchmark dataset comprising 160 categories and 20,000 image pairs, expanding the semantic space while rigorously correcting annotation noise to ensure data integrity. Furthermore, we propose a Score Purification Fusion (SPF) strategy that effectively suppresses noise interference and enhances the utilization efficiency of multimodal features. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across multiple evaluation benchmarks, substantiating the significant benefits of high-quality multimodal information for semantic segmentation.
📝 Abstract
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.
Problem

Research questions and friction points this paper is trying to address.

RGB-D semantic segmentation
benchmark dataset
annotation noise
category diversity
data scale
Innovation

Methods, ideas, or system contributions that make the work stand out.

RGB-D Semantic Segmentation
Large-Scale Benchmark
Score-Purified Fusion
High-Fidelity Annotation
Multimodal Information
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shaohua Dong
Shaohua Dong
University of North Texas
Computer Vision
Z
Zexuan Meng
University of North Texas, Denton, TX 76207, USA
H
Haiyan Sun
University of North Texas, Denton, TX 76207, USA
B
Bing Fan
University of North Texas, Denton, TX 76207, USA
C
Cuicui Zhang
University of North Texas, Denton, TX 76207, USA
D
Dylan Joseph
University of North Texas, Denton, TX 76207, USA
Kewei Sha
Kewei Sha
Associate Professor, University of North Texas
Security and PrivacyEdge ComputingBlockchainData Quality and Analytics
Yunhe Feng
Yunhe Feng
Assistant Professor at University of North Texas
Responsible AIEfficient Generative AIData Security and PrivacyApplied AI
Heng Fan
Heng Fan
Assistant Professor, University of North Texas
Computer VisionMachine LearningArtificial Intelligence