RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
Summary
This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.
View Cached Full Text
Cached at: 09/25/26, 07:47 PM
Paper page - RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
Source: https://huggingface.co/papers/2609.29028
Abstract
Inthispaper,weproposeRGBD20K,anoveldatasetforfacilitatingthedevelopmentofmorerobustandgeneralRGB-Dsemanticsegmentationbyencompassingabundantcategoriesandhigh-qualityannotations.RGBD20Kpossessesseveralattractiveproperties:(1)ExpandedSemanticSpace.Inparticular,itcovers160fine-grainedcategories,largelysurpassingthecategorydiversityofexistingpopularRGB-Dbenchmarks(e.g.,NYUv2with40classesandSUNRGB-Dwith37classes).Withsuchenrichedsemanticcoverage,weexpecttopromotethelearningofmoregeneralizablesegmentationmodels.(2)LargerScale.Comparedwithcurrentbenchmarks,RGBD20Koffers20,000RGB-Dimagepairs,providingasubstantiallylargertrainingresourcethatbenefitsthedevelopmentofmorepowerfuldeepmodels.(3)High-FidelityAnnotation.Weperformrigorousre-evaluationandcorrectionofexistinglabelstoresolvelong-standingannotationnoise,resultinginacleanandreliableground-truthfoundation.Furthermore,weproposeanovelscore-purifiedfusion(SPF)method,whichachievesstate-of-the-artperformanceacrossallevaluatedbenchmarks,demonstratingtheeffectivenessofourapproachinleveraginghigh-qualitymultimodalinformationforRGB-Dsemanticsegmentation.Thedatasetishere:https://github.com/ShaohuaDong2021/RGBD20K/.
View arXiv pageView PDFGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.29028
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.29028 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.29028 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.29028 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift
This paper introduces SULAND v2, a refined RGB surface landmine detection dataset and benchmark for UAV/UGV-based surveys, addressing annotation errors and domain-shift evaluation in object detection.
GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
Introduces GEOID-Flood, a large-scale multi-modal benchmark dataset for flood segmentation with over 14,000 tiles from 219 events across 65 countries, evaluating foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols.
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
RRS-10K is a benchmark dataset for evaluating vision-language models on rare remote sensing image interpretation, containing over 10,000 military-related images and multiple task formats. Evaluation of 52 models reveals moderate zero-shot performance and weaknesses in visual grounding and complex reasoning.
CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
This paper proposes COVER, a training-free method for converting 3D assets into sparse panoramic RGB-D-pose data with complete scene coverage and low redundancy, and introduces the CM-EVS dataset containing 36,373 curated frames from indoor and outdoor scenes.
Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
This paper presents a 3D-aware neural approach for RGB-NIR low-light imaging that fuses extremely noisy RGB observations with NIR cues in 3D space, eliminating the need for clean RGB supervision and improving robustness across different noise levels.