Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文指出标注不完整限制了目标检测基准质量,通过重新标注增加对象数量,并提出高召回率和不确定性感知的标注流程以提高数据质量。
📝 Abstract
While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.
Problem

Research questions and friction points this paper is trying to address.

Object Detection
Benchmark Quality
Annotation Incompleteness
Label Errors
Uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

high recall
soft labels
uncertainty-aware
annotation quality
label error detection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Sarina Penquitt
Institute of Computer Science, Osnabrück University, Osnabrück, Germany
J
Jonathan Klees
Institute of Computer Science, Osnabrück University, Osnabrück, Germany
A
Antonia van Betteray
Institute of Computer Science, Osnabrück University, Osnabrück, Germany
P
Parssa Jashnieh
Institute of Computer Science, Osnabrück University, Osnabrück, Germany
P
Peter Stehr
Department of Mathematics, University of Wuppertal, Wuppertal, Germany
Matthias Rottmann
Matthias Rottmann
Professor of Computer Science, Osnabrück University, Germany
Computer VisionDeep LearningSafe AIEfficient AINumerical Linear Algebra
L
Lars Schmarje
Wayve, London, United Kingdom