๐ค AI Summary
This work addresses the degradation of small-object detection performance in YOLO variants (v5โv11) across heterogeneous hardware platforms (CPU/GPU) and inference backends (ONNX Runtime, OpenVINO, TensorRT), specifically for objects occupying 1%โ5% of image area. We conduct a systematic benchmark evaluating accuracyโlatency trade-offs under realistic deployment conditions. This is the first cross-generational, horizontal sensitivity analysis of five YOLO versions to object scale, establishing a four-dimensional benchmark spanning hardware, model architecture, accuracy (mAP@0.5), and latency. Results show that YOLOv8/v10 achieve the best balance between CPU inference speed and small-object recall; YOLOv11 improves GPU mAP@0.5 by 3.2% but exhibits >40% miss rate for objects <2.5% image area. The study delivers a reproducible, cross-platform model selection decision map, providing empirical guidance for deploying small-object detectors in edge and cloud environments.
๐ Abstract
This paper provides an extensive evaluation of YOLO object detection models (v5, v8, v9, v10, v11) by com- paring their performance across various hardware platforms and optimization libraries. Our study investigates inference speed and detection accuracy on Intel and AMD CPUs using popular libraries such as ONNX and OpenVINO, as well as on GPUs through TensorRT and other GPU-optimized frameworks. Furthermore, we analyze the sensitivity of these YOLO models to object size within the image, examining performance when detecting objects that occupy 1%, 2.5%, and 5% of the total area of the image. By identifying the trade-offs in efficiency, accuracy, and object size adaptability, this paper offers insights for optimal model selection based on specific hardware constraints and detection requirements, aiding practitioners in deploying YOLO models effectively for real-world applications.