🤖 AI Summary
This paper systematically surveys the evolution and key technical challenges of the YOLO series (v1–v11) in real-time object detection. Addressing the persistent trade-off among inference speed, detection accuracy, and deployment efficiency, we employ a rigorous literature review coupled with multi-dimensional performance benchmarking across versions. Our analysis traces architectural innovations—including iterative improvements to backbone, neck, and head components—as well as core mechanisms such as attention modeling, model lightweighting, and end-to-end training, alongside multi-task extensions (e.g., instance segmentation, pose estimation, and tracking). We present the first unified technical taxonomy spanning all YOLO versions, precisely mapping critical advancements and standardized benchmark results. Furthermore, we offer a novel conceptual insight into the “triple balance” (speed–accuracy–efficiency) within unified detection frameworks, identifying domain-specific bottlenecks and scalability potential in healthcare, industrial automation, and beyond. The work delivers an authoritative, reproducible reference for developing next-generation real-time vision systems.
📝 Abstract
Over the past decade, object detection has advanced significantly, with the YOLO (You Only Look Once) family of models transforming the landscape of real-time vision applications through unified, end-to-end detection frameworks. From YOLOv1's pioneering regression-based detection to the latest YOLOv9, each version has systematically enhanced the balance between speed, accuracy, and deployment efficiency through continuous architectural and algorithmic advancements.. Beyond core object detection, modern YOLO architectures have expanded to support tasks such as instance segmentation, pose estimation, object tracking, and domain-specific applications including medical imaging and industrial automation. This paper offers a comprehensive review of the YOLO family, highlighting architectural innovations, performance benchmarks, extended capabilities, and real-world use cases. We critically analyze the evolution of YOLO models and discuss emerging research directions that extend their impact across diverse computer vision domains.