Score
Designs and implements lightweight deep-learning pipelines and models that locate and classify 2D objects with bounding boxes at real-time inference rates, including single-stage architectures and variants (e.g., YOLO family) and training procedures for small-object detection, objectness/saliency estimation, and combined detection/segmentation. Builds evaluation suites and robustness measures, trains models end-to-end, and uses visualization/explainability methods (e.g., Grad-CAM++) to analyze predictions and improve precision under varying conditions.
This paper systematically surveys the evolution and key technical challenges of the YOLO series (v1–v11) in real-time object detection. Addressing the persistent trade-off among inference speed, detection accuracy, and deployment efficiency, we employ a rigorous literature review coupled with multi-dimensional performance benchmarking across versions. Our analysis traces architectural innovations—including iterative improvements to backbone, neck, and head components—as well as core mechanisms such as attention modeling, model lightweighting, and end-to-end training, alongside multi-task extensions (e.g., instance segmentation, pose estimation, and tracking). We present the first unified technical taxonomy spanning all YOLO versions, precisely mapping critical advancements and standardized benchmark results. Furthermore, we offer a novel conceptual insight into the “triple balance” (speed–accuracy–efficiency) within unified detection frameworks, identifying domain-specific bottlenecks and scalability potential in healthcare, industrial automation, and beyond. The work delivers an authoritative, reproducible reference for developing next-generation real-time vision systems.
Real-time object detection faces persistent challenges in jointly optimizing latency, accuracy, and computational efficiency across evolving YOLO architectures. Method: This work introduces a novel reverse-temporal analytical framework—integrating systematic literature review, architectural evolution modeling, and multidimensional performance attribution analysis—to trace YOLO’s decade-long progression from v1 to v11. Contribution/Results: We construct the first comprehensive technical evolution atlas spanning all YOLO versions, uncovering paradigm shifts toward multimodal perception, contextual reasoning, and AGI integration. The study precisely identifies generational bottlenecks and pivotal breakthrough mechanisms (e.g., anchor-free design, transformer-based attention, dynamic head adaptation). Based on these insights, we propose a forward-looking roadmap for next-generation YOLO models—incorporating embodied intelligence and neuro-symbolic reasoning—thereby establishing a reusable methodology benchmark and practical engineering guideline for real-time vision system design and deployment.
This work addresses the degradation of small-object detection performance in YOLO variants (v5–v11) across heterogeneous hardware platforms (CPU/GPU) and inference backends (ONNX Runtime, OpenVINO, TensorRT), specifically for objects occupying 1%–5% of image area. We conduct a systematic benchmark evaluating accuracy–latency trade-offs under realistic deployment conditions. This is the first cross-generational, horizontal sensitivity analysis of five YOLO versions to object scale, establishing a four-dimensional benchmark spanning hardware, model architecture, accuracy (mAP@0.5), and latency. Results show that YOLOv8/v10 achieve the best balance between CPU inference speed and small-object recall; YOLOv11 improves GPU mAP@0.5 by 3.2% but exhibits >40% miss rate for objects <2.5% image area. The study delivers a reproducible, cross-platform model selection decision map, providing empirical guidance for deploying small-object detectors in edge and cloud environments.
To address insufficient multi-scale feature representation in real-time object detection, this paper proposes YOLO-MS—a lightweight, end-to-end trainable multi-scale collaborative feature enhancement framework. Methodologically, it introduces a multi-branch base module coupled with a multi-kernel convolutional fusion structure to reformulate cross-scale feature learning; incorporates an end-to-end joint optimization mechanism enabling training from scratch without ImageNet pretraining; and features plug-and-play compatibility for seamless integration into mainstream YOLO architectures. Experimentally, YOLO-MS-XS achieves 42.1% AP on COCO, outperforming RTMDet by 2.0 points. When deployed as a plug-in module, it elevates YOLOv8-N’s AP, APₗ, and APₘ to 20.3%, 55.1%, and 40.6%, respectively—while reducing both parameter count and FLOPs. These results demonstrate YOLO-MS’s effectiveness in enhancing multi-scale feature learning with minimal computational overhead.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
Existing tensor decomposition-based rank selection for embedded devices relies heavily on manual trial-and-error or incurs prohibitive computational overhead from automatic optimization. To address this, we propose a software-hardware co-designed real-time object detection framework. Our approach uniquely integrates Tensor Train (TT) decomposition with FPGA acceleration in a deeply coupled manner, enabling joint optimization of model compression ratio and hardware execution efficiency. Specifically, we apply TT decomposition to compress YOLOv5, design a custom FPGA accelerator, and perform software-hardware co-compiled optimizations. Evaluated on Jetson Nano and Xilinx Zynq FPGA platforms, the framework achieves 68% model size reduction, 3.2× inference speedup, and end-to-end latency under 32 ms—while preserving high detection accuracy. This work establishes a scalable, co-design paradigm for efficient, lightweight vision models at the edge.
To address the low FLOP efficiency and suboptimal accuracy–computation trade-off of YOLO-style models on embedded devices, this paper proposes LeYOLO—a FLOP-aware, scalable lightweight YOLO architecture. Methodologically, it introduces three key innovations: (1) an inverted-bottleneck backbone scaled under information bottleneck theory guidance; (2) a Fast Pyramid Feature Network (FPAN) for efficient multi-scale feature fusion; and (3) a decoupled Network-in-Network (DNiN) detection head. Leveraging joint scaling under strict FLOP constraints, LeYOLO achieves state-of-the-art FLOP–accuracy efficiency across a wide operating range. Specifically, LeYOLO-Small attains 38.2% mAP on COCO at just 4.5 GFLOPs—reducing computational cost by 42% versus YOLOv9-Tiny while maintaining superior accuracy. The full LeYOLO family spans 0.66–8.4 GFLOPs with corresponding mAPs of 25.2–41.0, establishing new SOTA in FLOP-per-mAP performance.
This work proposes an efficient object detection architecture based on YOLOv11 to address the challenges of weak small-object detection and insufficient feature representation in real-time applications. By integrating C3K2 modules into the backbone, SPPF structures into the neck, and C2PSA modules with spatial attention into the detection head, the model enhances multi-scale feature fusion and spatial detail perception. The proposed approach significantly improves mean average precision (mAP), particularly for small objects, while maintaining high inference speed. These advancements make the method well-suited for latency-sensitive real-world scenarios such as autonomous driving and video surveillance.
This work addresses the limitations of YOLO-family models in inference efficiency and multi-task performance on edge devices by presenting the first systematic analysis of the YOLO26 CNN architecture. The authors propose several key optimizations: eliminating the Distribution Focal Loss (DFL), introducing the ProgLoss loss function, adopting a small-object-aware label assignment strategy, employing the MuSGD optimizer, and designing a unified multi-task head that enables end-to-end inference without non-maximum suppression (NMS) while supporting instance segmentation, pose estimation, and oriented bounding box (OBB) detection. Experimental results demonstrate a 43% improvement in CPU inference speed over existing methods, achieving an optimal balance between real-time performance and multi-task accuracy, thereby offering an efficient solution for edge deployment.
This work addresses key limitations in current real-time visual detectors—namely, reliance on non-maximum suppression (NMS), redundant detection heads, low training efficiency, and inadequate annotation coverage for small objects. The authors propose YOLO26, a unified architecture featuring a novel lightweight, dual-head design without distribution focal loss (DFL), enabling end-to-end NMS-free inference. It further incorporates the MuSGD hybrid optimizer, Progressive Loss, and the STAL (Small-Target-Aware Labeling) assignment strategy to support unified multi-task modeling. An open-vocabulary extension, YOLOE-26, enhances generalization capability. On COCO, YOLO26 achieves 40.9–57.5 mAP with TensorRT latency of only 1.7–11.8 ms on an NVIDIA T4 GPU, outperforming all existing real-time detectors; YOLOE-26x attains 40.6 AP on LVIS.
This study systematically evaluates the performance and robustness of various YOLO models for object detection in robotic workspaces. By constructing a custom dataset tailored to robotic scenarios, integrating the COCO2017 benchmark, and incorporating image distortions to simulate real-world deployment conditions, the work presents the first comprehensive comparison of different YOLO variants in this specific context. The experimental results reveal significant differences among the models in terms of accuracy, inference speed, and resilience to visual perturbations. These findings provide empirical evidence and practical guidance for selecting appropriate YOLO architectures in robotic vision systems, balancing trade-offs between detection precision, computational efficiency, and robustness under realistic operating conditions.
This work addresses the challenge of achieving real-time performance in edge vision systems, where conventional multi-stage detection-classification pipelines suffer from fully GPU-serialized execution. The authors propose a five-step optimization methodology enabling zero-GPU-fallback INT8 deployment of classification models on NVIDIA Jetson Deep Learning Accelerators (DLAs), and construct a parallel inference pipeline with GPU-based detection and DLA-based classification. Key innovations include the first-ever DLA deployment workflow that entirely avoids GPU fallback, overcoming DLA operator limitations and quantization compatibility bottlenecks through techniques such as manual dynamic range calibration, quantization-aware training, and ONNX graph surgery. Evaluated on a Jetson Orin NX, the dual-head human attribute classifier operating in parallel with the detector incurs only a 0.8 FPS overhead (12.5 vs. 13.3 FPS) and supports cost-free scaling across dual DLAs.