real-time object detection

Designs and implements lightweight deep-learning pipelines and models that locate and classify 2D objects with bounding boxes at real-time inference rates, including single-stage architectures and variants (e.g., YOLO family) and training procedures for small-object detection, objectness/saliency estimation, and combined detection/segmentation. Builds evaluation suites and robustness measures, trains models end-to-end, and uses visualization/explainability methods (e.g., Grad-CAM++) to analyze predictions and improve precision under varying conditions.

real-timeobjectdetection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

YOLO11 to Its Genesis: A Decadal and Comprehensive Review of The You Only Look Once (YOLO) Series

Jun 12, 2024
RS
Ranjan Sapkota
🏛️ Cornell University | The University of Central Florida | Universidad de las Fuerzas Armadas | The University of Tennessee | Cooper Machine Company, Inc. | Indian Institute of Science Education and Research Thiruvananthapuram (IISER TVM) | ZenoRobotics, LLC | Texas A&M University | The Hong Kong Polytechnic University | City University of Hong Kong

Real-time object detection faces persistent challenges in jointly optimizing latency, accuracy, and computational efficiency across evolving YOLO architectures. Method: This work introduces a novel reverse-temporal analytical framework—integrating systematic literature review, architectural evolution modeling, and multidimensional performance attribution analysis—to trace YOLO’s decade-long progression from v1 to v11. Contribution/Results: We construct the first comprehensive technical evolution atlas spanning all YOLO versions, uncovering paradigm shifts toward multimodal perception, contextual reasoning, and AGI integration. The study precisely identifies generational bottlenecks and pivotal breakthrough mechanisms (e.g., anchor-free design, transformer-based attention, dynamic head adaptation). Based on these insights, we propose a forward-looking roadmap for next-generation YOLO models—incorporating embodied intelligence and neuro-symbolic reasoning—thereby establishing a reusable methodology benchmark and practical engineering guideline for real-time vision system design and deployment.

AI ApplicationsObject DetectionYOLO Series

Must-Read Papers

Most classic and influential ideas
View more

Small Object Detection with YOLO: A Performance Analysis Across Model Versions and Hardware

Apr 14, 2025
MF
Muhammad Fasih Tariq
🏛️ University of Management and Technology

This work addresses the degradation of small-object detection performance in YOLO variants (v5–v11) across heterogeneous hardware platforms (CPU/GPU) and inference backends (ONNX Runtime, OpenVINO, TensorRT), specifically for objects occupying 1%–5% of image area. We conduct a systematic benchmark evaluating accuracy–latency trade-offs under realistic deployment conditions. This is the first cross-generational, horizontal sensitivity analysis of five YOLO versions to object scale, establishing a four-dimensional benchmark spanning hardware, model architecture, accuracy (mAP@0.5), and latency. Results show that YOLOv8/v10 achieve the best balance between CPU inference speed and small-object recall; YOLOv11 improves GPU mAP@0.5 by 3.2% but exhibits >40% miss rate for objects <2.5% image area. The study delivers a reproducible, cross-platform model selection decision map, providing empirical guidance for deploying small-object detectors in edge and cloud environments.

Analyzing inference speed and accuracy on CPUs and GPUs with optimization librariesAssessing YOLO sensitivity to object size for optimal model selectionEvaluating YOLO model performance across versions and hardware platforms

To address insufficient multi-scale feature representation in real-time object detection, this paper proposes YOLO-MS—a lightweight, end-to-end trainable multi-scale collaborative feature enhancement framework. Methodologically, it introduces a multi-branch base module coupled with a multi-kernel convolutional fusion structure to reformulate cross-scale feature learning; incorporates an end-to-end joint optimization mechanism enabling training from scratch without ImageNet pretraining; and features plug-and-play compatibility for seamless integration into mainstream YOLO architectures. Experimentally, YOLO-MS-XS achieves 42.1% AP on COCO, outperforming RTMDet by 2.0 points. When deployed as a plug-in module, it elevates YOLOv8-N’s AP, APₗ, and APₘ to 20.3%, 55.1%, and 40.6%, respectively—while reducing both parameter count and FLOPs. These results demonstrate YOLO-MS’s effectiveness in enhancing multi-scale feature learning with minimal computational overhead.

Enhance multi-scale feature representationsImprove real-time object detectionOutperform existing YOLO models

Deep Learning and Machine Learning - Object Detection and Semantic Segmentation: From Theory to Applications

Oct 21, 2024
JR
Jintao Ren
🏛️ Aarhus University | Indiana University | Kyoto University | AppCubic | Rutgers University | University of Wisconsin-Madison | Georgia Institute of Technology | National Taiwan Normal University | University of Hawaii | Xi'an Jiaotong-Liverpool University | Zhejiang University | Purdue University

To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.

Bridging traditional methods with modern AI for large-scale detection tasksExploring object detection and semantic segmentation from theory to applicationsReviewing state-of-the-art deep learning architectures for computer vision

An Efficient Real-Time Object Detection Framework on Resource-Constricted Hardware Devices via Software and Hardware Co-design

Jul 01, 2021
SL
Shiyi Luo
🏛️ San Diego State University | University of California Irvine | University of California, Davis | California State University Fullerton

Existing tensor decomposition-based rank selection for embedded devices relies heavily on manual trial-and-error or incurs prohibitive computational overhead from automatic optimization. To address this, we propose a software-hardware co-designed real-time object detection framework. Our approach uniquely integrates Tensor Train (TT) decomposition with FPGA acceleration in a deeply coupled manner, enabling joint optimization of model compression ratio and hardware execution efficiency. Specifically, we apply TT decomposition to compress YOLOv5, design a custom FPGA accelerator, and perform software-hardware co-compiled optimizations. Evaluated on Jetson Nano and Xilinx Zynq FPGA platforms, the framework achieves 68% model size reduction, 3.2× inference speedup, and end-to-end latency under 32 ms—while preserving high detection accuracy. This work establishes a scalable, co-design paradigm for efficient, lightweight vision models at the edge.

Automating tensor rank selection for neural network compressionBalancing compression efficiency with minimal accuracy lossReducing computational complexity in model compression methods

LeYOLO, New Scalable and Efficient CNN Architecture for Object Detection

Jun 20, 2024
LH
Lilian Hollard
🏛️ Université de Reims Champagne-Ardenne | CEA | INRAE

To address the low FLOP efficiency and suboptimal accuracy–computation trade-off of YOLO-style models on embedded devices, this paper proposes LeYOLO—a FLOP-aware, scalable lightweight YOLO architecture. Methodologically, it introduces three key innovations: (1) an inverted-bottleneck backbone scaled under information bottleneck theory guidance; (2) a Fast Pyramid Feature Network (FPAN) for efficient multi-scale feature fusion; and (3) a decoupled Network-in-Network (DNiN) detection head. Leveraging joint scaling under strict FLOP constraints, LeYOLO achieves state-of-the-art FLOP–accuracy efficiency across a wide operating range. Specifically, LeYOLO-Small attains 38.2% mAP on COCO at just 4.5 GFLOPs—reducing computational cost by 42% versus YOLOv9-Tiny while maintaining superior accuracy. The full LeYOLO family spans 0.66–8.4 GFLOPs with corresponding mAPs of 25.2–41.0, establishing new SOTA in FLOP-per-mAP performance.

Achieving YOLO-like accuracy with MobileNet-level compactness for embedded systemsBridging performance gap between YOLO and SSDLite for low-resource devicesImproving parameter and FLOP efficiency in lightweight object detection models

Latest Papers

What's happening recently
View more

This work proposes an efficient object detection architecture based on YOLOv11 to address the challenges of weak small-object detection and insufficient feature representation in real-time applications. By integrating C3K2 modules into the backbone, SPPF structures into the neck, and C2PSA modules with spatial attention into the detection head, the model enhances multi-scale feature fusion and spatial detail perception. The proposed approach significantly improves mean average precision (mAP), particularly for small objects, while maintaining high inference speed. These advancements make the method well-suited for latency-sensitive real-world scenarios such as autonomous driving and video surveillance.

accuracyobject detectionreal-time performance

This work addresses the limitations of YOLO-family models in inference efficiency and multi-task performance on edge devices by presenting the first systematic analysis of the YOLO26 CNN architecture. The authors propose several key optimizations: eliminating the Distribution Focal Loss (DFL), introducing the ProgLoss loss function, adopting a small-object-aware label assignment strategy, employing the MuSGD optimizer, and designing a unified multi-task head that enables end-to-end inference without non-maximum suppression (NMS) while supporting instance segmentation, pose estimation, and oriented bounding box (OBB) detection. Experimental results demonstrate a 43% improvement in CPU inference speed over existing methods, achieving an optimal balance between real-time performance and multi-task accuracy, thereby offering an efficient solution for edge deployment.

architecture understandingcomputer visiondeep learning model

This work addresses key limitations in current real-time visual detectors—namely, reliance on non-maximum suppression (NMS), redundant detection heads, low training efficiency, and inadequate annotation coverage for small objects. The authors propose YOLO26, a unified architecture featuring a novel lightweight, dual-head design without distribution focal loss (DFL), enabling end-to-end NMS-free inference. It further incorporates the MuSGD hybrid optimizer, Progressive Loss, and the STAL (Small-Target-Aware Labeling) assignment strategy to support unified multi-task modeling. An open-vocabulary extension, YOLOE-26, enhances generalization capability. On COCO, YOLO26 achieves 40.9–57.5 mAP with TensorRT latency of only 1.7–11.8 ms on an NVIDIA T4 GPU, outperforming all existing real-time detectors; YOLOE-26x attains 40.6 AP on LVIS.

Distribution Focal Lossnon-maximum suppressionpositive label assignment

This study systematically evaluates the performance and robustness of various YOLO models for object detection in robotic workspaces. By constructing a custom dataset tailored to robotic scenarios, integrating the COCO2017 benchmark, and incorporating image distortions to simulate real-world deployment conditions, the work presents the first comprehensive comparison of different YOLO variants in this specific context. The experimental results reveal significant differences among the models in terms of accuracy, inference speed, and resilience to visual perturbations. These findings provide empirical evidence and practical guidance for selecting appropriate YOLO architectures in robotic vision systems, balancing trade-offs between detection precision, computational efficiency, and robustness under realistic operating conditions.

model evaluationobject detectionrobotics

This work addresses the challenge of achieving real-time performance in edge vision systems, where conventional multi-stage detection-classification pipelines suffer from fully GPU-serialized execution. The authors propose a five-step optimization methodology enabling zero-GPU-fallback INT8 deployment of classification models on NVIDIA Jetson Deep Learning Accelerators (DLAs), and construct a parallel inference pipeline with GPU-based detection and DLA-based classification. Key innovations include the first-ever DLA deployment workflow that entirely avoids GPU fallback, overcoming DLA operator limitations and quantization compatibility bottlenecks through techniques such as manual dynamic range calibration, quantization-aware training, and ONNX graph surgery. Evaluated on a Jetson Orin NX, the dual-head human attribute classifier operating in parallel with the detector incurs only a 0.8 FPS overhead (12.5 vs. 13.3 FPS) and supports cost-free scaling across dual DLAs.

edge inferencehierarchical classificationmodel deployment

Hot Scholars

RS

Ranjan Sapkota

Cornell University
Artificial IntelligenceAgentic AIAgricultural AutomationAgricultural Robotics
MK

Manoj Karkee

Cornell University
Agricultural AutomationAgricultural RoboticsSmart FarmingDigital Agriculture
MM

Ming-Ming Cheng

Professor of Computer Science, Nankai University
Computer VisionComputer GraphicsVisual AttentionSaliency
KS

Klamer Schutte

TNO, Intelligent Imaging
Artificial intelligenceimage processingcomputer vision
SK

Shu Kong

Texas A&M University
Computer VisionMachine Learning