🤖 AI Summary
To address the excessive computational overhead introduced by attention mechanisms in real-time object detection—compromising the balance between speed and accuracy—this paper proposes YOLOv12. Methodologically, it introduces (1) a lightweight Area Attention module for efficient local-global self-attention modeling; (2) a Residual Efficient Layer Aggregation Network (RELAN) integrated with FlashAttention to enhance feature reuse and long-range dependency capture; and (3) co-optimized backbone and detection head design. On COCO, YOLOv12 achieves 54.1% mAP—3.2% higher than YOLOv8/v10—with 128 FPS inference speed on an NVIDIA Tesla A100 and 19% lower computational cost. It outperforms state-of-the-art YOLO and DETR-based models, establishing a new latency–accuracy trade-off paradigm for real-time detection.
📝 Abstract
The YOLO (You Only Look Once) series has been a leading framework in real-time object detection, consistently improving the balance between speed and accuracy. However, integrating attention mechanisms into YOLO has been challenging due to their high computational overhead. YOLOv12 introduces a novel approach that successfully incorporates attention-based enhancements while preserving real-time performance. This paper provides a comprehensive review of YOLOv12's architectural innovations, including Area Attention for computationally efficient self-attention, Residual Efficient Layer Aggregation Networks for improved feature aggregation, and FlashAttention for optimized memory access. Additionally, we benchmark YOLOv12 against prior YOLO versions and competing object detectors, analyzing its improvements in accuracy, inference speed, and computational efficiency. Through this analysis, we demonstrate how YOLOv12 advances real-time object detection by refining the latency-accuracy trade-off and optimizing computational resources.