🤖 AI Summary
This study addresses the challenges of object detection in satellite imagery under wide-area, high-resolution scenarios, where small targets suffer from detail loss and standard square input formats are ill-suited. The authors introduce a new public dataset comprising 1,307 native high-resolution images and 19,101 annotated bounding boxes, establishing the first compact detection benchmark tailored to land–sea transportation scenes. Comprehensive evaluations of mainstream detectors—including YOLO, RT-DETR, DETR, and Faster R-CNN—are conducted under unified COCO metrics while preserving the original aspect ratios, achieving mAP50 scores of 84.4–88.2. Furthermore, the work proposes SkyDet, an ultra-lightweight, anchor-free model with only 1.22 million parameters, which attains 60.5 mAP50 and 24.32 mAP50–95 at a model size of 4.90 MB and inference latency of 13.74 ms.
📝 Abstract
Satellite object detection is challenged by small targets and wide-format scenes that lose detail under standard square-input resizing. We introduce SkySeaLand, a public dataset of 1,307 high-resolution satellite images and 19,101 verified bounding boxes across airplane, boat, car, and ship classes in terrestrial and maritime scenes. Native COCO and YOLO annotations are provided. The collection is dominated by large source images and wide scene geometry: 84.5 percent exceed 3,836 pixels on the longest side and 73.1 percent are near a 3:1 aspect ratio. We evaluate twelve detectors from the YOLO, RT-DETR, DETR, and Faster R-CNN families using a common split and COCO metrics. The tested YOLO and RT-DETR variants obtain 84.4--88.2 mAP50, with no consistent accuracy gain from larger parameter counts under the reported model-specific recipes. We also report SkyDet, a 1.22 M parameter anchor-free baseline that obtains 60.5 mAP50 and 24.32 mAP50-95 in a 4.90 MB footprint, with 13.74 ms latency (72.8 FPS) on a Tesla T4. SkySeaLand provides a compact benchmark for mixed land--maritime transportation detection, while SkyDet establishes a documented low-footprint reference rather than a state-of-the-art accuracy claim.