🤖 AI Summary
This work addresses the limitations of existing drone detection datasets, which often fail to reflect real-world anti-drone scenarios due to image compression artifacts and the absence of critical factors such as camera ego-motion, extremely small target scales, and diverse optical configurations. To bridge this gap, the authors introduce an open-source, high-quality multimodal dataset that, for the first time, synchronously captures uncompressed RGB and event-based camera data under realistic and complex conditions, incorporating significant camera motion and varied lens setups. The project provides a unified data loader and multimodal baseline models, enabling systematic studies on the trade-off between field of view and detection range. Experimental results demonstrate that the proposed dataset substantially enhances detection performance for minuscule drones and validates the efficacy of multimodal fusion in challenging environments.
📝 Abstract
Detecting UAVs in air spaces has become increasingly important due to UAVs widespread availability and easy usage. However, due to their small size, they are typically difficult to detect at a sufficient range. For the training of optimized detection algorithms, datasets have been published, covering optical sensing methods ranging from infrared to regular RGB to event-sensor-based. However, these datasets often fail to reflect realistic counter-UAV scenarios, lacking critical factors such as camera ego-motion, extremely small target scales, and diverse lens configurations, and introduce compression artefacts on the frame images. To address this gap, we introduce SkyEV, an open-source dataset featuring highly synchronized uncompressed RGB and event-based data. SkyEV distinguishes itself by capturing complex real-world conditions, including significant camera motion and varied optical setups, which are essential for testing the fundamental trade-off between Field of View and detection range. Furthermore, we provide a unified data loader and establish an experimental baseline using a multi-modal architecture, demonstrating the dataset's efficacy in detecting challenging, small-scale targets.