🤖 AI Summary
This study addresses the challenge that existing bioacoustic classifiers struggle to precisely localize bird vocalizations in the time–frequency domain, thereby limiting ecological analysis. The work proposes a novel approach by formulating bird call detection as an object detection task on spectrograms and employs the YOLOv11 model to achieve high-precision spatiotemporal localization. Key contributions include the introduction of IoMin—an evaluation metric better aligned with acoustic boundaries—the development of an open-source web-based audio annotation tool, and strong empirical results: on complex tropical soundscapes from Singapore, the method achieves an IoMin@50 F1-score of 81.8%, nearly doubling baseline performance; it also significantly outperforms baselines on unseen Hawaiian data (58.6% vs. 48.6%).
📝 Abstract
Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate bird vocalization detection as an object detection task on spectrograms and train YOLO11 models to localize bird calls in dense tropical soundscapes from Singapore. We additionally introduce an open-source browser-based annotation tool and propose Intersection over Minimum (IoMin), an evaluation metric that better handles ambiguous acoustic boundaries than standard IoU and is better suited to the problem at hand. The best YOLO model nearly doubles baseline performance on in-distribution soundscapes from Singapore (81.8% vs. 42.1% IoMin@50 F1-score) while still outperforming the baseline on unseen out-of-distribution recordings from Hawaii (58.6% vs. 48.6%). These results suggest that object detection frameworks are a promising approach to time-frequency localization of animal vocalizations in complex soundscapes.