đ€ AI Summary
This study addresses the limitations of existing dolphin acoustic datasets, which are small, restricted in access, and lack standardized benchmarks, thereby hindering research into complex intraspecific communication structures. We release the first large-scale, publicly available longitudinal dataset of dolphin whistles, accompanied by a complete open-source processing pipeline, expert annotations, and standardized evaluation protocols. Methodologically, we employ the Wav2Vec 2.0 self-supervised pre-trained model to perform detection, segmentation, and classification of bioacoustic signals. Experimental results demonstrate that this approach significantly outperforms general-purpose baselines such as AVES on both detection and classification tasks, effectively learning fine-grained acoustic representations. This work bridges a critical gap in single-species fine-grained acoustic structure research.
đ Abstract
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species'communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (Tursiops truncatus) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.