Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

📅 2026-06-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of insufficient training data and poor cross-domain generalization in underwater acoustic machine learning, primarily due to the scarcity of publicly available labeled datasets. To this end, the authors construct a new underwater audio dataset comprising over one thousand annotated recordings spanning eight categories of biological and mechanical sound sources. They propose a lightweight CNN architecture incorporating boundary-aware loss and feature alignment to effectively mitigate class imbalance and domain shift. Evaluated within a newly established cross-domain framework, the method achieves an in-domain accuracy of 96.35% and demonstrates a 42.60% improvement in zero-shot ship detection performance on the ShipsEar dataset. The work also releases the data curation pipeline and reproducible benchmarks to support future research in underwater acoustics.
📝 Abstract
Machine learning for underwater acoustics is constrained by the scarcity of publicly available labeled datasets. In contrast to air-acoustic domains, where large benchmarks enable rapid model development, underwater datasets are typically small and limited in acoustic diversity, restricting robust model training and cross-domain generalization. To help address this gap, we introduce a curated underwater audio dataset derived from an open-source maritime sound archive. The dataset contains over one thousand labeled audio segments across eight biologically and mechanically relevant acoustic classes, providing an additional resource for training models in data-limited underwater environments. Additionally, we establish a lightweight Convolutional Neural Network (CNN) baseline and propose a margin-enhanced loss with feature alignment to mitigate class confusion arising from data imbalance, acoustic similarity, and cross-domain mismatch. While the baseline achieves 96.35% in-domain accuracy, evaluation on ShipsEar reveals substantial domain shift; the proposed feature alignment improve zero-shot ship detection by 42.60%, demonstrating stronger robustness under distribution mismatch. We further release a transparent curation pipeline and reproducible benchmark to support future research on imbalance mitigation, domain adaptation, and data-efficient underwater acoustic classification.
Problem

Research questions and friction points this paper is trying to address.

underwater acoustics
dataset scarcity
cross-domain generalization
labeled data
acoustic classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

underwater acoustic dataset
feature alignment
margin-enhanced loss
cross-domain generalization
lightweight CNN
Quoc Thinh Vo
Quoc Thinh Vo
Student, Drexel University
machine learningsignal processingsoftware engineeringcloud engineering
D
David K. Han
Department of Electrical and Computer Engineering, Drexel University, Philadelphia, PA, USA