🤖 AI Summary
To address low accuracy and poor generalizability of underwater acoustic source localization in complex, dynamic marine environments, this paper proposes a multi-branch deep network for joint distance estimation between moving sources and receivers. Methodologically, we introduce an adaptive gain control (AGC) layer—novel in this domain—to enhance robustness to signal-to-noise ratio variations; integrate CNN and Conformer architectures for synergistic spatial-temporal feature modeling; and employ dual-input features—Log-Mel spectrograms and GCC-PHAT—enabling efficient cross-domain fine-tuning with minimal target-domain data. Evaluated on real underwater array recordings, our method establishes a new state-of-the-art benchmark, significantly outperforming existing approaches. In cross-domain transfer scenarios, it reduces localization error by 23.6%, demonstrating superior generalizability and practical engineering applicability.
📝 Abstract
Localizing acoustic sound sources in the ocean is a challenging task due to the complex and dynamic nature of the environment. Factors such as high background noise, irregular underwater geometries, and varying acoustic properties make accurate localization difficult. To address these obstacles, we propose a multi-branch network architecture designed to accurately predict the distance between a moving acoustic source and a receiver, tested on real-world underwater signal arrays. The network leverages Convolutional Neural Networks (CNNs) for robust spatial feature extraction and integrates Conformers with self-attention mechanism to effectively capture temporal dependencies. Log-mel spectrogram and generalized cross-correlation with phase transform (GCC-PHAT) features are employed as input representations. To further enhance the model performance, we introduce an Adaptive Gain Control (AGC) layer, that adaptively adjusts the amplitude of input features, ensuring consistent energy levels across varying ranges, signal strengths, and noise conditions. We assess the model's generalization capability by training it in one domain and testing it in a different domain, using only a limited amount of data from the test domain for fine-tuning. Our proposed method outperforms state-of-the-art (SOTA) approaches in similar settings, establishing new benchmarks for underwater sound localization.