🤖 AI Summary
This study addresses the limitation of existing audio-language model evaluations, which are predominantly confined to species classification and thereby restrict their applicability to broader bioacoustic tasks. To overcome this, we propose the BEANS-Next benchmark and the ROOTS dataset, establishing a comprehensive evaluation framework encompassing four major task families, including perception and scene understanding. Methodologically, we augment training resources by integrating behavioral metadata with synthetic generation techniques and employ multimodal fusion strategies to enhance the training of large-scale audio-language models. Experimental results demonstrate that our approach significantly improves overall model performance across all task families, advancing these models toward functioning as general-purpose bioacoustic assistants. All associated datasets and code pipelines have been made publicly available.
📝 Abstract
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.