BEANS-Next and ROOTS: Broadening Audio-Language Capabilities for Bioacoustics

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing audio-language model evaluations, which are predominantly confined to species classification and thereby restrict their applicability to broader bioacoustic tasks. To overcome this, we propose the BEANS-Next benchmark and the ROOTS dataset, establishing a comprehensive evaluation framework encompassing four major task families, including perception and scene understanding. Methodologically, we augment training resources by integrating behavioral metadata with synthetic generation techniques and employ multimodal fusion strategies to enhance the training of large-scale audio-language models. Experimental results demonstrate that our approach significantly improves overall model performance across all task families, advancing these models toward functioning as general-purpose bioacoustic assistants. All associated datasets and code pipelines have been made publicly available.
📝 Abstract
Bioacoustics and ethology encompass a wide range of audio understanding tasks, many of which stand to benefit from recent advances in large audio-language models. However, progress in the field has so far been assessed on a narrow set of tasks, primarily centered on label-centric biological category recognition, such as species and call-type classification. In this work, we introduce BEANS-Next, a benchmark grounded in a taxonomy of bioacoustics tasks spanning acoustic perception, biological category recognition, scene understanding, and in-context learning. Using BEANS-Next, we show that existing models exhibit limited performance beyond the task families emphasized by existing evaluations, constraining their usefulness for broader bioacoustic applications. To support progress on this broader task space, we also introduce ROOTS, a large-scale training resource built from expanded curated real-world data and previously underused behavioral and acoustic metadata, supplemented by audio-derived information and scalable synthetic generation where labeling is insufficient. We demonstrate that training on this dataset yields substantial progress across all task groups of BEANS-Next, moving audio-language models closer to their potential as general-purpose assistants for bioacoustics and ethology. To accelerate progress in the field, we open-source our benchmark, dataset, and data pipelines.
Problem

Research questions and friction points this paper is trying to address.

Bioacoustics
Audio-Language Models
Benchmark
Task Generalization
Ethology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bioacoustics
Audio-Language Models
Benchmark
Large-scale Dataset
Synthetic Data Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Christos Plachouras
Earth Species Project
David Robinson
David Robinson
Earth Species Project
Marius Miron
Marius Miron
Earth Species Project
Artificial IntelligenceDigital Signal ProcessingBioacousticsMusic Information RetrievalMachine Learning
G
Gagan Narula
Earth Species Project
P
Paul Laisné
Earth Species Project
A
Anthony L. T. Fine
Earth Species Project
Benno Weck
Benno Weck
Music Technology Group (Universitat Pompeu Fabra)
Music Information RetrievalNatural Language Processing
E
Ellen Gilsenan-McMahon
Earth Species Project
D
Diane Kim
Earth Species Project
L
Laura Hay Mack
Earth Species Project
M
Maddie Cusimano
Earth Species Project
S
Sara Keen
Earth Species Project
Lukas Rauch
Lukas Rauch
University of Kassel
Deep LearningSelf-Supervised LearningActive LearningBioacoustics
Benjamin Hoffman
Benjamin Hoffman
Earth Species Project
Animal communicationmachine learning
Emmanuel Chemla
Emmanuel Chemla
LSCP, ENS, Paris
Emmanouil Benetos
Emmanouil Benetos
Queen Mary University of London
Machine listeningAudio signal processingMusic information retrievalMachine learning
Johan Pauwels
Johan Pauwels
Queen Mary University of London
Music Information RetrievalAutomatic Key and Chord Estimation
Milad Alizadeh
Milad Alizadeh
University of Oxford
Machine LearningLarge Language Models
Matthieu Geist
Matthieu Geist
Earth Species Project (ex-google, ex-cohere, on leave of Professor, Université de Lorraine)
reinforcement learningmachine learning
Olivier Pietquin
Olivier Pietquin
Earth Species Project | ex Google DeepMind (On leave - Professor at University of Lille)
Machine LearningSpeech and Language ProcessingSignal ProcessingDialog systems