Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing large audio-language models in inferring high-order human activities from atomic acoustic events, highlighting the absence of systematic evaluations for their capacity to understand sound combinations. To bridge this gap, this work proposes a principled evaluation framework grounded in typicality and distractor similarity, and constructs a dedicated benchmark that integrates audio perception, linguistic reasoning, and systematic variant testing to quantitatively analyze the compositional reasoning capabilities of these models. The findings reveal an inherent limitation of current models, demonstrating their inability to reliably perform compositional reasoning based solely on audio inputs. To facilitate future research in enhancing complex auditory reasoning, the associated datasets and evaluation code have been fully open-sourced, providing a critical foundation for advancing the compositional understanding capabilities of large audio-language models.
📝 Abstract
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/
Problem

Research questions and friction points this paper is trying to address.

Large Audio-Language Models
Compositional Understanding
Sounding Actions
Higher-level Human Activities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Audio-Language Models
Compositional Understanding
Sounding Actions
Acoustic Perception
Benchmark Evaluation
🔎 Similar Papers
2024-09-15arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
M
Michel Olvera
LTCI, Télécom Paris, Institut Polytechnique de Paris
P
Paraskevas Stamatiadis
LTCI, Télécom Paris, Institut Polytechnique de Paris
C
Changhong Wang
LTCI, Télécom Paris, Institut Polytechnique de Paris
Gaël Richard
Gaël Richard
Professor, Télécom Paris, Institut polytechnique de Paris
Audio signal processingMachine listeningMusic ProcessingMusic Information RetrievalSound and music computing