When Capabilities Fail to Compose: Diagnosing the Compositionality Gap in Large Audio-Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic deficiencies of Large Audio-Language Models (LALMs) in composing independent capabilities. We propose a dual-sentence input diagnostic framework that conducts controlled composition experiments by combining attribute recognition and segment selection with downstream tasks such as ASR and QA. By integrating chain-of-thought analysis and positional bias probing techniques, this work systematically quantifies performance degradation patterns during multimodal capability integration. Our findings reveal significant compositional bottlenecks: performance deterioration occurs in 39 out of 40 experimental configurations, yielding an average accuracy decline of approximately 26.7% or a 28.5% increase in word error rate. These results provide critical empirical evidence for understanding and improving the compositional generalization capacity of LALMs.
📝 Abstract
Large audio-language models (LALMs) perform strongly on individual audio tasks, but whether these capabilities can be reliably composed remains underexplored. We conduct a controlled diagnostic study of capability composition in LALMs, requiring models to integrate audio-attribute recognition, cue-conditioned segment selection, and downstream ASR or question answering. We construct two-utterance inputs with distinct acoustic cues to evaluate composition across environmental sound, gender, and emotion cues, with ASR, Math QA, and Factual QA as downstream tasks. Across four open-source LALMs, compositional QA accuracy decreases in 39 of 40 model-task-cue settings, by an average of 26.7 percentage points. ASR exhibits a similarly consistent degradation, with WER increasing in 39 of 40 settings by an average of 28.5 percentage points, while the magnitude of degradation varies across models, cue types, and cue salience. We further probe these failures through output format, positional preference, and chain-of-thought (CoT) analyses. Our study reveals a systematic gap between possessing individual audio capabilities and reliably composing them.
Problem

Research questions and friction points this paper is trying to address.

Large Audio-Language Models
Compositionality Gap
Capability Composition
Audio Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Audio-Language Models
Compositionality Gap
Capability Composition
Chain-of-Thought
Diagnostic Study
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chien-Feng Liu
National Taiwan University
Chih-Kai Yang
Chih-Kai Yang
National Taiwan University
Deep LearningSpeech ProcessingNatural Language ProcessingMachine Learning
B
Bo-Han Feng
National Taiwan University
Y
Yu-Hsuan Li Liang
National Taiwan University
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing
C
Cheng-Fu Chou
National Taiwan University