🤖 AI Summary
This study addresses the high retraining costs and lack of interpretability in large audio-language models by proposing an interpretable, fine-tuning-free pipeline. The proposed method employs expert encoders to map audio inputs onto explicit causal semantic tree nodes, integrating automatic speech recognition transcripts with a frozen large language model for reasoning. Zero-shot cross-domain adaptation is achieved using only a lightweight head network and semantic graph retrieval, eliminating the need for extensive parameter updates. Evaluated on the SAKURA benchmark, this approach attains near state-of-the-art performance while requiring minimal parameters and training data. Furthermore, the framework supports seamless, lossless performance improvements when integrated with more capable large language models, offering a scalable and transparent alternative to conventional end-to-end retraining paradigms for audio-language understanding.
📝 Abstract
Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is a post-hoc account. We propose Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read, and the nodes are causal: replacing the deciding node with a distractor overturns 78% of correct answers on SAKURA. With a 7B reader, L2R outperforms all LALMs we compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, despite training about 1,400x fewer parameters on orders of magnitude less audio data. However, because any LLM can serve as the reader, we show that a stronger reader narrows this gap without retraining any audio component. A new domain is added with one small head: with five labelled clips per species, it outperforms QLoRA fine-tuning of an LALM on the same clips by 13-26 points.