🤖 AI Summary
This study addresses the inherent lack of audio understanding in large language models (LLMs) and the textual performance degradation and prohibitive costs associated with fine-tuning. To overcome these limitations, we propose a symbiotic architecture that bypasses the LLM backbone by directly writing audio features into the key-value cache via an audio-conditioned vector injector. This approach endows frozen LLMs with audio processing capabilities without modifying any model weights, structurally preserving the original textual representation space to prevent catastrophic forgetting while significantly reducing training overhead. Experimental results demonstrate that the proposed architecture surpasses existing frozen-parameter methods across multiple audio tasks, approaching the performance of full fine-tuning while perfectly retaining the LLM's original textual proficiency.
📝 Abstract
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.