Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent lack of audio understanding in large language models (LLMs) and the textual performance degradation and prohibitive costs associated with fine-tuning. To overcome these limitations, we propose a symbiotic architecture that bypasses the LLM backbone by directly writing audio features into the key-value cache via an audio-conditioned vector injector. This approach endows frozen LLMs with audio processing capabilities without modifying any model weights, structurally preserving the original textual representation space to prevent catastrophic forgetting while significantly reducing training overhead. Experimental results demonstrate that the proposed architecture surpasses existing frozen-parameter methods across multiple audio tasks, approaching the performance of full fine-tuning while perfectly retaining the LLM's original textual proficiency.
📝 Abstract
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Audio Understanding
Frozen LLM
Scalability
Audio Language Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Symbiotic Architecture
KV Cache Injection
Frozen Language Models
Audio Language Model
Parameter-Efficient
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.