DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prompt overfitting problem in large speech models fine-tuned exclusively on ASR instructions, which hinders generalization to novel tasks such as speech translation. To overcome this limitation, the work proposes an end-to-end alignment framework that freezes the LLM embedding matrix and employs a distance-based CTC loss to generate aligned speech embeddings. Theoretically, it demonstrates that explicit geometric regression is unnecessary, as the modified CTC loss inherently provides implicit geometric grounding, thereby significantly simplifying the architecture. Trained solely on LibriSpeech, the proposed approach outperforms cascaded systems and achieves zero-shot generalization to translation and emotion recognition tasks. Furthermore, its performance closely approaches the upper bound of cascaded pipelines while exhibiting consistent improvements with model scaling.
📝 Abstract
Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.
Problem

Research questions and friction points this paper is trying to address.

Speech-LLMs
prompt overfitting
instruction following
zero-shot generalization
automatic speech recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech-LLMs
Prompt Overfitting
End-to-End Framework
CTC Loss
Zero-shot Generalization
🔎 Similar Papers