🤖 AI Summary
This study addresses the prompt overfitting problem in large speech models fine-tuned exclusively on ASR instructions, which hinders generalization to novel tasks such as speech translation. To overcome this limitation, the work proposes an end-to-end alignment framework that freezes the LLM embedding matrix and employs a distance-based CTC loss to generate aligned speech embeddings. Theoretically, it demonstrates that explicit geometric regression is unnecessary, as the modified CTC loss inherently provides implicit geometric grounding, thereby significantly simplifying the architecture. Trained solely on LibriSpeech, the proposed approach outperforms cascaded systems and achieves zero-shot generalization to translation and emotion recognition tasks. Furthermore, its performance closely approaches the upper bound of cascaded pipelines while exhibiting consistent improvements with model scaling.
📝 Abstract
Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.