In-Context Adaptation of Encoder-Decoder Models in Speech Recognition

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether encoder-decoder models possess an intrinsic capacity for automatic speech recognition (ASR) context adaptation. Through controlled experiments across six architectures, the authors compare organized and interleaved demonstration paradigms while decoupling the contributions of lexical and speaker information. The findings confirm that context adaptation constitutes a general emergent capability of such models, functioning effectively out-of-the-box without fine-tuning. Moreover, organized demonstrations are shown to be more stable and effective than their interleaved counterparts. The proposed approach achieves consistent context adaptation performance, yielding relative improvements of up to 30% and 23% under oracle and initial hypothesis settings, respectively.
📝 Abstract
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.
Problem

Research questions and friction points this paper is trying to address.

In-context adaptation
Automatic speech recognition
Encoder-decoder models
Demonstration
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context learning
Encoder-decoder models
Speech recognition
Collated demonstration
Domain adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yen Meng
The Centre for Speech Technology Research, University of Edinburgh, United Kingdom
S
Sharon Goldwater
The Centre for Speech Technology Research, University of Edinburgh, United Kingdom
Hao Tang
Hao Tang
University of Edinburgh
Speech and Language ProcessingSpeech Recognition