Score
Designs and implements algorithms, pipelines, and evaluation protocols to adapt automatic speech recognition (ASR) systems to new speakers when only limited or no labeled per‑speaker data is available. Work includes few‑shot and zero‑shot methods, online/real‑time adaptation pipelines, measuring performance as a function of per‑speaker data size, and benchmarking adaptation approaches under low‑resource budgets.
To address the severe scarcity of training data for automatic speech recognition (ASR) in low-resource languages (e.g., Vatlongos, Nashta), this paper proposes three lightweight, text-only data augmentation methods: rule-based lexical substitution, random token replacement, and large language model–driven text generation—each followed by high-fidelity speech synthesis via text-to-speech (TTS). Crucially, all methods require no external annotations or parallel corpora, ensuring strong generalizability and deployment simplicity. When integrated into fine-tuning of Wav2Vec2-XLSR-53, the approach yields substantial multilingual ASR improvements, achieving a 14.3 percentage-point absolute reduction in word error rate (WER) for Nashta. Empirical evaluation confirms robust performance across both extremely low-resource and higher-resource languages. This work establishes a concise, universal, and high-performance paradigm for data expansion in low-resource ASR.
Target-speaker automatic speech recognition (ASR) in streaming multi-speaker scenarios—particularly under full speaker overlap—remains challenging without explicit speaker queries (e.g., enrollment utterances or embeddings). Method: This paper proposes an adaptive speaker-directed framework featuring a speaker-aware voice activity detection (VAD) module that generates dynamic kernels, which are injected into the ASR encoder layers to enable fine-grained, temporally consistent speaker adaptation—without external speaker priors. Contribution/Results: To our knowledge, this is the first work to tightly integrate dynamic kernel injection with streaming ASR, enabling fully query-free online speaker tracking and recognition. Experiments demonstrate state-of-the-art performance in both offline and streaming settings, with substantial improvements in robustness and word error rate reduction under extreme speaker overlap conditions.
This work addresses the modality mismatch between speech and text that arises when adapting large language models (LLMs) for automatic speech recognition using only textual data. To bridge the gap between the speech encoder and the LLM, the authors propose a hybrid batch training strategy that jointly leverages a small amount of target-domain speech data—less than four hours—and abundant text data. Remarkably, with only 10% of the target-domain speech data, the proposed approach significantly outperforms text-only adaptation on both in-domain and out-of-domain evaluations, achieving word error rates comparable to or even better than those obtained by conventional fine-tuning with full speech datasets. These results demonstrate the method’s high efficiency and practical applicability in low-resource domain adaptation scenarios.
This work addresses the speaker and language adaptation challenge in lightweight cross-lingual text-to-speech (TTS) systems when no speech recordings are available for the target language. We propose an adapter-based parameter-efficient fine-tuning method that decouples multilingual phonetic modeling from speaker representation. Inspired by second-language acquisition theory, we further introduce an objective accent evaluation metric to systematically analyze the impact of adapter placement, architecture, and number of training speakers on synthesis performance. Experiments demonstrate that our approach efficiently acquires novel language and speaker characteristics without target-language speech data, significantly mitigating catastrophic forgetting. Both subjective listening tests and objective evaluations confirm superior speech naturalness and accent fidelity over baseline methods. The proposed framework provides a scalable, interpretable, and lightweight solution for low-resource cross-lingual TTS.
To address the limited cross-lingual transferability and severe domain mismatch of self-supervised learning (SSL) pre-trained models in automatic speech recognition (ASR) for low-resource languages, this paper proposes a lightweight adapter method with *intermediate warm-start*. Under frozen SSL backbone constraints, only 1–5% of parameters are fine-tuned. A two-stage progressive adaptation jointly optimizes adapter architecture and downstream model initialization. The novel intermediate warm-start mechanism mitigates speech feature distribution shift, substantially improving generalization to unseen languages. Evaluated on the ML-SUPERB benchmark, our approach achieves up to 28% relative reduction in character/phone error rates over standard efficient fine-tuning, significantly alleviating the bottleneck in low-resource cross-lingual ASR adaptation.
This work addresses the challenge of effectively leveraging plain text data to enhance the performance of encoder-centric end-to-end automatic speech recognition (ASR) systems. The authors propose a novel approach that integrates modality alignment and dynamic downsampling to enable the encoder to directly produce token-level representations, replacing the conventional large decoder with a compact “large-encoder, small-decoder” architecture. Key innovations include simple yet effective strategies such as stochastic duration modeling. Evaluated on LibriSpeech, the method achieves substantial gains in both recognition accuracy and inference speed, matching or surpassing more complex state-of-the-art systems while significantly streamlining the training pipeline and overall model design. All code and training recipes are publicly released.
Human infants acquire phonemic units from just hundreds of hours of speech, whereas current self-supervised speech models require orders-of-magnitude more data—revealing a critical data-efficiency gap. This paper introduces a general-purpose speech representation learning framework for rapid low-resource language adaptation. We propose Multi-task Adaptive Pre-Training (MAdaPT) and a First-Order Bilevel Optimization (FOBLO) algorithm, enhanced by interleaved supervised initialization to improve meta-training stability. The approach is architecture-agnostic and biologically inspired, requiring less than one hour of unlabeled target-language speech to learn highly discriminative representations. Evaluated on ABX, sWUGGY, sBLIMP, and tSC benchmarks, our method consistently outperforms prior models: one-hour adaptation achieves performance comparable to standard training with 100× more data. Code and pretrained models are publicly released.
Low-resource automatic speech recognition (ASR) faces dual challenges: scarcity of labeled training data and prohibitive computational costs associated with large-scale models. To address these, we propose an efficient multilingual ASR framework comprising three key components: (1) construction of a cross-lingual unsupervised corpus to enable targeted continual pretraining using unlabeled data from linguistically related languages; (2) integration of morphology-aware subword tokenization to enhance subword modeling for low-resource languages; and (3) design of a scalable, reproducible data curation pipeline. Our 300M-parameter multilingual model achieves state-of-the-art Persian ASR performance—surpassing Whisper Large v3 despite using only 20% of its parameters and less supervised data—while matching its performance on Arabic and Urdu. Crucially, our work demonstrates that data relevance and training strategy outweigh sheer model scale, establishing a lightweight, reproducible, and cost-effective pathway for low-resource ASR.
This work addresses the performance degradation in elderly speech recognition caused by speaker variability, particularly the challenge of zero-shot real-time adaptation to unseen speakers. The authors propose an online speaker adaptation method based on cross-utterance audio-text prompts, which dynamically fuses acoustic and textual embeddings from both current and historical utterances during decoding to generate compact and consistent speaker representations. This approach introduces a cross-modal prompting mechanism into online adaptation for the first time, requiring no prior data from the target speaker. Experiments demonstrate absolute reductions of 0.61% and 1.22% (relative improvements of 2.99% and 4.48%) in WER and CER on the DementiaBank Pitt and JCCOCC MoCA datasets, respectively, with inference speeds up to 9.83× faster than offline batch processing, significantly outperforming conventional i/x-vector and ECAPA-TDNN features.