asr fine-tuning

Designs, implements, and evaluates fine‑tuning pipelines that adapt automatic speech recognition (ASR) models to individual speakers or speaker groups; this includes initializing target models from pre‑trained sources, applying speaker‑adaptive fine‑tuning and personalization techniques, and fine‑tuning on augmented or low‑resource speech data. Builds and analyzes methods to adjust acoustic and semantic representations and tuning strategies to improve speaker‑specific recognition and dialect discrimination accuracy.

asrfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited effectiveness of direct speaker-specific fine-tuning (SS-FT) of general-purpose pre-trained automatic speech recognition (ASR) models on highly variable disordered speech, such as that associated with dysarthria or aphasia. To overcome this challenge, the authors propose a two-stage adaptation framework: first performing speaker-independent fine-tuning (SI-FT) on multi-speaker disordered speech data, followed by speaker-specific fine-tuning. This study provides the first systematic validation of SI-FT as an effective initialization strategy for personalization. Evaluated on disordered speech benchmarks including AphasiaBank and UA-Speech, the approach significantly improves recognition accuracy, while incurring only controlled performance degradation on out-of-domain canonical speech datasets such as TED-LIUM v3 and FLEURS. Experiments using Whisper-Large-v3 and Qwen3-ASR consistently demonstrate the superiority of the two-stage strategy over direct SS-FT.

automatic speech recognitiondysarthric speechnon-normative speech

Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches

Jun 01, 2025
DF
Dena F. Mujtaba
🏛️ Michigan State University

Stuttered speech—characterized by disfluencies such as blocks, prolongations, and repetitions—severely degrades automatic speech recognition (ASR) performance, leading to substantially higher word error rates (WER); further challenges include high inter- and intra-speaker variability and scarcity of annotated stuttering data, limiting model robustness and inclusivity. To address this, we conduct the first systematic comparison of personalized (speaker-specific fine-tuning) versus generalized (multi-speaker joint fine-tuning) paradigms on Whisper and Wav2Vec 2.0, incorporating multi-scenario stuttering speech augmentation and speaker-adaptive feature normalization. Results show that personalized models reduce WER by 32.7% in spontaneous speech contexts—significantly outperforming generalized models—and deliver more accurate, reliable real-time transcription in practical applications such as virtual assistants and video interviews. This work advances fairness and practical utility of ASR for people who stutter.

ASR systems misinterpret stuttered speech, increasing word errorsLimited annotated stuttered speech data complicates ASR trainingPersonalized vs. generalized ASR models for stuttering accuracy

Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition

May 19, 2025
DW
Dominik Wagner
🏛️ Technische Hochschule Nürnberg Georg Simon Ohm | Korea Advanced Institute of Science & Technology | Friedrich-Alexander-Universität Erlangen-Nürnberg

Automatic speech recognition (ASR) for dysarthric speech suffers from low accuracy due to high acoustic variability and limited labeled data. Method: This paper proposes an LLM-driven personalized ASR optimization framework, introducing a novel integrated paradigm of “controllable speech synthesis + speaker adaptation + parameter-efficient fine-tuning”: (i) an LLM generates target text and guides Parler-TTS to synthesize high-fidelity, content-controlled dysarthric speech; (ii) x-vectors model speaker-specific characteristics; and (iii) AdaLoRA performs lightweight fine-tuning in the wav2vec 2.0 feature space, decoupling linguistic content from individual acoustic traits. Results: The method reduces relative word error rate (WER) by 23% compared to full-parameter fine-tuning; incorporating synthetic data yields an additional 7% relative WER reduction, achieving over 30% total relative WER improvement. It significantly enhances both personalization efficiency and ASR robustness for dysarthric speech.

Enhancing performance via AdaLoRA adapters and wav2vec 2.0Improving dysarthric speech recognition using synthetic speechPersonalizing ASR systems with x-vectors to reduce WER

Dysarthric speech recognition faces significant challenges, including substantial inter-speaker variability, pronounced acoustic-phonetic deviations from healthy speech, and scarcity of speaker-specific annotated data—leading to overfitting. To address these issues, this paper proposes a cross-speaker joint fine-tuning strategy based on a pre-trained automatic speech recognition (ASR) model, simultaneously fine-tuning on dysarthric speech from multiple speakers in the CDSD corpus. Unlike conventional speaker-isolated fine-tuning paradigms, our approach leverages shared representation learning across pathological speech to enhance generalization to diverse articulatory impairments, thereby substantially reducing reliance on per-speaker labeled data. Experimental results demonstrate that the proposed method achieves up to a 13.15% absolute reduction in word error rate (WER) on target speakers compared to speaker-specific fine-tuning, with marked improvements in recognition accuracy. This work provides an efficient and practical solution for low-resource dysarthric ASR.

Addresses dysarthric speech recognition challenges from severity variationsImproves individual speech pattern recognition via multi-speaker fine-tuningReduces word error rate and data dependency through cross-learning

How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario

Nov 27, 2024
SW
Shih-Heng Wang
🏛️ National Taiwan University | Carnegie Mellon University | The University of Texas at Austin | FAIR

To address the limited cross-lingual transferability and severe domain mismatch of self-supervised learning (SSL) pre-trained models in automatic speech recognition (ASR) for low-resource languages, this paper proposes a lightweight adapter method with *intermediate warm-start*. Under frozen SSL backbone constraints, only 1–5% of parameters are fine-tuned. A two-stage progressive adaptation jointly optimizes adapter architecture and downstream model initialization. The novel intermediate warm-start mechanism mitigates speech feature distribution shift, substantially improving generalization to unseen languages. Evaluated on the ML-SUPERB benchmark, our approach achieves up to 28% relative reduction in character/phone error rates over standard efficient fine-tuning, significantly alleviating the bottleneck in low-resource cross-lingual ASR adaptation.

Automatic Speech Recognition (ASR)Low-Resource LanguagesSpeech Self-Supervised Learning (SSL)

Latest Papers

What's happening recently
View more

This work addresses the limitations of frozen large language models (LLMs) in speech synthesis, which struggle to capture speaker-specific acoustic and perceptual characteristics, resulting in insufficient voice consistency and speech quality. To overcome this, the authors propose an efficient fine-tuning approach based on Low-Rank Adaptation (LoRA) applied to the Qwen-0.5B backbone, combined with training data exhibiting high acoustic diversity. This method significantly enhances voice cloning performance in terms of naturalness, speaker similarity, and signal-to-noise ratio. Experimental results demonstrate improvements of up to 0.42 points in DNS-MOS scores and a 34% increase in signal-to-noise ratio, confirming that LoRA serves not only as a parameter-efficient fine-tuning strategy but also as a critical mechanism for enabling speaker adaptation in compact LLM-based text-to-speech systems.

Data DiversityLarge Language ModelsSpeaker Adaptation

This study challenges the prevailing assumption that performance gains from supervised fine-tuning (SFT) in downstream tasks of speech foundation models stem primarily from methodological improvements. Instead, it systematically evaluates eight SFT variants across nine pretrained checkpoints of wav2vec 2.0, HuBERT, and WavLM on three SUPERB classification tasks, incorporating multiple random seeds to assess stability and transferability. The findings reveal that SFT’s apparent advantages are highly contingent on specific pretrained instances and random seeds, with optimal configurations showing little consistency or generalizability across checkpoints. These results suggest that most reported gains arise from favorable instance–seed matching rather than genuine improvements in model capacity or upper-bound performance, thereby questioning the universality of SFT enhancements.

downstream performanceinstance dependencypretrained checkpoint

This study addresses the significant performance degradation of general-purpose automatic speech recognition (ASR) systems on dysarthric speech, which hinders effective daily communication for affected individuals. The authors propose a personalized fine-tuning approach that leverages as little as 1.4 hours of user-specific read speech and mobile-device-collected correction feedback to fully fine-tune the Whisper base model, while also benchmarking alternatives such as LoRA adaptation and Qwen3-ASR. Experimental results demonstrate that the proposed method achieves a word error rate (WER) of 15.8% with only 1.4 hours of data, further improving to 9.7% when utilizing the full dataset of 92 hours of read speech and 8.8 hours of correction data. These results substantially outperform existing adaptation strategies, confirming the feasibility and deployment potential of highly effective personalized ASR under low-resource conditions.

ASR adaptationautomatic speech recognitiondysarthric speech

To address performance degradation in ASR models—particularly LLM-based ASR—under domain adaptation due to data mismatch and training complexity, this paper proposes a metric-driven fine-tuning framework. Methodologically, it introduces the first learning-rate scheduler adaptively calibrated via WER feedback, integrated with domain-aware data augmentation, multi-scale temporal transformations, and an anti-overfitting fine-tuning protocol. The framework unifies support for both conventional ASR and LLM-based ASR (e.g., Whisper, Qwen2-Audio) across architectures. Experiments on diverse multi-domain, multilingual, and variable-length benchmarks demonstrate that Whisper-Turbo achieves an average 23.6% relative WER reduction, while Qwen2-Audio exhibits markedly improved generalization and training stability. The core contributions are a metric-guided dynamic optimization paradigm and a generic, architecture-agnostic domain adaptation framework.

Adapting ASR models to specialized domains effectivelyOptimizing fine-tuning for large-scale LLM-based ASR systemsPreventing overfitting while improving domain-specific performance

This study investigates how self-supervised speech models (S3Ms) learn and encode speaker demographic attributes—such as gender, age, dialect, ethnicity, and native-language status—and how different fine-tuning strategies influence the retention or amplification of such information. By integrating S3Ms with speaker identification (SID) and automatic speech recognition (ASR) fine-tuning, alongside fairness-enhancing algorithms, and employing hierarchical representation probing and embedding subspace analysis, the work provides the first systematic characterization of the layered encoding mechanisms of demographic information in these models. The findings reveal that SID fine-tuning amplifies global acoustic variation features, whereas ASR fine-tuning selectively preserves semantics-related variation while suppressing acoustic variation. Furthermore, fairness interventions primarily attenuate acoustic variation with limited impact on semantic variation, offering theoretical grounding for developing equitable ASR systems.

fairness in ASRself-supervised speech recognitionspeaker group encoding

Hot Scholars

MH

Mark Hasegawa-Johnson

Professor of Electrical and Computer Engineering, University of Illinois
SpeechAudioNatural Language Processing
QW

Qianli Wang

DFKI & TU Berlin
ExplainabilityNatural Language Processing
BA

Busayo Awobade

Research Scientist, MLCollective
Speech processingMultilinguality.
KD

Kunal Dhawan

Research Scientist, NVIDIA
Machine LearningDeep LearningSpeech ProcessingNatural Language Processing