Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of dialectal speech data, which constrains the performance of speech language models (SLMs), and the reliance of conventional text-to-speech (TTS) augmentation on authentic dialectal resources. To overcome these limitations, this work proposes a zero-resource pseudo-dialect augmentation framework that leverages large language models to generate dialectal texts, subsequently synthesizing pseudo-dialectal speech via standard TTS systems. Furthermore, an intermediate standard-text prediction task is introduced as auxiliary supervision to achieve cross-dialectal semantic normalization, effectively bridging the semantic gap between dialects and standard languages. Notably, this approach enables multilingual extension without requiring any authentic dialectal speech. Extensive experiments demonstrate significant performance improvements for SLMs across Japanese, German, and Chinese dialect understanding tasks, validating the efficacy of the proposed framework in low-resource dialectal scenarios.
📝 Abstract
Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.
Problem

Research questions and friction points this paper is trying to address.

Speech Language Model
Dialect Robustness
Data Scarcity
Pseudo-Dialect Augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech Language Model
Pseudo-Dialect Augmentation
Intermediate Standard-Text Prediction
Dialect-Robustness
Data Scarcity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shunsuke Mitsumori
SB Intuitions, Waseda University, Tokyo, Japan
T
Tomoya Mizumoto
SB Intuitions
Yusuke Fujita
Yusuke Fujita
SB Intuitions
Automatic Speech RecognitionSpeech SeparationSpeaker Diarization