Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of annotated data and underrepresentation of patient populations in clinical speech tasks under low-resource language settings by proposing a data augmentation strategy that leverages voice cloning while preserving paralinguistic information. It presents the first systematic evaluation of eight voice cloning models with respect to their ability to retain paralinguistic cues indicative of conditions such as depression and anxiety, and further demonstrates cross-lingual voice cloning and transfer learning from English to Japanese. Experimental results show that most cloning models effectively preserve paralinguistic features with only minor performance degradation. Moreover, models trained on cloned data significantly outperform those using direct cross-lingual transfer when evaluated on real Japanese clinical speech, confirming the effectiveness and novelty of the proposed approach for low-resource clinical speech analysis.
📝 Abstract
Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear whether these preserve the paralinguistic signal such tasks depend on. We benchmark eight voice cloning models on five paralinguistic tasks across public and clinical datasets, showing most preserve signal with modest degradation. We then clone English clinical speech into Japanese and find that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

paralinguistic
voice cloning
synthetic speech
clinical speech
data augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

voice cloning
paralinguistic preservation
cross-lingual augmentation
clinical speech
synthetic data augmentation
🔎 Similar Papers
No similar papers found.