Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the speaker and language adaptation challenge in lightweight cross-lingual text-to-speech (TTS) systems when no speech recordings are available for the target language. We propose an adapter-based parameter-efficient fine-tuning method that decouples multilingual phonetic modeling from speaker representation. Inspired by second-language acquisition theory, we further introduce an objective accent evaluation metric to systematically analyze the impact of adapter placement, architecture, and number of training speakers on synthesis performance. Experiments demonstrate that our approach efficiently acquires novel language and speaker characteristics without target-language speech data, significantly mitigating catastrophic forgetting. Both subjective listening tests and objective evaluations confirm superior speech naturalness and accent fidelity over baseline methods. The proposed framework provides a scalable, interpretable, and lightweight solution for low-resource cross-lingual TTS.

Technology Category

Natural Language Processing: SpeechMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Adaptive Behavior

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
Problem

Research questions and friction points this paper is trying to address.

Adapting TTS to unseen speakers and languages
Synthesizing target voice without target language recordings
Avoiding catastrophic forgetting in lightweight TTS systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapters enable cross-lingual TTS synthesis
Adapters prevent catastrophic forgetting in adaptation
Objective metric measures accent nativeness objectively
🔎 Similar Papers
No similar papers found.
A
Alessio Falai
Amazon AGI
Z
Ziyao Zhang
Amazon AGI
A
Akos Gangoly
Amazon AGI