From Reliable Text to Real Voices: Trust-Aware Progressive Adaptation for Low-Resource TTS

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对低资源TTS适应问题,提出了一种信任感知的渐进式适应方法,通过合成到真实语音的转换,并结合伪标签可靠性加权,以提高内容准确性、自然度和说话人相似性。
📝 Abstract
Low-resource text-to-speech (TTS) adaptation is constrained by scarce paired data and costly manual transcription. Existing fixed-voice TTS systems can provide relatively accurate pronunciation, but their synthetic speech offers limited speaker diversity and may exhibit flat prosody. Real recordings provide natural prosody and diverse voices, yet their automatic speech recognition (ASR) pseudo-labels may contain transcription errors. We find that supervision order affects content accuracy and speaker similarity. We propose trust-aware progressive adaptation: synthetic-to-real adaptation first establishes text-speech correspondences, then restores reference-speaker control using real speech. Transcript-agreement weighting uses agreement between two fixed ASR systems as a proxy for pseudo-label reliability to limit noisy supervision. Experiments with FireRedTTS3 on Burmese and Lao and OmniVoice on Burmese show improved content accuracy with high naturalness and competitive speaker similarity. Jointly considering supervision order and pseudo-label reliability when combining synthetic and real speech offers a practical path to zero-shot voice cloning in low-resource languages with less manual transcription. Audio demos are available at https://insiderx-pro.github.io/S2R-Adaptation-TTS/
Problem

Research questions and friction points this paper is trying to address.

low-resource TTS
scarce paired data
manual transcription
synthetic speech
speaker diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

trust-aware progressive adaptation
transcript-agreement weighting
low-resource TTS
🔎 Similar Papers
No similar papers found.