DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the entanglement of speaker identity and accent in reference audio for zero-shot text-to-speech (TTS) synthesis by proposing a training-free decoupled control method. Building upon the F5-TTS architecture, the approach integrates LoRA-based parameter-efficient fine-tuning with an exemplar encoder to disentangle speaker and accent representations. Furthermore, it enables continuous and dynamic modulation of accent intensity through guidance weights at inference time, generalizing effectively to unseen out-of-domain accents. Experimental results demonstrate that the proposed method significantly improves accent probe accuracy while preserving high speaker similarity, achieving performance comparable to that of cascaded models.
📝 Abstract
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot TTS
Accent Control
Speaker Identity Disentanglement
Exemplar-Guided
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot TTS
Accent Control
Speaker-Accent Decoupling
Exemplar-Guided
LoRA Adaptation
🔎 Similar Papers