RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of reward alignment in discrete diffusion-based voice cloning, where complex sampling trajectories hinder effective optimization. To overcome this, we propose a ratio-free advantage-weighted denoising method that integrates group relative advantage estimation with masked token reconstruction. This approach enables efficient alignment without requiring reverse trajectory likelihoods or target audio, thereby optimizing speaker rewards for zero-shot text-to-speech synthesis. Experimental results demonstrate that the proposed method reduces the word error rate to 2.42% and increases speaker similarity to 0.755, significantly improving both synthesized speech quality and identity consistency.
📝 Abstract
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
Problem

Research questions and friction points this paper is trying to address.

zero-shot text-to-speech
voice cloning
discrete diffusion
reward alignment
speaker identity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Diffusion
Reward Alignment
Voice Cloning
Advantage-Weighted Denoising
Zero-shot TTS
🔎 Similar Papers
No similar papers found.