Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the asymmetric scaling laws governing inference steps and model depth with respect to intelligibility and speaker identity in masked diffusion text-to-speech (TTS) systems. By training masked diffusion codec TTS models of varying depths and employing a Best-of-K search strategy alongside automatic speech recognition and speaker verification evaluations, we quantify how test-time compute influences synthesis performance. Our analysis reveals a decoupled paradigm wherein refinement steps primarily enhance intelligibility while search strategies optimize speaker identity, demonstrating that depth and step count are non-interchangeable. Experiments show that the refinement process bridges 86.2% of the intelligibility gap but only 46.4% of the identity gap. Furthermore, we find that 62% of the identity bottleneck originates from the codec rather than the generator, providing new theoretical foundations for efficient speech synthesis.
📝 Abstract
Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweep refinement steps T in [1,16] at inference, measuring zero-shot synthesis via ASR word error rate (intelligibility) and speaker verification (identity) on 174 held-out speakers. Against measured floors, refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range - a 1.86x asymmetry robust across multiple error metrics. Retraining at 3x and 6x schedule attenuates but does not reverse this gap (1.84 to 1.36 to 1.23x), because intelligibility saturates with steps while identity continues improving. Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. Depth and steps are not interchangeable: separable B(d)B(T) fits significantly better (Delta AICc=+69.3) than substitution models. Analysis shows 62% of remaining identity deficit lies in the codec, not the generator. We conclude that refinement and depth target different bottlenecks and should be optimized separately.
Problem

Research questions and friction points this paper is trying to address.

Masked-Diffusion TTS
Test-Time Compute
Intelligibility
Speaker Identity
Scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked-Diffusion TTS
Test-Time Compute
Refinement Steps
Speaker Identity
Best-of-K Search