🤖 AI Summary
This study addresses the failure modes of neural text-to-speech (TTS) models—specifically looping, truncation, and counting errors—when processing repetitive phrases. We introduce a novel analytical framework that decouples textual repetitiveness from sequence length. Through paired controlled-set testing, multi-architecture benchmarking, and cross-validation across multiple automatic speech recognition (ASR) systems, we demonstrate that textual repetitiveness, rather than length, is the primary driver of model failure, further revealing a smoothing effect of periodicity on accuracy. Our analysis quantifies a precipitous drop in accuracy to 18.2% under highly repetitive conditions and shows that the proposed framework can predict the performance of emerging architectures with an error margin within one percentage point.
📝 Abstract
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.