🤖 AI Summary
This study addresses the issue of word omission and premature stopping in long-form text-to-speech (TTS) synthesis caused by short-utterance training. To isolate the effect of training sequence length on synthesis performance, we construct the first 189-hour Russian long-audio corpus supporting continuous generation. Methodologically, building upon mainstream TTS models such as CosyVoice3, we design a fine-grained alignment strategy to conduct comparative fine-tuning experiments between short- and long-window configurations. Experimental results demonstrate that long-target fine-tuning significantly mitigates generation degradation; notably, the word error rate of CosyVoice3 decreases from 99.9% to 47.3%. These findings effectively validate the critical role of long-sequence training in enhancing continuous generation capabilities for TTS systems.
📝 Abstract
Long-form text-to-speech must retain requested words over extended generations, yet sentence-level training and evaluation can hide omissions and early stops. We introduce Balalaika-Longform, an open Russian corpus of 189 hours in continuous units of 30 seconds to 15 minutes. Long units and matched short windows support fine-tuning comparisons on the same source recordings, and the accompanying evaluation retains every synthesis attempt. We fine-tune CosyVoice3, Qwen3-TTS, VoxCPM2 and F5-TTS on either view and test continuous synthesis on 50 text-voice pairs with voices unseen in fine-tuning, at about 75, 300 and 1,200 words. At 1,200 words, CosyVoice3 WER falls from 99.9% after short-window fine-tuning to 47.3% with long targets, and to 16.6% with punctuated, format-matched transcripts; paired intervals favor long targets for Qwen3-TTS and VoxCPM2, and by a small margin for F5-TTS, whose absolute WER stays above 90%. The results isolate the effect of training sequence length under a fixed continuous-generation protocol.