Balalaika-Longform: A Russian Speech Corpus for Continuous Long-Form Text-to-Speech

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of word omission and premature stopping in long-form text-to-speech (TTS) synthesis caused by short-utterance training. To isolate the effect of training sequence length on synthesis performance, we construct the first 189-hour Russian long-audio corpus supporting continuous generation. Methodologically, building upon mainstream TTS models such as CosyVoice3, we design a fine-grained alignment strategy to conduct comparative fine-tuning experiments between short- and long-window configurations. Experimental results demonstrate that long-target fine-tuning significantly mitigates generation degradation; notably, the word error rate of CosyVoice3 decreases from 99.9% to 47.3%. These findings effectively validate the critical role of long-sequence training in enhancing continuous generation capabilities for TTS systems.
📝 Abstract
Long-form text-to-speech must retain requested words over extended generations, yet sentence-level training and evaluation can hide omissions and early stops. We introduce Balalaika-Longform, an open Russian corpus of 189 hours in continuous units of 30 seconds to 15 minutes. Long units and matched short windows support fine-tuning comparisons on the same source recordings, and the accompanying evaluation retains every synthesis attempt. We fine-tune CosyVoice3, Qwen3-TTS, VoxCPM2 and F5-TTS on either view and test continuous synthesis on 50 text-voice pairs with voices unseen in fine-tuning, at about 75, 300 and 1,200 words. At 1,200 words, CosyVoice3 WER falls from 99.9% after short-window fine-tuning to 47.3% with long targets, and to 16.6% with punctuated, format-matched transcripts; paired intervals favor long targets for Qwen3-TTS and VoxCPM2, and by a small margin for F5-TTS, whose absolute WER stays above 90%. The results isolate the effect of training sequence length under a fixed continuous-generation protocol.
Problem

Research questions and friction points this paper is trying to address.

long-form text-to-speech
continuous speech synthesis
Russian speech corpus
word omission
early stopping
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-form text-to-speech
speech corpus
training sequence length
continuous generation
fine-tuning
N
Nikita Vasiliev
BitmanagerAI, Dubai, UAE; lab260, Yerevan, Armenia
Kirill Borodin
Kirill Borodin
MTUCI
deep learning for audiogen AIsafe AI
V
Vasilii Kudryavtsev
BitmanagerAI, Dubai, UAE; lab260, Yerevan, Armenia
M
Maxim Maslov
lab260, Yerevan, Armenia
Grach Mkrtchian
Grach Mkrtchian
MTUCI
Artificial IntelligenceAlgorithmsData Structures