BFA: Real-time Multilingual Text-to-speech Forced Alignment

📅 2025-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of real-time, multilingual text-to-speech forced alignment. We propose an efficient, fine-grained alignment method that explicitly models inter-phoneme gaps and silence segments. Its core innovation is a hierarchical decoding framework integrating a context-agnostic universal phoneme encoder (CUPE) with a Connectionist Temporal Classification (CTC) decoder, enabling simultaneous prediction of phoneme onset and offset boundaries—thereby significantly enhancing temporal structure modeling. Compared to conventional approaches, our system achieves boundary recall rates exceeding 92% on the TIMIT and Buckeye corpora—comparable to the Montreal Forced Aligner—while operating at 240× real-time speed. To our knowledge, this is the first method to achieve truly faster-than-real-time multilingual forced alignment. The approach thus enables high-accuracy, ultra-low-latency alignment essential for interactive speech applications.

Technology Category

Natural Language Processing: SpeechCognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
We present Bournemouth Forced Aligner (BFA), a system that combines a Contextless Universal Phoneme Encoder (CUPE) with a connectionist temporal classification (CTC)based decoder. BFA introduces explicit modelling of inter-phoneme gaps and silences and hierarchical decoding strategies, enabling fine-grained boundary prediction. Evaluations on TIMIT and Buckeye corpora show that BFA achieves competitive recall relative to Montreal Forced Aligner at relaxed tolerance levels, while predicting both onset and offset boundaries for richer temporal structure. BFA processes speech up to 240x faster than MFA, enabling faster than real-time alignment. This combination of speed and silence-aware alignment opens opportunities for interactive speech applications previously constrained by slow aligners.
Problem

Research questions and friction points this paper is trying to address.

Develops real-time multilingual speech-text alignment system
Improves phoneme boundary prediction with gap modeling
Enables interactive applications through faster-than-real-time processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines CUPE encoder with CTC decoder
Models inter-phoneme gaps and silences
Uses hierarchical decoding for boundary prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Abdul Rehman
Faculty of Media, Science and Technology, Bournemouth University, Bournemouth, United Kingdom
J
Jingyao Cai
Faculty of Media, Science and Technology, Bournemouth University, Bournemouth, United Kingdom
J
Jian-Jun Zhang
Faculty of Media, Science and Technology, Bournemouth University, Bournemouth, United Kingdom
X
Xiaosong Yang
Faculty of Media, Science and Technology, Bournemouth University, Bournemouth, United Kingdom