Score
Designs and implements systems that align speech audio with textual or phonetic annotations and compare resulting segmentations, producing time‑aligned labels (phoneme/phone boundaries, word boundaries, TextGrid-style intervals) from audio and transcripts across languages and dialects. Builds forced‑alignment algorithms, distance metrics, scaling and visualization or auditory tools for segmentation/alignment evaluation and error analysis (for example reporting mean boundary error, overlays, and support for large multilingual corpora).
Existing forced alignment tools produce only point estimates of segment boundaries without quantifying uncertainty. To address this limitation, we propose the first confidence interval estimation method for forced alignment based on model ensembling and order statistics: ten independently trained segment classification neural networks are aggregated; boundary predictions are centered at the median, and 97.85% confidence intervals are constructed via order statistics. The method supports Praat TextGrid point-tier output and provides an interpretable boundary diagnostic table. This work is the first to integrate neural network ensembling with order statistics for uncertainty modeling in forced alignment. Evaluated on the Buckeye and TIMIT corpora, our approach achieves marginally higher boundary accuracy than single-model baselines while enabling uncertainty-aware linguistic analysis. It has been deployed in real-world speech processing pipelines, facilitating robust, uncertainty-informed phonetic and phonological modeling.
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
This work addresses the challenge of insufficient word-level forced alignment accuracy in low-resource and unseen languages by proposing a multilingual alignment approach that integrates self-supervised speech representations with learnable dynamic time warping. The method introduces, for the first time, a learnable dynamic programming framework for word boundary inference, jointly leveraging representations from the MMS model and the UnSupSeg boundary detector. An iterative training mechanism further enhances alignment performance. Notably, the approach generalizes to over 1,100 languages without additional training and outperforms baseline systems—including the Montreal Forced Aligner and MMS—on the TIMIT and Buckeye datasets. It also achieves comparable or superior results on unseen languages such as Dutch, German, and Hebrew.
This work addresses the challenge of real-time, multilingual text-to-speech forced alignment. We propose an efficient, fine-grained alignment method that explicitly models inter-phoneme gaps and silence segments. Its core innovation is a hierarchical decoding framework integrating a context-agnostic universal phoneme encoder (CUPE) with a Connectionist Temporal Classification (CTC) decoder, enabling simultaneous prediction of phoneme onset and offset boundaries—thereby significantly enhancing temporal structure modeling. Compared to conventional approaches, our system achieves boundary recall rates exceeding 92% on the TIMIT and Buckeye corpora—comparable to the Montreal Forced Aligner—while operating at 240× real-time speed. To our knowledge, this is the first method to achieve truly faster-than-real-time multilingual forced alignment. The approach thus enables high-accuracy, ultra-low-latency alignment essential for interactive speech applications.
This work addresses the limited parallelizability of the classical dynamic time warping (DTW) algorithm, which suffers from quadratic time and memory complexity. The authors propose Segmental DTW, a novel approach that decomposes global sequence alignment into local subsequence DTW computations that can be executed in parallel, followed by a segment-level dynamic programming step to integrate the partial alignments. This method achieves near-full parallelism while preserving alignment accuracy comparable to standard DTW. Theoretical analysis and empirical evaluation on Chopin Mazurka audio alignment tasks demonstrate that one variant of the proposed method outperforms existing approaches in both computational efficiency and alignment performance.
This study addresses the challenge of accurate phoneme alignment in low-resource dialects by focusing on Chengdu Mandarin. Leveraging 17 hours of speech data and a custom phonetic lexicon, the authors propose a bootstrapped training pipeline: first, a text-dependent GMM-HMM aligner (Chengdu-MFA) is developed, whose outputs are then used as pseudo-labels to fine-tune a pretrained audio encoder for text-independent, frame-level classification-based alignment (Chengdu-FC). Evaluated on an expert-annotated test set, Chengdu-MFA reduces phoneme boundary error by 31.8% relative to a Mandarin baseline, and Chengdu-FC further improves this reduction to 61.2%. These results demonstrate the effectiveness and novelty of the proposed approach for phoneme alignment in low-resource dialectal settings.