🤖 AI Summary
This study addresses the limited accuracy of existing forced alignment methods under long audio, complex acoustic conditions, and ASR transcription errors by introducing AlignBench, a dedicated evaluation benchmark, and FuseAlign, a Transformer-based model. FuseAlign performs joint audio-visual contextual modeling with millisecond-level boundary refinement. It incorporates an online label correction mechanism via exponential moving average (EMA) snapshots to detect missing words without relying on lexicons or Viterbi decoding. Furthermore, convolutional upsampling and large-scale pseudo-labeled training are employed to enhance robustness. Experimental results demonstrate that FuseAlign significantly outperforms baselines on AlignBench and maintains robust performance in real-world ASR transcription scenarios, validating the critical contribution of each proposed module.
📝 Abstract
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.