FuseAlign: Forced Alignment in the Wild

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited accuracy of existing forced alignment methods under long audio, complex acoustic conditions, and ASR transcription errors by introducing AlignBench, a dedicated evaluation benchmark, and FuseAlign, a Transformer-based model. FuseAlign performs joint audio-visual contextual modeling with millisecond-level boundary refinement. It incorporates an online label correction mechanism via exponential moving average (EMA) snapshots to detect missing words without relying on lexicons or Viterbi decoding. Furthermore, convolutional upsampling and large-scale pseudo-labeled training are employed to enhance robustness. Experimental results demonstrate that FuseAlign significantly outperforms baselines on AlignBench and maintains robust performance in real-world ASR transcription scenarios, validating the critical contribution of each proposed module.
📝 Abstract
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.
Problem

Research questions and friction points this paper is trying to address.

forced alignment
word-level alignment
ASR errors
evaluation benchmark
real-world speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

Forced Alignment
FuseAlign
Online Label Correction
Joint Contextualization
AlignBench
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mithilesh Vaidya
Descript, Inc.
S
Stephen Bailey
Descript, Inc.
S
Sumukh Badam
Descript, Inc.
M
Matthew Bendel
Descript, Inc.
X
Xingzhe He
Descript, Inc.