FuseAlign: Forced Alignment in the Wild
This study addresses the limited accuracy of existing forced alignment methods under long audio, complex acoustic conditions, and ASR transcription errors by introducing AlignBench, a dedicated evaluation benchmark, and FuseAlign, a Transformer-based model. FuseAlign performs joint audio-visual contextual modeling with millisecond-level boundary refinement. It incorporates an online label correction mechanism via exponential moving average (EMA) snapshots to detect missing words without relying on lexicons or Viterbi decoding. Furthermore, convolutional upsampling and large-scale pseudo-labeled training are employed to enhance robustness. Experimental results demonstrate that FuseAlign significantly outperforms baselines on AlignBench and maintains robust performance in real-world ASR transcription scenarios, validating the critical contribution of each proposed module.