FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy Conditions

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of cross-study comparability in forced alignment evaluation caused by inconsistent data partitioning, text normalization, and scoring criteria by constructing a unified, open-source evaluation framework. Methodologically, it introduces a tolerance-based F1 metric to mitigate artificially inflated MAE scores and conducts dual-track evaluations of 21 models under both clean and noisy conditions. Furthermore, it analyzes errors stratified by positional and adjacent-word states while performing fine-grained boundary detection. The findings reveal systematic temporal biases, such as Whisper’s consistent 150ms anticipation. By providing reproducible code and a dynamically updated benchmark, this work establishes a standardized paradigm for forced alignment research.
πŸ“ Abstract
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the code, splits, phone mapping, text normalization and scoring script, with results published periodically. Track 1 gives every aligner the reference transcript and Track 2 gives it a recognizer's output, on the same audio, clean and degraded four ways, with 21 open models and 9 commercial APIs under a unified protocol. We score every boundary of an utterance and check the two labels beside it, so a word the recognizer missed or invented is charged. Using a tolerance-based F1 as our primary metric eliminates the 9% to 14% score inflation that standard MAE causes on recognition-dependent systems in conversational speech. We then group boundaries by their position and how many adjacent words were recognized correctly, which shows where a system lost the score. We discovered systematic bias in how current systems time words, with Whisper about 150 ms early and several commercial ASR APIs over 50 ms late. Code and results are at https://github.com/olewave/fa-bench
Problem

Research questions and friction points this paper is trying to address.

Forced Alignment
ASR Timestamps
Benchmark
Evaluation Metric
Speech Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Forced Alignment
Benchmark
ASR Timestamps
Tolerance-based F1
Systematic Bias
πŸ”Ž Similar Papers
No similar papers found.