Scaling Forced Alignment to End-User Devices

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对长音频序列对齐问题,通过应用Hirschberg算法和将语音文本对齐建模为受限随机游走的方法,优化了内存使用并提升了处理速度。
📝 Abstract
The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech and text as a constrained random walk, allowing us to prune the search space with arbitrary confidence while accounting for transcription errors. The Hirschberg optimization reduces memory usage from 140 GB to 5 MB for three-hour inputs while producing identical alignments in one-third the time of torchaudio when both run on a CPU. We achieve an additional 2x speedup with pruning on inputs longer than 20 minutes while preserving alignment accuracy in more than 98% of tested cases.
Problem

Research questions and friction points this paper is trying to address.

Forced Alignment
Viterbi Algorithm
Time and Space Complexity
End-User Devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hirschberg algorithm
constrained random walk
pruning
💼 Related Jobs
No related jobs found.
L
Lawry Sorenson
Brigham Young University, Computer Science Department, Provo, UT, United States
Michael Crandall
Michael Crandall
Brigham Young University, Computer Science Department, Provo, UT, United States
E
Eric K. Ringger
Brigham Young University, Computer Science Department, Provo, UT, United States
Stephen D. Richardson
Stephen D. Richardson
Brigham Young University, Computer Science Department, Provo, UT, United States