arXiv · 2609.21145
Scaling Forced Alignment to End-User Devices
Abstract
The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling poorly to long input sequences. We propose two optimizations to address this issue. First, we apply the Hirschberg algorithm to perform the alignment in place using linear memory. Second, we model the alignment between speech and text as a constrained random walk, allowing us to prune the search space with arbitrary confidence while accounting for transcription errors. The Hirschberg optimization reduces memory usage from 140 GB to 5 MB for three-hour inputs while producing identical alignments in one-third the time of torchaudio when both run on a CPU. We achieve an additional 2x speedup with pruning on inputs longer than 20 minutes while preserving alignment accuracy in more than 98% of tested cases.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lawry Sorenson, Michael Crandall, Eric K. Ringger, Stephen D. Richardson. 2026-09-17. Scaling Forced Alignment to End-User Devices. https://arxiv.org/abs/2609.21145
Cite the original work for its findings. Save a collection to share your selection of sources.