arXiv · 2609.21084
The Hidden Cost of Digits: Number Normalization and WER in ASR Systems
Abstract
Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts and proper processing of reference transcripts. Popular approaches often reduce text normalization to lowercase and remove punctuation, with no additional normalization applied to languages other than English. In this work, we analyze the impact of normalization of numerical expressions in the evaluation of ASR systems in various languages, using Polish as an example of a highly inflective language. We perform experiments on VoxPopuli and The Polish Parliamentary speech datasets and estimate word error rate (WER) differences for different text normalization approaches. We show that the difference due to the lack of number normalization in WER may be substantial - more than 2 percentage points, and often higher than the differences between systems in popular multilingual benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Stanisław Kacprzak, Mieszko Fraś. 2026-09-17. The Hidden Cost of Digits: Number Normalization and WER in ASR Systems. https://arxiv.org/abs/2609.21084
Cite the original work for its findings. Save a collection to share your selection of sources.