arXiv · 2608.26697
Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study
Abstract
Synthetic speech offers scalable supervision for automatic speech recognition (ASR), but its benefit depends on text selection, reference speech, and augmentation scale. We present a phoneme-based TTS-to-ASR pipeline using a single TTS model jointly trained from scratch on Arabic, French, Italian, and Portuguese with the F5-TTS architecture and language-monolingual ASR systems cover 13 test sets. Across the synthesis-scale sweep, random augmentation improves over matched real-only continuation on 11 sets. In the selection comparison, PFGS improves over real-only training on 12 sets and over random selection on nine, with a maximum relative WER reduction of 19.3% against random selection. With target texts and synthesis counts fixed, reference-speech filtering reduces absolute WER by 0.29 and 0.59 points on Italian and French Common Voice, respectively. These findings support treating TTS augmentation as a synthetic-corpus construction problem, rather than merely a question of generation scale.
Explore related subjects
Keep this discovery
Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang, Wei Xu. 2026-08-27. Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study. https://arxiv.org/abs/2608.26697
Cite the original work for its findings. Save a collection to share your selection of sources.