arXiv · cs/9811031
Speech Synthesis with Neural Networks
Abstract
Text-to-speech conversion has traditionally been performed either by concatenating short samples of speech or by using rule-based systems to convert a phonetic representation of speech into an acoustic representation, which is then converted into speech. This paper describes a system that uses a time-delay neural network (TDNN) to perform this phonetic-to-acoustic mapping, with another neural network to control the timing of the generated speech. The neural network system requires less memory than a concatenation system, and performed well in tests comparing it to commercial systems using other technologies.
Explore related subjects
Keep this discovery
Orhan Karaali, Gerald Corrigan, Ira Gerson. 1998-11-24. Speech Synthesis with Neural Networks. https://arxiv.org/abs/cs/9811031
Cite the original work for its findings. Save a collection to share your selection of sources.