arXiv · 2306.11662
Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer
Abstract
Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing and transferring details and variations in prosody is a challenge. We introduce phrase-level cross-lingual prosody transfer for expressive multi-lingual machine dubbing. The proposed phrase-level prosody transfer delivers a significant 6.2% MUSHRA score increase over a baseline with utterance-level global prosody transfer, thereby closing the gap between the baseline and expressive human dubbing by 23.2%, while preserving intelligibility of the synthesised speech.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Giuseppe Coccia, Patrick Lumban Tobing, Ravichander Vipperla, Viacheslav Klimkov, Vincent Pollet. 2023-06-20. Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer. https://arxiv.org/abs/2306.11662
Cite the original work for its findings. Save a collection to share your selection of sources.