arXiv · 2012.15178
Synthetic Source Language Augmentation for Colloquial Neural Machine Translation
Abstract
Neural machine translation (NMT) is typically domain-dependent and style-dependent, and it requires lots of training data. State-of-the-art NMT models often fall short in handling colloquial variations of its source language and the lack of parallel data in this regard is a challenging hurdle in systematically improving the existing models. In this work, we develop a novel colloquial Indonesian-English test-set collected from YouTube transcript and Twitter. We perform synthetic style augmentation to the source of formal Indonesian language and show that it improves the baseline Id-En models (in BLEU) over the new test data.
Explore related subjects
Keep this discovery
Asrul Sani Ariesandy, Mukhlis Amien, Alham Fikri Aji, Radityo Eko Prasojo. 2020-12-30. Synthetic Source Language Augmentation for Colloquial Neural Machine Translation. https://arxiv.org/abs/2012.15178
Cite the original work for its findings. Save a collection to share your selection of sources.