arXiv · 1911.04997
Character-based NMT with Transformer
Abstract
Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. In particular, our experiments on EN-DE show that character-based Transformer models are more robust than their BPE counterpart, both when translating noisy text, and when translating text from a different domain. To obtain comparable BLEU scores in clean, in-domain data and close the gap with BPE-based models we use known techniques to train deeper Transformer models.
Explore related subjects
Keep this discovery
Rohit Gupta, Laurent Besacier, Marc Dymetman, Matthias Gallé. 2019-11-12. Character-based NMT with Transformer. https://arxiv.org/abs/1911.04997
Cite the original work for its findings. Save a collection to share your selection of sources.