arXiv · 2404.12299
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
Abstract
In Simultaneous Machine Translation (SiMT) systems, training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency systems. However, it is very challenging to curate such a corpus due to limitations in the abilities of annotators, and hence, existing SI corpora are limited. Therefore, we propose a method to convert existing speech translation corpora into interpretation-style data, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus). We demonstrate that fine-tuning SiMT models in text-to-text and speech-to-text settings with the LLM-SI-Corpus reduces latencies while maintaining the same level of quality as the models trained with offline datasets. The LLM-SI-Corpus is available at \url{https://github.com/yusuke1997/LLM-SI-Corpus}.
Explore related subjects
Keep this discovery
Yusuke Sakai, Mana Makinae, Hidetaka Kamigaito, Taro Watanabe. 2024-04-18. Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair. https://arxiv.org/abs/2404.12299
Cite the original work for its findings. Save a collection to share your selection of sources.