arXiv · 2311.12405
IndoRobusta: Towards Robustness Against Diverse Code-Mixed Indonesian Local Languages
Abstract
Significant progress has been made on Indonesian NLP. Nevertheless, exploration of the code-mixing phenomenon in Indonesian is limited, despite many languages being frequently mixed with Indonesian in daily conversation. In this work, we explore code-mixing in Indonesian with four embedded languages, i.e., English, Sundanese, Javanese, and Malay; and introduce IndoRobusta, a framework to evaluate and improve the code-mixing robustness. Our analysis shows that the pre-training corpus bias affects the model's ability to better handle Indonesian-English code-mixing when compared to other local languages, despite having higher language diversity.
Explore related subjects
Keep this discovery
Muhammad Farid Adilazuarda, Samuel Cahyawijaya, Genta Indra Winata, Pascale Fung, Ayu Purwarianti. 2023-11-21. IndoRobusta: Towards Robustness Against Diverse Code-Mixed Indonesian Local Languages. https://arxiv.org/abs/2311.12405
Cite the original work for its findings. Save a collection to share your selection of sources.