arXiv · 2505.15722
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
Abstract
We present the first comprehensive study of Memorization in Multilingual Large Language Models (MLLMs), analyzing 95 languages using models across diverse model scales, architectures, and memorization definitions. As MLLMs are increasingly deployed, understanding their memorization behavior has become critical. Yet prior work has focused primarily on monolingual models, leaving multilingual memorization underexplored, despite the inherently long-tailed nature of training corpora. We find that the prevailing assumption, that memorization is highly correlated with training data availability, fails to fully explain memorization patterns in MLLMs. We hypothesize that the conventional focus on monolingual settings, effectively treating languages in isolation, may obscure the true patterns of memorization. To address this, we propose a novel graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization. Our analysis reveals that among similar languages, those with fewer training tokens tend to exhibit higher memorization, a trend that only emerges when cross-lingual relationships are explicitly modeled. These findings underscore the importance of a \textit{language-aware} perspective in evaluating and mitigating memorization vulnerabilities in MLLMs. This also constitutes empirical evidence that language similarity both explains Memorization in MLLMs and underpins Cross-lingual Transferability, with broad implications for multilingual NLP.
Explore related subjects
Keep this discovery
Xiaoyu Luo, Yiyi Chen, Johannes Bjerva, Qiongxiu Li. 2025-05-21. Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities. https://doi.org/10.18653/v1/2025.emnlp-main.978
Cite the original work for its findings. Save a collection to share your selection of sources.