arXiv · 2609.30535
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
Abstract
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kianté Brantley. 2026-09-24. Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence. https://arxiv.org/abs/2609.30535
Cite the original work for its findings. Save a collection to share your selection of sources.