π€ AI Summary
This study addresses the significant challenge of cross-lingual representation alignment across diverse writing systems in multilingual models. We propose a code-switching curriculum learning strategy that leverages large language models to generate word-level and sentence-level code-switched text, constructing the BabyBabelLM dataset. Using this resource, we conduct progressive pre-training on decoder-only Transformer architectures, transitioning from mixed-language corpora to monolingual data. This approach effectively enhances cross-lingual alignment among English, Dutch, and Chinese. Empirical evaluations demonstrate that our method significantly outperforms baseline models on the BabyLM evaluation suite. Ultimately, this work establishes an efficient data augmentation paradigm for multilingual pre-training.
π Abstract
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.