Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient cognitive grounding of code-switching mechanisms and the unclear efficiency of cross-lingual alignment in bilingual model pre-training. We systematically investigate, for the first time, how structural parameters and dynamic switching rates influence synthetic code-switching training. By controlling the structural positions and dynamic switching rates of code-switching, we synthesize code-switched data across two language pairs and employ a multi-stage dynamic training strategy for pre-training and cross-lingual alignment evaluation. Our findings demonstrate that appropriately regulating the structural and dynamic properties of code-switching significantly enhances cross-lingual alignment performance for typologically similar languages. This work provides a novel cognitively inspired paradigm for the pre-training of computational bilingual models.
📝 Abstract
During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.
Problem

Research questions and friction points this paper is trying to address.

Code-Switching
Bilingual Model
Cross-lingual Alignment
Pretraining
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code-Switching
Bilingual Model Training
Cross-lingual Alignment
Cognitively-Inspired
Synthetic Data
🔎 Similar Papers
No similar papers found.
Z
Zhuojing Huang
University of Göttingen, Germany
L
Luise Pohlmann
University of Göttingen, Germany
Lisa Beinborn
Lisa Beinborn
Human-Centered Data Science, University of Göttingen