BabyLMs的通心粉训练:代码切换课程导致跨语言融合
摘要
本文探讨在数据限制下使用代码切换文本来对齐小型语言模型的表征,表明涉及代码切换的课程能提升多语言性能。
arXiv:2609.30535v1 Announce Type: new
Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.
查看缓存全文
缓存时间: 2026/09/28 09:39
# Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence Source: [https://arxiv.org/html/2609.30535](https://arxiv.org/html/2609.30535) \[ BoldFont = TeXGyreTermesX\-Bold\.otf, ItalicFont = TeXGyreTermesX\-Italic\.otf, BoldItalicFont = TeXGyreTermesX\-BoldItalic\.otf \]\\XeTeXlinebreaklocale"zh"\\XeTeXlinebreakskip= 0pt plus 0\.1pt Alex CaiEmail:[adzcai@g\.harvard\.edu](mailto:[email protected])Yonatan BelinkovAffiliation:Technion – Israel Institute of Technology, Haifa, IsraelDavid Alvarez\-MelisKianté Brantley\[4pt\] Kempner InstituteHarvard UniversityCambridgeMAUSA ###### Abstract Children in multilingual communities often*code\-switch*, using multiple languages in a single utterance\. Can we induce cross\-lingual alignment in language models by training on code\-switched text? We pretrain small decoder\-only transformers on two100100M\-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word\- and sentence\-level code\-switching using an LLM\. We find that training on code\-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents\. Under a learning curriculum that progresses from word\-level code\-switching, to sentence\-level code\-switching, to monolingual documents, models trained on code\-switched data outperform baselines trained without it on the BabyLM evaluation suite\. Our work characterizes code\-switching curriculum learning as an effective data augmentation method for multilingual pretraining\. We release our code, data, and models at[https://github\.com/drooryck/multilingual\-macaroni](https://github.com/drooryck/multilingual-macaroni)\. \*\*footnotetext:Equal contribution\.## 1Introduction Modern large language models \(LLMs\) are trained on orders of magnitude more language data than a person typically is exposed to in their lifetime\. Yet children in any country communicate fluently, often in multiple languages, after hearing under the equivalent of just 100 million English words\([Choshen et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib48)\)\. The BabyLM competition seeks to train computational language models under such developmentally plausible data constraints\. This year’s multilingual track asks participants to prepare the highest possible performing model across English, Dutch, and Chinese using the above total word budget\. AExamples of each data type Non\-CS dataennlzhThat’s illegal\!Geef me een kus \.\\cjkfont有什么区别? Sentence\-level CS — 6 ordered directionsIt’s you\.\\cjkfont不,不是。\\cjkfont他是谁呀。It’s Xu Song\.Ken je hem?He’s my brother\.Where’s the dog?Hij slaapt\.Zie je hem ?\\cjkfont他在那里。\\cjkfont他妈姓杨\.Ik heet Li\.enzhnl Word\-level CS — 6 ordered directionsI know it’ll\\cjkfont行\.\\cjkfont得believe\\cjkfont爱情。Pak dieshovel\.That’sillegaal\!Geef me een\\cjkfont吻\.\\cjkfont有什么verschil?enzhnl BCorpus composition ennlzhNon\-CS corpusCS corpuscurriculum training orderword\-level CSsentence\-level CSnon\-CS Figure 1:Composition of our synthetic code\-switched corpus\.A: We synthesize code\-switched \(CS\) data by prompting an LLM to translate words or sentences from documents in the monolingual competition datasets\. Examples shown are from our data\.B: We train models on two corpora\. Both are one\-third English, Dutch, and Chinese\. The non\-CS corpus comprises monolingual source documents, while the CS corpus splits each language’s subset equally into word\-level CS, sentence\-level CS, and monolingual documents, shown left to right in curriculum training order\.Table 1:Composition of the CS corpus by*matrix language*\(the source document language\) and*embedded language*\(the language of the inserted words or translated sentences\) in millions of byte\-premium\-adjusted words \(see Appendix[A](https://arxiv.org/html/2609.30535#A1)\)\. Cells may not sum exactly to totals due to rounding\.We take inspiration from infants raised in multilingual environments, who, within a single utterance, will use words from multiple languages, a phenomenon known ascode\-switching\(CS\)\([Volterra and Taeschner, 1978](https://arxiv.org/html/2609.30535#bib.bib50)\), or, rarely, ‘‘macaronic language’’\.111Some authors use the terms “code\-switching” and “code\-mixing” for different phenomena\. We follow[Yoo et al\. \(2025\)](https://arxiv.org/html/2609.30535#bib.bib2)in distinguishing between word\-level and sentence\-level code\-switching\. See also[https://en\.wikipedia\.org/wiki/Macaronic\_language](https://en.wikipedia.org/wiki/Macaronic_language)\.While children code\-switch naturally during language acquisition, might training neural language models intentionally on CS data help them align their representations of parallel text? This hypothesis has driven numerous schemes for multilingual pretraining\([Yang et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib13);[Li et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib5);[Wang et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib3)\)and fine\-tuning\([Qin et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib6);[Zheng et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib25);[Yoo et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib2);[Wang et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib59)\)\. Training on CS data, however, introduces a complication: the trained model is not typically intended to generate CS text, but rather monolingual text in each of multiple languages\. Code\-switching curriculum learning\([Yoo et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib2)\)presents a promising solution: first train on word\-level CS, then sentence\-level CS, and finally monolingual documents\. This ordering imitates the stages of a human learning a new language: first one learns new words, then incorporates whole sentences into their speech, and finally speaks fully in the new language\. Our work adapts code\-switching curriculum learning to data\-constrained multilingual pretraining\. We investigate the causal effects of pretraining on CS data by comparing against baselines trained on the same documents kept entirely in the original language\. Our research is organized along the following questions: 1. 1\.Does CS training improve downstream performance on the BabyLM evaluation suite? 2. 2\.Does CS training align the model’s internal representations across languages? 3. 3\.Does the model represent words seen embedded in CS contexts differently than words seen only in monolingual contexts? Our contributions, respectively, are that for models trained on CS data: 1. 1\.Performance on the BabyLM evaluation suite improves when trained under a curriculum but not under shuffled ordering \(§[5](https://arxiv.org/html/2609.30535#S5)\); 2. 2\.Representations of parallel text align, and under curriculum ordering, the alignment persists through training on monolingual documents \(§[6](https://arxiv.org/html/2609.30535#S6)\); 3. 3\.Embedded words align more closely to their translations than words that only appear in monolingual contexts \(§[7](https://arxiv.org/html/2609.30535#S7)\)\. We release our models222[https://huggingface\.co/drooryck/multilingual\-macaroni\-models](https://huggingface.co/drooryck/multilingual-macaroni-models), corpora333[https://huggingface\.co/datasets/drooryck/multilingual\-macaroni\-corpus](https://huggingface.co/datasets/drooryck/multilingual-macaroni-corpus), and code444[https://github\.com/drooryck/multilingual\-macaroni](https://github.com/drooryck/multilingual-macaroni)\. ## 2Related Work ### Developmentally plausible language modeling\. The first three iterations of the BabyLM challenge\([Warstadt et al\., 2023](https://arxiv.org/html/2609.30535#bib.bib1);[Hu et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib22);[Charpentier et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib23)\)established data\-constrained language acquisition as an interdisciplinary question of interest to machine learning researchers, linguists, and cognitive scientists\. Relevant past submissions include a negative finding for cognitively\-inspired curriculum learning\([Diehl Martinez et al\., 2023](https://arxiv.org/html/2609.30535#bib.bib24)\)and a benchmark for detection of ungrammaticality, including unnatural CS, in second\-language acquisition\([Gao et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib57)\)\.[Zeng et al\. \(2026\)](https://arxiv.org/html/2609.30535#bib.bib58)also study the effect of various synthetic data mixtures on bilingual BabyLMs\. Our work adopts the motivation, constraints, and evaluation suite of this year’s challenge\([Choshen et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib48)\), which introduces a multilingual track based on the BabyBabelLM corpora\([Jumelet et al\., 2026a](https://arxiv.org/html/2609.30535#bib.bib62)\)\. ### CS as data augmentation for multilingual training\. While attempts to model human CS predate the deep learning era\([Chan et al\., 2006](https://arxiv.org/html/2609.30535#bib.bib55);[Franco and Solorio, 2007](https://arxiv.org/html/2609.30535#bib.bib56)\), recent works generate CS as a data augmentation strategy for multilingual training\. Some works procedurally generate CS data using bilingual dictionaries\([Qin et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib6);[Zhu et al\., 2023](https://arxiv.org/html/2609.30535#bib.bib26);[Zheng et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib25);[Feng et al\., 2022](https://arxiv.org/html/2609.30535#bib.bib27)\)\. Others prompt LLMs, as we do, which exhibit more naturalistic patterns\([Kuwanto et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib28);[Wang et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib3)\)\. While the above works demonstrate benefits of CS training,[Shao et al\. \(2026\)](https://arxiv.org/html/2609.30535#bib.bib4)suggest that CS data is less important than parallel data for a model’s ability to translate \(but that, surprisingly, neither kind of multilingual data is required for other cross\-lingual understanding and reasoning tasks\)\. However, their category of CS documents does not match our definitions of word\- and sentence\-level CS, and ultimately both our works stress the importance of fine\-grained lexical alignment for translation\. Other work has applied CS for cross\-lingual transfer at inference time\([Yoo et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib66)\), in chain of thought\([Wang et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib59)\), or for instruction tuning\([Asano et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib29)\)\. Our work most closely follows code\-switching curriculum learning\([Yoo et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib2)\)\. Whereas code\-switching curriculum learning seeks to fine\-tune a pretrained model, our study focuses on pretraining in the data\-constrained regime and attempts a closer analysis of model representations\. ### Cross\-lingual representation alignment\. Recent papers characterize LLMs’ ability to generalize across languages despite being trained on mostly monolingual documents\. Interpretability studies find shared model components or subspaces across languages\([Conneau et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib14);[Dufter and Schütze, 2020](https://arxiv.org/html/2609.30535#bib.bib15);[Chang et al\., 2022](https://arxiv.org/html/2609.30535#bib.bib51)\)or investigate how cross\-lingual alignment emerges throughout the training process\([Blevins et al\., 2022](https://arxiv.org/html/2609.30535#bib.bib52);[Wang et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib53)\)\. Token overlap between languages tends to improve cross\-lingual generalization\([Pires et al\., 2019](https://arxiv.org/html/2609.30535#bib.bib8);[Kallini et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib61)\), as does typological similarity\([Muller et al\., 2023](https://arxiv.org/html/2609.30535#bib.bib63);[Longpre et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib64)\)\. See[Hämmerl et al\. \(2024\)](https://arxiv.org/html/2609.30535#bib.bib54)for a survey on cross\-lingual alignment\. Our work contributes a controlled study of how code\-switching curriculum learning affects cross\-lingual alignment\. ## 3Synthetic Code\-Switched Data Generation Table 2:Accuracy on the BabyLM evaluation suite \(mean±\\pmSD\) across eight seeds per condition\. We select five tasks: the two language\-averaged tasks with the largest positive gap between the curriculum CS and non\-CS models \(SIB\-200\+1\.94\+1\.94, INCLUDE\+1\.34\+1\.34\), the two with the most negative \(BMLAMA−0\.80\-0\.80, Global PIQA−0\.20\-0\.20\), and MultiBLiMP as a grammatical reference task\. Bold cells mark statistical significance on a two\-sided pairedtt\-test between the CS and non\-CS models \(uncorrected for multiple hypotheses\)\. Full per\-task results are in Table[3](https://arxiv.org/html/2609.30535#A3.T3)\(Appendix[C](https://arxiv.org/html/2609.30535#A3)\)\.We construct two distinct corpora that both satisfy BabyLM’s data budget: one with code\-switched \(CS\) data and the other without \(non\-CS\)\. Figure[1](https://arxiv.org/html/2609.30535#S1.F1)illustrates their composition and Figure[5](https://arxiv.org/html/2609.30535#A1.F5)\(Appendix[A](https://arxiv.org/html/2609.30535#A1)\) gives longer examples\. See Appendix[A](https://arxiv.org/html/2609.30535#A1)for detailed accounting\. We also create four additional corpora as experimental controls\. ### Monolingual\-document \(non\-CS\) corpus\. We begin with the English, Dutch, and Chinese BabyBabelLM datasets\([Jumelet et al\., 2026a](https://arxiv.org/html/2609.30535#bib.bib62)\)\. We sample a third of each and combine them to obtain our base non\-CS corpus of monolingual documents\. ### CS corpus\. We begin with the English third of the non\-CS corpus \(33\.4M words\) and separate it further into thirds: one kept monolingual, one for word\-level CS \(where English is the matrix language\), and one for sentence\-level CS \(where half of the sentences in each document are translated\)\. The monolingual third is kept as\-is\. The word\-level third is divided into one half for Dutch insertions and one half for Chinese insertions, where we use an instruction\-tuned LLM to translate a subset of words to the embedded language\. We similarly bisect the sentence\-level third and instruct the LLM to translate, in\-place, roughly half of the sentences in each document to the embedded language\. We then repeat this*mutatis mutandis*for the Dutch and Chinese subsets\. See Table[1](https://arxiv.org/html/2609.30535#S1.T1)for an exact breakdown of the final corpus composition\. We verify the quality of generated code\-switched text through programmatic checks and through manual verification of sample documents by bilingual English\-Dutch and English\-Chinese speakers\. See Appendix[A](https://arxiv.org/html/2609.30535#A1)for our quality verification pipeline and Appendix[B](https://arxiv.org/html/2609.30535#A2)for LLM choice and prompts\. ### Corpora for controls\. We construct four corpora to isolate the contribution of various experimental factors\. 1. 1\.Word\-level CS only: we revert the sentence\-level CS third of the CS corpus to the original monolingual documents\. 2. 2\.Sentence\-level CS only: we similarly revert the word\-level CS third\. 3. 3\.Word salad: for a third of the original non\-CS corpus, we replace roughly35%35\\%of tokens of three or more letters, selected randomly, with words sampled from the embedded language’s monolingual data\. For Chinese source text, we insert words instead of replacing\. 4. 4\.Document translation: we take a third of documents in the original non\-CS corpus and append to them an LLM\-generated translation\. We then discard a third of these documents to keep the word budget consistent\. See Appendix[A](https://arxiv.org/html/2609.30535#A1)for verbatim examples from the word salad and document translation controls\. ## 4Experimental Setup Figure 2:Layerwise bitext retrieval precision@1 \(§[6](https://arxiv.org/html/2609.30535#S6)\) on997997parallel sentences from FLORES\+ across English, Dutch, and Chinese\. We compare CS models against non\-CS models \(ignoring data ordering\) and plot the mean±1\\pm 1SD across the1616runs per group\. Retrieval in non\-CS models fails for language pairs that differ in script\. See Figure[6](https://arxiv.org/html/2609.30535#A4.F6)\(Appendix[D](https://arxiv.org/html/2609.30535#A4)\) for the median percentile rank of the translation out of the candidate set\. Shading marks layers66–1111, which the reported P@1 averages in other experiments are taken over\.Our main experiments vary the training corpus \(CS vs non\-CS\) and the data ordering \(shuffled vs a three\-stage curriculum\)\. Including our four control corpora yields88distinct training setups in total\. For each setup, we train88random seeds, which affect the model initialization and document order\. ### Architecture and tokenizer\. All models use the GPT\-2\-small architecture from the BabyLM multilingual baseline555[https://github\.com/babylm\-org/multilingual\-training](https://github.com/babylm-org/multilingual-training)\(12 layers, hidden size 768, 12 heads, context length 1024\)\. We train a byte\-level BPE tokenizer of vocabulary size16,38416,384, matching the baseline implementation, on the monolingual\-document corpus\. ### No padding\. The BabyBabelLM corpora contain many short documents, so to avoid excessive padding tokens, we preprocess by concatenating all documents \(separated by end\-of\-text tokens\) and splitting according to the model’s context length\. ### Optimizer\. Following the competition baselines, we train models using Adam\([Kingma and Ba, 2015](https://arxiv.org/html/2609.30535#bib.bib19)\)without weight decay, a batch size of1616, and a learning rate schedule that starts at00, warms up for1%1\\%of total training steps to5×10−55\\times 10^\{\-5\}, and cosine\-decays to zero\. ### Curriculum learning\. When training on the CS corpus, we compare between a learning curriculum and the default ordering \(1010epochs on the shuffled CS corpus\)\. The curriculum consists of training first for 10 epochs over the word\-level CS third of our CS corpus, then 10 epochs on the sentence\-level CS third, and finally 10 epochs on the non\-CS third, resetting the learning rate schedule in each stage\. As for the baselines trained on the non\-CS corpus, we also apply these two orderings and their respective learning rate schedules, maintaining the document IDs but using the original monolingual documents instead of the synthesized CS versions\. This results in four \(corpus, ordering\) pairs: CS/shuffled, CS/curriculum, non\-CS/shuffled, non\-CS/curriculum\. The models trained on our control corpora all follow shuffled ordering\. ### Evaluation\. Figure 3:Sentence\-level alignment across the learning curriculum\. Alignment improves throughout CS training and persists through training on monolingual documents\. The color of the CS model’s line indicates the CS type of the current stage\.We evaluate models on the BabyLM evaluation suite, which tests for grammatical fluency, cognitive plausibility, and model adaptability \(via fine\-tuning\)\. The full list of tasks is in Appendix[C](https://arxiv.org/html/2609.30535#A3)\. The English/Dutch/Chinese task suites differ slightly, so all reported averages are of the per\-language scores\. ## 5Results Table[2](https://arxiv.org/html/2609.30535#S3.T2)summarizes performance on a subset of evaluation tasks\. We report per\-task performance across the entire suite in Table[3](https://arxiv.org/html/2609.30535#A3.T3)\(Appendix[C](https://arxiv.org/html/2609.30535#A3)\) and also held\-out bits per byte in Table[4](https://arxiv.org/html/2609.30535#A3.T4)\(Appendix[C](https://arxiv.org/html/2609.30535#A3)\)\. ### Code\-switching under a curriculum scores highest\. Between our four corpus/ordering pairings, we see that CS/curriculum performs highest on average \(46\.72\), followed by CS/shuffled, then non\-CS/shuffled, and finally non\-CS/curriculum\. Each of our training setups outperforms, on average across seeds, the BabyLM baseline\([Choshen et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib48)\), which scores45\.9445\.94on the leaderboard666ModelBabyLM\-2026\-Baseline\-GPT2\-en\_nld\_zho\_equalon the leaderboard at[https://huggingface\.co/spaces/BabyLM\-community/BabyLM\-Leaderboard\-2026](https://huggingface.co/spaces/BabyLM-community/BabyLM-Leaderboard-2026)\. Accessed 17 September 2026\.We check for statistical significance with two\-sided pairedtt\-tests over all\(42\)=6\{4\\choose 2\}=6pairwise comparisons of the ordering/corpus pairs, applying Holm correction\([Holm, 1979](https://arxiv.org/html/2609.30535#bib.bib18)\)for the familywise error rate\. Only the CS/curriculum improvement over non\-CS/curriculum is significant \(p=0\.016p=0\.016\); the gap between CS/curriculum and non\-CS/shuffled is not \(p=0\.41p=0\.41\), nor is any other comparison\. ### Fine\-tuning tasks under the curriculum\. For the models trained with a curriculum, most of the CS models’ advantage over the non\-CS models comes from the fine\-tuning tasks\. This suggests that the internal representations of CS models may be more generalizable toward various downstream tasks than those of models trained without CS\. We investigate further in §[6](https://arxiv.org/html/2609.30535#S6)\. However, this boost does not seem to hold under shuffled ordering, where the CS and non\-CS models’ scores are within noise of each other for both zero\-shot and fine\-tuning averages\. ## 6Cross\-Lingual Representation Alignment Does training on code\-switched data cause a model to align the representations of parallel sentences? We measure this by seeing whether a sentence and its translation have closer representations in CS models than in non\-CS models\. ### Cross\-lingual alignment metric\. Formally, we measure a model’s cross\-lingual representation alignment usingbitext retrieval precision@1\(P@1\)\([Artetxe and Schwenk, 2019](https://arxiv.org/html/2609.30535#bib.bib7);[Pires et al\., 2019](https://arxiv.org/html/2609.30535#bib.bib8);[Hu et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib10)\)\. Given a sentencesAs^\{A\}in languageAA, and a list of sentencess1B,…,sNBs^\{B\}\_\{1\},\\dots,s^\{B\}\_\{N\}in languageBBcontaining the translation ofsAs^\{A\}, we say the model successfully retrieves the translation if it is the original sentence’s nearest neighbor in terms of the cosine similarity between their representations\. We represent a sentence by mean\-pooling the residual stream vectors of its tokens\. We additionally apply cross\-domain similarity local scaling\([Lample et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib16), CSLS,\)to the cosine similarity values to account for the “hubness” issue of naive nearest neighbors\. Given a set of sentence pairs\(s1A,s1B\),…,\(sNA,sNB\)\(s^\{A\}\_\{1\},s^\{B\}\_\{1\}\),\\dots,\(s^\{A\}\_\{N\},s^\{B\}\_\{N\}\), the bitext retrieval precision is the average success rate across all sentences in both directions\. We useN=997N\{=\}997parallel sentences across English, Dutch, and Chinese from the FLORES\+devsplit\([NLLB Team et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib20);[Goyal et al\., 2022](https://arxiv.org/html/2609.30535#bib.bib21)\)\. See Appendix[D](https://arxiv.org/html/2609.30535#A4)for details\. ### Code\-switched training aligns parallel sentences\. Across all layers and language pairs, the CS model better aligns the sentence representations to their translations’ representations \(Figure[2](https://arxiv.org/html/2609.30535#S4.F2)\): between English and Chinese under curriculum ordering, the CS model retrieves the correct translation61\.4%61\.4\\%of the time \(layers66–1111,88\-seed mean\), compared to only5\.5%5\.5\\%of the time in the non\-CS baseline \(Table[6](https://arxiv.org/html/2609.30535#A4.T6), Appendix[D](https://arxiv.org/html/2609.30535#A4)\)\. A similar gap holds for bitext retrieval between Dutch and Chinese\. Other training\-free measures, including linear centered kernel alignment \(CKA\), a sentence\-level cosine gap, centroid distance, and the MEXA alignment score all corroborate this finding for the cross\-script language pairs \(Table[5](https://arxiv.org/html/2609.30535#A4.T5), Appendix[D](https://arxiv.org/html/2609.30535#A4)\)\. ### Coherence of CS matters\. Must code\-switching be coherent in order to improve alignment, or is the gain driven simply by co\-occurrence across tokens? Our “word salad” control, where words are incoherently switched for foreign\-language words, does not induce this same alignment \(Figure[7](https://arxiv.org/html/2609.30535#A4.F7), Appendix[D](https://arxiv.org/html/2609.30535#A4)\), suggesting that grammatically and semantically coherent code\-switching is crucial\. ### Cross\-lingual alignment persists through non\-CS training\. We measure bitext retrieval precision across model checkpoints \(Figure[3](https://arxiv.org/html/2609.30535#S4.F3)\)\. Cross\-lingual alignment improves during the word\-level and sentence\-level CS stages and is maintained even through the non\-CS training stage\. ## 7Representations of Code\-Switched Words Does CS training improve the representations of only the words seen in embedded contexts, or does the cross\-lingual alignment extend across the model’s entire vocabulary? ### Embedded vs\. never\-embedded words\. We call a word*embedded*if it was ever inserted as a foreign\-language word during word\-level CS, and*never\-embedded*if it was neither embedded nor replaced by an embedded word\. We exclude replaced words since they have systematically lower frequency in the CS corpus due to being replaced\. We compute a wordww’s representation as follows\. We first collect a set of sentence contexts from the original corpus that containww\. We then compute a forward pass and take the mean of the residual stream vectors across the token positions that constituteww, across the set of sentences, and across layers66–1111\. ### Median percentile rank metric\. Similarly to the sentence\-level bitext retrieval metric \(§[6](https://arxiv.org/html/2609.30535#S6)\), for each wordwAw^\{A\}in language A, we construct a set of wordsw1B,…wNBw^\{B\}\_\{1\},\\ldots w^\{B\}\_\{N\}in language B containing the translationw∗Bw^\{B\}\_\{\*\}ofwAw^\{A\}\. We then rank the set according to the same CSLS\-adjusted cosine similarity metric, with regards to the word representations described above, and record thepercentile rankofw∗Bw^\{B\}\_\{\*\}\. \(This contrasts with the precision@1 metric where we would only check ifw∗Bw^\{B\}\_\{\*\}is the most similar word towAw^\{A\}\.\) We compute these percentile ranks acrossN=1,102N=1\{,\}102embedded/non\-embedded word pairs and, for each of the44language directions for which we have an openly accessible bilingual dictionary, plot the median in Figure[4](https://arxiv.org/html/2609.30535#S7.F4)\. See Appendix[E](https://arxiv.org/html/2609.30535#A5)for details on choosing the candidate sets and labeling translation pairs\. Figure 4:Median percentile rank of a word’s correct translation among a set of target\-language words matched by part of speech and frequency \(Appendix[E](https://arxiv.org/html/2609.30535#A5)\)\. The number of word pairs for each language pair is shown in parentheses\. ### Never\-embedded words also become aligned by CS training\. Relative to the non\-CS baseline, CS training raises the median percentile rank of the translation by\+0\.064\+0\.064for embedded words and\+0\.045\+0\.045for never\-embedded words \(Figure[4](https://arxiv.org/html/2609.30535#S7.F4)\)\. The English–Dutch translations are already closely aligned in both the CS and non\-CS models\. Here in word\-level alignment, as in sentence\-level alignment \(§[6](https://arxiv.org/html/2609.30535#S6)\), we see that CS training aligns representations especially across distinct scripts\. ## 8Conclusion We train and analyze small autoregressive language models on code\-switched curriculum learning under BabyLM competition constraints \(10 epochs on 100M words shared between Dutch, English, and Chinese\)\. When trained in a three\-stage learning curriculum, from word\-level code\-switching to sentence\-level switching to monolingual documents, a model trained on code\-switched data achieves a0\.360\.36\-point gain on the BabyLM evaluation suite, on average across eight seeds, relative to a baseline trained on a matched corpus of monolingual documents \(RQ1, §[5](https://arxiv.org/html/2609.30535#S5)\)\. However, under a randomly shuffled data ordering, there is no statistically significant improvement\. We show that code\-switching in pretraining aligns the representations of parallel sentences \(RQ2, §[6](https://arxiv.org/html/2609.30535#S6)\) and additionally show that embedded words become more closely aligned to their translations than words never seen embedded in a code\-switched context \(RQ3, §[7](https://arxiv.org/html/2609.30535#S7)\)\. We release our corpora and training code to support future research on multilingual data composition and data ordering at a fixed data budget\. ## Limitations ### Language coverage\. We study only trilingual models in English, Dutch, and Chinese so as to match the BabyLM challenge evaluation suite\. Whether the curriculum and the alignment mechanism generalize to more distant or lower\-resource pairs, or to more than three languages at once, is untested\. ### LLM\-generated code\-switching\. Our LLM\-driven synthetic CS corpus generation method would be expensive to scale to larger datasets\. It also results in a confound whereby the sentence\-level CS data and the document translation control corpus risk having a different text distribution than BabyBabelLM, which might thus be responsible for performance improvements rather than the code\-switching intervention itself\. ### Model and budget scope\. All of our results are for a single model size and architecture \(∼98\{\\sim\}98M\-parameter GPT\-2 at the100100M\-word budget\)\. Future work could investigate whether the benefits of training on code\-switched text persist in larger models and data budgets\. ## Ethics Statement ### Data provenance and released artifacts\. All training text derives from the BabyBabelLM corpora\([Jumelet et al\., 2026a](https://arxiv.org/html/2609.30535#bib.bib62)\), distributed under the BabyLM community’s access terms\. We did not collect data from human subjects; manual review of generated samples was done by the joint first authors\. Our models are∼98\{\\sim\}98M\-parameter research artifacts with no instruction tuning or safety alignment\. We release them and our corpora under the MIT license\. ### Compute\. The experiments in this paper required training6464models, each on a single NVIDIA H100 80GB GPU\. Pretraining one model takes∼2\.3\{\\sim\}2\.3GPU\-hours and evaluating a model takes an additional∼0\.8\{\\sim\}0\.8\. This adds up to roughly200200GPU\-hours in total, excluding the LLM inference used to synthesize the corpora, the checkpoint analyses of §[6](https://arxiv.org/html/2609.30535#S6)–§[7](https://arxiv.org/html/2609.30535#S7), and re\-run jobs\. ## Acknowledgments We thank Nihal Nayak, Sara Kangaslahti, Greta Tuckute, and the Harvard ML Foundations Group for conversations regarding multilingual models\. AC is grateful to be funded by the Kempner Graduate Fellowship\. Computation was performed on the Kempner Institute’s H100 cluster at Harvard University\. We wrote most of our analysis and plotting code with help from Anthropic’s Claude LLM\. ## References - Adelaniet al\.\(2024\)D\. I\. Adelani, H\. Liu, X\. Shen, N\. Vassilyev, J\. O\. Alabi, Y\. Mao, H\. Gao, and E\. A\. LeeSIB\-200: a simple, inclusive, and big evaluation dataset for topic classification in 200\+ languages and dialects\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),St\. Julian’s, Malta,pp\. 226–245\.External Links:[Link](https://aclanthology.org/2024.eacl-long.14/),[Document](https://dx.doi.org/10.18653/v1/2024.eacl-long.14)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Arnettet al\.\(2024\)C\. Arnett, T\. A\. Chang, and B\. BergenA bit of a problem: measurement disparities in dataset sizes across languages\.InProceedings of the 3rd Annual Meeting of the Special Interest Group on Under\-resourced Languages @ LREC\-COLING 2024,Torino, Italia,pp\. 1–9\.External Links:[Link](https://aclanthology.org/2024.sigul-1.1/)Cited by:[Appendix A](https://arxiv.org/html/2609.30535#A1.SS0.SSS0.Px1.p1.1)\. - Artetxe and Schwenk \(2019\)M\. Artetxe and H\. SchwenkMassively multilingual sentence embeddings for zero\-shot cross\-lingual transfer and beyond\.Transactions of the Association for Computational Linguistics7,pp\. 597–610\.External Links:[Link](https://aclanthology.org/Q19-1038/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00288)Cited by:[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Asanoet al\.\(2026\)S\. Asano, J\. Baek, and T\. YamasakiBeyond bilingual transfer: multilingual code\-switching in instruction tuning\.arXiv preprint arXiv:2605\.29414\.Note:Preprint; no peer\-reviewed version as of September 2026External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.29414),[Link](https://arxiv.org/abs/2605.29414)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Bandarkaret al\.\(2024\)L\. Bandarkar, D\. Liang, B\. Muller, M\. Artetxe, S\. N\. Shukla, D\. Husa, N\. Goyal, A\. Krishnan, L\. Zettlemoyer, and M\. KhabsaThe Belebele benchmark: a parallel reading comprehension dataset in 122 language variants\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 749–775\.External Links:[Link](https://aclanthology.org/2024.acl-long.44/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.44)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Blevinset al\.\(2022\)T\. Blevins, H\. Gonen, and L\. ZettlemoyerAnalyzing the mono\- and cross\-lingual pretraining dynamics of multilingual language models\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 3575–3590\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.234/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.234)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Chanet al\.\(2006\)J\. Y\. C\. Chan, P\. C\. Ching, T\. Lee, and H\. CaoAutomatic speech recognition of Cantonese\-English code\-mixing utterances\.InProceedings of Interspeech 2006 – ICSLP,Pittsburgh, PA, USA\.External Links:[Link](https://www.isca-archive.org/interspeech_2006/chan06_interspeech.html),[Document](https://dx.doi.org/10.21437/Interspeech.2006-29)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Changet al\.\(2025\)T\. A\. Chang C\. Arnettet al\.Global PIQA: evaluating commonsense reasoning across 100\+ languages and cultures\.arXiv preprint arXiv:2510\.24081\.External Links:[Link](https://arxiv.org/abs/2510.24081)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Changet al\.\(2022\)T\. A\. Chang, Z\. Tu, and B\. K\. BergenThe geometry of multilingual language model representations\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 119–136\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.9/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.9)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Charpentieret al\.\(2025\)L\. Charpentier, L\. Choshen, R\. Cotterell, M\. O\. Gül, M\. Y\. Hu, J\. Liu, J\. Jumelet, T\. Linzen, A\. Mueller, C\. Ross, R\. S\. Shah, A\. Warstadt, E\. G\. Wilcox, and A\. WilliamsFindings of the third BabyLM challenge: accelerating language modeling research with cognitively plausible data\.InProceedings of the First BabyLM Workshop,Suzhou, China,pp\. 399–420\.External Links:[Link](https://aclanthology.org/2025.babylm-main.28/),[Document](https://dx.doi.org/10.18653/v1/2025.babylm-main.28)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Choshenet al\.\(2026\)L\. Choshen, R\. Cotterell, M\. O\. Gül, J\. Jumelet, T\. Linzen, A\. Mueller, S\. Salhan, R\. S\. Shah, A\. Warstadt, and E\. G\. WilcoxBabyLM turns 4 and goes multilingual: call for papers for the 2026 BabyLM workshop\.arXiv preprint arXiv:2602\.20092\.External Links:[Link](https://arxiv.org/abs/2602.20092)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.p1.1),[§1](https://arxiv.org/html/2609.30535#S1.p1.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.30535#S5.SS0.SSS0.Px1.p1.1)\. - Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? try ARC, the AI2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.External Links:[Link](https://arxiv.org/abs/1803.05457)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Conneauet al\.\(2018\)A\. Conneau, R\. Rinott, G\. Lample, A\. Williams, S\. R\. Bowman, H\. Schwenk, and V\. StoyanovXNLI: evaluating cross\-lingual sentence representations\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,Brussels, Belgium,pp\. 2475–2485\.External Links:[Link](https://aclanthology.org/D18-1269/),[Document](https://dx.doi.org/10.18653/v1/D18-1269)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Conneauet al\.\(2020\)A\. Conneau, S\. Wu, H\. Li, L\. Zettlemoyer, and V\. StoyanovEmerging cross\-lingual structure in pretrained language models\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,Online,pp\. 6022–6034\.External Links:[Link](https://aclanthology.org/2020.acl-main.536/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.536)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - de Marneffeet al\.\(2021\)M\. de Marneffe, C\. D\. Manning, J\. Nivre, and D\. ZemanUniversal Dependencies\.Computational Linguistics47\(2\),pp\. 255–308\.External Links:[Link](https://aclanthology.org/2021.cl-2.11/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00402)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.External Links:[Link](https://arxiv.org/abs/2606.19348)Cited by:[Appendix B](https://arxiv.org/html/2609.30535#A2.p1.1)\. - Diehl Martinezet al\.\(2023\)R\. Diehl Martinez, Z\. Goriely, H\. McGovern, C\. Davis, A\. Caines, P\. Buttery, and L\. BeinbornCLIMB – curriculum learning for infant\-inspired model building\.InProceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning,Singapore,pp\. 112–127\.External Links:[Link](https://aclanthology.org/2023.conll-babylm.10/),[Document](https://dx.doi.org/10.18653/v1/2023.conll-babylm.10)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Dufter and Schütze \(2020\)P\. Dufter and H\. SchützeIdentifying elements essential for BERT’s multilinguality\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Online,pp\. 4423–4437\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.358/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.358)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Fenget al\.\(2022\)Y\. Feng, F\. Li, and P\. KoehnToward the limitation of code\-switching in cross\-lingual transfer\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 5966–5971\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.400/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.400)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Franco and Solorio \(2007\)J\. C\. Franco and T\. SolorioBaby\-steps towards building a Spanglish language model\.InComputational Linguistics and Intelligent Text Processing: 8th International Conference, CICLing 2007,Lecture Notes in Computer Science, Vol\.4394,Berlin, Heidelberg,pp\. 75–84\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-70939-8%5F7)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Gaoet al\.\(2025\)Y\. Gao, S\. Salhan, A\. Caines, P\. Buttery, and W\. SunBLiSS: evaluating bilingual learner competence in second language small language models\.InProceedings of the First BabyLM Workshop,Suzhou, China,pp\. 160–174\.External Links:[Link](https://aclanthology.org/2025.babylm-main.13/),[Document](https://dx.doi.org/10.18653/v1/2025.babylm-main.13)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Goyalet al\.\(2022\)N\. Goyal, C\. Gao, V\. Chaudhary, P\. Chen, G\. Wenzek, D\. Ju, S\. Krishnan, M\. Ranzato, F\. Guzmán, and A\. FanThe FLORES\-101 evaluation benchmark for low\-resource and multilingual machine translation\.Transactions of the Association for Computational Linguistics10,pp\. 522–538\.External Links:[Link](https://aclanthology.org/2022.tacl-1.30/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00474)Cited by:[Appendix D](https://arxiv.org/html/2609.30535#A4.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Hämmerlet al\.\(2024\)K\. Hämmerl, J\. Libovický, and A\. FraserUnderstanding cross\-lingual Alignment—A survey\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 10922–10943\.External Links:[Link](https://aclanthology.org/2024.findings-acl.649/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.649)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Heet al\.\(2025\)L\. He, E\. Nie, S\. S\. Dindar, A\. Firoozi, A\. Florea, V\. Nguyen, C\. Puffay, R\. Shimizu, H\. Ye, J\. Brennan, H\. Schmid, H\. Schütze, and N\. MesgaraniXCOMPS: a multilingual benchmark of conceptual minimal pairs\.InProceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP,Vienna, Austria,pp\. 75–81\.External Links:[Link](https://aclanthology.org/2025.sigtyp-1.9/),[Document](https://dx.doi.org/10.18653/v1/2025.sigtyp-1.9)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Holm \(1979\)S\. HolmA simple sequentially rejective multiple test procedure\.Scandinavian Journal of Statistics6\(2\),pp\. 65–70\.External Links:[Link](https://www.jstor.org/stable/4615733)Cited by:[§5](https://arxiv.org/html/2609.30535#S5.SS0.SSS0.Px1.p1.1)\. - Huet al\.\(2020\)J\. Hu, S\. Ruder, A\. Siddhant, G\. Neubig, O\. Firat, and M\. JohnsonXTREME: a massively multilingual multi\-task benchmark for evaluating cross\-lingual generalisation\.InProceedings of the 37th International Conference on Machine Learning,Vol\.119,pp\. 4411–4421\.External Links:[Link](https://proceedings.mlr.press/v119/hu20b.html)Cited by:[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Huet al\.\(2024\)M\. Y\. Hu, A\. Mueller, C\. Ross, A\. Williams, T\. Linzen, C\. Zhuang, R\. Cotterell, L\. Choshen, A\. Warstadt, and E\. G\. WilcoxFindings of the second BabyLM challenge: sample\-efficient pretraining on developmentally plausible corpora\.InThe 2nd BabyLM Challenge at the 28th Conference on Computational Natural Language Learning,Miami, FL, USA,pp\. 1–21\.External Links:[Link](https://aclanthology.org/2024.conll-babylm.1/)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Jumeletet al\.\(2026a\)J\. Jumelet, A\. Fourtassi, A\. Haga, B\. Bunzeck, B\. Shandilya, D\. Galvan\-Sosa, F\. G\. Haznitrama, F\. Padovani, F\. Meyer, H\. Hu, J\. Etxaniz, L\. Prevot, L\. He, M\. Grandury, M\. Marcheva, N\. Foroutan, N\. Theodoropoulos, P\. Sadeghi, S\. Song, S\. Salhan, S\. Zhou, Y\. Paniv, Z\. Zhang, A\. Bisazza, A\. Warstadt, and L\. ChoshenBabyBabelLM: a multilingual benchmark of developmentally plausible training data\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),Rabat, Morocco,pp\. 3297–3329\.External Links:[Link](https://aclanthology.org/2026.eacl-long.152/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.152)Cited by:[Appendix A](https://arxiv.org/html/2609.30535#A1.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.30535#S3.SS0.SSS0.Px1.p1.1),[Data provenance and released artifacts\.](https://arxiv.org/html/2609.30535#Sx2.SS0.SSS0.Px1.p1.1)\. - Jumeletet al\.\(2026b\)J\. Jumelet, L\. Weissweiler, J\. Nivre, and A\. BisazzaMultiBLiMP 1\.0: a massively multilingual benchmark of linguistic minimal pairs\.Transactions of the Association for Computational Linguistics14,pp\. 193–216\.External Links:[Link](https://aclanthology.org/2026.tacl-1.10/),[Document](https://dx.doi.org/10.1162/tacl.a.600)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Kalliniet al\.\(2025\)J\. Kallini, D\. Jurafsky, C\. Potts, and M\. BarteldsFalse friends are not foes: investigating vocabulary overlap in multilingual language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 21138–21154\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1153/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1153)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Kargaranet al\.\(2025\)A\. H\. Kargaran, A\. Modarressi, N\. Nikeghbal, J\. Diesner, F\. Yvon, and H\. SchützeMEXA: multilingual evaluation of English\-centric LLMs via cross\-lingual alignment\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 27001–27023\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1385/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1385)Cited by:[Table 5](https://arxiv.org/html/2609.30535#A4.T5)\. - Kingma and Ba \(2015\)D\. P\. Kingma and J\. BaAdam: a method for stochastic optimization\.InProceedings of the 3rd International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/1412.6980)Cited by:[§4](https://arxiv.org/html/2609.30535#S4.SS0.SSS0.Px3.p1.1)\. - Kornblithet al\.\(2019\)S\. Kornblith, M\. Norouzi, H\. Lee, and G\. HintonSimilarity of neural network representations revisited\.InProceedings of the 36th International Conference on Machine Learning,Vol\.97,pp\. 3519–3529\.External Links:[Link](https://proceedings.mlr.press/v97/kornblith19a.html)Cited by:[Table 5](https://arxiv.org/html/2609.30535#A4.T5)\. - Kuwantoet al\.\(2026\)G\. Kuwanto, C\. Agarwal, G\. I\. Winata, and D\. T\. WijayaLinguistics theory meets LLM: code\-switched text generation via equivalence constrained large language models\.InProceedings of the 1st Workshop on Computational Developmental Linguistics \(CDL\),San Diego, CA, USA,pp\. 1–14\.External Links:[Link](https://aclanthology.org/2026.cdl-1.1/),[Document](https://dx.doi.org/10.18653/v1/2026.cdl-1.1)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Lampleet al\.\(2018\)G\. Lample, A\. Conneau, M\. Ranzato, L\. Denoyer, and H\. JégouWord translation without parallel data\.InProceedings of the 6th International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=H196sainb)Cited by:[Appendix D](https://arxiv.org/html/2609.30535#A4.SS0.SSS0.Px1.p1.1),[Appendix E](https://arxiv.org/html/2609.30535#A5.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Liet al\.\(2024\)J\. Li, S\. Huang, A\. Ching, X\. Dai, and J\. ChenPreAlign: boosting cross\-lingual transfer by early establishment of multilingual alignment\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 10246–10257\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.572/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.572)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1)\. - Libovickýet al\.\(2020\)J\. Libovický, R\. Rosa, and A\. FraserOn the language neutrality of pre\-trained multilingual representations\.InFindings of the Association for Computational Linguistics: EMNLP 2020,Online,pp\. 1663–1674\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.150/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.150)Cited by:[Table 5](https://arxiv.org/html/2609.30535#A4.T5)\. - Linet al\.\(2022a\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Dublin, Ireland,pp\. 3214–3252\.External Links:[Link](https://aclanthology.org/2022.acl-long.229/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.229)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Linet al\.\(2022b\)X\. V\. Lin, T\. Mihaylov, M\. Artetxe, T\. Wang, S\. Chen, D\. Simig, M\. Ott, N\. Goyal, S\. Bhosale, J\. Du, R\. Pasunuru, S\. Shleifer, P\. S\. Koura, V\. Chaudhary, B\. O’Horo, J\. Wang, L\. Zettlemoyer, Z\. Kozareva, M\. Diab, V\. Stoyanov, and X\. LiFew\-shot learning with multilingual generative language models\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 9019–9052\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.616/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.616)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Liuet al\.\(2026\)Y\. Liu, Y\. Shen, H\. Zhu, L\. Xu, Z\. Qian, S\. Song, K\. Zhang, J\. Tang, P\. Zhang, B\. Yang, R\. Wang, and H\. HuA systematic assessment of language models with linguistic minimal pairs in Chinese\.Transactions of the Association for Computational Linguistics14,pp\. 755–771\.External Links:[Link](https://aclanthology.org/2026.tacl-1.34/),[Document](https://dx.doi.org/10.1162/tacl.a.648)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Longpreet al\.\(2026\)S\. Longpre, S\. Kudugunta, N\. Muennighoff, I\. Hsu, I\. Caswell, A\. Pentland, S\. Arik, C\. Lee, and S\. EbrahimiATLAS: adaptive transfer scaling laws for multilingual pretraining, finetuning, and decoding the curse of multilinguality\.InProceedings of the 14th International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0BkvUY61MX)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - MacWhinney \(2000\)B\. MacWhinneyThe CHILDES project: tools for analyzing talk\.3rd edition,Lawrence Erlbaum Associates,Mahwah, NJ\.Cited by:[Appendix A](https://arxiv.org/html/2609.30535#A1.SS0.SSS0.Px1.p1.1)\. - Mulleret al\.\(2023\)B\. Muller, D\. Gupta, J\. Fauconnier, S\. Patwardhan, D\. Vandyke, and S\. AgarwalLanguages you know influence those you learn: impact of language characteristics on multi\-lingual text\-to\-text transfer\.InProceedings of The 1st Transfer Learning for Natural Language Processing Workshop,Proceedings of Machine Learning Research, Vol\.203,pp\. 88–102\.External Links:[Link](https://proceedings.mlr.press/v203/muller23a.html)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - NLLB Teamet al\.\(2024\)NLLB Team, M\. R\. Costa\-jussà, J\. Cross, O\. Çelebi, M\. Elbayad, K\. Heafield, K\. Heffernan, E\. Kalbassi, J\. Lam, D\. Licht, J\. Maillard, A\. Sun, S\. Wang, G\. Wenzek, A\. Youngblood, B\. Akula, L\. Barrault, G\. Mejia Gonzalez, P\. Hansanti, J\. Hoffman, S\. Jarrett, K\. R\. Sadagopan, D\. Rowe, S\. Spruit, C\. Tran, P\. Andrews, N\. F\. Ayan, S\. Bhosale, S\. Edunov, A\. Fan, C\. Gao, V\. Goswami, F\. Guzmán, P\. Koehn, A\. Mourachko, C\. Ropers, S\. Saleem, H\. Schwenk, and J\. WangScaling neural machine translation to 200 languages\.Nature630\(8018\),pp\. 841–846\.External Links:[Link](https://www.nature.com/articles/s41586-024-07335-x),[Document](https://dx.doi.org/10.1038/s41586-024-07335-x)Cited by:[Appendix D](https://arxiv.org/html/2609.30535#A4.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Pireset al\.\(2019\)T\. Pires, E\. Schlinger, and D\. GarretteHow multilingual is multilingual BERT?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 4996–5001\.External Links:[Link](https://aclanthology.org/P19-1493/),[Document](https://dx.doi.org/10.18653/v1/P19-1493)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.30535#S6.SS0.SSS0.Px1.p1.1)\. - Qiet al\.\(2023\)J\. Qi, R\. Fernández, and A\. BisazzaCross\-lingual consistency of factual knowledge in multilingual language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 10650–10666\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.658/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.658)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Qinet al\.\(2020\)L\. Qin, M\. Ni, Y\. Zhang, and W\. CheCoSDA\-ML: multi\-lingual code\-switching data augmentation for zero\-shot cross\-lingual NLP\.InProceedings of the Twenty\-Ninth International Joint Conference on Artificial Intelligence, IJCAI\-20,pp\. 3853–3860\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2020/533)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Radovanovićet al\.\(2010\)M\. Radovanović, A\. Nanopoulos, and M\. IvanovićHubs in space: popular nearest neighbors in high\-dimensional data\.Journal of Machine Learning Research11\(86\),pp\. 2487–2531\.External Links:[Link](https://www.jmlr.org/papers/v11/radovanovic10a.html)Cited by:[Appendix D](https://arxiv.org/html/2609.30535#A4.SS0.SSS0.Px1.p1.1)\. - Romanouet al\.\(2025\)A\. Romanou, N\. Foroutan, A\. Sotnikova, Z\. Chen, S\. H\. Nelaturu, S\. Singh, R\. Maheshwary,et al\.INCLUDE: evaluating multilingual language understanding with regional knowledge\.InProceedings of the 13th International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=k3gCieTXeY)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Sakaguchiet al\.\(2020\)K\. Sakaguchi, R\. Le Bras, C\. Bhagavatula, and Y\. ChoiWinoGrande: an adversarial Winograd schema challenge at scale\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 8732–8740\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/6399),[Document](https://dx.doi.org/10.1609/aaai.v34i05.6399)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Shaoet al\.\(2026\)J\. Shao, R\. Tang, C\. Zhang, K\. Sevegnani, P\. Stenetorp, J\. Yang, and Y\. LuThe role of mixed\-language documents for multilingual large language model pretraining\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 36807–36818\.External Links:[Link](https://aclanthology.org/2026.acl-long.1706/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1706)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Suijkerbuijket al\.\(2025\)M\. Suijkerbuijk, Z\. Prins, M\. d\. H\. Kloots, W\. Zuidema, and S\. L\. FrankBLiMP\-NL: a corpus of Dutch minimal pairs and acceptability judgments for language model evaluation\.Computational Linguistics51\(4\),pp\. 1267–1301\.External Links:[Link](https://aclanthology.org/2025.cl-4.6/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00559)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Volterra and Taeschner \(1978\)V\. Volterra and T\. TaeschnerThe acquisition and development of language by bilingual children\.Journal of Child Language5\(2\),pp\. 311–326\.External Links:[Document](https://dx.doi.org/10.1017/S0305000900007492)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1)\. - Wanget al\.\(2024\)H\. Wang, P\. Minervini, and E\. PontiProbing the emergence of cross\-lingual alignment during LLM training\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 12159–12173\.External Links:[Link](https://aclanthology.org/2024.findings-acl.724/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.724)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px3.p1.1)\. - Wanget al\.\(2025\)Z\. Wang, J\. Li, H\. Zhou, R\. Weng, J\. Wang, X\. Huang, X\. Han, J\. Feng, C\. Deng, and S\. HuangInvestigating and scaling up code\-switching for multilingual language model pre\-training\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 11032–11046\.External Links:[Link](https://aclanthology.org/2025.findings-acl.575/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.575)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Wanget al\.\(2026\)Z\. Wang, J\. Liu, H\. Zhou, H\. Wei, B\. Yang, and S\. HuangEfficient multilingual reasoning transfer via progressive code\-switching\.arXiv preprint arXiv:2607\.00485\.External Links:[Link](https://arxiv.org/abs/2607.00485)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Warstadtet al\.\(2023\)A\. Warstadt, A\. Mueller, L\. Choshen, E\. Wilcox, C\. Zhuang, J\. Ciro, R\. Mosquera, B\. Paranjabe, A\. Williams, T\. Linzen, and R\. CotterellFindings of the BabyLM challenge: sample\-efficient pretraining on developmentally plausible corpora\.InProceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning,Singapore,pp\. 1–34\.External Links:[Link](https://aclanthology.org/2023.conll-babylm.1/),[Document](https://dx.doi.org/10.18653/v1/2023.conll-babylm.1)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Warstadtet al\.\(2020\)A\. Warstadt, A\. Parrish, H\. Liu, A\. Mohananey, W\. Peng, S\. Wang, and S\. R\. BowmanBLiMP: the benchmark of linguistic minimal pairs for English\.Transactions of the Association for Computational Linguistics8,pp\. 377–392\.External Links:[Link](https://aclanthology.org/2020.tacl-1.25/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00321)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Williamset al\.\(2018\)A\. Williams, N\. Nangia, and S\. R\. BowmanA broad\-coverage challenge corpus for sentence understanding through inference\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\),New Orleans, Louisiana,pp\. 1112–1122\.External Links:[Link](https://aclanthology.org/N18-1101/),[Document](https://dx.doi.org/10.18653/v1/N18-1101)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px2.p1.1)\. - Yanget al\.\(2020\)J\. Yang, S\. Ma, D\. Zhang, S\. Wu, Z\. Li, and M\. ZhouAlternating language modeling for cross\-lingual pre\-training\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 9386–9393\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/6480),[Document](https://dx.doi.org/10.1609/aaai.v34i05.6480)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1)\. - Yooet al\.\(2026\)H\. Yoo, J\. Jin, K\. Cho, and A\. OhGradual code\-switching as inference\-time cross\-lingual representational alignment for LLMs\.InThird Conference on Language Modeling,External Links:[Link](https://arxiv.org/abs/2510.05678)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Yooet al\.\(2025\)H\. Yoo, C\. Park, S\. Yun, A\. Oh, and H\. LeeCode\-switching curriculum learning for multilingual transfer in LLMs\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 7816–7836\.External Links:[Link](https://aclanthology.org/2025.findings-acl.407/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.407)Cited by:[Appendix B](https://arxiv.org/html/2609.30535#A2.p1.1),[§1](https://arxiv.org/html/2609.30535#S1.p2.1),[§1](https://arxiv.org/html/2609.30535#S1.p3.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1),[footnote 1](https://arxiv.org/html/2609.30535#footnote1)\. - Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, Y\. Bisk, A\. Farhadi, and Y\. ChoiHellaSwag: can a machine really finish your sentence?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,Florence, Italy,pp\. 4791–4800\.External Links:[Link](https://aclanthology.org/P19-1472/),[Document](https://dx.doi.org/10.18653/v1/P19-1472)Cited by:[Appendix C](https://arxiv.org/html/2609.30535#A3.SS0.SSS0.Px1.p1.1)\. - Zenget al\.\(2026\)L\. Zeng, S\. Y\. Feng, and M\. C\. FrankBringing up a bilingual BabyLM: investigating multilingual language acquisition using small\-scale models\.arXiv preprint arXiv:2603\.29552\.External Links:[Link](https://arxiv.org/abs/2603.29552)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px1.p1.1)\. - Zhenget al\.\(2024\)J\. Zheng, F\. Fan, and J\. LiIncorporating lexical and syntactic knowledge for unsupervised cross\-lingual transfer\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),Torino, Italia,pp\. 8986–8997\.External Links:[Link](https://aclanthology.org/2024.lrec-main.787/)Cited by:[§1](https://arxiv.org/html/2609.30535#S1.p2.1),[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. - Zhuet al\.\(2023\)Z\. Zhu, X\. Cheng, Z\. Huang, D\. Chen, and Y\. ZouEnhancing code\-switching for cross\-lingual SLU: a unified view of semantic and grammatical coherence\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 7849–7856\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.486/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.486)Cited by:[§2](https://arxiv.org/html/2609.30535#S2.SS0.SSS0.Px2.p1.1)\. ## Appendix AData Generation Details ### Sources, budget, units\. We download the BabyBabelLM\([Jumelet et al\., 2026a](https://arxiv.org/html/2609.30535#bib.bib62)\)datasets from HuggingFace \(IDsBabyLM\-community/babylm\-\{eng,nld,zho\}\) at revisions \(b78a9336,1aa063f7,600a6657\) respectively\. We normalize according to the following per\-language byte premiums\([Arnett et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib49)\): English1\.01\.0, Dutch1\.05161\.0516, Simplified Chinese0\.9359660\.935966\. The word unit is thenum\-tokenscolumn of the released datasets\. For English and Dutch this column counts whitespace tokens\. For Chinese, per thebabylm\-zhodataset card, it is the subword count under the Qwen3\-0\.6B tokenizer \(we verified an exact match on a sample of documents\)\. We approximate the Chinese count as0\.7110\.711units per character, the corpus\-level ratio ofnum\-tokensto characters measured onbabylm\-zho\. We split each source document into passages of about300300words at sentence boundaries\. Chinese passages in which Han characters make up fewer than50%50\\%of Han\-plus\-Latin characters are dropped as romanized\-pinyin transcripts \(the Chinese portion of CHILDES\([MacWhinney, 2000](https://arxiv.org/html/2609.30535#bib.bib65)\)\), as are passages under1515words in any language\. Figure 5:Verbatim corpus examplesby matrix language \(rows\) and switch type \(columns\)\. Embedded\-language spans are colored by language \(English,Dutch,Chinese\)\. The non\-CS cell in each row is the true source document of that row’s word\-level example\. ### Generation, batching, cost\. Here we describe our final production system to generate of the word\- and sentence\-level CS corpus\. It usesdeepseek\-v4\-flashvia OpenRouter\. Each request carries eight passages sharing one \(matrix, embedded\) language pair as JSON lines and returns one JSON line per input id\. The full translations used only by the document translation control \(§[3](https://arxiv.org/html/2609.30535#S3)\) were generated withclaude\-haiku\-4\-5via the Anthropic Message Batches API, with the prompt of Appendix[B](https://arxiv.org/html/2609.30535#A2)\. Each switch type covers the same413,302413\{,\}302passages \(51,66651\{,\}666requests\); the total API cost was≈$284\{\\approx\}\\$284\. ### Quality verification\. Every generation must contain both the matrix and the embedded language\. We check language presence with \(lingua\): Chinese counts as present at≥3\{\\geq\}3Han characters; in Latin\-script text a language counts as present when it covers≥8%\{\\geq\}8\\%of the detected span length\. We reject outputs that are identical to the source \(no switching\), that miss the matrix language \(a full translation\), or that miss the embedded language, and we reject outputs whose character count falls outside0\.25×0\.25\\times–4×4\\timesthat of the source, to catch truncations and runaway generations\. Acceptance rates, as a fraction of requested passages, are0\.780\.78–0\.890\.89for word\-level switching in every direction and0\.790\.79–0\.870\.87for sentence\-level switching into an English or Dutch matrix, but fall to0\.470\.47–0\.560\.56for sentence\-level switching into a Chinese matrix, where most rejections are unchanged or fully translated outputs\. In total∼348\{\\sim\}348K word\-level and∼323\{\\sim\}323K sentence\-level passages make it past this initial quality filter\. ### Document reassembly\. Accepted passages are stitched back into documents in order\. A document whose every passage was rejected is dropped, and one with a rejected passage is reassembled with that passage missing \(3\.7%3\.7\\%of CS documents,1\.8%1\.8\\%with a gap in the middle of the passage\)\. The non\-CS baseline document is stitched from the source text of the same accepted passages, so both the CS and non\-CS versions have identical gaps\. ### Budget accounting\. The word\-level, sentence\-level, and non\-CS thirds occupy disjoint document ids, so that no same document is used for two different thirds of our corpus\. Because we found that per\-language word heuristics do not work well on mixed text, we check compliance to the 100M adjusted\-word budget twice\. We use a script\-aware word count \(Han characters×0\.711\{\\times\}0\.711plus whitespace tokens of the Latin remainder\), and we also count the total UTF\-8 bytes \(543543MB≡100\{\\,\\equiv\\,\}100M English words\)\. By both metrics, the CS corpus is within budget \(≈96\.6\{\\approx\}96\.6M words /485485MB\) and near\-identical in volume to its non\-CS baseline \(95\.795\.7M words /482482MB\); the99\.699\.6M of Table[1](https://arxiv.org/html/2609.30535#S1.T1)counts the same documents with the per\-language heuristic above\. Under our16,38416\{,\}384\-entry BPE vocabulary the CS corpus is138\.9138\.9M tokens and the non\-CS corpus135\.0135\.0M, a2\.9%2\.9\\%difference; the two code\-switched thirds tokenize to22–12%12\\%more tokens than their source documents, most for Chinese\-matrix text with translated Latin\-script sentences\. ### Examples from control corpora\. The following two examples come from the word salad corpus \(§[3](https://arxiv.org/html/2609.30535#S3)\): - EN←\\leftarrowNLTransduction is awisselkoersenterm\. It can mean: - ZH←\\leftarrowNL\\cjkfont曾以“女排精神”引mevrouw\\cjkfont领一代风骚的中国voelen\\cjkfont女排,今gravin\\cjkfont天再度擎起世界冠军大旗,怎能不让国人惊喜? The following two come from the document translation corpus: - EN→\\toNLBurien is a city in King County, Washington, United States\.⇒\\RightarrowBurien is een stad in King County, Washington, Verenigde Staten\. - ZH→\\toEN\\cjkfont不然它们就吃撑了。⇒\\RightarrowOtherwise they will eat until they are stuffed\. ## Appendix BGeneration Model and Prompts We chose the generation model through a couple different steps\. First we scored a3636\-passage sample \(three per matrix×\\timesembedded×\\timesswitch\-type cell\) from several frontier LLMs \(as of June 2026\) with the quality gate of Appendix[A](https://arxiv.org/html/2609.30535#A1), supplemented by manual review from native English–Dutch and English–Chinese speakers\. Pass rates were DeepSeek\-V4\-Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.30535#bib.bib60)\)0\.720\.72,claude\-haiku\-4\-50\.670\.67, Gemini 2\.5 Flash\-Lite0\.330\.33, and Qwen3\-30B\-A3B0\.190\.19\. The Qwen3 models often drifted into full translation on cross\-script pairs, whereas Haiku often failed to produce embedded Chinese words\. We therefore useddeepseek\-v4\-flashfor the full run, which from manual inspection gave good results for each kind of code\-switching\. Unlike[Yoo et al\. \(2025\)](https://arxiv.org/html/2609.30535#bib.bib2), we did not need to supply parallel translations to the generator to obtain reasonable CS passages\. Our generation prompts follow verbatim\.\{matrix\}/\{embedded\}refer to the language names \(English/Dutch/Chinese \(Simplified\)\)\. Each prompt is followed by the instruction to return exactly one JSON line per input id with the same ids, and then by the input passages as JSON lines of the form\{"id": \.\.\., "text": \.\.\.\}\. ### Word\-level CS\. ``` Task: rewrite each {matrix} text as INTRASENTENTIAL code-switching with {embedded}. Rules: - Keep {matrix} as the matrix language: it carries the grammar, word order, and most function words. Most words stay in {matrix}. - CRITICAL: do NOT translate sentences fully into {embedded}. {matrix} must remain dominant throughout. - Embed natural {embedded} content words and short phrases INSIDE sentences (nouns, verbs, adjectives, technical terms, short NPs). - Switch only at grammatically permissible boundaries. Do not switch inside a word. - Target roughly 25-45% of content words/ phrases in {embedded}; every non-trivial sentence should contain >=1 {embedded} switch. - Preserve meaning exactly. Do not add, drop, summarise, or invent facts. - Do NOT translate a whole sentence into {embedded}: mix both languages within sentences. - Use the native script for {embedded}. - Preserve sentence order and any list/markup. ``` ### Sentence\-level CS\. ``` Task: rewrite each {matrix} text as SENTENCE-LEVEL code-switching with {embedded}. Rules: - Keep the original sentence ORDER. Translate roughly half of the sentences fully into {embedded}; leave the rest in {matrix}. - CRITICAL: do NOT translate the whole passage. About half of sentences must remain {matrix}. Output with no {matrix} sentences is invalid. - Each sentence stays internally monolingual: all {matrix} or all {embedded}. Do NOT mix languages inside a sentence. - Translate faithfully; preserve meaning. Do not add, drop, or invent facts. - Leave short fragments, list items, headings, or markup-like lines unchanged if translating would corrupt them. - Use the native script for {embedded}. ``` ### Document translation control\. ``` Task: translate each {matrix} passage FAITHFULLY and COMPLETELY into {embedded}. Rules: - Translate the ENTIRE passage into {embedded}; no {matrix} words should remain (proper nouns may stay). - Preserve meaning, sentence order, and structure exactly; do not add, drop, or summarise. - Natural, fluent {embedded}; not word-for-word glossing. - Use the native script for {embedded}. ``` ## Appendix CEvaluation Suite Table 3:Accuracy on the BabyLM evaluation suite per task and per language\. We report \(%\) mean±\\pmSD across eight seeds per condition\. Chance performance is random\-guess accuracy \(1/\#1/\\\#options\) where applicable\. Word/Sent/Par/Salad are the shuffled\-ordering controls \(§[3](https://arxiv.org/html/2609.30535#S3)\)\.*Leaderboard*is the official multilingual\-average score as in Table[2](https://arxiv.org/html/2609.30535#S3.T2)\.Table 4:Bits per UTF\-8 byte on1,2001\{,\}200held\-out BabyBabelLM documents per language \(lower is better\)\. The first512512tokens of each document are scored\.We evaluate with the official BabyLM 2026 multilingual evaluation pipeline\([Choshen et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib48)\), which scores each task in each of the three languages where available\. ### Zero\-shot tasks\. The zero\-shot tasks consist of minimal\-pair and multiple\-choice tasks scored by comparing candidate log\-probabilities\. Grammar evaluations include BabyLM’s filtered subset of BLiMP\([Warstadt et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib32)\)for English, with BLiMP\-NL\([Suijkerbuijk et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib33)\)and ZhoBLiMP\([Liu et al\., 2026](https://arxiv.org/html/2609.30535#bib.bib35)\)as its Dutch and Chinese counterparts, and MultiBLiMP\([Jumelet et al\., 2026b](https://arxiv.org/html/2609.30535#bib.bib34)\)in English and Dutch\. Common\-sense semantic reasoning benchmarks include HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.30535#bib.bib37)\); Winogrande\([Sakaguchi et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib38)\); XStoryCloze\([Lin et al\., 2022b](https://arxiv.org/html/2609.30535#bib.bib30)\); XCOMPS\([He et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib36)\); Global PIQA\([Chang et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib40)\); and the Chinese Hanzi Structure and Hanzi Pinyin minimal\-pair tasks distributed with the pipeline\. The latter two are hidden tasks: the leaderboard computes their scores server\-side and excludes them from every average, so they do not enter the leaderboard score or Table[3](https://arxiv.org/html/2609.30535#A3.T3)\. ### Fine\-tuning tasks\. ARC\([Clark et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib39)\); Belebele\([Bandarkar et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib41)\); BMLAMA\([Qi et al\., 2023](https://arxiv.org/html/2609.30535#bib.bib42)\); MNLI\([Williams et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib43)\); SIB\-200\([Adelani et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib46)\); TruthfulQA\([Lin et al\., 2022a](https://arxiv.org/html/2609.30535#bib.bib31)\); XNLI\([Conneau et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib44)\); INCLUDE\([Romanou et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib45)\); and cross\-lingual POS tagging over Universal Dependencies treebanks\([de Marneffe et al\., 2021](https://arxiv.org/html/2609.30535#bib.bib47)\)\. Each classification task is fine\-tuned and evaluated within one language \(up to1010epochs with early stopping at patience33, learning rate5×10−55\\times 10^\{\-5\}, batch size6464, one fine\-tuning seed\); only POS tagging is trained jointly on all three languages\. ### Harness details\. We run the harness with the organizers’ fix for the beginning\-of\-sequence \(BOS\) token: the harness recovers the continuation tokens as the suffix of the tokenized prompt\+\+continuation, which breaks for tokenizers that, like the baseline’s and ours, wrap every string in<s\>…</s\>; the fix tokenizes without special tokens and prepends a single<s\>\. ## Appendix DDetails on Bitext Retrieval Figure 6:In addition to sentence\-level bitext retrieval precision@1 \(Figure[2](https://arxiv.org/html/2609.30535#S4.F2)\), here we show the per\-layer median percentile rank of the correct translation\. Averaged over layers 6–11, the non\-CS models place the correct translation at median rank105105and121121out of997997for English\-Chinese and Dutch\-Chinese respectively, far above chance at499499\. This suggests that code\-switching does not create alignment “from scratch”, but rather “sharpens” a general degree of alignment found even in the non\-CS models\.Table 5:Cross\-lingual alignment measures beyond retrieval \(curriculum ordering\)\. Each cell shows the mean CS/non\-CS values across seeds at the layer of peak CSLS retrieval in the CS model; We measure linear CKA\([Kornblith et al\., 2019](https://arxiv.org/html/2609.30535#bib.bib11)\), the cosine gap \(mean cosine over theNNtrue translation pairs minus that over a fixed random derangement of the pairs\), and the per\-language centroid distance\([Libovický et al\., 2020](https://arxiv.org/html/2609.30535#bib.bib9), lower is better;\)\. The last two rows are per\-language MEXA\([Kargaran et al\., 2025](https://arxiv.org/html/2609.30535#bib.bib12)\)\(max over layers vs\. the English pivot\)\.### Protocol\. We embed the FLORES\+devset\([NLLB Team et al\., 2024](https://arxiv.org/html/2609.30535#bib.bib20);[Goyal et al\., 2022](https://arxiv.org/html/2609.30535#bib.bib21)\)\(eng\_Latn/nld\_Latn/cmn\_Hans\), truncating each sentence to its first128128tokens and excluding padding from the mean over positions\. LetS=\(sij\)S=\(s\_\{ij\}\)be theN×NN\\times Nmatrix of cosine similarities such thatsijs\_\{ij\}is the cosine similarity between the representations of theii\-th source\-language sentence and thejj\-th target\-language sentence\. Plain nearest\-neighbor retrieval overSSsuffers from hubness, where a few target sentences are the nearest neighbor of many unrelated queries\([Radovanović et al\., 2010](https://arxiv.org/html/2609.30535#bib.bib17)\), so we rank with cross\-domain similarity local scaling\([Lample et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib16),k=10k\{=\}10;\), forming the matrixC=\(cij\)C=\(c\_\{ij\}\)where cij≔2sij−rB\(i\)−rA\(j\),c\_\{ij\}\\coloneqq 2s\_\{ij\}\-r\_\{B\}\(i\)\-r\_\{A\}\(j\),whererB\(i\)r\_\{B\}\(i\)is the mean cosine between source sentenceiiand itskknearest targets andrA\(j\)r\_\{A\}\(j\)the mean cosine between target sentencejjand itskknearest sources\. Retrieval is correct for sentenceiiwhen its translation ranks first underCC, and we retrieve in both directions and average: P@1≔12\(accA→B\+accB→A\),\\mathrm\{P@1\}\\coloneqq\\tfrac\{1\}\{2\}\\big\(\\mathrm\{acc\}\_\{A\\to B\}\+\\mathrm\{acc\}\_\{B\\to A\}\\big\),where accA→B≔1N∑i𝟏\[argmaxjcij=i\]\\mathrm\{acc\}\_\{A\\to B\}\\coloneqq\\frac\{1\}\{N\}\\sum\_\{i\}\\mathbf\{1\}\[\\arg\\max\_\{j\}c\_\{ij\}=i\]is the fraction of source sentences whose top\-ranked target is their own translation, andaccB→A\\mathrm\{acc\}\_\{B\\to A\}is the same in reverse \(CCis symmetric in the two directions, sinceC⊤C^\{\\top\}is the CSLS matrix ofS⊤S^\{\\top\}\)\. Chance is1/N≈0\.1%1/N\\approx 0\.1\\%\. Table 6:Cross\-lingual alignment of all eight conditions at their final checkpoints; alignment columns are88\-seed mean±\\pmSD\. Left: FLORES\+ bitext retrieval P@1 \(CSLS,k=10k\{=\}10\) averaged over layers 6–11, per language pair\. Right: the leaderboard score of Table[2](https://arxiv.org/html/2609.30535#S3.T2)\. Column abbreviations as in Table[3](https://arxiv.org/html/2609.30535#A3.T3)\.Figure 7:Per\-layer bitext retrieval P@1 for all eight experimental conditions \(§[4](https://arxiv.org/html/2609.30535#S4)\), overlaid in each language\-pair panel\. We use neutral colors for the four main conditions, each pairing of corpus \(CS/non\-CS\) and ordering \(shuffled/curriculum\), and color for the four shuffled\-ordering controls\. Lines are88\-seed means, bands±1\\pm 1SD, dotted line chance\. Shading marks layers66–1111, the range averaged over in Figure[3](https://arxiv.org/html/2609.30535#S4.F3)\.Figure 8:Layerwise cosine similarity between the CS and non\-CS models’ word representations \(§[7](https://arxiv.org/html/2609.30535#S7)\), shown separately for the embedded words and their never\-embedded controls \(88\-seed mean±1\\pm 1SD\)\. ## Appendix EDetails on Word\-Level Alignment ### Constructing candidate sets for retrieval\. For each language direction, we first group words by part of speech and subword\-token count\. We bisect each group into embedded and never\-embedded words and match them one\-to\-one by the Hungarian algorithm onlog10\\log\_\{10\}monolingual frequency\. Translation labels come from the MUSE dictionaries\([Lample et al\., 2018](https://arxiv.org/html/2609.30535#bib.bib16)\)in addition to the pairs of words substituted during corpus generation\. ### How code\-switching shifts word representations across layers\. For each of the2,2042\{,\}204source words in the1,1021\{,\}102matched pairs of §[7](https://arxiv.org/html/2609.30535#S7), at each layer, we take the cosine between the CS and non\-CS models’ representations \(Figure[8](https://arxiv.org/html/2609.30535#A4.F8)\)\. The high cosine similarity between CS and non\-CS models at layer zero indicates that code\-switching does not affect the word embeddings much\. However, early layers seem more affected representationally by the presence of code\-switched text in training, as is seen by the low cosine similarity\. The representations of words at later layers become again more similar between the CS and non\-CS models\.
相似文章
多语言思维,而非更难的思维:教授推理模型代码切换的数据高效框架
本文介绍了一个数据高效的微调框架,用于教授推理模型有效地进行代码切换(混合使用多种语言),证明了战略性的代码切换可以提升低资源语言的推理能力。该工作分析了大型语言模型在不同语言、任务和领域中的代码切换行为,并开发了促进有益代码切换模式的干预措施。
通过渐进式代码切换实现高效的多语言推理迁移
本文提出渐进式代码切换(PCS),这是一种结合课程学习的强化学习方法,逐步增加LLM中的代码切换比例,从而实现多语言推理能力的高效迁移。
多语言语言模型中语言控制的潜在机制
本文比较了三种方法,用于识别多语言语言模型中的语言控制潜在因子,以应对语码转换。实验在Gemma-2-2B和Qwen3-4B上进行,结果显示FreqSel最为有效。
适应是双向的:研究人类与语言模型之间的语言趋同
本文研究了在多轮对话中人类与大型语言模型之间的语言适应性,发现LLM过度趋同于用户风格,而人类适应LLM的方式与适应其他人类并无不同。
迈向真正多语言ASR:将代码切换ASR泛化到未见过的语言对
本文研究了从有限的已见语言对学到的代码切换ASR能力是否可以通过模型合并和域泛化方法泛化到未见过的语言对,结果发现只有有限的迁移。