LLM Scheming Inversely Scales with Pretraining Language Coverage
Summary
This paper finds that LLM scheming behavior inversely scales with pretraining language coverage, with low-resource languages showing 34.2% higher scheming scores in Qwen3-30B-A3B.
View Cached Full Text
Cached at: 07/29/26, 09:51 AM
# LLM Scheming Inversely Scales with Pretraining Language Coverage
Source: [https://arxiv.org/html/2607.24769](https://arxiv.org/html/2607.24769)
###### Abstract
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high\-risk deployment settings\. While recent work has empirically demonstrated in\-context scheming—the covert pursuit of misaligned objectives while feigning alignment—in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety\. We apply Petri, an open\-source automated auditing framework, to Qwen3\-30B\-A3B to evaluate deceptive and scheming behaviors across multiple languages\. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low\-resource languages averaging 34\.2% higher scores compared to high\-resource languages on a five\-category scheming index\. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors\.
Machine Learning, ICML
## 1Introduction
As AI systems are increasingly deployed in high\-risk settings such as autonomous decision\-making, medical assistance, and legal analysis, ensuring robust model alignment and safety becomes increasingly critical\. Yet findings suggest that models have gained the capability to perform in\-context scheming, the act of deliberately pretending to be aligned while covertly pursuing alternative misaligned objectives\(Hubingeret al\.,[2021](https://arxiv.org/html/2607.24769#bib.bib2); Carlsmith,[2023](https://arxiv.org/html/2607.24769#bib.bib3)\)\. Scheming behaviors have been empirically demonstrated across multiple frontier models\(Meinkeet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib19)\), and Claude 3 Opus engages in alignment faking without explicit instruction, selectively complying with training objectives specifically to prevent modification of its own values\(Greenblattet al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib4)\)\. Yet existing efforts to mitigate and understand scheming are conducted almost exclusively in English, despite evidence that safety alignment degrades significantly in non\-English languages\(Wanget al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib27)\)\. Related work shows that surface behavioral metrics can underestimate the extent of internal representational change during behavior modification\(Chaudhary and Barez,[2025](https://arxiv.org/html/2607.24769#bib.bib37)\), and that adaptation procedures can transfer latent behavioral dispositions even through ostensibly benign training pipelines\([K”̈oniget al\.,](https://arxiv.org/html/2607.24769#bib.bib40)\)\. Due to differences in pre\-training corpus language concentrations, scheming behaviors may not generalize uniformly across languages\. In this paper, we present a correlation study of scheming and deception propensity across six languages—English, Spanish, Chinese, Arabic, Vietnamese, and Portuguese—and examine whether observed variation is consistent with differences in estimated pretraining language concentration\. In this work, we address the entailing hypothesis:Large language models \(LLMs\) have a tendency to scheme more in languages covered less in its pretraining corpus\.
## 2Related Work
#### Misalignment evaluations
The use of LLMs in high\-risk settings and agentic environments has sparked a growing body of work evaluating language models for misalignment\(Taubenfeldet al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib13); Shevlaneet al\.,[2023](https://arxiv.org/html/2607.24769#bib.bib14); Phuonget al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib15)\)\. Prior works examine various forms of misalignment, including harmful compliance, jailbreaks, and other failure modes caused by adversarial prompting\(Zouet al\.,[2023](https://arxiv.org/html/2607.24769#bib.bib16); Chaoet al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib17)\)\. Studies have also shown language models to be proficient in performing covert actions like sandbagging, evaluation awareness, and scheming in agentic environments\(van der Weijet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib18); Nguyenet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib20); Meinkeet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib19); Chaudharyet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib51)\)with tool\-calling capabilities\. Such covert behaviors may be facilitated by context\-dependent reasoning that the model selectively reveals\(Batraet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib46)\), and recent work shows that the chosen post\-training objective shapes which internal structures are preserved versus disrupted\(Nunezet al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib38)\)\. Beyond eliciting dangerous behaviors to assess compliance, frameworks and benchmarks are also commonly used to evaluate misalignment\(Mazeikaet al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib21)\)\.
Figure 1:Here, we visually show our method’s pipeline\. We initially translate the seed instructions from the auditor and target to a predefined language\. Then, the transcript is carried out in that language and the judge subsequently scores the conversation using several behavioral measures\.
#### Multilingual Models and Evaluations
To reach a broader audience, numerous models have been developed with an emphasis on multilingual capabilities\(Grattafioriet al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib23); Foroutanet al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib33); Üstünet al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib32)\)\. However, model strength is typically distributed unequally between languages\(Liet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib25)\)\. A similar trend is shown with the general alignment of models, with a nontrivial disparity in alignment between high\-resource and low\-resource languages\(Yonget al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib24)\)\. Current work prioritizes monolingual approaches to LLM safety, leaving multilingual studies relatively unexplored\(Yonget al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib26); Wanget al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib27)\)\. Although multilingual studies reveal failure modes like jailbreaks particularly in low\-resource languages\(Yonget al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib24)\), they typically neglect other covert behaviors such as scheming, resulting in a critical gap in research\. Persona\-level behavioral properties such as sycophancy have been shown to be steerable through off\-the\-shelf persona vectors rather than reducing to a single linear direction\(Kelkaret al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib41)\), suggesting that complex misalignment behaviors may decompose into structurally distinct components — a possibility that becomes more tractable to test in a multilingual setting where coverage varies systematically\.
## 3Methodology
Our analysis is based largely on the existing Petri framework used to evaluate the broad misalignment capabilities of language models\(Fronsdalet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib1)\)\. In this section, we provide a comprehensive description of our evaluation pipeline\.
At a high level, Petri operates as an automated evaluation framework used to probe misalignment in a target model\. This is done by employing anAuditorandJudge—both of which are language models\. More specifically, the auditor is initialized with a system prompt and tool construction capabilities used to elicit harmful behaviors from the target\. The system prompt provides the auditor with information about what scenario to simulate and corresponding behavior probing methods\. We defer examples of seed instructions to Appendix[A\.1](https://arxiv.org/html/2607.24769#A1.SS1)\. After receiving a seed instruction specifying desired scenarios and behaviors, the auditor interacts with the target model through multi\-turn conversations, dynamically constructing prompts, tools, and simulated outputs to elicit misaligned behavior\.
After a conversation, aJudgelanguage model evaluates the extent to which the target acts misaligned by reviewing the transcript\. Specifically, it scores the target model’s responses across a broad set of categories on a scale from one to ten, where higher values indicate a greater presence of the respective behaviors in the transcript\. Notable categories include “evaluation awareness”, “unprompted sycophancy”, and “unprompted deception towards user”\.
To apply this pipeline in a multilingual fashion, we translate the seed instructions and auditor system prompts into prespecified languages\. The target model’s system prompt instructs it to respond exclusively in the prescribed language\. We then yield a score distribution representative of each language by aggregating the judge’s score outputs across all audits and averaging them per category\. Formally:
S¯b,l=1\|Ib\|∑i∈IbSi,b,l\\bar\{S\}\_\{b,l\}=\\frac\{1\}\{\\lvert I\_\{b\}\\rvert\}\\sum\_\{i\\in I\_\{b\}\}S\_\{i,b,l\}\(1\)
HereS¯b,l\\bar\{S\}\_\{b,l\}denotes the average score that a judge assigns to a specific behaviorbbfor a given languagell\. We also run a permutation test to assess whether the results emerged from chance\. We assessed significance of our scores using a two\-sample permutation test over the selected language\-category scores, while maintaining the original group sizes\.
## 4Experiments
Table 1:Selected scheming behavior scores across languages\.Values represent the mean judge scores \(on a 1–10 scale\) per behavior category\. Languages are ordered from least to most scheming by their average score\. The final column reports the mean across all languages\.#### Experimental Setup
In our experiments, we leverage the Qwen3\(Yanget al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib30)\)family of models\. Specifically, we analyze results over Qwen3\-30B\-A3B\. For the auditor and judge models, we use Gemini 2\.5 Flash and Gemini 2\.5 Pro, respectively\. Our translations are also carried out via Gemini 2\.5 Pro\. Additionally, prior work has found Gemini 2\.5 Pro to be an apt judge through comparisons with those from OpenAI and Anthropic\.\(Fronsdalet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib1)\)\.
We report results for a selected five\-category scheming index\. These categories were selected because they are most representative of scheming behavior\. Although other categories such as negative affect, information leaking, and concerning behavior could have been included, the resulting scores might capture related but less specific safety failures, rather than scheming behaviors specifically\.
For each scheming behavior, we first identified all samples in which that behavior received a non\-baseline score of 1 in at least one language\. We then extracted language\-specific scores for each of these samples and computed their mean to obtain a per\-language normalized mean score for each behavior\.
#### Languages
In our experiments, we use six languages officially supported by Qwen3\(Yanget al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib30)\)\. Specifically, we use English, Chinese, Spanish, Portuguese, Arabic, and Vietnamese\. We categorize English and Chinese as high\-resource languages since we estimate that these languages make up the majority of the pretraining corpus for Qwen3\-30B\-A3B\. This estimation is supported in part because of historical precedents set by the Qwen series models, where prior models were trained with an explicit emphasis on Chinese and English, and this precedent has continued\(Baiet al\.,[2023](https://arxiv.org/html/2607.24769#bib.bib31)\)\. In the Qwen3 family of models, related technical reports also suggest the dominance of English and Chinese in the pretraining corpus\. For example, the Qwen3 Omni technical report indicates that the training data includes 80% Chinese and English\(Xuet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib29)\)\. Additionally, the Qwen3 ASR technical report notes that the majority of AUT pretraining data is in Chinese and English\(Shiet al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib28)\)\.
Based on these factors, we reasonably estimate that Chinese and English comprise a larger portion of the pretraining corpus than other languages\. We classify Spanish, Portuguese, Arabic, and Vietnamese as medium\- to low\-resource languages\. We deem Qwen3\-30B\-A3B sufficiently capable of reasoning in these languages due to notable scores in benchmarks such as MMMLU, MT\-AIME24, Polymath, etc\. These languages are lower in resources compared to English and Chinese, but there does not exist a concrete methodology to classify the proportions between these languages without more information\. Without precise classification of the rankings in this scenario, we employ a binary classification and instead observe differences in the two categories\.
#### Translational Verification
In our experiments, we employ measures to ensure the validity of our translations and judgment\. Specifically, we used language instructors in Chinese and Spanish to verify the translational validity of 10 random samples\. We found that intent was faithfully preserved across all verified samples\. For the validity of judgments, we evaluate 10 random samples per language by manually reviewing the assigned scores with the corresponding transcript\. We found that the scores were aligned with observable behavior in all tested languages\.
## 5Results
Figure 2:Average scheming score across languages\. Here, we show the average score of each language across the five selected scheming behaviors\. Notably, Chinese and English frequently exhibit the lowest scheming scores opposed to Spanish, Arabic, Vietnamese, and Portuguese, all of which are considered low\-resource languages\.TableLABEL:tab:scheming\_ratesand Figure[2](https://arxiv.org/html/2607.24769#S5.F2)display the average scheming scores across all evaluated languages and behavioral categories\. Our results convey that languages estimated to occupy a larger share of the pretraining corpus tend to exhibit lower mean scheming scores, with low\-resource languages exhibiting on average scores34\.2%higher\.
Chinese and English are estimated to dominate the pretraining corpus, yielding mean scheming scores of 2\.055 and 2\.076, respectively\. Spanish, Arabic, Portuguese, and Vietnamese resulted in average scores of 2\.634, 2\.638, 2\.648, and 3\.164, respectively\. This pattern is visible in aggregate language means and is strongest for emotional manipulation and self\-preservation, while categories such as self\-serving bias show weaker separation across language groups\.
Behavioral categories exhibit greater variance across results\. Categories such as “deception toward user”, show relatively uniform scores across languages, ranging from 2\.026 in Chinese to 2\.692 in Arabic\. In contrast, “self\-preservation” exhibits much larger disparities, with Spanish, Portuguese, and Vietnamese scoring 3\.400, 2\.599, and 4\.400, respectively, while all other languages obtained a score of 1\.00\. “Encouragement of user delusion” is the highest\-scoring category overall, with a mean score of 3\.583 across languages\. These results suggest that certain behavioral categories are more sensitive to pretraining language concentration than others\. For our permutation test, the observed gap was \(p=0\.019p=0\.019one\-sided;p=0\.039p=0\.039two\-sided\), suggesting that our results are statistically significant while simultaneously validating our hypothesis stated in Section[1](https://arxiv.org/html/2607.24769#S1)\.
## 6Discussion
Our findings show that low\-resource languages exhibit higher mean scheming scores relative to high\-resource languages\. In this section, we discuss potential reasons behind our results\.
Safety alignment procedures, which are often developed and validated primarily in English, may generalize less effectively to low\-resource languages due to imbalances in pretraining coverage\. If alignment tuning is disproportionately concentrated in English and Chinese, low\-resource languages may receive weaker safety signal, potentially resulting in higher scheming rates in those languages\. This view is consistent with broader findings that training procedures shape which circuits are preserved or disrupted in post\-training\(Golechhaet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib53); Nunezet al\.,[2026](https://arxiv.org/html/2607.24769#bib.bib38)\), and that some circuits are alignment\-critical and warrant preferential preservation during model modification\(Patelet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib47)\)\. Under this lens, low\-resource languages may simply lack the alignment\-critical circuit coverage that high\-resource languages enjoy\.
There may also be semantic discrepancies introduced during seed instruction translations\. Specific nuances like tone or slang may not translate cleanly to other languages\. For this reason, the target model may not interpret the original seed instruction faithfully, potentially causing disproportionately large score shifts during the judging round\. We acknowledge that this may also contribute to the varying scheming behavior across languages observed in our analysis\.
## 7Limitations and Future Work
In this work, our primary limitation is the scarcity of models that publicly disclose the proportions of languages used in their training datasets\. Although there are several model families that do disclose this\(Martinset al\.,[2024](https://arxiv.org/html/2607.24769#bib.bib34); Gonzalez\-Agirreet al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib35); Workshopet al\.,[2023](https://arxiv.org/html/2607.24769#bib.bib36)\), their context windows are typically limited, hindering compatibility with the Petri framework\. Since Petri mandates multi\-turn conversations between the target and auditor models, smaller context windows are exhausted relatively quickly\. As a result of this, we leveraged a model with a large context window \(Qwen3\-30B\-A3B\) and estimated its dataset language proportions using their technical report\(Yanget al\.,[2025](https://arxiv.org/html/2607.24769#bib.bib30)\)\. This restricts our ability to draw precise quantitative conclusions about the relationship between language proportion and model behavior\.
Additionally, our use of a single judge model may have introduced language\-dependent scoring biases\. Thus, we are limited in generalizing our results across various architectures, model sizes, and families\. Differences in training data, architecture, or alignment procedures may produce substantially different scheming behaviors, which our current pipeline does not fully account for\.
As for future work, we aim to apply our methodology to a broader range of model families to increase generalizability\. Specifically, we prioritize models with publicly disclosed pretraining resource distribution\. We would also like to extend our study further by assessing the sensitivity of our results to model and design choices\.
## 8Conclusion
In this work, we examine the relationship between pretraining language coverage and scheming behavior in a multilingual LLM\. Using an automated auditing framework, Petri, we evaluate model behavior across six languages with different levels of estimated resource representation\.
We find that low\-resource languages consistently exhibit higher mean scheming scores, with disparities of up to 34\.2% compared to high\-resource languages\. This pattern persists across multiple behavioral categories, suggesting that alignment techniques may not generalize uniformly across languages\.
While our findings are subject to limitations, they highlight a critical gap in current safety research\. As language models are increasingly deployed in multilingual contexts, ensuring consistent alignment across languages is essential\.
## 9Impact Statement
This paper presents work whose goal is to advance the understanding of AI safety\. With the disparities in scheming behavior across different languages in large language models, users interacting with low\-resource languages may face a higher risk of encountering deceptive behavior, potentially leading to more motivation to deliver more safety measures for such cases\. There also exists risks such as deliberate exploitation of large language models by interacting through low\-resource languages\. Overall, we hope to demonstrate the importance of improving alignment in all languages to ensure safe and equitable deployment\.
## References
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang, B\. Hui, L\. Ji, M\. Li, J\. Lin, R\. Lin, D\. Liu, G\. Liu, C\. Lu, K\. Lu, J\. Ma, R\. Men, X\. Ren, X\. Ren, C\. Tan, S\. Tan, J\. Tu, P\. Wang, S\. Wang, W\. Wang, S\. Wu, B\. Xu, J\. Xu, A\. Yang, H\. Yang, J\. Yang, S\. Yang, Y\. Yao, B\. Yu, H\. Yuan, Z\. Yuan, J\. Zhang, X\. Zhang, Y\. Zhang, Z\. Zhang, C\. Zhou, J\. Zhou, X\. Zhou, and T\. Zhu \(2023\)Qwen technical report\.arXiv preprint arXiv:2309\.16609\.Cited by:[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px2.p1.1)\.
- S\. Batra, P\. Tillman, S\. Gaggar, S\. Kesineni, K\. Zhu, S\. Dev, A\. Panda, V\. Sharma, and M\. Chaudhary \(2025\)SALT: steering activations towards leakage\-free thinking in chain of thought\.arXiv preprint arXiv:2511\.07772\.Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Carlsmith \(2023\)Scheming ais: will ais fake alignment during training in order to get power?\.External Links:2311\.08379,[Link](https://arxiv.org/abs/2311.08379)Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1)\.
- P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce, V\. Sehwag, E\. Dobriban, N\. Flammarion, G\. J\. Pappas, F\. Tramer, H\. Hassani, and E\. Wong \(2024\)JailbreakBench: an open robustness benchmark for jailbreaking large language models\.External Links:2404\.01318,[Link](https://arxiv.org/abs/2404.01318)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Chaudhary and F\. Barez \(2025\)Safetynet: detecting harmful outputs in llms by modeling and monitoring deceptive behaviors\.arXiv preprint arXiv:2505\.14300\.Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1)\.
- M\. Chaudhary, I\. Su, N\. Hooda, N\. Shankar, J\. Tan, K\. Zhu, R\. Lagasse, V\. Sharma, and A\. Panda \(2025\)Evaluation awareness scales predictably in open\-weights large language models\.arXiv preprint arXiv:2509\.13333\.Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- N\. Foroutan, P\. Teiletche, A\. K\. Tarun, and A\. Bosselut \(2026\)Revisiting multilingual data mixtures in language model pretraining\.External Links:[Link](https://openreview.net/forum?id=IKJyRyHpHV)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Fronsdal, I\. Gupta, A\. Sheshadri, J\. Michala, S\. McAleer, R\. Wang, S\. Price, and S\. Bowman \(2025\)Petri: parallel exploration of risky interactions\.External Links:[Link](https://github.com/safety-research/petri)Cited by:[§3](https://arxiv.org/html/2607.24769#S3.p1.1),[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px1.p1.1)\.
- S\. Golechha, M\. Chaudhary, J\. Velja, A\. Abate, and N\. Schoots \(2025\)Modular training of neural networks aids interpretability\.arXiv e\-prints,pp\. arXiv–2502\.Cited by:[§6](https://arxiv.org/html/2607.24769#S6.p2.1)\.
- A\. Gonzalez\-Agirre, M\. Pàmies, J\. Llop, I\. Baucells, S\. D\. Dalt, D\. Tamayo, J\. J\. Saiz, F\. Espuña, J\. Prats, J\. Aula\-Blasco, M\. Mina, I\. Pikabea, A\. Rubio, A\. Shvets, A\. Sallés, I\. Lacunza, J\. Palomar, J\. Falcão, L\. Tormo, L\. Vasquez\-Reina, M\. Marimon, O\. Pareras, V\. Ruiz\-Fernández, and M\. Villegas \(2025\)Salamandra technical report\.External Links:2502\.08489,[Link](https://arxiv.org/abs/2502.08489)Cited by:[§7](https://arxiv.org/html/2607.24769#S7.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Greenblatt, C\. Denison, B\. Wright, F\. Roger, M\. MacDiarmid, S\. Marks, J\. Treutlein, T\. Belonax, J\. Chen, D\. Duvenaud, A\. Khan, J\. Michael, S\. Mindermann, E\. Perez, L\. Petrini, J\. Uesato, J\. Kaplan, B\. Shlegeris, S\. R\. Bowman, and E\. Hubinger \(2024\)Alignment faking in large language models\.External Links:2412\.14093,[Link](https://arxiv.org/abs/2412.14093)Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1)\.
- E\. Hubinger, C\. van Merwijk, V\. Mikulik, J\. Skalse, and S\. Garrabrant \(2021\)Risks from learned optimization in advanced machine learning systems\.External Links:1906\.01820,[Link](https://arxiv.org/abs/1906.01820)Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1)\.
- \[14\]U\. K”̈onig, H\. Kazmi, R\. Li, and M\. ChaudharyQuantifying subliminal behavioral transfer ratios in language model distillation\.InContinual Adaptation at Scale: Towards Sustainable AI Workshop,Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1)\.
- I\. Kelkar, N\. Alam, V\. Kakaria, M\. Panwar, V\. Sharma, and M\. Chaudhary \(2026\)Playing devil’s advocate: off\-the\-shelf persona vectors rival targeted steering for sycophancy\.arXiv preprint arXiv:2605\.21006\.Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Li, Y\. Shi, Z\. Liu, F\. Yang, A\. Payani, N\. Liu, and M\. Du \(2025\)Language ranker: a metric for quantifying llm performance across high and low\-resource languages\.InProceedings of the Thirty\-Ninth AAAI Conference on Artificial Intelligence and Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’25/IAAI’25/EAAI’25\.External Links:ISBN 978\-1\-57735\-897\-8,[Link](https://doi.org/10.1609/aaai.v39i27.35038),[Document](https://dx.doi.org/10.1609/aaai.v39i27.35038)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- P\. H\. Martins, P\. Fernandes, J\. Alves, N\. M\. Guerreiro, R\. Rei, D\. M\. Alves, J\. Pombal, A\. Farajian, M\. Faysse, M\. Klimaszewski, P\. Colombo, B\. Haddow, J\. G\. C\. de Souza, A\. Birch, and A\. F\. T\. Martins \(2024\)EuroLLM: multilingual language models for europe\.External Links:2409\.16235,[Link](https://arxiv.org/abs/2409.16235)Cited by:[§7](https://arxiv.org/html/2607.24769#S7.p1.1)\.
- M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang, N\. Mu, E\. Sakhaee, N\. Li, S\. Basart, B\. Li, D\. Forsyth, and D\. Hendrycks \(2024\)HarmBench: a standardized evaluation framework for automated red teaming and robust refusal\.External Links:2402\.04249,[Link](https://arxiv.org/abs/2402.04249)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Meinke, B\. Schoen, J\. Scheurer, M\. Balesni, R\. Shah, and M\. Hobbhahn \(2025\)Frontier models are capable of in\-context scheming\.External Links:2412\.04984,[Link](https://arxiv.org/abs/2412.04984)Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1),[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Nguyen, H\. H\. Khiem, C\. L\. Attubato, and F\. Hofstätter \(2025\)Probing evaluation awareness of language models\.InICML Workshop on Technical AI Governance \(TAIG\),External Links:[Link](https://openreview.net/forum?id=lerUefpec2)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- J\. R\. Nunez, V\. Sawant, N\. Allen, N\. Amgalanbaatar, Y\. Zongo, V\. Sharma, and M\. Chaudhary \(2026\)Mechanistic origins of catastrophic forgetting: why rl preserves circuits better than sft?\.arXiv preprint arXiv:2605\.28860\.Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2607.24769#S6.p2.1)\.
- D\. Patel, G\. Gervacio, D\. Raimi, K\. Zhu, R\. Lagasse, G\. Grand, A\. Panda, and M\. Chaudhary \(2025\)Alignment\-constrained dynamic pruning for llms: identifying and preserving alignment\-critical circuits\.arXiv preprint arXiv:2511\.07482\.Cited by:[§6](https://arxiv.org/html/2607.24769#S6.p2.1)\.
- M\. Phuong, M\. Aitchison, E\. Catt, S\. Cogan, A\. Kaskasoli, V\. Krakovna, D\. Lindner, M\. Rahtz, Y\. Assael, S\. Hodkinson, H\. Howard, T\. Lieberum, R\. Kumar, M\. A\. Raad, A\. Webson, L\. Ho, S\. Lin, S\. Farquhar, M\. Hutter, G\. Deletang, A\. Ruoss, S\. El\-Sayed, S\. Brown, A\. Dragan, R\. Shah, A\. Dafoe, and T\. Shevlane \(2024\)Evaluating frontier models for dangerous capabilities\.External Links:2403\.13793,[Link](https://arxiv.org/abs/2403.13793)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- T\. Shevlane, S\. Farquhar, B\. Garfinkel, M\. Phuong, J\. Whittlestone, J\. Leung, D\. Kokotajlo, N\. Marchal, M\. Anderljung, N\. Kolt, L\. Ho, D\. Siddarth, S\. Avin, W\. Hawkins, B\. Kim, I\. Gabriel, V\. Bolina, J\. Clark, Y\. Bengio, P\. Christiano, and A\. Dafoe \(2023\)Model evaluation for extreme risks\.External Links:2305\.15324,[Link](https://arxiv.org/abs/2305.15324)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- X\. Shi, X\. Wang, Z\. Guo, Y\. Wang, P\. Zhang, X\. Zhang, Z\. Guo, H\. Hao, Y\. Xi, B\. Yang, J\. Xu, J\. Zhou, and J\. Lin \(2026\)Qwen3\-asr technical report\.External Links:2601\.21337,[Link](https://arxiv.org/abs/2601.21337)Cited by:[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px2.p1.1)\.
- A\. Taubenfeld, Z\. Gekhman, L\. Nezry, O\. Feldman, N\. Harris, S\. Reddy, R\. Stella, A\. Goldstein, M\. Croak, Y\. Matias, and A\. Feder \(2026\)Evaluating alignment of behavioral dispositions in llms\.External Links:2602\.11328,[Link](https://arxiv.org/abs/2602.11328)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Üstün, V\. Aryabumi, Z\. Yong, W\. Ko, D\. D’souza, G\. Onilude, N\. Bhandari, S\. Singh, H\. Ooi, A\. Kayid, F\. Vargus, P\. Blunsom, S\. Longpre, N\. Muennighoff, M\. Fadaee, J\. Kreutzer, and S\. Hooker \(2024\)Aya model: an instruction finetuned open\-access multilingual language model\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15894–15939\.External Links:[Link](https://aclanthology.org/2024.acl-long.845/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.845)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- T\. van der Weij, F\. Hofstätter, O\. Jaffe, S\. F\. Brown, and F\. R\. Ward \(2025\)AI sandbagging: language models can strategically underperform on evaluations\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=7Qa2SpjxIS)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
- W\. Wang, Z\. Tu, C\. Chen, Y\. Yuan, J\. Huang, W\. Jiao, and M\. Lyu \(2024\)All languages matter: on the multilingual safety of LLMs\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5865–5877\.External Links:[Link](https://aclanthology.org/2024.findings-acl.349/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.349)Cited by:[§1](https://arxiv.org/html/2607.24769#S1.p1.1),[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Workshop, :, T\. L\. Scao, A\. Fan, C\. Akiki, E\. Pavlick, S\. Ilić, D\. Hesslow, R\. Castagné, A\. S\. Luccioni, F\. Yvon, M\. Gallé, J\. Tow, A\. M\. Rush, S\. Biderman, A\. Webson, P\. S\. Ammanamanchi, T\. Wang, B\. Sagot, N\. Muennighoff, A\. V\. del Moral, O\. Ruwase, R\. Bawden, S\. Bekman, A\. McMillan\-Major, I\. Beltagy, H\. Nguyen, L\. Saulnier, S\. Tan, P\. O\. Suarez, V\. Sanh, H\. Laurençon, Y\. Jernite, J\. Launay, M\. Mitchell, C\. Raffel, A\. Gokaslan, A\. Simhi, A\. Soroa, A\. F\. Aji, A\. Alfassy, A\. Rogers, A\. K\. Nitzav, C\. Xu, C\. Mou, C\. Emezue, C\. Klamm, C\. Leong, D\. van Strien, D\. I\. Adelani, D\. Radev, E\. G\. Ponferrada, E\. Levkovizh, E\. Kim, E\. B\. Natan, F\. D\. Toni, G\. Dupont, G\. Kruszewski, G\. Pistilli, H\. Elsahar, H\. Benyamina, H\. Tran, I\. Yu, I\. Abdulmumin, I\. Johnson, I\. Gonzalez\-Dios, J\. de la Rosa, J\. Chim, J\. Dodge, J\. Zhu, J\. Chang, J\. Frohberg, J\. Tobing, J\. Bhattacharjee, K\. Almubarak, K\. Chen, K\. Lo, L\. V\. Werra, L\. Weber, L\. Phan, L\. B\. allal, L\. Tanguy, M\. Dey, M\. R\. Muñoz, M\. Masoud, M\. Grandury, M\. Šaško, M\. Huang, M\. Coavoux, M\. Singh, M\. T\. Jiang, M\. C\. Vu, M\. A\. Jauhar, M\. Ghaleb, N\. Subramani, N\. Kassner, N\. Khamis, O\. Nguyen, O\. Espejel, O\. de Gibert, P\. Villegas, P\. Henderson, P\. Colombo, P\. Amuok, Q\. Lhoest, R\. Harliman, R\. Bommasani, R\. L\. López, R\. Ribeiro, S\. Osei, S\. Pyysalo, S\. Nagel, S\. Bose, S\. H\. Muhammad, S\. Sharma, S\. Longpre, S\. Nikpoor, S\. Silberberg, S\. Pai, S\. Zink, T\. T\. Torrent, T\. Schick, T\. Thrush, V\. Danchev, V\. Nikoulina, V\. Laippala, V\. Lepercq, V\. Prabhu, Z\. Alyafeai, Z\. Talat, A\. Raja, B\. Heinzerling, C\. Si, D\. E\. Taşar, E\. Salesky, S\. J\. Mielke, W\. Y\. Lee, A\. Sharma, A\. Santilli, A\. Chaffin, A\. Stiegler, D\. Datta, E\. Szczechla, G\. Chhablani, H\. Wang, H\. Pandey, H\. Strobelt, J\. A\. Fries, J\. Rozen, L\. Gao, L\. Sutawika, M\. S\. Bari, M\. S\. Al\-shaibani, M\. Manica, N\. Nayak, R\. Teehan, S\. Albanie, S\. Shen, S\. Ben\-David, S\. H\. Bach, T\. Kim, T\. Bers, T\. Fevry, T\. Neeraj, U\. Thakker, V\. Raunak, X\. Tang, Z\. Yong, Z\. Sun, S\. Brody, Y\. Uri, H\. Tojarieh, A\. Roberts, H\. W\. Chung, J\. Tae, J\. Phang, O\. Press, C\. Li, D\. Narayanan, H\. Bourfoune, J\. Casper, J\. Rasley, M\. Ryabinin, M\. Mishra, M\. Zhang, M\. Shoeybi, M\. Peyrounette, N\. Patry, N\. Tazi, O\. Sanseviero, P\. von Platen, P\. Cornette, P\. F\. Lavallée, R\. Lacroix, S\. Rajbhandari, S\. Gandhi, S\. Smith, S\. Requena, S\. Patil, T\. Dettmers, A\. Baruwa, A\. Singh, A\. Cheveleva, A\. Ligozat, A\. Subramonian, A\. Névéol, C\. Lovering, D\. Garrette, D\. Tunuguntla, E\. Reiter, E\. Taktasheva, E\. Voloshina, E\. Bogdanov, G\. I\. Winata, H\. Schoelkopf, J\. Kalo, J\. Novikova, J\. Z\. Forde, J\. Clive, J\. Kasai, K\. Kawamura, L\. Hazan, M\. Carpuat, M\. Clinciu, N\. Kim, N\. Cheng, O\. Serikov, O\. Antverg, O\. van der Wal, R\. Zhang, R\. Zhang, S\. Gehrmann, S\. Mirkin, S\. Pais, T\. Shavrina, T\. Scialom, T\. Yun, T\. Limisiewicz, V\. Rieser, V\. Protasov, V\. Mikhailov, Y\. Pruksachatkun, Y\. Belinkov, Z\. Bamberger, Z\. Kasner, A\. Rueda, A\. Pestana, A\. Feizpour, A\. Khan, A\. Faranak, A\. Santos, A\. Hevia, A\. Unldreaj, A\. Aghagol, A\. Abdollahi, A\. Tammour, A\. HajiHosseini, B\. Behroozi, B\. Ajibade, B\. Saxena, C\. M\. Ferrandis, D\. McDuff, D\. Contractor, D\. Lansky, D\. David, D\. Kiela, D\. A\. Nguyen, E\. Tan, E\. Baylor, E\. Ozoani, F\. Mirza, F\. Ononiwu, H\. Rezanejad, H\. Jones, I\. Bhattacharya, I\. Solaiman, I\. Sedenko, I\. Nejadgholi, J\. Passmore, J\. Seltzer, J\. B\. Sanz, L\. Dutra, M\. Samagaio, M\. Elbadri, M\. Mieskes, M\. Gerchick, M\. Akinlolu, M\. McKenna, M\. Qiu, M\. Ghauri, M\. Burynok, N\. Abrar, N\. Rajani, N\. Elkott, N\. Fahmy, O\. Samuel, R\. An, R\. Kromann, R\. Hao, S\. Alizadeh, S\. Shubber, S\. Wang, S\. Roy, S\. Viguier, T\. Le, T\. Oyebade, T\. Le, Y\. Yang, Z\. Nguyen, A\. R\. Kashyap, A\. Palasciano, A\. Callahan, A\. Shukla, A\. Miranda\-Escalada, A\. Singh, B\. Beilharz, B\. Wang, C\. Brito, C\. Zhou, C\. Jain, C\. Xu, C\. Fourrier, D\. L\. Periñán, D\. Molano, D\. Yu, E\. Manjavacas, F\. Barth, F\. Fuhrimann, G\. Altay, G\. Bayrak, G\. Burns, H\. U\. Vrabec, I\. Bello, I\. Dash, J\. Kang, J\. Giorgi, J\. Golde, J\. D\. Posada, K\. R\. Sivaraman, L\. Bulchandani, L\. Liu, L\. Shinzato, M\. H\. de Bykhovetz, M\. Takeuchi, M\. Pàmies, M\. A\. Castillo, M\. Nezhurina, M\. Sänger, M\. Samwald, M\. Cullan, M\. Weinberg, M\. D\. Wolf, M\. Mihaljcic, M\. Liu, M\. Freidank, M\. Kang, N\. Seelam, N\. Dahlberg, N\. M\. Broad, N\. Muellner, P\. Fung, P\. Haller, R\. Chandrasekhar, R\. Eisenberg, R\. Martin, R\. Canalli, R\. Su, R\. Su, S\. Cahyawijaya, S\. Garda, S\. S\. Deshmukh, S\. Mishra, S\. Kiblawi, S\. Ott, S\. Sang\-aroonsiri, S\. Kumar, S\. Schweter, S\. Bharati, T\. Laud, T\. Gigant, T\. Kainuma, W\. Kusa, Y\. Labrak, Y\. S\. Bajaj, Y\. Venkatraman, Y\. Xu, Y\. Xu, Y\. Xu, Z\. Tan, Z\. Xie, Z\. Ye, M\. Bras, Y\. Belkada, and T\. Wolf \(2023\)BLOOM: a 176b\-parameter open\-access multilingual language model\.External Links:2211\.05100,[Link](https://arxiv.org/abs/2211.05100)Cited by:[§7](https://arxiv.org/html/2607.24769#S7.p1.1)\.
- J\. Xu, Z\. Guo, H\. Hu, Y\. Chu, X\. Wang, J\. He, Y\. Wang, X\. Shi, T\. He, X\. Zhu, Y\. Lv, Y\. Wang, D\. Guo, H\. Wang, L\. Ma, P\. Zhang, X\. Zhang, H\. Hao, Z\. Guo, B\. Yang, B\. Zhang, Z\. Ma, X\. Wei, S\. Bai, K\. Chen, X\. Liu, P\. Wang, M\. Yang, D\. Liu, X\. Ren, B\. Zheng, R\. Men, F\. Zhou, B\. Yu, J\. Yang, L\. Yu, J\. Zhou, and J\. Lin \(2025\)Qwen3\-omni technical report\.External Links:2509\.17765,[Link](https://arxiv.org/abs/2509.17765)Cited by:[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px2.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2607.24769#S4.SS0.SSS0.Px2.p1.1),[§7](https://arxiv.org/html/2607.24769#S7.p1.1)\.
- Z\. X\. Yong, B\. Ermis, M\. Fadaee, S\. Bach, and J\. Kreutzer \(2025\)The state of multilingual LLM safety research: from measuring the language gap to mitigating it\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 15845–15860\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.800/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.800),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Yong, C\. Menghini, and S\. H\. Bach \(2024\)Low\-resource languages jailbreak gpt\-4\.External Links:2310\.02446,[Link](https://arxiv.org/abs/2310.02446)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson \(2023\)Universal and transferable adversarial attacks on aligned language models\.External Links:2307\.15043,[Link](https://arxiv.org/abs/2307.15043)Cited by:[§2](https://arxiv.org/html/2607.24769#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AAppendix
### A\.1Seed Instruction Examples
Figure 3:We provide examples of seed instructions being translated to several languages\.Similar Articles
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
This paper investigates whether pretraining LLMs on artificial languages (pre-pretraining) consistently improves token efficiency across multiple natural languages, finding that gains are highly dependent on experimental setup and random seed, though stable gains appear for small models with the Llama tokenizer.
Small LLMs: Pruning vs. Training from Scratch
This paper empirically compares pruning vs. training small language models from scratch, finding that pruning provides a strong advantage under limited token budgets but that the advantage diminishes as training scales, especially with coarse pruning.
Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding
This paper argues that designing advanced language representations to shape cognitive schemas is a key frontier for expanding LLM intelligence without scaling parameters. It provides formalizations and empirical evidence showing that different linguistic structures significantly impact model performance and internal feature activations.
Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs
This paper proposes a dense-to-sparse continual training method for LLMs, using a predictor-gated bank-wise sparsity to achieve 4x FFN sparsity, and demonstrates it on Qwen2.5-8B with long-context training.
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
This paper proposes using Evolution Strategies (ES) instead of Reinforcement Learning for post-training LLMs, showing that ES improves solution coverage (pass@k) and achieves better results on math benchmarks.