Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
Summary
This study compares probing techniques for identifying latent language in multilingual LLMs, finding that different methods yield inconsistent results, indicating they expose distinct aspects of multilingual processing rather than a single internal lingua franca.
View Cached Full Text
Cached at: 09/02/26, 05:46 AM
# Lingua Franca or Probing Artifact?Rethinking Latent Language in Multilingual LLMs
Source: [https://arxiv.org/html/2609.00155](https://arxiv.org/html/2609.00155)
Badr AlKhamissiAntoine BosselutAffiliation:EPFLAffiliation:Correspondence:\{deniz\.bayazit,badr\.alkhamissi,antoine\.bosselut\}@epfl\.ch
###### Abstract
Latent language identification is often used to argue that multilingual language models route computation through language\-specific states, such as English pivots\. However, existing probes infer latent language from different signals, such as the geometry of hidden states or what can be decoded from intermediate representations\. Since such claims shape conclusions about how models share and route information across languages, we ask whether these probes measure the same phenomenon or expose distinct aspects of multilingual computation\. We study this question across model families, training regimes, domains, tasks, checkpoints, and up to 27 languages\. We find that identification probes systematically disagree: the GMM\-based representation probe, which draws evidence from hidden state geometry, shows earlier cross\-lingual mixing, whereas decoding\-based probes, which rely on output\-space decodability, retain sharper language\-specific and more English\-biased signals\. These differences track model multilinguality and training progression, but are comparatively stable across domains\. Our results suggest a more cautious interpretation of latent language identification, where current probes expose different aspects of multilingual processing, rather than directly revealing a single internallingua franca\.111The code is available at: [https://github\.com/bayazitdeniz/latent\-lid](https://github.com/bayazitdeniz/latent-lid)
## 1Introduction
Figure 1:Do multilingual LLMs work in English? It depends on how you ask\.We probe the same intermediate hidden states of a multilingual LLM with three probes: logitlens, tuned lens, and a representation\-based GMM, and recover three different latent language distributions over the task’s languages \(tr,fr\), and other candidate languages like English \(en\)\. We study when and why these probes disagree across models, training regimes, and tasks\.Multilingual language models are often assumed to share information across languages by mapping them into common or partially language\-neutral representations\([Foroutan et al\., 2022](https://arxiv.org/html/2609.00155#bib.bib9);[Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25)\)\. Yet the form of this sharing remains poorly understood\. One influential hypothesis is that multilingual models rely on a*latent language*: an intermediate language\-like regime, sometimes described as an internal pivot orlingua franca, through which the model routes meaning before producing an output\.[Wendler et al\. \(2024\)](https://arxiv.org/html/2609.00155#bib.bib21)provide evidence for such behavior in English\-dominated models like LLaMA\-2\([Touvron et al\., 2023](https://arxiv.org/html/2609.00155#bib.bib17)\), but newer models trained on more balanced multilingual data may behave differently\([Üstün et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib22);[Foroutan et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib18);[Schut et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib2)\)\. This phenomenon raises a broader question:what part of a latent language estimate reflects the model’s internal computation, and what part reflects the assumptions of the diagnostic method?
Figure 2:Two families of latent language identification \(LLID\) probes\.Given an intermediate hidden statehh,decoding\-basedmethods \(left\) projecthhthrough the model’s unembedding \(optionally via a tuned lens\) and derive language scores from the resulting vocabulary distribution\.Representation\-basedmethods \(right\) instead comparehhwith language\-conditioned activation distributions, here a per\-layer GMM with one component per candidate language\. The two families operate on the samehhbut extract different evidence: output\-space decodability versus activation\-space geometry\.Interpreting such probes is challenging because latent language identification \(LLID\) is inferred rather than directly observed\. As illustrated in Fig\.[1](https://arxiv.org/html/2609.00155#S1.F1), a latent language estimate depends on what kind of evidence is extracted from the model\. Prior work relies on two broad families \(Fig\.[2](https://arxiv.org/html/2609.00155#S1.F2)\):representation\-basedprobes, which infer language from the geometry of hidden states\([Shani and Basirat, 2025](https://arxiv.org/html/2609.00155#bib.bib19)\), anddecoding\-basedprobes, which infer language by projecting intermediate hidden states through the model’s output head\([Wendler et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib21);[Zhong et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib20)\)\. These families make different assumptions about what it means for a model to internally use a language, and it is unclear whether they identify a common mechanism or expose different processes in multilingual encoding and decoding\.
In this work, we treat LLID as a measurement problem\. Instead of asking whether multilingual LLMs “think in English,” we ask whether latent language estimates reflect model properties, such as multilinguality and training progression, or diagnostic choices, such as representation geometry versus output\-space decodability\. To answer this, we compare a GMM\-based representation probe and decoding probes across controlled and open\-ended multilingual tasks\. We first test the internal consistency and generalization of decoding\-based LLID across controlled and open\-ended settings \(§[4](https://arxiv.org/html/2609.00155#S4)–[5](https://arxiv.org/html/2609.00155#S5)\), before asking whether representation\-based LLID recovers the same patterns under matched conditions \(§[6](https://arxiv.org/html/2609.00155#S6)\)\. We then trace how these estimates change across domains, training regimes, and checkpoints\.
Our analysis shows that LLID probes should not be treated as interchangeable measurements of a single phenomenon\. The GMM representation probe reveals earlier cross\-lingual mixing and weaker English dominance, while decoding\-based probes retain sharper language\-specific and more English\-biased signals\. This pattern recurs across our analyses, but the layerwise behavior of each probe is model\-dependent\. For a given probe, where these effects emerge in the layer stack and how they change across layers vary with model multilinguality and over the course of model pretraining, although multilinguality alone does not determine them\. By contrast, domain and language\-inventory changes shift LLID estimates only modestly\. Rather than supporting strong claims about an internallingua franca, each diagnostic provides a partial view of multilingual computation, shaped by the evidence it extracts from hidden states or decoded outputs\.
## 2A Framework for Comparing Latent Language Probes
### 2\.1Task Definition: Latent Language Identification
Letℒ\\mathcal\{L\}be a finite set of candidate languages, and leth∈ℝdh\\in\\mathbb\{R\}^\{d\}denote a hidden state extracted from a given model layer and sequence positions\. We represent latent language with a variableZ∈ℒZ\\in\\mathcal\{L\}and estimate its conditional distribution givenhhas:
q\(ℓ\)≈Pr\(Z=ℓ∣h\),ℓ∈ℒq\(\\ell\)\\approx\\Pr\(Z=\\ell\\mid h\),\\quad\\ell\\in\\mathcal\{L\}As the variableZZis not directly observed, we treatqqas a diagnostic estimate, not as ground\-truth evidence of an internal language variable\. Different LLID estimators may therefore produce different distributions for the same prompt and layer\. In practice, each estimator constructsqqby selecting one or more token positions and aggregating the resulting language evidence \(§[3](https://arxiv.org/html/2609.00155#S3)\)\.
### 2\.2Two Kinds of Evidence for LLID and Their Limits
Existing LLID methods differ mainly in the evidence they use to constructqq\. We distinguish two families\. Representation\-based methods infer language from where a hidden state lies in activation space relative to language\-conditioned structure learned from multilingual data\. Decoding\-based methods infer language from what can be read out from that hidden state through the model’s output vocabulary\.
#### Representation\-based LLID\.
This approach compares a hidden statehhto language\-conditioned structure learned from multilingual data, such as GMM posteriors\([Shani and Basirat, 2025](https://arxiv.org/html/2609.00155#bib.bib19)\), language subspaces\([Zhao et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib23)\), or neuron\-level indicators\([Zeng et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib26)\)\. This comparison yields per\-language scores, which are then aggregated into a layerwise distributionqq\.
In our experiments, we fit a per\-layer GMM with one language\-conditioned component per candidate language using hidden states from examples in each language, and use the resulting component posteriors as the layerwise language distributionqq\.
A common concern for this approach is that learned structure may conflate language identity with correlated factors such as script, tokenization, or domain\. In our experiments, however, we find consistent behavior across QA domains such as STEM and Arts & Humanities \(§[6](https://arxiv.org/html/2609.00155#S6.SS0.SSS0.Px1)\)\.
#### Decoding\-based LLID\.
These approaches estimate latent language by asking what language can be decoded from an intermediate layer\. Let𝒱\\mathcal\{V\}be the model’s vocabulary and letWU∈ℝ\|𝒱\|×dW\_\{U\}\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}be the model’s unembedding matrix\. Given any hidden stateh∈ℝdh\\in\\mathbb\{R\}^\{d\}, after final layer normalization, we form an intermediate vocabulary distribution:
p:=softmax\(WUh\)p:=\\mathrm\{softmax\}\(W\_\{U\}h\)Applying the model’s unmodified output head directly to an intermediate hidden state is known as the logitlens\([nostalgebraist, 2020](https://arxiv.org/html/2609.00155#bib.bib28)\), and we refer to this direct projection as the raw logitlens\. Because intermediate representations may not align with those expected by the final output head, we also consider the tuned lens\([Belrose et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib5)\), which learns a per\-layer affine transformation to predict the final\-layer hidden state from an intermediate one before applying the output head\.
Prior work mapsppto language scoresqqin several ways, such as summing probability over valid first\-token prefixes of language\-specific target words\([Wendler et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib21);[Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25)\), scoring the probability of multi\-token language\-specific answer sequences\([Zhong et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib20)\), or decoding one or more tokens fromppand classifying the result with an external language identification \(LID\) model\([Ozaki et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib27)\)\. In our experiments, for tasks where the completion is known, we sum the probability over valid first\-tokens, whereas if the task is open\-ended, we use two rollout\-based variants \(argmax and top\-pp\) where short continuations decoded from intermediate layers are passed to a LID classifier\([Kargaran et al\., 2023](https://arxiv.org/html/2609.00155#bib.bib3)\)\.
Decoding\-based LLID methods assume that the final unembedding remains meaningful at intermediate layers, which may not hold in earlier layers\. A further concern is that token\-level heuristics such as target\-string matching require non\-trivial design choices \(e\.g\., gathering valid completions\) to yield a meaningful distribution overℒ\\mathcal\{L\}\.
### 2\.3Metrics to Compare LLID Estimates
LLID estimates vary along several dimensions that cannot be captured by a single metric: distributions can be confident or diffuse, while their dominant language can match the task\-relevant language or pivot to another one\. Furthermore, two estimators can agree or disagree\. We therefore use a small set of complementary metrics\. Unless otherwise stated, each is computed per estimator and averaged over prompts at each layer\.
#### Entropy\.
Entropy,H=−∑ℓ∈ℒq\(ℓ\)logq\(ℓ\)H=\-\\sum\_\{\\ell\\in\\mathcal\{L\}\}q\(\\ell\)\\log q\(\\ell\), measures how diffuse the LLID distribution is; lower values indicate a peaky estimate, while higher values indicate mixing across languages\.
#### Dominance\.
We measure dominance withD=maxℓ∈ℒq\(ℓ\)D=\\max\_\{\\ell\\in\\mathcal\{L\}\}q\(\\ell\), which is high when most probability mass is assigned to a single language\.
#### Pivot rate\.
Pivot rate is the fraction of prompts whose dominant estimated language differs from the task\-relevant languageℓx\\ell\_\{x\}\(e\.g\., the source or target language in translation\):
Px=𝕀\[argmaxℓ∈ℒq\(ℓ\)≠ℓx\]P\_\{x\}=\\mathbb\{I\}\\big\[\\arg\\max\_\{\\ell\\in\\mathcal\{L\}\}q\(\\ell\)\\neq\\ell\_\{x\}\\big\]
#### Agreement\.
Given two LLID estimatorsa\{a\}andb\{b\}, agreement measures how often the estimators assign the same dominant language:
Ax=𝕀\[argmaxℓ∈ℒqa\(ℓ\)=argmaxℓ∈ℒqb\(ℓ\)\]A\_\{x\}=\\mathbb\{I\}\\big\[\\arg\\max\_\{\\ell\\in\\mathcal\{L\}\}q\_\{a\}\(\\ell\)=\\arg\\max\_\{\\ell\\in\\mathcal\{L\}\}q\_\{b\}\(\\ell\)\\big\]
In our experiments, this is used to compare decoding\-based and representation\-based LLIDs\.
## 3Experimental Setup
### 3\.1Models & Measuring Multilinguality
We evaluate up to 12 autoregressive language models spanning English\-centric baselines, multilingual base models, and instruction\-tuned variants\. Because training mixtures are not fully available, we consider the following multilinguality proxies: zero\-shot multilingual QA accuracy, cross\-lingual perplexity, and byte\-normalized likelihood\. Unless otherwise stated, cross\-model plots order models by increasing zero\-shot macro\-accuracy onINCLUDE\([Romanou et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib24)\)\. Model details, scores, and rankings are in Appendix[B](https://arxiv.org/html/2609.00155#A2)\.
### 3\.2Evaluation Regimes
We evaluate LLID in two settings: \(1\) a controlled regime where ground\-truth completions are available, and \(2\) an open\-ended regime that tests whether the same patterns hold in natural text\. The exact language inventory,i\.e\., the set of candidate languages included in each setting, is listed in Table[1](https://arxiv.org/html/2609.00155#A0.T1)in Appendix[A](https://arxiv.org/html/2609.00155#A1)\.
#### Controlled Target\-String Evaluation\.
We use three synthetic task formats: copy, cloze, and translation\. Copy asks the model to reproduce a concept word in the same language, cloze asks the model to complete a masked context, and translation asks for the concept word in a specified target language\. These examples build on prior target\-word settings for latent language analysis\([Wendler et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib21);[Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25);[Zhong et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib20)\)\. For each promptxxand candidate languageℓ\\ell, letwℓ,xw\_\{\\ell,x\}be the language\-ℓ\\ellanswer\. Following[Wendler et al\.](https://arxiv.org/html/2609.00155#bib.bib21), letStartτ\(wℓ,x\)⊆V\\textsc\{Start\}\_\{\\tau\}\(w\_\{\\ell,x\}\)\\subseteq Vdenote the set of vocabulary tokens under tokenizerτ\\tauthat can beginwℓ,xw\_\{\\ell,x\}\. For simplicity, we fixxx,ℓ\\ell, andτ\\tauand writeStart\(w\)\\textsc\{Start\}\(w\)\. The controlled decoding\-based language score is then defined as:
sx\(ℓ\)=∑t∈Start\(w\)Pθ\(xn\+1=t∣hn\),s\_\{x\}\(\\ell\)=\\sum\_\{\\mathclap\{t\\in\\textsc\{Start\}\(w\)\}\}P\_\{\\theta\}\(x\_\{n\+1\}=t\\mid h\_\{n\}\),
wherehnh\_\{n\}is the hidden state at the final prompt token\. We refer to this summed mass as theStart\(w\)\\textsc\{Start\}\(w\)probability\. When a language distribution is needed, we normalizesxs\_\{x\}over the candidate setℒ\\mathcal\{L\}to obtainqxq\_\{x\}\. Construction details forStart\(w\)\\textsc\{Start\}\(w\)and candidate language completions are in Appendix[A\.2](https://arxiv.org/html/2609.00155#A1.SS2)\.
#### Open\-Ended Evaluation\.
For natural\-language prompts without constrained completions, we summarize both representation\-based and decoding\-based LLID at the midpoint of the prompt and aggregate over a forward window of five positions\. We evaluate on PUD sentence data\([Zeman et al\., 2017](https://arxiv.org/html/2609.00155#bib.bib1)\)and INCLUDE question\-answering prompts\([Romanou et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib24)\), using PUD variants to vary the language inventory and INCLUDE domains to test domain sensitivity\. Exact dataset composition and additional evaluation details are in Appendix[A\.3](https://arxiv.org/html/2609.00155#A1.SS3)\.
Figure 3:Latent language probability rises in later layers in controlled translation; English is a stronger competing latent language in Llama\-2 than in Aya\-23 and Apertus\.For each translation prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same translation targets\. English is excluded as a source or target, but retained as a candidate latent language\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.
### 3\.3LLID Estimators
For open\-ended evaluation, all estimators produce a layerwise distributionqqover candidate languages\. We compare:
- •repr: a representation\-based GMM posterior fit to multilingual hidden states;
- •dec\-rmax: decoding\-based LLID using deterministic argmax rollouts from intermediate layers;
- •dec\-topp: decoding\-based LLID using top\-pprollouts\.
Fordec\-rmaxanddec\-topp, we generate short continuations from intermediate layers, pass them through GlotLID\([Kargaran et al\., 2023](https://arxiv.org/html/2609.00155#bib.bib3)\), and aggregate the resulting per\-language scores intoqq\. We evaluate both the raw logitlens and tuned lens defined in §[2\.2](https://arxiv.org/html/2609.00155#S2.SS2), separating the rollout procedure from the projection used to map intermediate hidden states into the output vocabulary\. Fitting details for GMMs, tuned lenses, and GlotLID are in Appendix[C](https://arxiv.org/html/2609.00155#A3)\.
Figure 4:Later\-layer latent language estimates depend strongly on the probe: raw logitlens assigns more probability to English, while tuned lens and GMM estimates favor task\-relevant languages\.Rows aggregate the 50–75% and 75–100% layer windows\.Left:translation prompts using normalizedStart\(w\)\\textsc\{Start\}\(w\)\{\}probabilities\.Right:open\-ended INCLUDE prompts using raw logitlens top\-pp, tuned lens top\-pp, and representation\-based GMM LLID\. Bars pool prompt–layer–language probabilities within each category before averaging; diamonds show the average maximum probability within the task\-relevant or other\-language category\. Error bars show SE\. Cases with English as source or target language are excluded\.
## 4The Controlled Case: Latent Language Under Known Completions
We begin by testing decoding\-based LLID in the controlled setting, where the set of valid completions is known across languages and we can directly track when language\-specific completions become decodable from intermediate layers\. We track the summedStart\(w\)\\textsc\{Start\}\(w\)\{\}probability assigned to language\-specific tokens of the same concept across layers for three models ordered by increasing multilingual capability: Llama\-2\-7B, Aya\-23\-8B, and Apertus\-8B\. Consistent with prior observations on Llama\-2\-7B, we find that across all model families the decoded language signal is weak in early layers and becomes most visible in later layers across translation \(Fig\.[3](https://arxiv.org/html/2609.00155#S3.F3)\), copy \(Fig\.[7](https://arxiv.org/html/2609.00155#A2.F7)\), and cloze tasks \(Fig\.[8](https://arxiv.org/html/2609.00155#A4.F8)\)\. However, the sharpness of this signal differs substantially across models with different degrees of multilinguality\. The effect is also not necessarily uniform across tasks and languages within a model\.
Across models, the copy task yields sharper and higher late\-layer probabilities for the observed language\. This is expected: copying can rely on reproducing a surface form that is already present in the context\. Cloze, on the other hand, requires retrieving a concept and lexicalizing it in the appropriate language222Expressing an underlying concept using the words/forms of a particular language\.\. Compared with copy, the cloze plots show lower peaks, more competing languages, and more diffuse probability mass, especially for Llama\-2 and Aya\-23\.
Even when English is excluded from the translation direction \(Fig\.[3](https://arxiv.org/html/2609.00155#S3.F3)\), Llama\-2 frequently maintains relatively high probability mass on English lexicalizations across layers, suggesting a stronger English\-centered bias in its internal representations\. This effect is weaker in Aya and further reduced in Apertus, where the target language probability rises relatively earlier and becomes more consistently dominant in later layers\.
Taken together, these results show thatStart\(w\)\\textsc\{Start\}\(w\)\{\}recovers a meaningful late\-layer language signal in the controlled setting\. The diagnostic is strongest when the task makes the relevant surface form directly available, as in copy prompts, and weaker when the model must retrieve and lexicalize a concept, as in cloze prompts\. The diagnosis also varies systematically with model multilinguality: less multilingual models more often retain high\-probability English alternatives, while the more multilingual Apertus\-8B shifts probability toward the intended non\-English target earlier\.
Figure 5:In open\-ended INCLUDE, latent language dominance and pivoting vary across models and probes\.Representation\-based estimates exhibit model\-specific changes across the layer stack, whereas decoding\-based estimates generally shift later, depending on model multilinguality\. For each prompt and layer, dominance is the highest probability assigned to a candidate latent language; a pivot occurs when the dominant language differs from the prompt language\. Rows are LLID estimators and columns are models\. Pivot rates are separated into English and other languages; the dashed line shows their sum\. INCLUDE contains no English prompts\.
## 5Generalizing Decoding\-based Estimates to Open\-Ended Generation
The controlled setting above provides a favorable testbed for theStart\(w\)\\textsc\{Start\}\(w\)\{\}measurement as the relevant cross\-lingual completions are known in advance, so the diagnostic can compare a fixed set of lexical alternatives\. However, forStart\(w\)\\textsc\{Start\}\(w\)\{\}to support claims about latent language or multilingual pivoting, its signal should generalize beyond controlled settings with predefined candidate answers\. We therefore ask whether similar language preferences can be recovered in open\-ended next\-token prediction, where the model generates continuations in context rather than selecting among predefined translations\.
We extend the decoding diagnostic to open\-ended generation by generating short continuations from intermediate hidden states under either the raw logitlens or tuned lens, using argmax or top\-ppdecoding, and then classifying the continuation language\. Fig\.[4](https://arxiv.org/html/2609.00155#S3.F4)shows that the projection method changes how probability mass is allocated across languages\. Raw logitlens decoding assigns substantially more mass to English than the controlledStart\(w\)\\textsc\{Start\}\(w\)\{\}setting, suggesting that English lexicalizations remain accessible when the model is not restricted to known translations\. For some models, raw logitlens outputs also become difficult to interpret, given the representational drift between intermediate and final\-layer hidden states\([Belrose et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib5)\)\. The tuned lens shifts average probability toward the task\-relevant language while average probability assigned to other languages remains low\. This suggests that the tuned lens may partly collapse intermediate language evidence toward the final decoded language, consistent with the concern raised by[Wendler et al\. \(2024\)](https://arxiv.org/html/2609.00155#bib.bib21)\.
Figure 6:Training dynamics of latent language estimates across pretraining on PUD21\.Rows are model families and columns are metrics\. Colors denote early, middle, and late layer bins; solid and dashed lines denoterepranddec\-topp, respectively\. Dominance is the average top\-language probability; confident pivot rate is the fraction of prompts where a non\-task language is dominant with probability at least 0\.6; agreement measures whether the two probes select the same dominant language\. Thexx\-axis is normalized training progress, spanning∼\{\\sim\}210B to 4T training tokens\. Layer bins use normalized layer indexj/Jj/J: earlyj/J<0\.25j/J<0\.25, middle0\.25≤j/J<0\.750\.25\\leq j/J<0\.75, and latej/J≥0\.75j/J\\geq 0\.75\. Most checkpoint\-dependent changes occur in late layers, while representation–decoding agreement remains low across training\.
## 6Representation vs\. Decoding: Different Evidence, Different LLID Behaviors
Having established how decoding\-based estimates vary across projections, rollouts, and task settings, we now compare them with representation\-based LLID on the same hidden states\. Representation\-based LLID provides a complementary view of the same intermediate states examined by decoding\-based probes\. Instead of asking what language can be decoded from a hidden state, it asks which language\-conditioned region of activation space the hidden state resembles\. This distinction matters because a hidden state may be geometrically closer to one language’s cluster without being decodable as that language, and conversely, an intermediate state may decode into a language whose representation\-space signature is not dominant\. Disagreement between the two estimators should therefore be treated as measurement information rather than as immediate diagnostic failure\. Such disagreement provides evidence that the probes capture different aspects of multilingual processing\.
Comparing the rightmost column of Fig\.[4](https://arxiv.org/html/2609.00155#S3.F4)\(Repr\-GMM\) with the decoding\-based columns on the same open\-ended INCLUDE prompts makes this contrast visible\. The representation probe assigns substantially more probability mass to the task\-relevant language and correspondingly less to English than raw logitlens top\-ppdecoding does, particularly in the 50–75% layer window\. This gap is most pronounced for models with moderate multilinguality \(e\.g\., Llama\-3\.1\), where decoding\-based estimates still show a strong English bar while the representation\-based estimate already favors the task languages\. For the most multilingual models \(e\.g\., Apertus, Apertus\-Instruct\), the average probabilities suggest closer agreement as both probes assign relatively little probability to English and substantial probability to the task\-relevant language in late layers\. However, the maximum\-language estimates also reveal that, under the representation probe, a single non\-task, non\-English language can still receive more probability than the task\-relevant language\. These patterns suggest that the two estimator families are sensitive to different aspects of the same hidden states\. The GMM representation probe reflects the geometric neighborhood of the input language’s activation, while decoding\-based estimators reflect the accessibility of language\-specific tokens through the unembedding matrix, which retains a stronger English bias even when the representation probe assigns substantially less probability to English\.
The precise layerwise patterns are also model\-dependent \(Fig\.[5](https://arxiv.org/html/2609.00155#S4.F5)\)\. For the representation probe, Llama\-2 exhibits a mixture of English\- and other\-language pivots, Aya\-23 shows a pronounced but temporary increase in English\-directed pivots, while Apertus\-8B pivots are almost entirely to non\-English languages\. OLMo\-2 shows another distinct pattern, maintaining high representation\-probe dominance across nearly the entire layer stack\. By contrast, dominance drops sharply near the start and recovers through the middle\-to\-late layers for Llama\-2 and Aya\-23, while Apertus\-8B remains low until a late\-layer rise\.
These patterns reflect again the different evidence used by the two probe families\. The representation probe is fit directly to input\-language activation distributions, so it is most confident where hidden states preserve a strong language\-conditioned geometric signature\. As states become more mixed or task\-oriented, they are less cleanly assigned to a single fitted language distribution, producing the early dominance drop in Llama\-2 and Aya\-23 and the low early\-to\-middle dominance in Apertus\. Their later recovery indicates that the states become concentrated again in a language\-conditioned region\. OLMo\-2 is the exception, as its high dominance suggests a persistent separation of language\-conditioned activation regions rather than a pronounced early\-to\-mid layer mixing phase\. Decoding\-based estimators, on the other hand, generally change much later in the network\.
#### Across Domains\.
We next ask whether the model and estimator\-specific trajectories in Fig\.[5](https://arxiv.org/html/2609.00155#S4.F5)persist across language inventories and QA subject domains\. Fig\.[16](https://arxiv.org/html/2609.00155#A4.F16)repeats the dominance and pivot\-rate analyses across PUD language\-set variants and INCLUDE domains\. For a given model and LLID estimator, the curves largely preserve the same layerwise shape across settings\. Thus, domain and language inventory changes affect absolute levels less than model or diagnostic choice\. The clearest systematic exception is candidate\-inventory size: PUD21 generally yields lower dominance and higher pivot rates than PUD9, as expected when probability mass is distributed over more possible languages\.
#### Across Pretraining\.
We ask whether latent language estimates appear suddenly, consolidate gradually, or drift as the model trains\. Fig\.[6](https://arxiv.org/html/2609.00155#S5.F6)compares OLMo\-2\-7B and Apertus\-8B across checkpoints \(see Table[3](https://arxiv.org/html/2609.00155#A2.T3)for checkpoint alignment\), tracking representation and decoding dominance, confident pivot rate, and estimator agreement in early, middle, and late layer bins\. We use confident pivot rate to avoid treating small fluctuations in diffuse LLID distributions as meaningful pivots, retaining only cases where a non\-task language is dominant with probability at least 0\.6\. The two models show qualitatively different trajectories\. For OLMo, representation\-based dominance is high and stable from the earliest checkpoint, while decoding dominance consolidates over the first∼\{\\sim\}20–30% of training; confident pivoting is low forreprbut moderate and slowly increasing fordec\-topp, consistent with English\-centric training where the representation space separates languages early but the decoding pathway retains an English bias\. For Apertus, both estimators undergo a sharp transition between roughly 20% and 40% of training, after which late\-layer dominance reaches∼\{\\sim\}0\.9\. Representation–decoding agreement remains low throughout training for both models, and is especially low for Apertus\. These dynamics suggest that most of the checkpoint\-dependent changes occur in the late\-layer bin for both models\. Overall, the checkpoint analysis shows that representation\- and decoding\-based estimates of LLID behavior can change at different points in training\.
#### Base vs\. Instruct\.
Finally, we ask whether instruction tuning increases English dominance in LLID estimates\. This hypothesis is plausible because instruction data is often less multilingual than pretraining data, so instruction tuning might reorient intermediate states or decoded continuations toward English\. In Fig\.[4](https://arxiv.org/html/2609.00155#S3.F4), base and instruction\-tuned variants show only modest differences in their average LLID probabilities\. The layerwise comparisons in the Appendix \(Figs\.[20](https://arxiv.org/html/2609.00155#A4.F20)–[21](https://arxiv.org/html/2609.00155#A4.F21)\) likewise show broadly similar pivot and entropy trajectories, although some model\-specific deviations remain\. Contrary to our initial expectation, instruction tuning produces little systematic shift toward English in LLID estimates, and its overall effects are considerably smaller than the changes observed across pretraining checkpoints\.
## 7Related Work
#### Shared and language\-specific multilingualstructure\.
Multilingual models maintain both shared and language\-specific structure in hidden states, through related but distinguishable subspaces\([Chang et al\., 2022](https://arxiv.org/html/2609.00155#bib.bib8)\), language\-neutral subnetworks\([Foroutan et al\., 2022](https://arxiv.org/html/2609.00155#bib.bib9)\), and language\-specific neurons or components\([Zeng et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib26);[Zhao et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib23)\)\. These findings provide a basis for representation\-based LLID, while highlighting that hidden states can simultaneously encode shared and language\-specific structure\.
#### Latent pivots and output language\.
A second line of work studies whether multilingual models internally route computation through a pivot language\. Decoding\-based probes find English\-pivot behavior in English\-dominated models and different pivoting patterns under non\-English\-centric training\([Wendler et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib21);[Zhong et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib20)\)\. Causal interventions further separate output language from conceptual content, suggesting that language form and meaning can be partially disentangled in the residual stream\([Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25)\)\. Our work differs by asking whether the probes used to infer such latent language behavior agree on the same hidden states\.
#### Training dynamics of multilingual representations\.
Recent work studies how multilingual abilities emerge during training rather than only at the final checkpoint, showing that cross\-lingual abilities\([Blevins et al\., 2022](https://arxiv.org/html/2609.00155#bib.bib6)\)and internal linguistic features\([Bayazit et al\., 2026](https://arxiv.org/html/2609.00155#bib.bib7)\)develop, stabilize, or change over pretraining\. Accordingly, we examine LLID across training checkpoints: if latent language estimates reflect model properties rather than diagnostic artifacts, they should vary systematically with training progression\.
## 8Conclusion
We reframed latent language identification as a measurement problem, comparing representation\-based and decoding\-based probes across models, training regimes, and tasks\. The GMM representation probe and decoding\-based probes disagree systematically: the former identifies language structure earlier and with weaker English bias, while decoding pathways retain sharper, more English\-favored signals\. These differences vary systematically with multilinguality and training progression, but their exact layerwise trajectories remain model and lens\-dependent and cannot be predicted from multilinguality alone\. Domain and language\-inventory changes, by contrast, shift LLID estimates only modestly\. Rather than confirming a single internallingua franca, existing probes expose complementary aspects of multilingual computation, showing that the choice of probe is itself part of the claim\.
## Acknowledgements
We thank Clara Meister, Negar Foroutan, Angelika Romanou, and Ayush K\. Tarun for their helpful discussions and feedback on our manuscript\. We also gratefully acknowledge the support of the Swiss National Science Foundation \(No\. 215390\), the AI2050 program at Schmidt Sciences \(Grant \#G\-25\-69783\), Sony Group Corporation, and the Swiss National Supercomputing Center \(CSCS\) in the form of an infrastructure engineering and development project\. This research is with support from Google\.org and the Google Cloud Research Credits program for the Gemini Academic Program\. This project is also funded by the European Union \(ERC, RESPECT\-LM, 101222478\)\. Views and opinions expressed are however those of the author\(s\) only and do not necessarily reflect those of the European Union or the European Research Council\. Neither the European Union nor the granting authority can be held responsible for them\.
## Limitations
Our findings come with several limitations\. We evaluate decoder\-only models in the 7–9B parameter range, and our checkpoint analysis is necessarily restricted to OLMo\-2 and Apertus, the two model families that release intermediate training checkpoints publicly\. Whether probe disagreement persists at other scales, or for non\-decoder architectures, remains an open question\. The precise layerwise LLID patterns also vary across model families, so the trajectories observed here should not be assumed to generalize unchanged to other models\. We compare only two probe families, of which the representation\-based family is instantiated only with the GMM estimator\([Shani and Basirat, 2025](https://arxiv.org/html/2609.00155#bib.bib19)\), and do not examine causal interventions\([Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25)\)or neuron\-level attribution\([Zeng et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib26)\), which may align with either family or expose a third aspect of multilingual processing\. Because latent language is inferred rather than observed \(i\.e\., no ground truth\), our work documents probe disagreement but cannot adjudicate which \(if any\) best reflects internal computation\. Probe\-specific choices \(e\.g\., GMM components, the tuned lens fitting corpus, the open\-ended anchor position, and the concept\-word construction\) may shift absolute LLID values, though we expect qualitative trends to be robust\. Finally, our 27\-language inventory excludes very low\-resource languages and formal languages such as code and mathematics\. We also leave open whether probe disagreement predicts downstream failure modes in cross\-lingual transfer or generation\.
## Ethics Statement
This work uses publicly available pretrained language models and multilingual evaluation datasets \(PUD/UD, INCLUDE, and Fineweb2\), used in accordance with their respective licenses\. Our analysis is diagnostic in nature and does not involve training new models, collecting human data, or deploying systems in user\-facing settings\. We see no direct ethical concerns arising from this work\.
## References
- Aryabumiet al\.\(2024\)V\. Aryabumi, J\. Dang, D\. Talupuru, S\. Dash, D\. Cairuz, H\. Lin, B\. Venkitesh, M\. Smith, K\. Marchisio, S\. Ruder, A\. Locatelli, J\. Kreutzer, N\. Frosst, P\. Blunsom, M\. Fadaee, A\. Üstün, and S\. HookerAya 23: open weight releases to further multilingual progress\.External Links:2405\.15032Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.9.1)\.
- Bayazitet al\.\(2026\)D\. Bayazit, A\. Mueller, and A\. BosselutCrosscoding through time: tracking emergence & consolidation of linguistic representations throughout LLM pretraining\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 1353–1377\.External Links:[Link](https://aclanthology.org/2026.acl-long.60/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.60),ISBN 979\-8\-89176\-390\-6Cited by:[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px3.p1.1)\.
- Belroseet al\.\(2025\)N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. SteinhardtEliciting latent predictions from transformers with the tuned lens\.External Links:2303\.08112,[Link](https://arxiv.org/abs/2303.08112)Cited by:[Appendix C](https://arxiv.org/html/2609.00155#A3.SS0.SSS0.Px3.p1.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p1.2),[§5](https://arxiv.org/html/2609.00155#S5.p2.1)\.
- Blevinset al\.\(2022\)T\. Blevins, H\. Gonen, and L\. ZettlemoyerAnalyzing the mono\- and cross\-lingual pretraining dynamics of multilingual language models\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 3575–3590\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.234/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.234)Cited by:[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px3.p1.1)\.
- Changet al\.\(2022\)T\. A\. Chang, Z\. Tu, and B\. K\. BergenThe geometry of multilingual language model representations\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 119–136\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.9/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.9)Cited by:[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px1.p1.1)\.
- Dumaset al\.\(2025\)C\. Dumas, C\. Wendler, V\. Veselovsky, G\. Monea, and R\. WestSeparating tongue from thought: activation patching reveals language\-agnostic concept representations in transformers\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 31822–31841\.External Links:[Link](https://aclanthology.org/2025.acl-long.1536/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1536),ISBN 979\-8\-89176\-251\-0Cited by:[§A\.2](https://arxiv.org/html/2609.00155#A1.SS2.p1.1),[§1](https://arxiv.org/html/2609.00155#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p2.1),[§3\.2](https://arxiv.org/html/2609.00155#S3.SS2.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px2.p1.1),[Limitations](https://arxiv.org/html/2609.00155#Sx2.p1.1)\.
- Foroutanet al\.\(2022\)N\. Foroutan, M\. Banaei, R\. Lebret, A\. Bosselut, and K\. AbererDiscovering language\-neutral sub\-networks in multilingual language models\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 7560–7575\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.513/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.513)Cited by:[§1](https://arxiv.org/html/2609.00155#S1.p1.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px1.p1.1)\.
- Foroutanet al\.\(2025\)N\. Foroutan, P\. Teiletche, A\. K\. Tarun, and A\. BosselutRevisiting multilingual data mixtures in language model pretraining\.External Links:2510\.25947,[Link](https://arxiv.org/abs/2510.25947)Cited by:[§1](https://arxiv.org/html/2609.00155#S1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.6.1)\.
- Hernández\-Canoet al\.\(2026\)A\. Hernández\-Cano, A\. Hägele, A\. H\. Huang, A\. Romanou, A\. Solergibert, B\. Pásztor, B\. Messmer, D\. Garbaya, E\. F\. Ďurech, I\. Hakimi, J\. G\. Giraldo, M\. Ismayilzada, N\. Foroutan, S\. Moalla, T\. Chen, V\. Sabolčec, Y\. Xu, M\. Aerni, B\. AlKhamissi, I\. A\. Marinas, M\. H\. Amani, M\. Ansaripour, I\. Badanin, H\. Benoit, E\. Boros, N\. J\. Browning, F\. Bösch, M\. Böther, N\. Canova, C\. Challier, C\. Charmillot, J\. Coles, J\. M\. Deriu, A\. Devos, L\. Drescher, D\. Dzenhaliou, M\. Ehrmann, D\. Fan, S\. Fan, S\. Gao, M\. Gila, M\. Grandury, D\. Hashemi, A\. M\. Hoyle, J\. Jiang, M\. Klein, A\. Kucharavy, A\. Kucherenko, F\. Lübeck, R\. Machacek, T\. I\. Manitaras, A\. Marfurt, K\. Matoba, S\. Matrenok, H\. Mendonça, F\. R\. Mohamed, S\. Montariol, L\. Mouchel, S\. Najem\-Meyer, J\. Ni, G\. Oliva, M\. Pagliardini, E\. Palme, A\. Panferov, L\. Paoletti, M\. Passerini, I\. Pavlov, A\. Poiroux, K\. Ponkshe, N\. Ranchin, J\. Rando, M\. Sauser, J\. Saydaliev, M\. Sayfiddinov, M\. Schneider, S\. Schuppli, M\. Scialanga, A\. Semenov, K\. Shridhar, R\. Singhal, A\. Sotnikova, A\. Sternfeld, A\. K\. Tarun, P\. Teiletche, J\. Vamvas, X\. Yao, H\. Zhao, A\. Ilic, A\. Klimovic, A\. Krause, C\. Gulcehre, D\. Rosenthal, E\. Ash, F\. Tramèr, J\. VandeVondele, L\. Veraldi, M\. Rajman, T\. C\. Schulthess, T\. Hoefler, A\. Bosselut, M\. Jaggi, and I\. SchlagApertus: democratizing open and compliant LLMs for global language environments\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 46877–46955\.External Links:[Link](https://aclanthology.org/2026.acl-long.2172/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2172),ISBN 979\-8\-89176\-390\-6Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.3.1)\.
- Kargaranet al\.\(2023\)A\. H\. Kargaran, A\. Imani, F\. Yvon, and H\. SchuetzeGlotLID: language identification for low\-resource languages\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6155–6218\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.410/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.410)Cited by:[§A\.3](https://arxiv.org/html/2609.00155#A1.SS3.p1.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p2.1),[§3\.3](https://arxiv.org/html/2609.00155#S3.SS3.p1.2)\.
- Martinset al\.\(2025\)P\. H\. Martins, J\. Alves, P\. Fernandes, N\. M\. Guerreiro, R\. Rei, A\. Farajian, M\. Klimaszewski, D\. M\. Alves, J\. Pombal, N\. Boizard, M\. Faysse, P\. Colombo, F\. Yvon, B\. Haddow, J\. G\. C\. de Souza, A\. Birch, and A\. F\. T\. MartinsEuroLLM\-9b: technical report\.External Links:2506\.04079,[Link](https://arxiv.org/abs/2506.04079)Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.8.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting GPT: the logit lens\. lesswrong, 2020\.\.External Links:[Link](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p1.2)\.
- Ozakiet al\.\(2025\)S\. Ozaki, T\. Hiraoka, H\. Otake, H\. Ouchi, M\. Isonuma, B\. Heinzerling, K\. Inui, T\. Watanabe, Y\. Miyao, Y\. Oseki, and Y\. TakagiDo llms need to think in one language? correlation between latent language and task performance\.External Links:2505\.21458,[Link](https://arxiv.org/abs/2505.21458)Cited by:[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p2.1)\.
- Penedoet al\.\(2025\)G\. Penedo, H\. Kydlíček, V\. Sabolčec, B\. Messmer, N\. Foroutan, A\. H\. Kargaran, C\. Raffel, M\. Jaggi, L\. V\. Werra, and T\. WolfFineWeb2: one pipeline to scale them all — adapting pre\-training data processing to every language\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=jnRBe6zatP)Cited by:[Appendix C](https://arxiv.org/html/2609.00155#A3.SS0.SSS0.Px3.p1.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.12.1)\.
- Romanouet al\.\(2025\)A\. Romanou, N\. Foroutan, A\. Sotnikova, S\. H\. Nelaturu, S\. Singh, R\. Maheshwary, M\. Altomare, Z\. Chen, M\. Haggag, S\. A, A\. Amayuelas, A\. H\. Amirudin, D\. Boiko, M\. Chang, J\. Chim, G\. Cohen, A\. K\. Dalmia, A\. Diress, S\. Duwal, D\. Dzenhaliou, D\. Florez, F\. Farestam, J\. M\. Imperial, S\. Islam, P\. Isotalo, M\. Jabbarishiviari, B\. F\. Karlsson, E\. Khalilov, C\. Klamm, F\. Koto, D\. Krzemiński, G\. de Melo, S\. Montariol, Y\. Nan, J\. Niklaus, J\. Novikova, J\. S\. Obando Ceron, D\. Paul, E\. Ploeger, J\. Purbey, S\. Rajwal, S\. S\. Ravi, S\. Rydell, R\. Santhosh, D\. Sharma, M\. Prifti Skenduli, A\. Soltani Moakhar, B\. moakhar, A\. Tarun, A\. T\. Wasi, T\. Weerasinghe, S\. Yilmaz, M\. Zhang, I\. Schlag, M\. Fadaee, S\. Hooker, and A\. BosselutINCLUDE: evaluating multilingual language understanding with regional knowledge\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 83291–83322\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf)Cited by:[§3\.1](https://arxiv.org/html/2609.00155#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.00155#S3.SS2.SSS0.Px2.p1.1)\.
- Schutet al\.\(2025\)L\. Schut, Y\. Gal, and S\. FarquharDo multilingual LLMs think in english?\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,External Links:[Link](https://openreview.net/forum?id=I8BOtOPcOv)Cited by:[§1](https://arxiv.org/html/2609.00155#S1.p1.1)\.
- Shani and Basirat \(2025\)N\. Shani and A\. BasiratLanguage dominance in multilingual large language models\.InProceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Y\. Belinkov, A\. Mueller, N\. Kim, H\. Mohebbi, H\. Chen, D\. Arad, and G\. Sarti \(Eds\.\),Suzhou, China,pp\. 137–148\.External Links:[Link](https://aclanthology.org/2025.blackboxnlp-1.7/),[Document](https://dx.doi.org/10.18653/v1/2025.blackboxnlp-1.7),ISBN 979\-8\-89176\-346\-3Cited by:[§A\.3](https://arxiv.org/html/2609.00155#A1.SS3.p1.1),[§1](https://arxiv.org/html/2609.00155#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px1.p1.1),[Limitations](https://arxiv.org/html/2609.00155#Sx2.p1.1)\.
- team \(2024\)M\. A\. teamMistral nemo\.Note:https://mistral\.ai/news/mistral\-nemoAccessed: 2026Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.5.1)\.
- Touvronet al\.\(2023\)H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale, D\. Bikel, L\. Blecher, C\. C\. Ferrer, M\. Chen, G\. Cucurull, D\. Esiobu, J\. Fernandes, J\. Fu, W\. Fu, B\. Fuller, C\. Gao, V\. Goswami, N\. Goyal, A\. Hartshorn, S\. Hosseini, R\. Hou, H\. Inan, M\. Kardas, V\. Kerkez, M\. Khabsa, I\. Kloumann, A\. Korenev, P\. S\. Koura, M\. Lachaux, T\. Lavril, J\. Lee, D\. Liskovich, Y\. Lu, Y\. Mao, X\. Martinet, T\. Mihaylov, P\. Mishra, I\. Molybog, Y\. Nie, A\. Poulton, J\. Reizenstein, R\. Rungta, K\. Saladi, A\. Schelten, R\. Silva, E\. M\. Smith, R\. Subramanian, X\. E\. Tan, B\. Tang, R\. Taylor, A\. Williams, J\. X\. Kuan, P\. Xu, Z\. Yan, I\. Zarov, Y\. Zhang, A\. Fan, M\. Kambadur, S\. Narang, A\. Rodriguez, R\. Stojnic, S\. Edunov, and T\. ScialomLlama 2: open foundation and fine\-tuned chat models\.External Links:2307\.09288,[Link](https://arxiv.org/abs/2307.09288)Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.11.1),[§1](https://arxiv.org/html/2609.00155#S1.p1.1)\.
- Üstünet al\.\(2024\)A\. Üstün, V\. Aryabumi, Z\. Yong, W\. Ko, D\. D’souza, G\. Onilude, N\. Bhandari, S\. Singh, H\. Ooi, A\. Kayid, F\. Vargus, P\. Blunsom, S\. Longpre, N\. Muennighoff, M\. Fadaee, J\. Kreutzer, and S\. HookerAya model: an instruction finetuned open\-access multilingual language model\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15894–15939\.External Links:[Link](https://aclanthology.org/2024.acl-long.845/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.845)Cited by:[§1](https://arxiv.org/html/2609.00155#S1.p1.1)\.
- Walshet al\.\(2025\)E\. P\. Walsh, L\. Soldaini, D\. Groeneveld, K\. Lo, S\. Arora, A\. Bhagia, Y\. Gu, S\. Huang, M\. Jordan, N\. Lambert, D\. Schwenk, O\. Tafjord, T\. Anderson, D\. Atkinson, F\. Brahman, C\. Clark, P\. Dasigi, N\. Dziri, A\. Ettinger, M\. Guerquin, D\. Heineman, H\. Ivison, P\. W\. Koh, J\. Liu, S\. Malik, W\. Merrill, L\. J\. V\. Miranda, J\. Morrison, T\. Murray, C\. Nam, J\. Poznanski, V\. Pyatkin, A\. Rangapur, M\. Schmitz, S\. Skjonsberg, D\. Wadden, C\. Wilhelm, M\. Wilson, L\. Zettlemoyer, A\. Farhadi, N\. A\. Smith, and H\. Hajishirzi2 OLMo 2 furious \(COLM’s version\)\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=2ezugTT9kU)Cited by:[Table 2](https://arxiv.org/html/2609.00155#A2.T2.2.1.10.1)\.
- Wendleret al\.\(2024\)C\. Wendler, V\. Veselovsky, G\. Monea, and R\. WestDo llamas work in English? on the latent language of multilingual transformers\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15366–15394\.External Links:[Link](https://aclanthology.org/2024.acl-long.820/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.820)Cited by:[§A\.2](https://arxiv.org/html/2609.00155#A1.SS2.p1.1),[§1](https://arxiv.org/html/2609.00155#S1.p1.1),[§1](https://arxiv.org/html/2609.00155#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p2.1),[§3\.2](https://arxiv.org/html/2609.00155#S3.SS2.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.00155#S5.p2.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px2.p1.1)\.
- Zemanet al\.\(2017\)D\. Zeman, M\. Popel, M\. Straka, J\. Hajič, J\. Nivre, F\. Ginter, J\. Luotolahti, S\. Pyysalo, S\. Petrov, M\. Potthast, F\. Tyers, E\. Badmaeva, M\. Gokirmak, A\. Nedoluzhko, S\. Cinková, J\. Hajič jr\., J\. Hlaváčová, V\. Kettnerová, Z\. Urešová, J\. Kanerva, S\. Ojala, A\. Missilä, C\. D\. Manning, S\. Schuster, S\. Reddy, D\. Taji, N\. Habash, H\. Leung, M\. de Marneffe, M\. Sanguinetti, M\. Simi, H\. Kanayama, V\. de Paiva, K\. Droganova, H\. Martínez Alonso, Ç\. Çöltekin, U\. Sulubacak, H\. Uszkoreit, V\. Macketanz, A\. Burchardt, K\. Harris, K\. Marheinecke, G\. Rehm, T\. Kayadelen, M\. Attia, A\. Elkahky, Z\. Yu, E\. Pitler, S\. Lertpradit, M\. Mandl, J\. Kirchner, H\. F\. Alcalde, J\. Strnadová, E\. Banerjee, R\. Manurung, A\. Stella, A\. Shimada, S\. Kwak, G\. Mendonça, T\. Lando, R\. Nitisaroj, and J\. LiCoNLL 2017 shared task: multilingual parsing from raw text to Universal Dependencies\.InProceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies,J\. Hajič and D\. Zeman \(Eds\.\),Vancouver, Canada,pp\. 1–19\.External Links:[Link](https://aclanthology.org/K17-3001/),[Document](https://dx.doi.org/10.18653/v1/K17-3001)Cited by:[§3\.2](https://arxiv.org/html/2609.00155#S3.SS2.SSS0.Px2.p1.1)\.
- Zenget al\.\(2025\)H\. Zeng, S\. Han, L\. Chen, and K\. YuConverging to a lingua franca: evolution of linguistic regions and semantics alignment in multilingual large language models\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 10602–10617\.External Links:[Link](https://aclanthology.org/2025.coling-main.707/)Cited by:[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px1.p1.1),[Limitations](https://arxiv.org/html/2609.00155#Sx2.p1.1)\.
- Zhaoet al\.\(2025\)W\. Zhao, J\. Guo, Y\. Deng, T\. Wu, W\. Zhang, Y\. Hu, X\. Sui, Y\. Zhao, W\. Che, B\. Qin, T\. Chua, and T\. LiuWhen less language is more: language\-reasoning disentanglement makes llms better multilingual reasoners\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 38608–38642\.External Links:[Document](https://dx.doi.org/10.52202/085713-1290),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/372bd0e47f2d5bceca7e300e1446849c-Paper-Conference.pdf)Cited by:[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px1.p1.1)\.
- Zhonget al\.\(2025\)C\. Zhong, Q\. Liu, F\. Cheng, J\. Jiang, Z\. Wan, C\. Chu, Y\. Murawaki, and S\. KurohashiWhat language do non\-English\-centric large language models think in?\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 26333–26346\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1350/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1350),ISBN 979\-8\-89176\-256\-5Cited by:[§A\.2](https://arxiv.org/html/2609.00155#A1.SS2.p1.1),[§1](https://arxiv.org/html/2609.00155#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.00155#S2.SS2.SSS0.Px2.p2.1),[§3\.2](https://arxiv.org/html/2609.00155#S3.SS2.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.00155#S7.SS0.SSS0.Px2.p1.1)\.
Setting\# Langs\# EntriesLanguagesPUD99900 train \+ 900 held\-outArabic \(ar\), Czech \(cs\), English \(en\), French \(fr\), Hindi \(hi\), Icelandic \(is\), Indonesian \(id\), Portuguese \(pt\), Spanish \(es\)PUD21212,100 train \+ 2,100 held\-outArabic \(ar\), Russian \(ru\), Hindi \(hi\), Chinese \(zh\), Korean \(ko\), Japanese \(ja\), Indonesian \(id\), German \(de\), English \(en\), Icelandic \(is\), Swedish \(sv\), Spanish \(es\), French \(fr\), Galician \(gl\), Italian \(it\), Portuguese \(pt\), Czech \(cs\), Polish \(pl\), Turkish \(tr\), Finnish \(fi\), Thai \(th\)UD6 extension6600 train \+ 600 held\-outUkrainian \(uk\), Bulgarian \(bg\), Serbian \(sr\), Urdu \(ur\), Persian \(fa\), Marathi \(mr\)INCLUDE\-1010882 promptsArabic \(90\), Spanish \(90\), Finnish \(90\), French \(81\), Hindi \(90\), Indonesian \(90\), Portuguese \(90\), Russian \(90\), Turkish \(90\), Chinese \(81\)Copy/cloze6414 prompts per taskArabic \(ar\), Hindi \(hi\), Chinese \(zh\), Russian \(ru\), English \(en\), French \(fr\)Translation916,146 promptsArabic \(ar\), Spanish \(es\), Finnish \(fi\), French \(fr\), Hindi \(hi\), Indonesian \(id\), Russian \(ru\), Turkish \(tr\), Chinese \(zh\)
Table 1:Language inventories used in the experiments\.PUD and UD rows are balanced at 100 train and 100 held\-out prompts per language\. PUD9\+UD6 and PUD21\+UD6 add the UD6 extension to the corresponding PUD inventory\. The INCLUDE subset is capped at 30 prompts per language\-domain cell, but French and Chinese have fewer total prompts because their STEM cells contain only 21 selected examples\. Target\-string copy and cloze counts are for the six\-language main\-plot subset; translation counts aggregate the nine target\-specific runs\.## Appendix ADataset Details
In this section we describe the datasets and the preprocessing steps\. Note that we did not collect these datasets ourselves; they may contain personally identifying or offensive content\.
### A\.1Language Inventory
Table[1](https://arxiv.org/html/2609.00155#A0.T1)gives the language inventories used across the controlled target\-string runs, PUD runs, and INCLUDE runs\. In the main text, we use the shorthand PUD9, PUD21, PUD9\+UD6, PUD21\+UD6 to avoid repeatedly listing these languages\. Entry counts refer to prepared prompt rows; for PUD and UD settings, we report the deterministic train and held\-out splits separately\. The PUD9\+UD6 and PUD21\+UD6 settings are formed by adding the UD6 extension row to PUD9 or PUD21\.
### A\.2Controlled Target\-String Construction
The controlled copy, cloze, and translation prompts are derived from the target\-word setup of prior latent language studies\([Wendler et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib21);[Dumas et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib25);[Zhong et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib20)\)\. Each example has a known concept completion in multiple languages, and the logitlens query is made at the final task token, before the answer begins\. For instruction\-tuned models, formatting may append assistant\-control tokens, so we anchor scoring to the end of the original task prefix rather than the final formatted token\.
We first build a common\-concept grid and then derive copy, cloze, and directed translation prompts from that grid\. Copy and translation prompts use four completed in\-context examples followed by a query line; cloze prompts use two completed cloze demonstrations followed by the query context\. For translation and cloze settings, word translations and cloze contexts are constructed jointly so that the context disambiguates homonyms\. We filter examples where the translated cloze answer does not align with the independently translated concept word\. We do not remove examples solely because the English and target language forms share a tokenizer prefix\.
For each remaining promptxx, candidate languageℓ\\ell, and model tokenizerτ\\tau, we constructStart\(w\)\\textsc\{Start\}\(w\), the set of vocabulary tokens that can begin the language\-ℓ\\ellanswerwℓ,xw\_\{\\ell,x\}, as defined in §[3\.2](https://arxiv.org/html/2609.00155#S3.SS2)\. This set includes vocabulary tokens with tokenizer\-specific prefixes and whitespace variants when applicable\. For tokenizers with byte\-fallback behavior, we also include valid byte\-level starting tokens\.
The corresponding scoresx\(ℓ\)s\_\{x\}\(\\ell\), also defined in §[3\.2](https://arxiv.org/html/2609.00155#S3.SS2), is the summed next\-token probability assigned to these valid starting tokens\. We use this raw summed mass for the probability curves\. When computing distributional metrics such as entropy, dominance, or pivot rate, we normalize these scores over the candidate language set:
qx\(ℓ\)=sx\(ℓ\)∑ℓ′∈ℒsx\(ℓ′\)q\_\{x\}\(\\ell\)=\\frac\{s\_\{x\}\(\\ell\)\}\{\\sum\_\{\\ell^\{\\prime\}\\in\\mathcal\{L\}\}s\_\{x\}\(\\ell^\{\\prime\}\)\}The language entropy is computed from this normalized distributionqxq\_\{x\}, while vocabulary entropy is computed over the model’s next\-token vocabulary distribution\.
### A\.3Open\-Ended Evaluation Details
For natural\-language data without constrained completions, both representation\-based and decoding\-based LLID are evaluated at a matched prompt anchor\. Unless otherwise stated, we use the midpoint of the prompt and aggregate over a forward window of five positions\. Decoding\-based open\-ended LLID uses either deterministic argmax rollouts or top\-pprollouts withp=0\.9p=0\.9and five samples\. We score each decoded continuation with GlotLID\([Kargaran et al\., 2023](https://arxiv.org/html/2609.00155#bib.bib3)\), collapse its labels to the candidate\-language inventory, and normalize the resulting distribution\. For top\-pprollouts, we average these distributions uniformly across samples to obtainqq\. We use PUD as sentence\-level natural\-language data\. PUD21 is the broadest setting, while PUD9 follows the language subset used by[Shani and Basirat \(2025\)](https://arxiv.org/html/2609.00155#bib.bib19)\. We also evaluate PUD9\+UD6 and PUD21\+UD6 by adding six Universal Dependencies languages selected for script or family overlap\. For domain analyses, we use INCLUDE question\-answering prompts in 10 languages and three domains: Social Science, Arts & Humanities, and STEM\.
## Appendix BModel Details
Table[2](https://arxiv.org/html/2609.00155#A2.T2)reports all evaluated models and multilinguality proxies\. Since direct pretraining mixtures are not fully available, we use zero\-shot QA accuracy on INCLUDE, cross\-lingual perplexity, and byte\-normalized likelihood as external proxies for multilingual capability\. For the training dynamics analysis, we align OLMo\-2 and Apertus checkpoints as shown in Table[3](https://arxiv.org/html/2609.00155#A2.T3)\.
The INCLUDE score is zero\-shot accuracy fromlm\-evaluation\-harnessruns oninclude\-base\-44, macro\-averaged over 44 language groups\. We use it to sort the models because it directly measures multilingual task performance\.
The PUD21 line NLL is a likelihood\-based auxiliary score on the PUD21 parallel sentence set\. We compute the autoregressive negative log\-likelihood of each teacher\-forced line, convert it to bits, and average across PUD\-only parallel rows in the 21 languages\. Lower is better\.
The PUD21\+UD6 BPB score is a byte\-normalized likelihood on the broader PUD21\+UD6 corpus\. We sum token\-level negative log\-likelihood over all evaluated lines, convert from nats to bits, and divide by the total number of UTF\-8 bytes:
BPB=∑x−log2pθ\(x\)∑x\|utf8\(x\)\|\\mathrm\{BPB\}=\\frac\{\\sum\_\{x\}\-\\log\_\{2\}p\_\{\\theta\}\(x\)\}\{\\sum\_\{x\}\|\\mathrm\{utf8\}\(x\)\|\}This reduces, but does not eliminate, tokenization and script effects; lower BPB is better\.
ModelINCLUDE0\-shot macro\(↑\\uparrow\)PUD21 lineNLL\(↓\\downarrow\)PUD21\+UD6BPB\(↓\\downarrow\)Apertus\-8B\-Instruct55\.0169\.81\.19Apertus\-8B\([Hernández\-Cano et al\., 2026](https://arxiv.org/html/2609.00155#bib.bib14)\)53\.1138\.30\.97Llama\-3\.1\-8B\-Instruct52\.8152\.91\.09Mistral\-Nemo\-Instruct\-2407\([team, 2024](https://arxiv.org/html/2609.00155#bib.bib13)\)52\.2166\.71\.20Llama\-3\.1\-8B\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib15)\)48\.9148\.81\.05EuroLLM\-9B\-Instruct47\.8156\.01\.17EuroLLM\-9B\([Martins et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib11)\)43\.3149\.21\.11Aya\-23\-8B\([Aryabumi et al\., 2024](https://arxiv.org/html/2609.00155#bib.bib12)\)40\.1172\.01\.26OLMo\-2\-1124\-7B\([Walsh et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib10)\)33\.8181\.11\.30Llama\-2\-7B\([Touvron et al\., 2023](https://arxiv.org/html/2609.00155#bib.bib17)\)27\.7170\.01\.22GPT\-2\([Radford et al\., 2019](https://arxiv.org/html/2609.00155#bib.bib16)\)26\.2365\.42\.62GPT\-2\-XL25\.9302\.32\.20
Table 2:Multilingual Performance\.Models are sorted by decreasing zero\-shot macro accuracy on INCLUDE\. Line NLL is an auxiliary likelihood score computed on a parallel corpus across languages\. BPB is likelihood normalized by UTF\-8 byte count, but it remains script\-sensitive because byte counts differ across writing systems\. Best is bold; second\-best is underlined\.Alignment IdxOLMo\-2 tokensApertus tokens1210B210B2625B630B31\.04T1\.05T41\.46T1\.47T51\.87T1\.89T62\.29T2\.31T72\.70T2\.73T83\.12T3\.15T93\.53T3\.57T103\.95T3\.99T
Table 3:Token\-aligned checkpoints,used in the OLMo\-2 and Apertus training dynamics comparison\.Figure 7:Copy prompts: latent language probability across model layers\.For each copy prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same prompt language\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.
## Appendix CTraining Details
#### Hardware, Packages & Artifacts\.
We run experiments on NVIDIA A100 80GB GPUs, using one GPU at a time\. Each run lasts between one to six hours depending on the dataset size\. We use Python 3\.10 and use transformers333[https://huggingface\.co/docs/transformers](https://huggingface.co/docs/transformers), nnsight444[https://nnsight\.net/](https://nnsight.net/), and datasets555[https://huggingface\.co/datasets](https://huggingface.co/datasets)packages\. We used AI\-assisted tools to support code development, and we reviewed and tested the resulting code\. For writing, AI tools were used to rephrase the wording of the manuscript\.
#### GMM Fitting\.
Representation\-based LLID requires a language\-conditioned model of activation space\. We fit one per\-layer Gaussian model per candidate language from hidden states collected on the train split\. Each GMM setup uses at most 100 training prompts per language, uniform language priors, layerwise PCA retaining 98% explained variance, and diagonal covariance\. Unless otherwise stated, we fit token\-level hidden states and use the surface\-token alignment from the dataset preprocessing to select comparable positions at evaluation time\. We fit separate GMMs per candidate\-language inventory:pud9,pud21,pud9\_ud6,pud21\_ud6, andinclude\_10lang\. For INCLUDE plots that contain English as a possible latent language, we fit with the same PUD calibration source but add English to the candidate\-language set\.
#### Tuned Lens Fitting\.
For decoding\-based LLID, we compare raw logitlens projections with tuned lens projections\([Belrose et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib5)\)\. Tuned lenses are fit on multilingual FineWeb2 data\([Penedo et al\., 2025](https://arxiv.org/html/2609.00155#bib.bib4)\), keeping the fitting corpus independent of PUD, INCLUDE, and controlled target\-string evaluation prompts\. Our main tuned lens artifacts use the 27\-language inventory\. For each model, FineWeb2 rows are tokenized into windows of length up to 2048 tokens, subject to the model’s supported context length\. Training examples are balanced by strict round\-robin over languages, dropping excess rows from higher\-resource languages\. We train for one epoch with at most 1000 optimization steps, AdamW with learning rate10−310^\{\-3\}and weight decay 0, temperature 1\.0, identity and bias regularization weights10−410^\{\-4\}, and bfloat16 autocast\. Validation uses a lightweight balanced subset, and checkpoints are written periodically\. We do not train a tuned lens translator for the final transformer layer; at runtime, the final layer uses the raw model\-head projection\.
## Appendix DAdditional Results
The appendix figures below are organized by analysis group\. For the controlled synthetic tasks, Figs\.[7](https://arxiv.org/html/2609.00155#A2.F7)–[11](https://arxiv.org/html/2609.00155#A4.F11)show layerwiseStart\(w\)\\textsc\{Start\}\(w\)\{\}probability curves, while Figs\.[12](https://arxiv.org/html/2609.00155#A4.F12)–[14](https://arxiv.org/html/2609.00155#A4.F14)summarize the same task family into language\-probability categories across layer windows\. For open\-ended language probabilities, Fig\.[15](https://arxiv.org/html/2609.00155#A4.F15)gives the PUD21\+UD6 counterpart\. For training analyses, Figs\.[17](https://arxiv.org/html/2609.00155#A4.F17)–[19](https://arxiv.org/html/2609.00155#A4.F19)show checkpoint dynamics and full\-layer trajectories, while Figs\.[20](https://arxiv.org/html/2609.00155#A4.F20)–[21](https://arxiv.org/html/2609.00155#A4.F21)compare base and instruction\-tuned model variants\.
Figure 8:Cloze prompts: latent language probability across model layers\.For each cloze prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same prompt language\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.Figure 9:Translation prompts: latent language probability across model layers, grouped by translation target\.For each translation prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same translation target\. English source and target cases are included\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.Figure 10:Translation prompts: latent language probability across model layers, grouped by translation source\.For each translation prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same translation source\. English source and target cases are included\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.Figure 11:Translation prompts: latent language probability across model layers, grouped by translation source\.For each translation prompt and model layer, we measure the summed next\-token probability of all tokens that can begin the corresponding answer in each candidate language \(Start\(w\)\\textsc\{Start\}\(w\)\{\}probability\)\. Rows are models and columns group prompts with the same translation source\. English is excluded as a source or target, but retained as a candidate latent language\. Upper strips show average latent language distribution entropy and average first\-token vocabulary entropy\.Figure 12:Copy prompts: latent language probabilities across layer windows\.Columns show the first 50%, 50–75%, and 75–100% of layers\. Bars pool individual prompt–layer–language probabilities within each category before averaging; diamonds show the average maximum probability within the task\-relevant or other\-language category\. English prompts are excluded, but English remains a candidate latent language\. Error bars show SE\.Figure 13:Cloze prompts: latent language probabilities across layer windows\.Same layout and aggregation as Fig\.[12](https://arxiv.org/html/2609.00155#A4.F12), evaluated on cloze prompts\. English prompt cases are excluded\.Figure 14:Translation prompts: latent language probabilities across layer windows\.Same layout and aggregation as Fig\.[12](https://arxiv.org/html/2609.00155#A4.F12)\. For translation, the task\-relevant category contains both the translation source and target language\. English is excluded as a source or target, but retained as a candidate latent language\.Figure 15:Open\-ended PUD21\+UD6: latent language probabilities across layer windows\.Rows are LLID estimators and columns are layer windows\. Bars, diamonds, and error bars use the same aggregation as Fig\.[12](https://arxiv.org/html/2609.00155#A4.F12)\. The task\-relevant category is the prompt language\. English prompts are excluded, but English remains a candidate latent language\.\(a\)Dominance\(b\)Pivot rate
Figure 16:Latent language dominance and pivoting are comparatively stable across domains and language inventories\.Rows are language models and columns are LLID methods\. For each prompt and layer, dominance is the highest probability assigned to a candidate latent language; a pivot occurs when the dominant language differs from the prompt language\. Curves largely cluster across PUD variants \(PUD9, PUD9\+UD6, PUD21, PUD21\+UD6\) and INCLUDE domains \(Arts/Humanities, Social Sciences, and STEM\) within each model–estimator panel, while the layerwise trajectories can differ substantially across models\.Figure 17:Training dynamics of latent language estimates across pretraining on PUD21 with four layer bins\.This is the quartile\-binned counterpart to Fig\.[6](https://arxiv.org/html/2609.00155#S5.F6), using the same rows, metrics, estimator overlays, and prompt\-level definitions\. Colors divide the normalized layer indexj/Jj/Jinto quartiles: Q1 hasj/J<0\.25j/J<0\.25, Q2 has0\.25≤j/J<0\.500\.25\\leq j/J<0\.50, Q3 has0\.50≤j/J<0\.750\.50\\leq j/J<0\.75, and Q4 hasj/J≥0\.75j/J\\geq 0\.75\.Figure 18:Representation\-based LLID checkpoint trajectories across layers on PUD21\.Rows are model families and columns are layerwise metrics\. Each curve is one token\-aligned checkpoint, colored by its number of training tokens\. Metrics are averaged over prompts at each layer\. Unlike Fig\.[6](https://arxiv.org/html/2609.00155#S5.F6), this view keeps the full layer axis instead of aggregating layers into bins\.Figure 19:Decoding\-based LLID checkpoint trajectories across layers on PUD21\.Same layout as Fig\.[18](https://arxiv.org/html/2609.00155#A4.F18), using the raw logitlens argmax decoding probe instead of the representation probe\.Figure 20:Base\-vs\-instruct LLID trajectories on PUD21\.Rows are model families\. Columns show pivot rate underreprand raw logitlensdec\-topp, entropy under the same two estimators, and their dominant language agreement\. Blue solid and orange dashed curves denote base and instruction\-tuned variants, respectively\. Agreement is the fraction of matched prompt–layer pairs for which the two probes select the same dominant language\.Figure 21:Base\-vs\-instruct LLID trajectories on INCLUDE\-10 with English as a candidate language\.Same layout as Fig\.[20](https://arxiv.org/html/2609.00155#A4.F20), evaluated on INCLUDE\-10\. INCLUDE contains no English prompts, but English is retained as a candidate latent language\.Similar Articles
The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
The paper proposes the interlingua hypothesis, suggesting that large language models perform translation by encoding source text into a latent task-agnostic feature space and decoding from it, supported by empirical evidence on variance, causal influence, and monolingual fine-tuning.
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
The study uses exploratory factor analysis to compare latent structures in human and LLM responses on assessments, revealing that LLMs rely on statistically opaque mechanisms unlike human reasoning.
Latent Mechanisms of Language Control in Multilingual Language Models
This paper compares three methods to identify language-controlling latents in multilingual language models to address code-switching, with experiments on Gemma-2-2B and Qwen3-4B showing FreqSel as the most effective.
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
The paper investigates how prompt and response languages affect LLM content generation, finding that prompt language significantly influences output length while maintaining semantic fidelity through conceptual paraphrasing.
Are you speaking my languages? On spoken language adherence in multimodal LLMs
This paper addresses the problem of spoken language adherence in multimodal LLMs for ASR, proposing a soft prompting approach and novel metric to quantify language violations. It evaluates three mitigation strategies—zero-shot prompting, supervised fine-tuning, and chain-of-thought reasoning—across multiple languages to improve transcription fidelity.