Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models
Summary
This paper uses causal interventions to investigate syntactic mechanisms in multilingual language models, revealing cross-lingual transfer that is graded based on typological similarity.
View Cached Full Text
Cached at: 09/01/26, 12:10 PM
# Causal Interventions Reveal Typologically OrganizedSyntactic Mechanisms in Multilingual Language Models
Source: [https://arxiv.org/html/2608.28924](https://arxiv.org/html/2608.28924)
rmTeXGyreTermesX \[russian\]rm\[Path=fonts/, BoldFont=FreeSerifBold\.otf\]FreeSerif\.otf \[\*arabic\]rm\[Path=fonts/\]NotoNaskhArabic\-Regular\.ttf \[Scale=0\.7\] \[Scale=0\.7\] \[Scale=0\.7\]
Toshiki Nakai††thanks:Work initiated while author was at Saarland University\.Affiliation:Leipzig University, ScaDS\.AI Dresden/LeipzigEmail:[kyle@utexas\.edu](mailto:)Kyle MahowaldAffiliation:The University of Texas at AustinEmail:[toshiki\.nakai@uni\-leipzig\.de](mailto:)Julius SteuerAffiliation:Heidelberg Institute for Theoretical StudiesEmail:[julius\.steuer@h\-its\.org](mailto:)
###### Abstract
Linguistic theory has long recognized cross\-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar mechanisms\. However, this hypothesis has been difficult to test empirically due to our lack of fine\-grained, manipulable access of human processing mechanisms\. In this work, we take advantage of techniques from mechanistic interpretability to study such a question in multilingual LMs\. We first isolate language\-internal mechanisms before attempting to transfer them cross\-lingually\. Across four models and three well\-studied constructions \(subject–verb number agreement, anaphoric pronoun gender agreement, and filler–gap object extraction\) we find consistent cross\-lingual mechanism transfer\. We further find transfer to be graded, with more transfer between more typologically similar languages\. We believe our work provides novel hypotheses about cross\-linguistic syntactic structures and multilingual processing, and more broadly shows how the study of language models can help inform linguistic theory\.
Figure 1:We localize LM’s syntactic mechanisms to language\-internal subspaces\. We then test such mechanism’s effects on other languages\. If the intervention is still efficacious, this indicates the model’s syntactic mechanisms are language\-agnostic\. This provides evidence of*cross\-lingual shared mechanisms*\.## 1Introduction
Linguistics has long accepted there are shared abstract syntactic structures across different languages — perhaps due to innate constraints\([Chomsky, 1965](https://arxiv.org/html/2608.28924#bib.bib9)\)or functional evolutionary pressures\([Greenberg et al\., 1963](https://arxiv.org/html/2608.28924#bib.bib18);[Comrie, 1989](https://arxiv.org/html/2608.28924#bib.bib10)\)\. It has often been hypothesized that these similar structures across different languages recruit similar machinery in human processing\([Hartsuiker et al\., 2004](https://arxiv.org/html/2608.28924#bib.bib20);[Hawkins, 2014](https://arxiv.org/html/2608.28924#bib.bib21);[Norcliffe et al\., 2015](https://arxiv.org/html/2608.28924#bib.bib46);[Malik\-Moraleda et al\., 2022](https://arxiv.org/html/2608.28924#bib.bib41)\)\. However, testing for such shared mechanisms in humans is difficult due to a lack of granular manipulable access needed to ask whether processing of two languages recruit the*same*internal mechanism\.
In recent years, neural Language Models \(LMs\) have shown remarkable syntactic competence — possessing an ability to produce and process utterances previously thought to require abstract linguistic representations, both mono\- and cross\-lingually\([Linzen et al\., 2016](https://arxiv.org/html/2608.28924#bib.bib37);[Wilcox et al\., 2018](https://arxiv.org/html/2608.28924#bib.bib63);[Manning et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib42);[Hu et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib26);[Warstadt et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib62);[Jumelet et al\., 2026](https://arxiv.org/html/2608.28924#bib.bib27)\)\. This competence, combined with privileged access to their internal computations has led to a host of work treating them as model organisms to formulate and test linguistic hypotheses\([Linzen and Baroni, 2021](https://arxiv.org/html/2608.28924#bib.bib36);[Warstadt and Bowman, 2022](https://arxiv.org/html/2608.28924#bib.bib61);[Wilcox et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib65);[Boguraev et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib6);[Boguraev and Mahowald, 2026](https://arxiv.org/html/2608.28924#bib.bib5);[Misra and Kim, 2026](https://arxiv.org/html/2608.28924#bib.bib44)\)\.
In this work, we take advantage of advances in mechanistic interpretability, particularly Causal Abstraction\([Geiger et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib14)\), to understand the internal organization of high\-level linguistic abstractions in multilingual LMs\. The basic idea of the Causal Abstraction framework is to identify internal structure in the neural model that controls downstream behavior: for our purposes, syntactic behavior\([Arora et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib1);[Boguraev et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib6)\)\. Our goal here is to identify causally relevant linguistic structures in one language and then test whether intervening on that same structure inanotherlanguage has a similar effect \([Figure1](https://arxiv.org/html/2608.28924#S0.F1)\)\. If so, we would take that to be evidence for shared mechanisms across languages—offering potential insights into the underlying similarity of structures across languages, as well as into how multilingual neural models organize syntactic information\.
In particular, we test for shared mechanisms across three structurally varied phenomena: subject–verb number agreement, anaphoric pronoun gender agreement, and filler–gap object extraction\. Across fifteen languages and four LMs spanning sizes, families, and training specs, we find LMs recruit shared, repurposed mechanisms for each phenomenon rather than language\-specific ones\. We further provide evidence that this representational sharing is driven by typological similarity, with transfer organized around typological ‘hubs’ — languages broadly similar to a large amount of the other studied languages\.
Taken together, our results characterize multilingual syntactic competence in LMs as the result of abstractions repurposed across languages\. We further believe our results provide novel hypotheses about multilingual linguistic organization which could potentially be tested in humans\.
## 2Cross\-Lingual LM Interpretability
When multilingual LMs with syntactic capabilities in multiple languages emerged, there was a flurry of interest in whether they learned shared representations across languages or whether they mostly learned each language independently\. Early work in this space took advantage of probing — i\.e\., training classifiers on the internal states of an LM to recover high\-level labels\.[Chi et al\. \(2020\)](https://arxiv.org/html/2608.28924#bib.bib8)used structural probes\([Hewitt and Manning, 2019](https://arxiv.org/html/2608.28924#bib.bib25)\)to show universal\-dependency relations could be localized to language\-agnostic subspaces, suggesting linguistic representations pattern similarly cross\-lingually and[Papadimitriou et al\. \(2021\)](https://arxiv.org/html/2608.28924#bib.bib50)demonstrated probes trained to recover high\-level semantic features like ‘subjecthood’ transferred in linguistically notable ways across languages\. Such results were evidence of language\-agnostic linguistic features in LMs\.
However, probing can be over\-expressive\([Hewitt and Liang, 2019](https://arxiv.org/html/2608.28924#bib.bib24);[Voita and Titov, 2020](https://arxiv.org/html/2608.28924#bib.bib60)\), prompting a shift to*causal methods*which ensure faithfulness by manipulating internal states and checking for behavioral effect\. Early iterations of these methods produced linguistic insights, finding both language\-neutral and language\-specific representations\([Gonen et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib16);[Srinivasan et al\., 2023](https://arxiv.org/html/2608.28924#bib.bib59)\), and finding cross\-lingual grammatical circuits in English and Spanish\([Ferrando and Costa\-jussà, 2024](https://arxiv.org/html/2608.28924#bib.bib12)\)\. However, such variants could not elucidate the ‘features’ utilized for computation, merely localizing behavior to components\.
To surface such features, sparse\-decomposition methods such as Sparse Auto\-Encoders \(SAEs\) and Cross\-Layer Transcoders \(CLTs\) were adopted\. These methods likewise found cross\-lingually shared feature\-circuits\([Brinkmann et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib7)\)and that language identity is encoded in later model layers\([Harrasse et al\., 2026](https://arxiv.org/html/2608.28924#bib.bib19)\)\. However, these methods find ‘interpretable features’ in untrained models\([Heap et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib22)\)and are unstable across training runs\([Paulo and Belrose, 2026](https://arxiv.org/html/2608.28924#bib.bib51);[Leask et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib34)\)\.
We instead use Distributed Alignment Search\([Geiger et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib15), DAS;\), to localize abstract features in multilingual LMs\. DAS is highly efficacious for syntactic interpretability\([Arora et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib1)\), and has been utilized to study the representational organization of syntactic features in English filler–gap constructions\([Boguraev et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib6)\)and their respective island constraints\([Boguraev and Mahowald, 2026](https://arxiv.org/html/2608.28924#bib.bib5)\)\. Our methodology allows precise tests of whether causally implicated syntactic mechanisms transfer cross\-lingually\.
## 3Methods
### 3\.1Data
Subject–Verb Number AgreementNP1PPNP2VPEnglishThe man / The mennearthe cabinetis / areItalianL’uomo / Gli uominivicino allatavolaè / sonoBulgarianМъжът / Мъжетедомасатае / саAnaphoric Pronoun Gender AgreementNameVPbecausePron\.EnglishJames / Maryliedbecausehe / sheItalianMario / Mariariseperchélui / leiBulgarianИван / Мариявиказащототой / тяFiller–Gap \(Object\)PrefixFillerNPVPGapEnglishI knowthat / whatshesawthe / \.ItalianSoche / cosaleivideil / \.BulgarianЗнам,че / каквотявидянего / \.Table 1:Exemplar minimal pairs for a representative subset of languages with a full set of exemplars are in[§A](https://arxiv.org/html/2608.28924#A1)\. Our gender agreement stimuli are violable:“Mary was powerful\. James lied becauseshetold him to”, although we do not find this to affect our LMs behaviorally\. Instead, we find our studied LMs to score at or above 90% accuracy on 90% \(115/128\) of templates and below 80% accuracy only once\.#### Stimuli
We study three constructions — subject–verb number agreement, anaphoric pronoun gender agreement, and object extraction from an embeddedwh\-question — across fifteen languages — English, French, German, Spanish, Italian, Portuguese, Russian, Bulgarian, Farsi, Dutch, Chinese, Japanese, Korean, Galician and Romanian\. We design data templates as in[Arora et al\. \(2024\)](https://arxiv.org/html/2608.28924#bib.bib1)to allow sampling of large sets of unique minimal pairs \(examples in[Table1](https://arxiv.org/html/2608.28924#S3.T1); full templates in[§A](https://arxiv.org/html/2608.28924#A1)\)\.111As German and Dutch place embedded verbs clause\-finally, their filler–gap templates cannot be perfectly aligned; we therefore exclude theVPposition for these languages and analyze it only across the five with a non\-final embedded verb\.Typological constraints prevent every language from being represented in each construction \(e\.g\., Chinese does not have grammatical number\)\. Nevertheless, there is a core set of seven languages across templates: English, German, Italian, Russian, Bulgarian, Dutch, and French\. Behavioral evaluation shows LMs correctly process the stimuli with details in[§C](https://arxiv.org/html/2608.28924#A3)\.
#### Controls
We design three controls\.Random Labelsmeasures the specificity of our interventions\. Per[Hewitt and Liang \(2019\)](https://arxiv.org/html/2608.28924#bib.bib24)this control keeps the critical context but pairs it with unrelated, ungrammatical labels\. An exemplar is in \(1\)\.
- Theman/menbeside the car→\\rightarrowdog/run
Interventions demonstrating strong specificity, as desired,*should not generalize*to this control\.Other phenomenaevaluates each core\-language intervention on the other constructions\. We again*expect no transfer*, as the constructions are syntactically distinct\. Finally, following[Kumon and Yanaka \(2026\)](https://arxiv.org/html/2608.28924#bib.bib30),OOD labelsconsists of templates with the same critical context but with novel grammatical labels\. We cannot extend this control to gender agreement, as there are no cross\-lingually robust alternate labels\. As such, we only report related results in[§E](https://arxiv.org/html/2608.28924#A5)\.
### 3\.2Models
We study four models:mGPT\([Shliazhko et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib57), 1\.3b params\.;\)andtiny\-aya\-base\([Salamanca et al\., 2026](https://arxiv.org/html/2608.28924#bib.bib56), 3\.4b params\.;\)both expressly multilingual, trained on 61 and 70\+ languages respectively, as well asLlama\-3\.2\-1bandLlama\-3\.2\-3b\([Grattafiori et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib17)\), both not expressly multilingual, trained on generic web data, but nevertheless possessing multilingual capabilities\. By evaluating four models, we can study effects of architectures and training factors \(e\.g\. model size, training data composition\) on representation sharing and ensure our findings reflect general trends rather than that of one model\.
### 3\.3Distributed Alignment Search
DAS is a supervised interpretability method which finds low\-rankdd\-dimensional subspaces in an LM in which a high\-level causal variable can be localized and manipulated\. Specifically, given a base valueb∈ℝnb\\in\\mathbb\{R\}^\{n\}— a tensor at a specific internal site when an LM processes some input — and corresponding source values∈ℝns\\in\\mathbb\{R\}^\{n\}— the tensor at the same site when the model processes a minimally\-paired input — DAS finds a subspace in which intervening froms→bs\\rightarrow b, whilst keeping all orthogonal components fixed, changes the output prediction correspondingly\. This is operationalized as
b\+\(sa⊤−ba⊤\)a\\textbf\{b\}\+\(\\textbf\{sa\}^\{\\top\}\-\\textbf\{ba\}^\{\\top\}\)\\textbf\{a\}
wherea∈ℝn×da\\in\\mathbb\{R\}^\{n\\times d\}is a rotation matrix learned by minimizing the cross\-entropy loss of the LM’s prediction under intervention through gradient descent\. Following[Arora et al\. \(2024\)](https://arxiv.org/html/2608.28924#bib.bib1);[Boguraev et al\. \(2025\)](https://arxiv.org/html/2608.28924#bib.bib6);[Boguraev and Mahowald \(2026\)](https://arxiv.org/html/2608.28924#bib.bib5)we learn 1D subspaces \(d=1d=1\) in the residual stream of our given LMs\. When the span we are intervening on contains multiple tokens, we intervene on the pooled representation across the span, and when two separate template positions combine into one token \(i\.e\., French elision producing the single tokenqu’ilfromque\+\+il\) we evaluate both corresponding interventions at the combined token\.
#### Training and Evaluation
We train interventions at each template position and model layer\. Following[Arora et al\. \(2024\)](https://arxiv.org/html/2608.28924#bib.bib1);[Boguraev et al\. \(2025\)](https://arxiv.org/html/2608.28924#bib.bib6), we train interventions for 100 steps with a batch size of 4, filtering training items to those on which the LM is behaviorally competent\. We useOddsto evaluate the interventions\. This metric measures the probability increase for the source label after intervention, relative to the probability decrease of the base label\. HigherOddsindicate higher causal efficacy\. In cases of aggregation, we report theMax Odds: the maximumOddsvalue across layers at a given position\. The use of these metrics is consistent with other work using DAS for linguistic evaluation\([Arora et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib1);[Boguraev et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib6);[Boguraev and Mahowald, 2026](https://arxiv.org/html/2608.28924#bib.bib5)\)\. We measure these metrics across a held\-out evaluation set of 100 unique minimal pairs\.
## 4Experiment 1: Do LMs Share Mechanisms Cross\-Lingually?
In our first experiment, we ask whether the mechanisms LMs use to process a given construction transfer cross\-lingually\. That is, we use DAS to find a subspace causally implicated in a particular syntactic phenomenon \(e\.g\., subject–verb agreement\) in Language A and see if that same subspace can causally control the same phenomenon in Language B\. In principle, this subspace could exploit the lexically\-matched templates to discover translation\-equivalent alignments\([Gonen et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib16), à la\), rather than abstract syntactic features\. We flag this alternative hypothesis here and return to distinguishing it from genuine syntactic abstraction in the discussions of[Section4](https://arxiv.org/html/2608.28924#S4)and[Section5](https://arxiv.org/html/2608.28924#S5)\.
#### Setup
We measure theOddsacross layers and positions for all learned interventions evaluated across the same construction multilingually, as well as on controls\. We report the performance across distinct groups:within\-language interventions— e\.g\., training on English subject–verb agreement and evaluating on English subject–verb agreement,cross\-lingual interventions— e\.g\., training on English subject–verb agreement and evaluating on subject–verb agreement in all other languages,within\-family interventions— e\.g\., training on English subject–verb agreement in English and evaluating on subject–verb agreement of all other Germanic languages, andcontrols—random labelsandother phenomenain the main text, withOOD labelsin[§E](https://arxiv.org/html/2608.28924#A5)\.
We compare these values across layers, positions, and models\. We also run statistical tests on theMax Oddsof transferred interventions normalized by within\-language performance\. This intuitively measures how much an LM’s mechanisms for a given language generalize to other languages with respect to its own, baseline, performance\.
#### Hypothesis
We hypothesize there are shared syntactic mechanisms in LMs\. That is, we expect cross\-lingual transfer to be significantly above controls\. We further expect higher transfer within a family than to the full set of evaluated languages\.
Figure 2:We calculate theOddswhen evaluating our trained interventions on a held\-out test set of the same language they are trained on \(red line\), all other languages \(green line\), all other languages within the same family \(yellow line\), and across two controls \(grey and black lines\)\. We find that across all positions, phenomena, and models, interventions consistently generalize across languages with pairwise t\-tests on theMax Oddsshowing significant differences between both the cross\-language and control conditions \(p<0\.05p<0\.05\)\.
#### Results
Our results are in[Figure2](https://arxiv.org/html/2608.28924#S4.F2)\. Across layers, positions, and LMs, the cross\-lingual generalizationOdds\(green lines\) are consistently higher than controls \(grey lines\), and at similar levels across phenomena\. We also see the information flow through the LMs: early layers possess causally efficacious information at early positions, with the information passing through the middle layers in middle positions before reaching later layers at the final position\. Pairwise t\-tests with family\-wise Holm\-Bonferroni corrections show significantly higher transferMax Oddscross\-lingually than to corresponding controls \(p<0\.05p<0\.05\)\.222For bar charts of theMax Oddssee[§D](https://arxiv.org/html/2608.28924#A4)We see qualitatively similar results forOOD labels, with both transfer condition’sMax Oddsbroadly significantly above controls\. This provides confidence that our results are not just overfit to the lexical items in our templates \(full results in[§E](https://arxiv.org/html/2608.28924#A5)\)\. All told, these results provide evidence of shared, causally efficacious, syntactic subspaces in multilingual LMs, suggesting the emergence of language\-agnostic mechanisms for these constructions\.
However, this generalization is not ubiquitous\.[Figure2](https://arxiv.org/html/2608.28924#S4.F2)visually suggests transfer is highest within families \(yellow lines are consistently above green lines\), but pairwise t\-tests onMax Oddsshow mixed results\. For number agreement, the within\-family and all\-language groups are significantly different atPPfor all models, atNP1fortiny\-aya\-baseandLlama\-3\.2\-3band atNP2forLlama\-3\.2\-3b\. For gender agreement, we find significant differences atVPandnameformGPT,Llama\-3\.2\-3bandtiny\-aya\-base\. Finally, for filler–gaps, we find no within\-family vs\. cross\-language difference\. This may reflect the low diversity of this language set or the high within\-languageOddsand consistent cross\-lingual transferOddsdeflating the normalizedMax Oddswe test with respect to the other constructions\.
#### Discussion
Our results show LMs converge to abstract cross\-lingual syntactic mechanisms\. Further, they suggest representational sharing is not without structure, but such structure is not as simple as, say, language family\. However, our current experimental set\-up does not lend insights into which factors are responsible for transfer\. This question is centered in[Section5](https://arxiv.org/html/2608.28924#S5)\.
We further find evidence against the alternative, lexical\-translation subspace hypothesis\. That is, our results in[§E](https://arxiv.org/html/2608.28924#A5)demonstrate interventions generalize across lexical items \(both within a language and cross\-lingually\)\. These results provide evidence our results are not merely lexically specific, and therefore the subspaces not just translation\-equivalent\. We revisit this hypothesis in[Section5](https://arxiv.org/html/2608.28924#S5), where the typological structure of transfer further bears negative evidence for it\.
Figure 3:We visualize our languages on the top principal components of the transfer matrix\. We find language family\-based clusters not to be pure\.
## 5Experiment 2: What Drives Representational Reuse?
We have found the syntactic mechanisms utilized by LMs can be transferred cross\-lingually\. However, exactly what governs this transfer is unclear\. Furthermore, we have not teased apart the effects of the LM itself\. Here, we analyze these factors\.
Number AgreementGender AgreementFiller–Gap \(Object\)TermNP1\(\.68\)PP\(\.71\)NP2\(\.85\)Name\(\.69\)VP\(\.75\)because\(\.76\)Filler\(\.73\)NP\(\.61\)Verb\(\.78\)Typological dist\.−0\.577∗∗∗\\mathbf\{\-0\.577\}^\{\*\*\*\}−0\.558∗∗∗\\mathbf\{\-0\.558\}^\{\*\*\*\}−0\.432∗∗∗\\mathbf\{\-0\.432\}^\{\*\*\*\}−0\.462∗∗∗\\mathbf\{\-0\.462\}^\{\*\*\*\}−0\.320∗∗∗\\mathbf\{\-0\.320\}^\{\*\*\*\}−0\.498∗∗∗\\mathbf\{\-0\.498\}^\{\*\*\*\}−0\.544∗∗∗\\mathbf\{\-0\.544\}^\{\*\*\*\}−0\.781∗∗∗\\mathbf\{\-0\.781\}^\{\*\*\*\}−1\.04∗∗∗\\mathbf\{\-1\.04\}^\{\*\*\*\}Model size0\.3630\.3630\.2680\.2680\.4810\.4810\.4000\.4000\.3790\.3790\.5500\.550−0\.013\-0\.013−0\.151\-0\.1510\.2240\.224Typological dist\.×\\timesmodel size−0\.043\-0\.043−0\.100∗∗∗\\mathbf\{\-0\.100\}^\{\*\*\*\}−0\.078∗∗\\mathbf\{\-0\.078\}^\{\*\*\}−0\.055\-0\.055−0\.158∗∗∗\\mathbf\{\-0\.158\}^\{\*\*\*\}−0\.063\-0\.0630\.0150\.0150\.0100\.010−0\.103\-0\.103
Table 2:LMEM coefficients predicting theMax Oddsfrom linguistic and architectural factors\. Estimates significant under the LRT arebolded\.∗p<\.05\{\}^\{\*\}p<\.05;p∗∗<\.01\{\}^\{\*\*\}p<\.01;∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001\.Rc2R^\{2\}\_\{c\}values in parentheses next to each position\. We find significant negative effects of typological distance, suggesting that transfer is stronger between typologically similar languages\. Interaction effects between typological distance and model size further suggest such effects grow stronger in larger models for the tasks of Number Agreement and Gender Agreement\.#### Set\-Up
For each model, phenomenon, and position, we create ann×nn\\times nmatrix,MM, wherennis the number of languages evaluated, and cellMi,jM\_\{i,j\}is theMax Oddsof the interventions trained on languageiiwhen evaluated on languagejj\. We then perform dimensionality reduction with PCA\. Visualizing where our languages fall in this space provides insights about what languages have similar transfer patterns — an implicit representation of which languages our LMs represent similarly\.
We also fit a linear mixed\-effects model \(LMEM\) predicting theMax Oddsat each position from linguistic and LM\-specific predictors\. In particular, our LMEM has the LM, training language, and evaluation language as random effects with the typological distance between two languages, model size and their interaction as fixed effects\. We operationalize typological distance as the cosine distance betweenlang2vecsyntax vectors\([Dryer and Haspelmath, 2013](https://arxiv.org/html/2608.28924#bib.bib11);[Littell et al\., 2017](https://arxiv.org/html/2608.28924#bib.bib39)\)\. For full LMEM details, see[§F](https://arxiv.org/html/2608.28924#A6)\.
#### Hypothesis
We expect more transfer between typologically similar languages\. We further expect smaller models to show more transfer as their smaller parameter count would explicitly promote representational reuse\.
#### Results
Our PCA analysis \([Figure3](https://arxiv.org/html/2608.28924#S4.F3)\) shows some clustering by language family: Romance, Slavic and Germanic languages generally cluster with themselves, with more typologically distant languages further afield\. However, closer analysis suggests language family alone does not explain clustering: the average silhouette score \(i\.e\., cluster purity\) across language families with more than one language demonstrates low cohesion \([Table3](https://arxiv.org/html/2608.28924#S5.T3)\)\.
To decipher the role of other factors, we turn to our LMEM \([Table2](https://arxiv.org/html/2608.28924#S5.T2)\)\. Across all phenomena and positions, typological distance is a strong, significant predictor of transfer, with less distance between two languages associated with more transfer\. That is, we consistently see stronger transfer between languages that are more syntactically similar — independent of construction or model\. We do not find a significant main effect of model size at any position — refuting our hypothesis that smaller models would show more transfer\. We do observe significant interactions of model size and typological distance at several positions across the agreement templates, with it consistently indicating that typological distance becomes a stronger predictor of transfer as a model gets larger, albeit weakly\.
However,lang2vecis a coarse representation of a language’s typological aggregate and does not tell us what lower\-level features may be governing such transfer\. We believe there are many factors which shape representational similarity during training, such as tokenizer overlap between languages, syntactic feature overlap, and number of cognates\. In order to further characterize linguistic similarity, we correlate such metrics with each other, and the previously used typological distance, to show that such metrics generally covary\.
Subject Verb Number AgreementNP1PPNP2mGPT\-0\.0800\.201\-0\.158Llama\-3\.2\-1b0\.0650\.029\-0\.093Llama\-3\.2\-3b\-0\.1790\.0970\.120Tiny\-Aya0\.2360\.039\-0\.229Anaphoric Pronoun Gender AgreementNameVPbecausemGPT0\.226\-0\.038\-0\.175Llama\-3\.2\-1b\-0\.0110\.015\-0\.181Llama\-3\.2\-3b\-0\.0930\.081\-0\.444Tiny\-Aya0\.087\-0\.104\-0\.322Filler–Gap \(Object\)FillerNPVerbmGPT0\.3870\.5430\.106Llama\-3\.2\-1b0\.220\-0\.019\-0\.003Llama\-3\.2\-3b0\.2970\.2320\.259Tiny\-Aya0\.3310\.1340\.277
Table 3:Silhouette scores measuring language\-family cluster purity in PCA space\. Each cell is the mean silhouette width for the given model and position\. Scores range from−1\-1\(languages closer to other families\) to\+1\+1\(tight within\-family clusters\)\.Figure 4:We measure the correlation between different operationalizations of typological similarity, pooled across positions and phenomenon\. All pairwise correlations are significantly positive \(all cellsp<\.01p<\.01\) and broadly strong, suggesting these factors are entangled in the model’s training data and learned representations\. Correlations with template overlap tend to be weaker, reflecting the hand\-designed nature of our templates\.In particular, we take four measures of linguistic overlap: \(1\)template\-level tokenizer overlap, calculated as the Jensen\-Shannon divergence between token distributions at each position of our templates, pooled across our phenomenon and positions \(2\)corpus\-level tokenizer overlap, calculated as the Jensen\-Shannon divergence between token distributions across all text in the FLORES\+ dataset\([NLLB Team et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib45), a dataset consisting of 2019 sentences translated from English into 227 different languages;\), \(3\)overlap in GramBank features\([Skirgård et al\., 2023](https://arxiv.org/html/2608.28924#bib.bib58)\), calculated as the Hamming Distance between binary feature vectors, and \(4\)proportion of cognates, calculated as the proportion of cognates in the LexiBank Indo\-European Cognate Relationships database\([List et al\., 2022](https://arxiv.org/html/2608.28924#bib.bib38);[Heggarty et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib23);[Blum et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib3);[Blum et al\., 2026](https://arxiv.org/html/2608.28924#bib.bib4)\)\. While there is broad coverage of these metrics across languages, we note there is not perfect coverage\. Accordingly, we subset correlations to languages with such data\.333Each metric’s language coverage is in[§G](https://arxiv.org/html/2608.28924#A7)\.
Correlations are in[Figure4](https://arxiv.org/html/2608.28924#S5.F4)\. All pairwise correlations are significant, and generally strongly positive\. Template overlap shows weaker correlation, likely due to our hand\-designed templates differing from the overall language distribution and the low token variance of positions with fixed inventories, such asfiller\(2 items\) andbecause\(1 item\)\. More broadly, these results suggest that typological similarity is broadly an effect of many lower\-level features\. Unfortunately, our data is not powered enough to perform regressions on individual features, leaving such analysis to future work\.
#### Discussion
Our results show LM representational reuse is governed by high\-level linguistic features, with mechanisms transferring more between linguistically similar languages\. We further demonstrate thatlang2vectypological distance is strongly correlated with a broader set of features, both surface\-level \(tokenizer overlap\) and linguistically driven \(Grambank Feature overlap\), which shape an LM’s representational development during training\. These results further provide evidence against the lexical\-translation hypothesis because if our learned subspaces were merely lexical translations, we would not expect there to be typologically\-graded cross\-lingual transfer\.
Our results also bear on the debate regarding the role of tokenizer overlap in cross\-lingual capabilities\. Early work demonstrated overlap aids zero\-shot transfer\([Pires et al\., 2019](https://arxiv.org/html/2608.28924#bib.bib53)\), but more recent work has suggested that language\-specific tokenizers had benefit over shared ones\([Rust et al\., 2021](https://arxiv.org/html/2608.28924#bib.bib55)\), and that overlap can even harm syntactic performance\([Limisiewicz et al\., 2023](https://arxiv.org/html/2608.28924#bib.bib35)\)\. More recently still,[Kallini et al\. \(2025\)](https://arxiv.org/html/2608.28924#bib.bib28)demonstrated that the semantic similarity of shared tokens, rather than overlap alone, drives cross\-lingual benefit\. While our results do not possess enough power to adjudicate on the debate, they suggest tokenizer overlap can be beneficial for cross\-lingual mechanism sharing\. However we also find above\-control transfer between languages with disjoint scripts, suggesting tokenizer overlap alone is not the full story\.
## 6Experiment 3: An Investigation Into the Effects of Training Data
In their work on English filler–gap constructions,[Boguraev et al\. \(2025\)](https://arxiv.org/html/2608.28924#bib.bib6)find that frequent constructions show more outward representational transfer, and infrequent constructions show more inward transfer\. In our final experiment, we investigate whether there are similar effects of training data magnitude on multilingual representation sharing\.
#### Set\-Up
We focus onmGPTwhich is the only open\-data model we study, trained on the multilingual Colossal Clean Crawled Corpus\([Raffel et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib54), mC4;\)and Wikipedia\. Due to computational constraints, we do not rerun the full data generation pipeline, instead using per\-language mC4 document counts as a proxy for training\-data magnitude\. To calculate inward and outward transfer, we follow[Boguraev et al\. \(2025\)](https://arxiv.org/html/2608.28924#bib.bib6): we compute the in\- and out\-degree of each language as a node in the transfer matrix across a sweep of edge\-thresholds, with final inward and outward transfer values the area under the sweep’s curve \(AUC\)\.
#### Hypothesis
We expect languages more prevalent in the training data to serve as stronger sources for transfer\. Conversely, we expect less represented languages to serve as stronger sinks for transfer\.
#### Results
Our results are in[Table4](https://arxiv.org/html/2608.28924#S6.T4)\. We consistently see a positive correlation between training\-data and out\-degree, but such a correlation is weak in two of our three phenomena\. Further, only the filler–gap phenomenon shows a negative correlation between training\-data and in\-degree, with such correlation being negligible\. Taken together, these results show little evidence for our hypothesis\.
There is, however, strong correlation between a language’s in\- and out\-degree AUCs \([Table4](https://arxiv.org/html/2608.28924#S6.T4)third row\)\. This result suggests multilingual transfer is not organized by ‘sources’ and ‘sinks’, but by ‘hubs’ which promote and receive large amounts of transfer\. We correlate a language’s ‘hub’\-ness \(mean in\- and out\-degree AUC\) with typological centrality \(average pairwise typological distance across languages\) and find the two to be highly correlated across phenomena \([Table4](https://arxiv.org/html/2608.28924#S6.T4)fourth row\)\. In comparison, ‘hub’\-ness is generally not correlated with training data amount \([Table4](https://arxiv.org/html/2608.28924#S6.T4)fifth row\)\.
NumberAgreementGenderAgreementFiller–Gap\(Object\)Training Datarr\(mC4, Out\-Degree\)0\.180\.630\.36rr\(mC4, In\-Degree\)0\.250\.77\-0\.06Typological ‘Hubs’rr\(In, Out\)0\.700\.890\.44rr\(Centr\., Hub Score\)0\.590\.850\.75rr\(mC4, Hub Score\)0\.230\.720\.20
Table 4:Pearson correlation,rr, between different data\-metrics\. ‘Hub’\-ness is most correlated with typological centrality, not by training data magnitude\.
#### Discussion
Our results refute our hypothesis: transfer is not strongly coupled to training\-data frequency\. Instead, we find specific languages act as ‘hubs’ in our transfer network, both receiving and promoting large amounts of representational transfer\. We further find that ‘hub’\-ness is strongly correlated with how similar a given language is to the other languages in the network\. Such results provide additional evidence that multilingual representational sharing is governed by abstract linguistic features\. We note these correlations are computed over few items \(n=7\-13n=7\\text\{\-\}13\) and are therefore underpowered and suggestive, rather than conclusive\.
## 7Conclusion
Linguists have long recognized seemingly parallel structure cross\-lingually\. These observations have led to many theories of cross\-lingual regularities in linguistic mechanisms\. However, thus far it has been hard to probe the degree to which internal processing overlaps in humans, and by which factors such overlap is governed\.
In this work, we take advantage of multilingual LMs to do just that\. First, we show surface\-level similarities are reflected in LM’s internal organization, finding abstract mechanisms to be learned and repurposed by LMs cross\-lingually\. Second, we show such transfer is gradient, modulated by typological similarity between languages\. Finally, we show there is little effect of training data magnitude on representation sharing, with transfer instead organized around typological ‘hubs’\.
Our findings point to linguistically interesting hypotheses regarding cross\-linguistic syntactic structures and human multilingual processing — hypotheses we imagine tested in psycholinguistic studies\. For instance, syntactic priming paradigms could study whether languages showing greater transfer in LMs also exhibit stronger cross\-linguistic priming in bilingual speakers\. More broadly, we believe our work shows how the study of Language Models can help inform linguistic theory\([Futrell and Mahowald, 2026](https://arxiv.org/html/2608.28924#bib.bib13)\)\.
## Limitations
This work demonstrates how Language Models can be used to study linguistically interesting questions\. However the relationship between linguistic processing in neural models and humans is hotly debated\(e\.g\.,[Wilcox et al\., 2020](https://arxiv.org/html/2608.28924#bib.bib64);[Kuribayashi et al\., 2021](https://arxiv.org/html/2608.28924#bib.bib31);[Oh and Schuler, 2023a](https://arxiv.org/html/2608.28924#bib.bib47);[Oh and Schuler, 2023b](https://arxiv.org/html/2608.28924#bib.bib48);[Oh et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib49);[Kuribayashi et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib32);[Piantadosi, 2024](https://arxiv.org/html/2608.28924#bib.bib52);[Katzir, 2023](https://arxiv.org/html/2608.28924#bib.bib29);[Mahowald et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib40),i\.a\.\)\. As such, our work should not be seen as providing definitive evidence about the nature of multilingual linguistic processing in humans, merely supporting evidence about the analyses general\-purpose learners converge on to process such phenomena\.
Further, while DAS enables us to identify the presence of abstract cross\-lingual mechanisms in a neural language model, we only have indirect access to the totality of the underlying mechanism via the design of the experimental stimuli: our results show that, e\.g\., nominal gender and verbal number features are causally related in the model’s activations\. DAS does not, however, reveal the complete mechanism the model uses to process the feature\. For example, English marks only singular and plural, while Arabic also marks dual\. Transferring an intervention on the number feature from English to Arabic might be disrupted by this extra dual category\. Our method could reflect the disruption — by comparing transfer quality across languages with and without a dual — but it wouldn’t be able to explicitly reveal dual number as part of the mechanism operating on number\. As such, more work explicitly reverse\-engineering complete mechanisms associated with linguistic features would be useful and elucidate this distinction\. However, this is outside of our current scope\.
Like much work on multilingual LMs, the range of languages we study here is a small subset of human languages and is skewed towards widely spoken ones\. A more typologically diverse sample could potentially increase the richness of the conclusions we can draw\.
Finally our work relies on templatically generated sentences to ensure large amounts of tight minimal pairs\. Such sentences are known to differ from naturalistic sentences in meaningful ways\. As such, extending this work to natural sentences would be a meaningful venture\.
## Acknowledgments
We thank Olaf de Rohan Wilner, Agnese Lombardi, Youn\-Gyu Park, Siyuan Song, Qing Yao, Leticia Hanada, Branimir Boguraev, Venus Shirazy, Maria Helena Fernandez Serrano, and Edoardo Giorgi for providing assistance with template design and judgments on resulting stimuli for the many different languages we study\. We acknowledge funding from NSF CAREER grant 2339729 to Kyle Mahowald\. Julius Steuer received funding from the Klaus Tschira Foundation, Heidelberg, Germany\.
## References
- Arora et al\. \(2024\)Aryaman Arora, Dan Jurafsky, and Christopher Potts\. 2024\.[CausalGym: Benchmarking causal interpretability methods on linguistic tasks](https://doi.org/10.18653/v1/2024.acl-long.785)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 14638–14663, Bangkok, Thailand\. Association for Computational Linguistics\.
- Barr et al\. \(2013\)Dale J\. Barr, R\. Levy, Christoph Scheepers, and Harry J\. Tily\. 2013\.[Random effects structure for confirmatory hypothesis testing: Keep it maximal](https://api.semanticscholar.org/CorpusID:6868055)\.*Journal of memory and language*, 68\.
- Blum et al\. \(2025\)Frederic Blum, Carlos Barrientos, Johannes Englisch, Robert Forkel, Simon J\. Greenhill, Christoph Rzymski, and Johann\-Mattis List\. 2025\.[Lexibank 2: pre\-computed features for large\-scale lexical data](https://doi.org/10.12688/openreseurope.20216.2)\.*Open Research Europe*, 5\(126\):1–27\.
- Blum et al\. \(2026\)Frederic Blum, Carlos Barrientos, Johannes Englisch, Robert Forkel, Simon J\. Greenhill, Christoph Rzymski, and Johann\-Mattis List\. 2026\.[Lexibank 2: pre\-computed features for large\-scale lexical data \(v2\.2\)](https://doi.org/10.5281/zenodo.20712996)\.
- Boguraev and Mahowald \(2026\)Sasha Boguraev and Kyle Mahowald\. 2026\.[Causal drawbridges: Characterizing gradient blocking of syntactic islands in transformer lms](https://arxiv.org/abs/2604.13950)\.*Preprint*, arXiv:2604\.13950\.
- Boguraev et al\. \(2025\)Sasha Boguraev, Christopher Potts, and Kyle Mahowald\. 2025\.[Causal interventions reveal shared structure across English filler–gap constructions](https://doi.org/10.18653/v1/2025.emnlp-main.1271)\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 25021–25042\. Association for Computational Linguistics\.
- Brinkmann et al\. \(2025\)Jannik Brinkmann, Chris Wendler, Christian Bartelt, and Aaron Mueller\. 2025\.[Large language models share representations of latent grammatical concepts across typologically diverse languages](https://doi.org/10.18653/v1/2025.naacl-long.312)\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pages 6131–6150\. Association for Computational Linguistics\.
- Chi et al\. \(2020\)Ethan A\. Chi, John Hewitt, and Christopher D\. Manning\. 2020\.[Finding universal grammatical relations in multilingual BERT](https://doi.org/10.18653/v1/2020.acl-main.493)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 5564–5577\. Association for Computational Linguistics\.
- Chomsky \(1965\)Noam Chomsky\. 1965\.*Aspects of the theory of syntax*\.MIT Press\.
- Comrie \(1989\)Bernard Comrie\. 1989\.*Language universals and linguistic typology: Syntax and morphology*\.University of Chicago press\.
- Dryer and Haspelmath \(2013\)Matthew S\. Dryer and Martin Haspelmath, editors\. 2013\.[*WALS Online \(v2020\.4\)*](https://doi.org/10.5281/zenodo.13950591)\.Zenodo\.
- Ferrando and Costa\-jussà \(2024\)Javier Ferrando and Marta R\. Costa\-jussà\. 2024\.[On the similarity of circuits across languages: a case study on the subject\-verb agreement task](https://doi.org/10.18653/v1/2024.findings-emnlp.591)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 10115–10125\. Association for Computational Linguistics\.
- Futrell and Mahowald \(2026\)Richard Futrell and Kyle Mahowald\. 2026\.[How linguistics learned to stop worrying and love the language models](https://doi.org/10.1017/S0140525X2510112X)\.*Behavioral and Brain Sciences*, 49:e198\.
- Geiger et al\. \(2025\)Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and 1 others\. 2025\.Causal abstraction: A theoretical foundation for mechanistic interpretability\.*Journal of Machine Learning Research*, 26\(83\):1–64\.
- Geiger et al\. \(2024\)Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman\. 2024\.[Finding alignments between interpretable causal variables and distributed neural representations](https://proceedings.mlr.press/v236/geiger24a.html)\.In*Proceedings of the Third Conference on Causal Learning and Reasoning*, volume 236 of*Proceedings of Machine Learning Research*, pages 160–187\. PMLR\.
- Gonen et al\. \(2020\)Hila Gonen, Shauli Ravfogel, Yanai Elazar, and Yoav Goldberg\. 2020\.[It’s not Greek to mBERT: Inducing word\-level translations from multilingual BERT](https://doi.org/10.18653/v1/2020.blackboxnlp-1.5)\.In*Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP*, pages 45–56\. Association for Computational Linguistics\.
- Grattafiori et al\. \(2024\)Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others\. 2024\.[The llama 3 herd of models](https://arxiv.org/abs/2407.21783)\.*Preprint*, arXiv:2407\.21783\.
- Greenberg et al\. \(1963\)Joseph H Greenberg and 1 others\. 1963\.Some universals of grammar with particular reference to the order of meaningful elements\.*Universals of language*, 2\(1\):73–113\.
- Harrasse et al\. \(2026\)Abir Harrasse, Florent Draye, Punya Syon Pandey, Zhijing Jin, and Bernhard Schölkopf\. 2026\.[Tracing multilingual representations in llms with cross\-layer transcoders](https://arxiv.org/abs/2511.10840)\.*Preprint*, arXiv:2511\.10840\.
- Hartsuiker et al\. \(2004\)Robert J Hartsuiker, Martin J Pickering, and Eline Veltkamp\. 2004\.Is syntax separate or shared between languages? cross\-linguistic syntactic priming in spanish\-english bilinguals\.*Psychological science*, 15\(6\):409–414\.
- Hawkins \(2014\)John Hawkins\. 2014\.*Cross\-Linguistic Variation and Efficiency*\.Oxford University Press\.
- Heap et al\. \(2025\)Thomas Heap, Tim Lawson, Lucy Farnik, and Laurence Aitchison\. 2025\.Sparse autoencoders can interpret randomly initialized transformers\.*arXiv e\-prints*, pages arXiv–2501\.
- Heggarty et al\. \(2024\)Paul Heggarty, Cormac Anderson, and Matthew Scarborough\. 2024\.Indo\-european cognate relationships database \(ie\-cor version 1\.1\)\.
- Hewitt and Liang \(2019\)John Hewitt and Percy Liang\. 2019\.[Designing and interpreting probes with control tasks](https://doi.org/10.18653/v1/D19-1275)\.In*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\)*, pages 2733–2743\. Association for Computational Linguistics\.
- Hewitt and Manning \(2019\)John Hewitt and Christopher D\. Manning\. 2019\.[A structural probe for finding syntax in word representations](https://doi.org/10.18653/v1/N19-1419)\.In*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pages 4129–4138\. Association for Computational Linguistics\.
- Hu et al\. \(2020\)Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy\. 2020\.[A systematic assessment of syntactic generalization in neural language models](https://doi.org/10.18653/v1/2020.acl-main.158)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 1725–1744\. Association for Computational Linguistics\.
- Jumelet et al\. \(2026\)Jaap Jumelet, Leonie Weissweiler, Joakim Nivre, and Arianna Bisazza\. 2026\.[MultiBLiMP 1\.0: A massively multilingual benchmark of linguistic minimal pairs](https://doi.org/10.1162/tacl.a.600)\.*Transactions of the Association for Computational Linguistics*, 14:193–216\.
- Kallini et al\. \(2025\)Julie Kallini, Dan Jurafsky, Christopher Potts, and Martijn Bartelds\. 2025\.[False Friends are not foes: Investigating vocabulary overlap in multilingual language models](https://doi.org/10.18653/v1/2025.findings-emnlp.1153)\.In*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 21138–21154\. Association for Computational Linguistics\.
- Katzir \(2023\)Roni Katzir\. 2023\.[Why large language models are poor theories of human linguistic cognition: A reply to Piantadosi](https://doi.org/10.5964/bioling.13153)\.*Biolinguistics*, 17:e13153\.
- Kumon and Yanaka \(2026\)Ryoma Kumon and Hitomi Yanaka\. 2026\.[Fine\-grained analysis of shared syntactic mechanisms in language models](https://doi.org/10.18653/v1/2026.acl-long.2078)\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 44866–44891, San Diego, California, United States\. Association for Computational Linguistics\.
- Kuribayashi et al\. \(2021\)Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, and Kentaro Inui\. 2021\.[Lower perplexity is not always human\-like](https://doi.org/10.18653/v1/2021.acl-long.405)\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 5203–5217\. Association for Computational Linguistics\.
- Kuribayashi et al\. \(2025\)Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin\. 2025\.[Large language models are human\-like internally](https://doi.org/10.1162/tacl.a.58)\.*Transactions of the Association for Computational Linguistics*, 13:1743–1766\.
- Kuznetsova et al\. \(2017\)Alexandra Kuznetsova, Per B\. Brockhoff, and Rune H\. B\. Christensen\. 2017\.[lmerTest package: Tests in linear mixed effects models](https://doi.org/10.18637/jss.v082.i13)\.*Journal of Statistical Software*, 82\(13\):1–26\.
- Leask et al\. \(2025\)Patrick Leask, Bart Bussmann, Michael Pearce, Joseph Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda\. 2025\.Sparse autoencoders do not find canonical units of analysis\.In*International Conference on Learning Representations*, volume 2025, pages 53617–53642\.
- Limisiewicz et al\. \(2023\)Tomasz Limisiewicz, Jiří Balhar, and David Mareček\. 2023\.[Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages](https://doi.org/10.18653/v1/2023.findings-acl.350)\.In*Findings of the Association for Computational Linguistics: ACL 2023*, pages 5661–5681\. Association for Computational Linguistics\.
- Linzen and Baroni \(2021\)Tal Linzen and Marco Baroni\. 2021\.Syntactic structure from deep learning\.*Annual Review of Linguistics*, 7\(1\):195–212\.
- Linzen et al\. \(2016\)Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg\. 2016\.[Assessing the ability of LSTMs to learn syntax\-sensitive dependencies](https://doi.org/10.1162/tacl_a_00115)\.*Transactions of the Association for Computational Linguistics*, 4:521–535\.
- List et al\. \(2022\)Johann\-Mattis List, Robert Forkel, Simon J\. Greenhill, Christopher Rzymski, Johannes Englisch, and Russell D\. Gray\. 2022\.[Lexibank, a public repository of standardized wordlists with computed phonological and lexical features](https://doi.org/10.1038/s41597-022-01432-0)\.*Scientific Data*, 9:316\.
- Littell et al\. \(2017\)Patrick Littell, David R\. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin\. 2017\.[URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors](https://aclanthology.org/E17-2002/)\.In*Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers*, pages 8–14\. Association for Computational Linguistics\.
- Mahowald et al\. \(2024\)Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko\. 2024\.Dissociating language and thought in large language models\.*Trends in cognitive sciences*, 28\(6\):517–540\.
- Malik\-Moraleda et al\. \(2022\)Saima Malik\-Moraleda, Dima Ayyash, Jeanne Gallée, Josef Affourtit, Malte Hoffmann, Zachary Mineroff, Olessia Jouravlev, and Evelina Fedorenko\. 2022\.An investigation across 45 languages and 12 language families reveals a universal language network\.*Nature neuroscience*, 25\(8\):1014–1019\.
- Manning et al\. \(2020\)Christopher D\. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy\. 2020\.[Emergent linguistic structure in artificial neural networks trained by self\-supervision](https://doi.org/10.1073/pnas.1907367117)\.*Proceedings of the National Academy of Sciences*, 117\(48\):30046–30054\.
- Misra \(2022\)Kanishka Misra\. 2022\.minicons: Enabling flexible behavioral and representational analyses of transformer language models\.*arXiv preprint arXiv:2203\.13112*\.
- Misra and Kim \(2026\)Kanishka Misra and Najoung Kim\. 2026\.[A systematic framework for generating novel experimental hypotheses from language models](https://arxiv.org/abs/2408.05086)\.*Preprint*, arXiv:2408\.05086\.
- NLLB Team et al\. \(2024\)NLLB Team, Marta R\. Costa\-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others\. 2024\.[Scaling neural machine translation to 200 languages](https://doi.org/10.1038/s41586-024-07335-x)\.*Nature*, 630\(8018\):841–846\.
- Norcliffe et al\. \(2015\)Elisabeth Norcliffe, Alice C Harris, and T Florian Jaeger\. 2015\.Cross\-linguistic psycholinguistics and its critical role in theory development: Early beginnings and recent advances\.*Language, Cognition and Neuroscience*, 30\(9\):1009–1032\.
- Oh and Schuler \(2023a\)Byung\-Doh Oh and William Schuler\. 2023a\.[Transformer\-based language model surprisal predicts human reading times best with about two billion training tokens](https://doi.org/10.18653/v1/2023.findings-emnlp.128)\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 1915–1921\. Association for Computational Linguistics\.
- Oh and Schuler \(2023b\)Byung\-Doh Oh and William Schuler\. 2023b\.[Why does surprisal from larger transformer\-based language models provide a poorer fit to human reading times?](https://doi.org/10.1162/tacl_a_00548)*Transactions of the Association for Computational Linguistics*, 11:336–350\.
- Oh et al\. \(2024\)Byung\-Doh Oh, Shisen Yue, and William Schuler\. 2024\.[Frequency explains the inverse correlation of large language models’ size, training data amount, and surprisal’s fit to reading times](https://doi.org/10.18653/v1/2024.eacl-long.162)\.In*Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 2644–2663\. Association for Computational Linguistics\.
- Papadimitriou et al\. \(2021\)Isabel Papadimitriou, Ethan A\. Chi, Richard Futrell, and Kyle Mahowald\. 2021\.[Deep subjecthood: Higher\-order grammatical features in multilingual BERT](https://doi.org/10.18653/v1/2021.eacl-main.215)\.In*Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume*, pages 2522–2532\. Association for Computational Linguistics\.
- Paulo and Belrose \(2026\)Gonçalo Paulo and Nora Belrose\. 2026\.[Sparse autoencoders trained on the same data learn different features](https://proceedings.iclr.cc/paper_files/paper/2026/file/3c1fe56b043848b211030c202764c6a7-Paper-Conference.pdf)\.In*International Conference on Learning Representations*, volume 2026, pages 35748–35760\.
- Piantadosi \(2024\)Steven T\. Piantadosi\. 2024\.[Modern language models refute Chomsky’s approach to language](https://doi.org/10.5281/ZENODO.11351540)\.In Edward Gibson and Moshe Poliak, editors,*From fieldwork to linguistic theory: A tribute to Dan Everett*\. Language Science Press\.
- Pires et al\. \(2019\)Telmo Pires, Eva Schlinger, and Dan Garrette\. 2019\.[How multilingual is multilingual BERT?](https://doi.org/10.18653/v1/P19-1493)In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pages 4996–5001\. Association for Computational Linguistics\.
- Raffel et al\. \(2020\)Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J\. Liu\. 2020\.[Exploring the limits of transfer learning with a unified text\-to\-text transformer](http://jmlr.org/papers/v21/20-074.html)\.*Journal of Machine Learning Research*, 21\(140\):1–67\.
- Rust et al\. \(2021\)Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych\. 2021\.[How good is your tokenizer? on the monolingual performance of multilingual language models](https://doi.org/10.18653/v1/2021.acl-long.243)\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 3118–3135\. Association for Computational Linguistics\.
- Salamanca et al\. \(2026\)Alejandro R\. Salamanca, Diana Abagyan, Daniel D’souza, Ammar Khairi, David Mora, Saurabh Dash, Viraat Aryabumi, Sara Rajaee, Mehrnaz Mofakhami, Ananya Sahu, Thomas Euyang, Brittawnya Prince, Madeline Smith, Hangyu Lin, Acyr Locatelli, Sara Hooker, Tom Kocmi, Aidan Gomez, Ivan Zhang, and 7 others\. 2026\.[Tiny aya: Bridging scale and multilingual depth](https://arxiv.org/abs/2603.11510)\.*Preprint*, arXiv:2603\.11510\.
- Shliazhko et al\. \(2024\)Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Anastasia Kozlova, Vladislav Mikhailov, and Tatiana Shavrina\. 2024\.mgpt: Few\-shot learners go multilingual\.*Transactions of the Association for Computational Linguistics*, 12:58–79\.
- Skirgård et al\. \(2023\)Hedvig Skirgård, Hannah J\. Haynie, Damián E\. Blasi, Harald Hammarström, Jeremy Collins, Jay J\. Latarche, Jakob Lesage, Tobias Weber, Alena Witzlack\-Makarevich, Sam Passmore, Angela Chira, Luke Maurits, Russell Dinnage, Michael Dunn, Ger Reesink, Ruth Singer, Claire Bowern, Patience Epps, Jane Hill, and 86 others\. 2023\.[Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss](https://doi.org/10.1126/sciadv.adg6175)\.*Science Advances*, 9\(16\)\.
- Srinivasan et al\. \(2023\)Anirudh Srinivasan, Venkata Subrahmanyan Govindarajan, and Kyle Mahowald\. 2023\.[Counterfactually probing language identity in multilingual models](https://doi.org/10.18653/v1/2023.mrl-1.3)\.In*Proceedings of the 3rd Workshop on Multi\-lingual Representation Learning \(MRL\)*, pages 24–36\. Association for Computational Linguistics\.
- Voita and Titov \(2020\)Elena Voita and Ivan Titov\. 2020\.[Information\-theoretic probing with minimum description length](https://doi.org/10.18653/v1/2020.emnlp-main.14)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 183–196\. Association for Computational Linguistics\.
- Warstadt and Bowman \(2022\)Alex Warstadt and Samuel R Bowman\. 2022\.What artificial neural networks can tell us about human language acquisition\.In*Algebraic structures in natural language*, pages 17–60\. CRC Press\.
- Warstadt et al\. \(2020\)Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng\-Fu Wang, and Samuel R\. Bowman\. 2020\.[BLiMP: The benchmark of linguistic minimal pairs for English](https://doi.org/10.1162/tacl_a_00321)\.*Transactions of the Association for Computational Linguistics*, 8:377–392\.
- Wilcox et al\. \(2018\)Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell\. 2018\.[What do RNN language models learn about filler–gap dependencies?](https://doi.org/10.18653/v1/W18-5423)In*Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP*, pages 211–221\. Association for Computational Linguistics\.
- Wilcox et al\. \(2020\)Ethan G\. Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger P\. Levy\. 2020\.[On the predictive power of neural language models for human real\-time comprehension behavior](https://escholarship.org/uc/item/738338tm)\.In*Proceedings of the Annual Meeting of the Cognitive Science Society*, volume 42\.
- Wilcox et al\. \(2024\)Ethan Gotlieb Wilcox, Richard Futrell, and Roger Levy\. 2024\.Using computational models to test syntactic learnability\.*Linguistic Inquiry*, 55\(4\):805–848\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others\. 2020\.[Transformers: State\-of\-the\-art natural language processing](https://doi.org/10.18653/v1/2020.emnlp-demos.6)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 38–45\. Association for Computational Linguistics\.
- Wu et al\. \(2024\)Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah Goodman, Christopher Manning, and Christopher Potts\. 2024\.[pyvene: A library for understanding and improving PyTorch models via interventions](https://doi.org/10.18653/v1/2024.naacl-demo.16)\.In*Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 3: System Demonstrations\)*, pages 158–165\. Association for Computational Linguistics\.
## Appendix ATemplates
Exemplar stimuli we utilize to train and evaluate our interventions can be seen in[Table5](https://arxiv.org/html/2608.28924#A2.T5)\. We also provide exemplar stimuli for our controls: exemplar stimuli forrandom labelscan be seen in[Table6](https://arxiv.org/html/2608.28924#A2.T6), and exemplar stimuli forOOD labelscan be seen in[Table7](https://arxiv.org/html/2608.28924#A2.T7)\. We publish the full templates on Zenodo\.444[Templates](https://doi.org/10.5281/zenodo.22025512)
## Appendix BIntervention Models
We publish all intervention models on Hugging Face\.555![[Uncaptioned image]](https://arxiv.org/html/2608.28924v1/hf-logo.png)[Intervention models](https://huggingface.co/sashaboguraev/causal-multi-interventions)The code used for model training, evaluation and analysis is available on GitHub\.666[Training and evaluation code](https://github.com/justeuer/multilingual-interventions)
LanguageNP1PPNP2VPEnglishThe man / The mennearthe cabinetis / areSpanishEl hombre / Los hombrescerca dela alacenaes / sonGermanDer Mann / Die Männernebendem Schrankist / sindDutchDe man / De mannennaastde kastis / zijnFrenchL’homme / Les hommesprès dela tableest / sontItalianL’uomo / Gli uominivicino allatavolaè / sonoPortugueseO homem / Os homensperto damesaé / sãoRomanianBărbatul / Bărbațiilângăfereastrăeste / suntBulgarianМъжът / Мъжетедомасатае / саRussianМужчина / Мужчиныпередшкафомявляется / являютсяFarsiمرد / مردانکنارکمداست / هستندGalicianO home / Os homespretoda fiestraé / son
\(a\)Subject–Verb Number Agreement
LanguageNameVPBecausePronounEnglishJames / Maryliedbecausehe / sheSpanishMiguel / Maríacorrióporqueél / ellaGermanThomas / Mariaaßweiler / sieDutchJan / Annaloogomdathij / zijFrenchPierre / Mariechuchotaparce qu’il / elleItalianMario / Mariariseperchélui / leiPortugueseJoão / Mariasussurraporqueele / elaRomanianIon / Mariaa râspentru căel / eaBulgarianИван / Мариявиказащототой / тяRussianИван / Марияходил / ходилапотому чтоон / онаKorean\\krfont철수 군이 / 지은 양이\\krfont떨렸다,\\krfont왜냐하면\\krfont그가 / 그녀가Japanese\\jpfontはるとくんが / さくらちゃんが\\jpfont来た、\\jpfontなぜなら\\jpfont彼が / 彼女がChinese\\zhfont何先生 / 袁夫人\\zhfont笑了,\\zhfont因为\\zhfont他 / 她
\(b\)Anaphoric Pronoun Gender Agreement
LanguagePrefixFillerNPVPGapEnglishI knowthat / whatshesawthe / \.FrenchJe saisqu’ / ce qu’ellea vuça / \.GermanIch weiß,dass / wassie—das / sahDutchIk weetdat / watzij—hem / zagItalianSoche / cosaleivideil / \.RussianЯ знаю,что / когоонавиделаэто / \.BulgarianЗнам,че / каквотявидянего / \.
\(c\)Filler–Gap \(Object\)\.
Table 5:Exemplar minimal pairs per language for each construction\. The minimal differences upstream \(denoted by the two options separated by a ‘/’\) in theNP1,Name, andFillerpositions each correspond to the given labels in theVP,Pronoun, andGappositions respectively\. Such minimal pairs are used to evaluate LM linguistic competence, and to ensure that our causal interventions are successful\. ‘\-’ indicates positions where the given construction has no item\. This appears only in German and Dutch Filler–Gap stimuli as they are V2 languages, mandating any embedded verbs appear clause\-final\. The contrast at the VP position for Russian gender agreement \(ходил/ходила\) reflects the past\-tense verb’s obligatory agreement with the subject, and is not an additional label\.LanguageNP1PPNP2VPEnglishThe man / The mennearthe cabinetdog / giveSpanishEl hombre / Los hombrescerca dela alacenaperro / darGermanDer Mann / Die Männernebendem SchrankHund / gebenDutchDe man / De mannennaastde kasthond / gevenFrenchL’homme / Les hommesprès dela tablechien / donnerItalianL’uomo / Gli uominivicino allatavolacane / darePortugueseO homem / Os homensperto damesacão / darRomanianBărbatul / Bărbațiilângăfereastrăcâine / daBulgarianМъжът / Мъжетедомасатакуче / даваRussianМужчина / Мужчиныпередшкафомсобака / датьFarsiمرد / مردانکنارکمدسگ / دادنGalicianO home / Os homespretoda fiestracan / dar
\(a\)Subject–Verb Number Agreement \(random labels\)
LanguageNameVPBecausePronounEnglishJames / Maryliedbecausedog / giveSpanishMiguel / Maríacorrióporqueperro / darGermanThomas / MariaaßweilHund / gebenDutchJan / Annaloogomdathond / gevenFrenchPierre / Mariechuchotaparce qu’chien / donnerItalianMario / Mariariseperchécane / darePortugueseJoão / Mariasussurraporquecão / darRomanianIon / Mariaa râspentru căcâine / daBulgarianИван / Мариявиказащотокуче / даваRussianИван / Марияходил / ходилапотому чтособака / датьKorean\\krfont철수 군이 / 지은 양이\\krfont떨렸다,\\krfont왜냐하면\\krfont개 / 주다Japanese\\jpfontはるとくんが / さくらちゃんが\\jpfont来た、\\jpfontなぜなら\\jpfont犬 / あげるChinese\\zhfont何先生 / 袁夫人\\zhfont笑了,\\zhfont因为\\zhfont狗 / 给
\(b\)Anaphoric Pronoun Gender Agreement \(random labels\)
LanguagePrefixFillerNPVPGapEnglishI knowthat / whatshesawdog / giveFrenchJe saisqu’ / ce qu’ellea vuchien / donnerGermanIch weiß,dass / wassie—Hund / gebenDutchIk weetdat / watzij—hond / gevenItalianSoche / cosaleividecane / dareRussianЯ знаю,что / когоонавиделасобака / датьBulgarianЗнам,че / каквотявидякуче / дава
\(c\)Filler–Gap \(Object,random labels\)\.
Table 6:Sampled exemplar minimal pairs per language for each construction forrandom labels: the stimulus is identical to the true construction, but the counterfactual target are arbitrary tokens\. We expect a well\-calibrated intervention to be*specific*and not generalize to these\.LanguageNP1PPNP2VPEnglishThe man / The mennearthe cabinethas / haveSpanishEl hombre / Los hombrescerca dela alacenatiene / tienenGermanDer Mann / Die Männernebendem Schrankhat / habenDutchDe man / De mannennaastde kastheeft / hebbenFrenchL’homme / Les hommesprès dela tablea / ontItalianL’uomo / Gli uominivicino allatavolaha / hannoPortugueseO homem / Os homensperto damesatem / têmRomanianBărbatul / Bărbațiilângăfereastrăare / auBulgarianМъжът / Мъжетедомасатаможе / могатRussianМужчина / Мужчиныпередшкафомимеет / имеютFarsiمرد / مردانکنارکمددارد / دارندGalicianO home / Os homespretoda fiestraestá / están
\(a\)Subject–Verb Number Agreement \(OOD labels\)
LanguagePrefixFillerNPVPGapEnglishI knowthat / whatshesawa / ,FrenchJe saisque / ce queellea vucela / ,GermanIch weiß,dass / wassie—das / fandDutchIk weetdat / watzij—hem / vondItalianSoche / cosaleivideun / ,RussianЯ знаю,что / когоонавиделато / ,BulgarianЗнам,че / каквотявидятова / ,
\(b\)Filler–Gap \(Object,OOD labels\)\.
Table 7:Sampled exemplar minimal pairs per language for each construction in theOOD labelscontrol\. As inspired by[Kumon and Yanaka \(2026\)](https://arxiv.org/html/2608.28924#bib.bib30): the stimulus is identical to the true construction, but the counterfactual target is a plausible yet out\-of\-distribution continuation\. We expect a well\-trained intervention to target the abstract feature of interest, and thus generalize to these\.
## Appendix CBehavioral Results
Before attempting to study the internals of the LMs on our three phenomena, we must first confirm that they are behaviorally competent with the given task\. We utilize a targeted syntactic evaluation paradigm inspired by the method developed by[Wilcox et al\. \(2018\)](https://arxiv.org/html/2608.28924#bib.bib63)to study RNN’s competence of filler–gap constructions\. The method is as follows: given two minimal pairs,bbandss, with corresponding labelsℓb\\ell\_\{b\}andℓs\\ell\_\{s\}, an LM should correctly predict that the surprisal,SS, of each given label is higher given the correct context than in the incorrect context\. Formally, the metric, termedaccuracy, is a binary measure that is calculated as:
accuracy=\{1ifS\(ℓb∣b\)<S\(ℓb∣s\)andS\(ℓs∣b\)\>S\(ℓs∣s\),0otherwise\.\\textsc\{accuracy\}=\\begin\{cases\}1&\\text\{if \}S\(\\ell\_\{b\}\\mid b\)<S\(\\ell\_\{b\}\\mid s\)\\\\ &\\text\{ and \}S\(\\ell\_\{s\}\\mid b\)\>S\(\\ell\_\{s\}\\mid s\),\\\\\[4\.0pt\] 0&\\text\{otherwise\.\}\\end\{cases\}The LMs behavioral results can be seen in[Figure5](https://arxiv.org/html/2608.28924#A3.F5)\. Specifically, for each construction we study \(column facets\) and each LM \(row facets\), we show the mean\-averaged performance over 160 randomly sampled minimal pairs\. We can see that our LMs are broadly competent at processing such constructions, scoring at or above 90% accuracy on 90% \(115/128\) of templates and below 80% accuracy only once \(Llama‑3\.2‑1bon Bulgarian number agreement\)\. This gives us evidence that the LMs are syntactically competent with these phenomena, thus licensing our next experiments\.
Digging deeper into these results, it is evident that the expressly multilingual LMs are generally the most performant on the templates, perhaps as expected\. Interestingly, however,Llama\-3\.2\-3bdoes not lag far behind them\.Llama\-3\.2\-1bis the least performant of all the models on the task, however still showing strong performance\. We do not see any systematic effects of language family, although we do note that the majority of these languages, and families, are high\-resource and thus likely not data scarce during training\.
Figure 5:Behavioral performance of each model on each of our templatic stimuli\. We broadly see strong performance from all of our models on all templates\. Generally multilingual LMs perform the best, but bothLlamamodels show strong performance as well\. All bars but thirteen are above 90% performance, with twelve of those above 80%\. The only bar below 80% isLlama‑3\.2‑1bon Bulgarian number agreement\.
## Appendix DMax\-Odds
The bar charts with the normalizedMax Oddsfor each of our five experimental groups in Experiment One can be seen in[Figure6](https://arxiv.org/html/2608.28924#A4.F6)\. We further provide the Holm\-Bonferroni corrected p\-values of all reported comparisons for the main results in[Table8](https://arxiv.org/html/2608.28924#A4.T8), and for theOOD labelsin[Table9](https://arxiv.org/html/2608.28924#A4.T9)\.
Filler–Gap \(Object\)Subject Verb Number AgreementAnaphoric Pronoun Gender AgreementModelFillerNPVerbNP1PPNP2NAMEVPbecauseOther phenomena\(Control:specificity\)Llama\-3\.2\-1b1\.2×𝟏𝟎−𝟗∗∗∗\\mathbf\{1\.2\\times 10^\{\-9\}\}^\{\*\*\*\}5\.2×𝟏𝟎−𝟕∗∗∗\\mathbf\{5\.2\\times 10^\{\-7\}\}^\{\*\*\*\}1\.5×𝟏𝟎−𝟕∗∗∗\\mathbf\{1\.5\\times 10^\{\-7\}\}^\{\*\*\*\}3\.0×𝟏𝟎−𝟒𝟎∗∗∗\\mathbf\{3\.0\\times 10^\{\-40\}\}^\{\*\*\*\}2\.2×𝟏𝟎−𝟐𝟖∗∗∗\\mathbf\{2\.2\\times 10^\{\-28\}\}^\{\*\*\*\}1\.9×𝟏𝟎−𝟑𝟒∗∗∗\\mathbf\{1\.9\\times 10^\{\-34\}\}^\{\*\*\*\}1\.3×𝟏𝟎−𝟑𝟐∗∗∗\\mathbf\{1\.3\\times 10^\{\-32\}\}^\{\*\*\*\}5\.6×𝟏𝟎−𝟑𝟎∗∗∗\\mathbf\{5\.6\\times 10^\{\-30\}\}^\{\*\*\*\}2\.7×𝟏𝟎−𝟐𝟐∗∗∗\\mathbf\{2\.7\\times 10^\{\-22\}\}^\{\*\*\*\}Llama\-3\.2\-3b1\.9×𝟏𝟎−𝟕∗∗∗\\mathbf\{1\.9\\times 10^\{\-7\}\}^\{\*\*\*\}4\.7×𝟏𝟎−𝟔∗∗∗\\mathbf\{4\.7\\times 10^\{\-6\}\}^\{\*\*\*\}5\.9×𝟏𝟎−𝟗∗∗∗\\mathbf\{5\.9\\times 10^\{\-9\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟒𝟏∗∗∗\\mathbf\{1\.2\\times 10^\{\-41\}\}^\{\*\*\*\}1\.7×𝟏𝟎−𝟑𝟓∗∗∗\\mathbf\{1\.7\\times 10^\{\-35\}\}^\{\*\*\*\}1\.4×𝟏𝟎−𝟒𝟎∗∗∗\\mathbf\{1\.4\\times 10^\{\-40\}\}^\{\*\*\*\}1\.3×𝟏𝟎−𝟑𝟏∗∗∗\\mathbf\{1\.3\\times 10^\{\-31\}\}^\{\*\*\*\}2\.4×𝟏𝟎−𝟑𝟓∗∗∗\\mathbf\{2\.4\\times 10^\{\-35\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟑𝟑∗∗∗\\mathbf\{1\.2\\times 10^\{\-33\}\}^\{\*\*\*\}mGPT1\.0×𝟏𝟎−𝟓∗∗∗\\mathbf\{1\.0\\times 10^\{\-5\}\}^\{\*\*\*\}6\.1×𝟏𝟎−𝟔∗∗∗\\mathbf\{6\.1\\times 10^\{\-6\}\}^\{\*\*\*\}3\.9×𝟏𝟎−𝟕∗∗∗\\mathbf\{3\.9\\times 10^\{\-7\}\}^\{\*\*\*\}2\.8×𝟏𝟎−𝟑𝟓∗∗∗\\mathbf\{2\.8\\times 10^\{\-35\}\}^\{\*\*\*\}8\.3×𝟏𝟎−𝟑𝟔∗∗∗\\mathbf\{8\.3\\times 10^\{\-36\}\}^\{\*\*\*\}2\.4×𝟏𝟎−𝟑𝟑∗∗∗\\mathbf\{2\.4\\times 10^\{\-33\}\}^\{\*\*\*\}3\.1×𝟏𝟎−𝟐𝟐∗∗∗\\mathbf\{3\.1\\times 10^\{\-22\}\}^\{\*\*\*\}8\.3×𝟏𝟎−𝟐𝟓∗∗∗\\mathbf\{8\.3\\times 10^\{\-25\}\}^\{\*\*\*\}1\.9×𝟏𝟎−𝟑𝟐∗∗∗\\mathbf\{1\.9\\times 10^\{\-32\}\}^\{\*\*\*\}tiny\-aya\-base5\.9×𝟏𝟎−𝟗∗∗∗\\mathbf\{5\.9\\times 10^\{\-9\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟗∗∗∗\\mathbf\{1\.2\\times 10^\{\-9\}\}^\{\*\*\*\}6\.9×𝟏𝟎−𝟗∗∗∗\\mathbf\{6\.9\\times 10^\{\-9\}\}^\{\*\*\*\}8\.4×𝟏𝟎−𝟓𝟖∗∗∗\\mathbf\{8\.4\\times 10^\{\-58\}\}^\{\*\*\*\}1\.3×𝟏𝟎−𝟑𝟗∗∗∗\\mathbf\{1\.3\\times 10^\{\-39\}\}^\{\*\*\*\}3\.1×𝟏𝟎−𝟓𝟐∗∗∗\\mathbf\{3\.1\\times 10^\{\-52\}\}^\{\*\*\*\}1\.0×𝟏𝟎−𝟒𝟒∗∗∗\\mathbf\{1\.0\\times 10^\{\-44\}\}^\{\*\*\*\}1\.6×𝟏𝟎−𝟒𝟎∗∗∗\\mathbf\{1\.6\\times 10^\{\-40\}\}^\{\*\*\*\}6\.6×𝟏𝟎−𝟒𝟗∗∗∗\\mathbf\{6\.6\\times 10^\{\-49\}\}^\{\*\*\*\}Random Labels\(Control:selectivity\)Llama\-3\.2\-1b2\.0×𝟏𝟎−𝟒∗∗∗\\mathbf\{2\.0\\times 10^\{\-4\}\}^\{\*\*\*\}3\.9×𝟏𝟎−𝟓∗∗∗\\mathbf\{3\.9\\times 10^\{\-5\}\}^\{\*\*\*\}1\.3×𝟏𝟎−𝟔∗∗∗\\mathbf\{1\.3\\times 10^\{\-6\}\}^\{\*\*\*\}1\.7×𝟏𝟎−𝟑𝟏∗∗∗\\mathbf\{1\.7\\times 10^\{\-31\}\}^\{\*\*\*\}8\.0×𝟏𝟎−𝟏𝟔∗∗∗\\mathbf\{8\.0\\times 10^\{\-16\}\}^\{\*\*\*\}5\.9×𝟏𝟎−𝟏𝟕∗∗∗\\mathbf\{5\.9\\times 10^\{\-17\}\}^\{\*\*\*\}4\.1×𝟏𝟎−𝟑𝟑∗∗∗\\mathbf\{4\.1\\times 10^\{\-33\}\}^\{\*\*\*\}2\.2×𝟏𝟎−𝟐𝟓∗∗∗\\mathbf\{2\.2\\times 10^\{\-25\}\}^\{\*\*\*\}1\.0×𝟏𝟎−𝟏𝟐∗∗∗\\mathbf\{1\.0\\times 10^\{\-12\}\}^\{\*\*\*\}Llama\-3\.2\-3b4\.1×𝟏𝟎−𝟒∗∗∗\\mathbf\{4\.1\\times 10^\{\-4\}\}^\{\*\*\*\}5\.4×𝟏𝟎−𝟒∗∗∗\\mathbf\{5\.4\\times 10^\{\-4\}\}^\{\*\*\*\}7\.4×𝟏𝟎−𝟖∗∗∗\\mathbf\{7\.4\\times 10^\{\-8\}\}^\{\*\*\*\}7\.0×𝟏𝟎−𝟑𝟐∗∗∗\\mathbf\{7\.0\\times 10^\{\-32\}\}^\{\*\*\*\}4\.1×𝟏𝟎−𝟐𝟓∗∗∗\\mathbf\{4\.1\\times 10^\{\-25\}\}^\{\*\*\*\}1\.4×𝟏𝟎−𝟐𝟓∗∗∗\\mathbf\{1\.4\\times 10^\{\-25\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟑𝟏∗∗∗\\mathbf\{1\.2\\times 10^\{\-31\}\}^\{\*\*\*\}8\.7×𝟏𝟎−𝟑𝟏∗∗∗\\mathbf\{8\.7\\times 10^\{\-31\}\}^\{\*\*\*\}3\.1×𝟏𝟎−𝟐𝟒∗∗∗\\mathbf\{3\.1\\times 10^\{\-24\}\}^\{\*\*\*\}mGPT0\.002∗∗\\mathbf\{0\.002^\{\*\*\}\}0\.004∗∗\\mathbf\{0\.004^\{\*\*\}\}2\.6×𝟏𝟎−𝟒∗∗∗\\mathbf\{2\.6\\times 10^\{\-4\}\}^\{\*\*\*\}1\.8×𝟏𝟎−𝟐𝟒∗∗∗\\mathbf\{1\.8\\times 10^\{\-24\}\}^\{\*\*\*\}7\.6×𝟏𝟎−𝟏𝟓∗∗∗\\mathbf\{7\.6\\times 10^\{\-15\}\}^\{\*\*\*\}7\.0×𝟏𝟎−𝟐𝟑∗∗∗\\mathbf\{7\.0\\times 10^\{\-23\}\}^\{\*\*\*\}1\.8×𝟏𝟎−𝟐𝟒∗∗∗\\mathbf\{1\.8\\times 10^\{\-24\}\}^\{\*\*\*\}6\.7×𝟏𝟎−𝟏𝟗∗∗∗\\mathbf\{6\.7\\times 10^\{\-19\}\}^\{\*\*\*\}1\.7×𝟏𝟎−𝟏𝟐∗∗∗\\mathbf\{1\.7\\times 10^\{\-12\}\}^\{\*\*\*\}tiny\-aya\-base4\.0×𝟏𝟎−𝟒∗∗∗\\mathbf\{4\.0\\times 10^\{\-4\}\}^\{\*\*\*\}4\.1×𝟏𝟎−𝟒∗∗∗\\mathbf\{4\.1\\times 10^\{\-4\}\}^\{\*\*\*\}5\.4×𝟏𝟎−𝟕∗∗∗\\mathbf\{5\.4\\times 10^\{\-7\}\}^\{\*\*\*\}1\.1×𝟏𝟎−𝟒𝟖∗∗∗\\mathbf\{1\.1\\times 10^\{\-48\}\}^\{\*\*\*\}3\.3×𝟏𝟎−𝟐𝟑∗∗∗\\mathbf\{3\.3\\times 10^\{\-23\}\}^\{\*\*\*\}2\.3×𝟏𝟎−𝟑𝟓∗∗∗\\mathbf\{2\.3\\times 10^\{\-35\}\}^\{\*\*\*\}6\.6×𝟏𝟎−𝟒𝟓∗∗∗\\mathbf\{6\.6\\times 10^\{\-45\}\}^\{\*\*\*\}2\.6×𝟏𝟎−𝟑𝟕∗∗∗\\mathbf\{2\.6\\times 10^\{\-37\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟑𝟕∗∗∗\\mathbf\{1\.2\\times 10^\{\-37\}\}^\{\*\*\*\}Cross Language \(Within\-Family\)\(Test:within\-family vs\. all\-language transfer\)Llama\-3\.2\-1b0\.8430\.8431\.0001\.0001\.0001\.0000\.081†0\.081^\{\\dagger\}0\.013∗\\mathbf\{0\.013^\{\*\}\}0\.2780\.2780\.066†0\.066^\{\\dagger\}0\.2080\.2080\.9410\.941Llama\-3\.2\-3b0\.8430\.8431\.0001\.0001\.0001\.0000\.039∗\\mathbf\{0\.039^\{\*\}\}0\.007∗∗\\mathbf\{0\.007^\{\*\*\}\}0\.001∗∗\\mathbf\{0\.001^\{\*\*\}\}0\.003∗∗\\mathbf\{0\.003^\{\*\*\}\}0\.039∗\\mathbf\{0\.039^\{\*\}\}0\.7380\.738mGPT0\.8430\.8431\.0001\.0001\.0001\.0000\.2170\.2170\.002∗∗\\mathbf\{0\.002^\{\*\*\}\}0\.4220\.4224\.7×𝟏𝟎−𝟓∗∗∗\\mathbf\{4\.7\\times 10^\{\-5\}\}^\{\*\*\*\}0\.002∗∗\\mathbf\{0\.002^\{\*\*\}\}0\.3820\.382tiny\-aya\-base0\.7990\.7991\.0001\.0001\.0001\.0000\.019∗\\mathbf\{0\.019^\{\*\}\}0\.003∗∗\\mathbf\{0\.003^\{\*\*\}\}0\.4220\.4221\.9×𝟏𝟎−𝟒∗∗∗\\mathbf\{1\.9\\times 10^\{\-4\}\}^\{\*\*\*\}0\.007∗∗\\mathbf\{0\.007^\{\*\*\}\}0\.8430\.843
Table 8:Holm\-bonferroni correctedpp\-values \(correction applied within each comparison type, pooled across phenomena/models/positions\) for all pairwisett\-tests comparing cross\-lingual transfer \(Cross Language \(all\)\) against our two selectivity controls and against family\-restricted transfer, for Experiment One \([Section4](https://arxiv.org/html/2608.28924#S4);[Figure6](https://arxiv.org/html/2608.28924#A4.F6)\)\.Random LabelsandOther phenomenaare our controls: cross\-lingual transfer should, and does, significantly exceed both\.Cross Language \(Within\-Family\)instead tests whether transfer is stronger among typologically related languages than to the full evaluated set; results here are mixed\.†p<\.1\{\}^\{\\dagger\}p<\.1;∗p<\.05\{\}^\{\*\}p<\.05;p∗∗<\.01\{\}^\{\*\*\}p<\.01;∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001; significant cells bolded\.Filler–Gap \(Object\)Subject Verb Number AgreementModelFillerNPVerbNP1PPNP2Other phenomena\(Control:specificity\)Llama\-3\.2\-1b3\.3×𝟏𝟎−𝟕∗∗∗\\mathbf\{3\.3\\times 10^\{\-7\}\}^\{\*\*\*\}2\.5×𝟏𝟎−𝟔∗∗∗\\mathbf\{2\.5\\times 10^\{\-6\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟔∗∗∗\\mathbf\{1\.2\\times 10^\{\-6\}\}^\{\*\*\*\}5\.2×𝟏𝟎−𝟑𝟗∗∗∗\\mathbf\{5\.2\\times 10^\{\-39\}\}^\{\*\*\*\}7\.2×𝟏𝟎−𝟑𝟐∗∗∗\\mathbf\{7\.2\\times 10^\{\-32\}\}^\{\*\*\*\}3\.3×𝟏𝟎−𝟒𝟎∗∗∗\\mathbf\{3\.3\\times 10^\{\-40\}\}^\{\*\*\*\}Llama\-3\.2\-3b4\.4×𝟏𝟎−𝟔∗∗∗\\mathbf\{4\.4\\times 10^\{\-6\}\}^\{\*\*\*\}3\.2×𝟏𝟎−𝟔∗∗∗\\mathbf\{3\.2\\times 10^\{\-6\}\}^\{\*\*\*\}6\.5×𝟏𝟎−𝟕∗∗∗\\mathbf\{6\.5\\times 10^\{\-7\}\}^\{\*\*\*\}3\.0×𝟏𝟎−𝟒𝟐∗∗∗\\mathbf\{3\.0\\times 10^\{\-42\}\}^\{\*\*\*\}2\.4×𝟏𝟎−𝟒𝟎∗∗∗\\mathbf\{2\.4\\times 10^\{\-40\}\}^\{\*\*\*\}1\.1×𝟏𝟎−𝟒𝟑∗∗∗\\mathbf\{1\.1\\times 10^\{\-43\}\}^\{\*\*\*\}mGPT0\.042∗\\mathbf\{0\.042^\{\*\}\}8\.9×𝟏𝟎−𝟔∗∗∗\\mathbf\{8\.9\\times 10^\{\-6\}\}^\{\*\*\*\}4\.3×𝟏𝟎−𝟔∗∗∗\\mathbf\{4\.3\\times 10^\{\-6\}\}^\{\*\*\*\}5\.2×𝟏𝟎−𝟑𝟎∗∗∗\\mathbf\{5\.2\\times 10^\{\-30\}\}^\{\*\*\*\}5\.3×𝟏𝟎−𝟒𝟒∗∗∗\\mathbf\{5\.3\\times 10^\{\-44\}\}^\{\*\*\*\}4\.3×𝟏𝟎−𝟓𝟓∗∗∗\\mathbf\{4\.3\\times 10^\{\-55\}\}^\{\*\*\*\}tiny\-aya\-base8\.9×𝟏𝟎−𝟔∗∗∗\\mathbf\{8\.9\\times 10^\{\-6\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟖∗∗∗\\mathbf\{1\.2\\times 10^\{\-8\}\}^\{\*\*\*\}1\.6×𝟏𝟎−𝟗∗∗∗\\mathbf\{1\.6\\times 10^\{\-9\}\}^\{\*\*\*\}1\.1×𝟏𝟎−𝟓𝟏∗∗∗\\mathbf\{1\.1\\times 10^\{\-51\}\}^\{\*\*\*\}5\.7×𝟏𝟎−𝟒𝟓∗∗∗\\mathbf\{5\.7\\times 10^\{\-45\}\}^\{\*\*\*\}2\.4×𝟏𝟎−𝟓𝟗∗∗∗\\mathbf\{2\.4\\times 10^\{\-59\}\}^\{\*\*\*\}Random Labels\(Control:selectivity\)Llama\-3\.2\-1b0\.4270\.4270\.011∗\\mathbf\{0\.011^\{\*\}\}7\.8×𝟏𝟎−𝟓∗∗∗\\mathbf\{7\.8\\times 10^\{\-5\}\}^\{\*\*\*\}1\.6×𝟏𝟎−𝟑𝟏∗∗∗\\mathbf\{1\.6\\times 10^\{\-31\}\}^\{\*\*\*\}7\.2×𝟏𝟎−𝟏𝟗∗∗∗\\mathbf\{7\.2\\times 10^\{\-19\}\}^\{\*\*\*\}2\.3×𝟏𝟎−𝟐𝟐∗∗∗\\mathbf\{2\.3\\times 10^\{\-22\}\}^\{\*\*\*\}Llama\-3\.2\-3b0\.4270\.4270\.4270\.4271\.8×𝟏𝟎−𝟒∗∗∗\\mathbf\{1\.8\\times 10^\{\-4\}\}^\{\*\*\*\}1\.2×𝟏𝟎−𝟑𝟐∗∗∗\\mathbf\{1\.2\\times 10^\{\-32\}\}^\{\*\*\*\}4\.0×𝟏𝟎−𝟑𝟎∗∗∗\\mathbf\{4\.0\\times 10^\{\-30\}\}^\{\*\*\*\}3\.3×𝟏𝟎−𝟐𝟖∗∗∗\\mathbf\{3\.3\\times 10^\{\-28\}\}^\{\*\*\*\}mGPT0\.4270\.4270\.4880\.4880\.1550\.1551\.2×𝟏𝟎−𝟏𝟗∗∗∗\\mathbf\{1\.2\\times 10^\{\-19\}\}^\{\*\*\*\}6\.3×𝟏𝟎−𝟐𝟐∗∗∗\\mathbf\{6\.3\\times 10^\{\-22\}\}^\{\*\*\*\}2\.9×𝟏𝟎−𝟒𝟐∗∗∗\\mathbf\{2\.9\\times 10^\{\-42\}\}^\{\*\*\*\}tiny\-aya\-base0\.3270\.3270\.4880\.4883\.4×𝟏𝟎−𝟓∗∗∗\\mathbf\{3\.4\\times 10^\{\-5\}\}^\{\*\*\*\}8\.7×𝟏𝟎−𝟒𝟐∗∗∗\\mathbf\{8\.7\\times 10^\{\-42\}\}^\{\*\*\*\}8\.9×𝟏𝟎−𝟑𝟎∗∗∗\\mathbf\{8\.9\\times 10^\{\-30\}\}^\{\*\*\*\}7\.6×𝟏𝟎−𝟒𝟐∗∗∗\\mathbf\{7\.6\\times 10^\{\-42\}\}^\{\*\*\*\}Cross Language \(Within\-Family\)\(Test:within\-family vs\. all\-language transfer\)Llama\-3\.2\-1b1\.0001\.0001\.0001\.0001\.0001\.0000\.7830\.7830\.083†0\.083^\{\\dagger\}0\.4420\.442Llama\-3\.2\-3b1\.0001\.0001\.0001\.0001\.0001\.0000\.3040\.3040\.1330\.1330\.1480\.148mGPT1\.0001\.0000\.8390\.8391\.0001\.0000\.3040\.3040\.023∗\\mathbf\{0\.023^\{\*\}\}0\.3110\.311tiny\-aya\-base1\.0001\.0001\.0001\.0001\.0001\.0000\.012∗\\mathbf\{0\.012^\{\*\}\}0\.082†0\.082^\{\\dagger\}0\.4420\.442
Table 9:Holm\-bonferroni correctedpp\-values \(correction applied within each comparison type, pooled across phenomena/models/positions\) for theOOD labels\([§E](https://arxiv.org/html/2608.28924#A5);[Figure7](https://arxiv.org/html/2608.28924#A5.F7)\)\. As in[Table8](https://arxiv.org/html/2608.28924#A4.T8),Random LabelsandOther phenomenaare selectivity controls compared againstCross Language \(all\)transfer, andCross Language \(Within\-Family\)tests family\-restricted vs\. all\-language transfer\.†p<\.1\{\}^\{\\dagger\}p<\.1;∗p<\.05\{\}^\{\*\}p<\.05;p∗∗<\.01\{\}^\{\*\*\}p<\.01;∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001; significant cells bolded\.Figure 6:We take theMax Oddsacross layers for each model at each position, before normalizing by the within\-language value \(red bar\)\. We see strong performance in both our cross\-language \(green bar\) and within\-family \(yellow bar\) transfer conditions, with both being significantly above both controls in all cases when tested with t\-tests with a Holm\-Bonferroni correction\. We further find the within\-family transfer to be significantly higher than the cross\-language bar in a select few comparisons, but for this to not hold more broadly\.
## Appendix EOut\-of\-Distribution Controls
As described in[Section3\.1](https://arxiv.org/html/2608.28924#S3.SS1), we also evaluate the generalization of our interventions to a set of lexically disjoint labels, which are still grammatically correct and exhibiting the same phenomenon we are testing\. As noted above, we can only do this for two of our three phenomena — subject–verb number agreement and filler–gap object extraction — because there is not an alternative pronoun we can replace the ‘he/she’ labels in our templates with robustly\.
Our results can be seen in[Figure7](https://arxiv.org/html/2608.28924#A5.F7)\. Broadly, we see the same pattern as in our main set of results\. Generalization across languages is consistently significantly higher than our across\-phenomenon control for both phenomena, and consistently significantly higher than transfer to our random\-label control in the case of subject–verb agreement\. In the filler–gap case, we find generalization to OOD labels directionally correct for all comparisons\. However, accompanying t\-tests are broadly not significantly at most positions — only showing significance at theVPposition for bothLlamamodels andtiny\-aya\-baseand theNPposition forLlama\-3\.2\-1b\. We attribute this to two main causes\. First, there is a relatively lesser amount of samples for this phenomenon, which can make similar effect sizes less statistically significant, and secondly the raw transferOddsare similar cross\-linguistically as in the other phenomena, but the raw within\-language transfer is much higher\. As such, the normalizedMax Oddsare artificially deflated, leading to closer comparisons\. These two effects likely cause the lack of significance we see\. We also find there to be little difference between the unselective cross\-lingual condition, and the family limited cross\-lingual condition, with there being significant differences only at theNP1position andPPposition fortiny\-aya\-baseandmGPTrespectively\.
These results, in broadly matching what we find on our critical stimuli, provide us confidence that our interventions are not merely overfit to the lexical items we utilize in our templates, and instead indicative of general syntactic mechanisms we have uncovered\.
Figure 7:We evaluate our interventions on syntactically appropriate but lexically disjoint labels from our training set\. We find our results to broadly hold in this setting\.
## Appendix FLinear\-Mixed Effects Modeling Details
We perform all regressions with thelmerTestpackage in R\([Kuznetsova et al\., 2017](https://arxiv.org/html/2608.28924#bib.bib33)\)\. We fit a linear mixed\-effects model at each position with our dependent variable as themax oddsat each training — evaluation\-set pair\. We treat the LM, training language, and evaluation language as crossed random effects\. Our fixed effects are the typological distance between the given languages, model size, and their interaction, with both scaled by z\-scoring before regression\. To estimate theβ\\betacoefficients, we fit our models with REML, and to estimate the significance of each term, we refit them with ML to calculate the likelihood\-ratio test for each fixed\-effect term against a nested model with that term removed\. We fit the full model as per[Barr et al\. \(2013\)](https://arxiv.org/html/2608.28924#bib.bib2), namely with maximal random effects structure, removing collinear slopes until our model converges — in this case this leaves us with an intercepts\-only model\. The final full model formula can be seen in[Equation1](https://arxiv.org/html/2608.28924#A6.E1)\.
max\_odds∼\\displaystyle\\text\{max\\\_odds\}\\sim\{\}typological\_distance∗model\_size\\displaystyle\\text\{typological\\\_distance\}\*\\text\{model\\\_size\}\(1\)\+\(1∣train\_lang\)\\displaystyle\+\(1\\mid\\text\{train\\\_lang\}\)\+\(1∣eval\_lang\)\\displaystyle\+\(1\\mid\\text\{eval\\\_lang\}\)\+\(1∣model\)\\displaystyle\+\(1\\mid\\text\{model\}\)
## Appendix GMetric Coverage
As mentioned in[Section5](https://arxiv.org/html/2608.28924#S5), the metrics we utilize to calculate typological distance do not have perfect coverage over our languages\. We report the coverage in[Table10](https://arxiv.org/html/2608.28924#A7.T10)\. Grambank does not have entries for German, Bulgarian, Spanish, and Romanian because it focuses on typological breadth and diversity\. Furthermore, the cognate data we utilize, Cognate Proportion \(IECOR\) is limited to Indo\-European languages, so it is unavailable for Galician, Chinese, Japanese, and Korean\.
MetricEnglish \(en\)
German \(de\)
Dutch \(nl\)
French \(fr\)
Italian \(it\)
Bulgarian \(bg\)
Russian \(ru\)
Spanish \(es\)
Portuguese \(pt\)
Romanian \(ro\)
Galician \(gl\)
Farsi \(fa\)
Chinese \(zh\)
Japanese \(ja\)
Korean \(ko\)
Template Overlap✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓Corpus Overlap✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓Feature Overlap \(Grambank\)✓✗✓✓✓✗✓✗✓✗✓✓✓✓✓Cognate Proportion✓✓✓✓✓✓✓✓✓✓✗✓✗✗✗Typological Dist\.✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓Table 10:Language coverage for each distance/overlap metric used in the typological\-distance correlation analysis\. Feature Overlap \(Grambank\) is missing for German, Bulgarian, Spanish, and Romanian because Grambank lacks entries for them\. The cognate data we utilize, Cognate Proportion \(IECOR\) is limited to Indo\-European languages, so it is unavailable for Galician, Chinese, Japanese, and Korean\.
## Appendix HComputational Details
We access all models studied through thetransformerspackage[Wolf et al\. \(2020\)](https://arxiv.org/html/2608.28924#bib.bib66), perform behavioral evaluations using theminiconspackage\([Misra, 2022](https://arxiv.org/html/2608.28924#bib.bib43)\)to compute surprisals and train DAS using thepyvenepackage\([Wu et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib67)\)\. Training and evaluation is done using the same hyperparameters found in the extensive sweep performed by[Arora et al\. \(2024\)](https://arxiv.org/html/2608.28924#bib.bib1)\. All experiments were ran on a cluster of 4 NVIDIA A40 \(46 GB each, driver 610\.43\.02\) GPUs, 64 CPU cores, 251 GB RAM, Linux/RHEL 8\. Experiments were run as independent single\-GPU jobs, up to 4 concurrently\. Computational estimates on this hardware can be seen in[Table11](https://arxiv.org/html/2608.28924#A8.T11)\.
TrainingEvaluationModel / PhenomenonRuns\# InterventionsSeconds per InterventionTime \(h\)Pairs\# InterventionsSeconds per InterventionTime \(h\)mGPT\(24 layers\)number agreement12728\.22\.0579721\.9422\.4gender agreement13728\.22\.1534722\.8330\.3filler–gap7728\.11\.1294725\.1430\.2total32––5\.31407––82\.9tiny\-aya\-base\(36 layers\)number agreement1210816\.96\.15791083\.0953\.6gender agreement1310816\.66\.55341083\.0649\.1filler–gap710816\.93\.52941084\.6440\.9total32––16\.11407––143\.6Llama\-3\.2\-1b\(16 layers\)number agreement12489\.21\.5579482\.3117\.9gender agreement134810\.61\.8534483\.2222\.9filler–gap7489\.50\.9294485\.5021\.6total32––4\.21407––62\.3Llama\-3\.2\-3b\(28 layers\)number agreement128416\.04\.5579842\.9940\.4gender agreement138420\.96\.3534843\.2840\.9filler–gap78415\.42\.5294846\.9347\.6total32––13\.31407––128\.9ALL128––≈\\approx395628––≈\\approx418
Table 11:Per\-model compute cost on a single NVIDIA A40 GPU\. A run is one language, with interventions at every position and layer\. As there are three positions evaluated for each phenomenon, this just totals three times the number of layers in each model\. Times reported reflect the median value of all runs performed during our experiments\.
## Appendix ILLM Usage Statement
LLMs were utilized in the development of code and analyses, particularly for the adaptation of related codebases\([Arora et al\., 2024](https://arxiv.org/html/2608.28924#bib.bib1);[Boguraev et al\., 2025](https://arxiv.org/html/2608.28924#bib.bib6), namely those of \)to our work, as well as generating initial versions of some templates before refinement with native speaker informants\. LLMs were further used to provide feedback on the clarity of our writing during the editing process\.Similar Articles
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
This paper tests whether increasing interventional data in pretraining improves LLMs' causal direction reasoning, using controlled Simpson's-paradox worlds. It finds that the training mixture does not govern interventional evidence use; instead the evidence type in the inference-time context is the decisive factor.
An In-Vitro Study on Cross-Lingual Generalization in Language Models
This paper introduces an in-vitro framework with two procedurally generated languages to study cross-lingual generalization in language models, finding that tokenization's preservation of reusable substructure is more critical than lexical similarity or data balance for transferring capabilities across languages.
Latent Mechanisms of Language Control in Multilingual Language Models
This paper compares three methods to identify language-controlling latents in multilingual language models to address code-switching, with experiments on Gemma-2-2B and Qwen3-4B showing FreqSel as the most effective.
When English Isn't the Best Teacher: Source Language Effects in Cross-Lingual In-Context Learning
This paper empirically studies cross-lingual transfer in in-context learning across seven tasks, six models, and typologically diverse languages, showing that fine-tuning based expectations do not consistently apply and offering new heuristics for source language selection.
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
This paper proposes a causal framework for probing internal visual representations in Multimodal Large Language Models, revealing differences in how entities and abstract concepts are encoded. The study highlights that increasing model depth is crucial for encoding abstract concepts and uncovers a disconnect between perception and reasoning in current MLLMs.