Vision-Language Models are Fragile Multilingual Associators
Summary
This paper introduces M2BIND, a benchmark to evaluate whether vision-language models maintain stable visual-linguistic associations across languages. It finds that binding is not language-invariant, with cross-family and cross-script settings causing significant performance collapse and weaker internal causal binding.
View Cached Full Text
Cached at: 08/14/26, 09:24 AM
# Vision-Language Models are Fragile Multilingual Associators
Source: [https://arxiv.org/html/2608.12333](https://arxiv.org/html/2608.12333)
Ritabrata Chakraborty1,4,5,Rajatsubhra Chakraborty2, Shivakumara Palaiahnakote3,Angelo Cangelosi4,Umapada Pal5
1Manipal University Jaipur, India2University of North Carolina Charlotte, USA 3University of Salford, UK4University of Manchester, UK 5Indian Statistical Institute Kolkata, India [https://ritabrata04\.github\.io/m2bind/](https://ritabrata04.github.io/m2bind/)
###### Abstract
Vision\-language models must associate visual entities with textual attributes\. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored\. We introduce M2BIND, a benchmark varying the language of the context and query across multiple languages\. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions\. We find that binding is not language\-invariant: cross\-family and cross\-script settings trigger significant binding collapse, with the model’s internal binding computation shifting to later layers and losing causal strength\. Closely related languages preserve associations comparatively better\. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation\.
Vision\-Language Models are Fragile Multilingual Associators
Ritabrata Chakraborty1,4,5, Rajatsubhra Chakraborty2,Shivakumara Palaiahnakote3,Angelo Cangelosi4,Umapada Pal51Manipal University Jaipur, India2University of North Carolina Charlotte, USA3University of Salford, UK4University of Manchester, UK5Indian Statistical Institute Kolkata, India[https://ritabrata04\.github\.io/m2bind/](https://ritabrata04.github.io/m2bind/)
## 1Introduction
Figure 1:Concept binding across languages\.\(Top\)The world map situates the languages in our study across continents, families, and scripts\.\(Bottom\)AFrenchspeaker who also readsEnglisheffortlessly mapsjauneonto the same concept asyellow, and so links the 1 image, the 2Frenchcontext, and the 3Englishquery to answercone→\\rightarrowI \(✓\)\. VLMs frequently fail this association \(✗\)\. We investigate whether such bindings survive when the context and query languages differ\.A reader who speaksFrenchandEnglishdoes not pause to translate when informed about something“yellow"and asked about something“jaune"\. Humans fuse what we see with what we read into a single concept, and we do so without effort no matter which of the world’s languages the words arrive in \([Fig\.˜1](https://arxiv.org/html/2608.12333#S1.F1)\)Bonner and Epstein \([2021](https://arxiv.org/html/2608.12333#bib.bib3)\); Xieet al\.\([2017](https://arxiv.org/html/2608.12333#bib.bib4)\)\. Thesemultilingual associationscome innately with our understanding of languages and how they describe \(here, in the visual sense\) the world around us\.
Vision\-language models \(VLMs\)OpenAIet al\.\([2024](https://arxiv.org/html/2608.12333#bib.bib9)\); Teamet al\.\([2025](https://arxiv.org/html/2608.12333#bib.bib8)\)now operate in exactly these settings, answering questions about images for users spread across the languages and scripts of the map in[Fig\.˜1](https://arxiv.org/html/2608.12333#S1.F1)\. The languages depicted \(top\) are chosen to span multiple scripts and language families across several continents, so our findings speak to globally diverse deployment rather than to any single region\. If the concept associations a VLM forms internally are tethered to the surface form ofEnglish, then every non\-Englishuser inherits a silent degradation that monolingual evaluation never exposes\. Such associations are governed by*binding*, which in its simplest causal form asks :given X:Y and Y:Z, can the model infer X:Z?\(see[Eq\.˜1](https://arxiv.org/html/2608.12333#S2.E1)\)\.Jiahai Feng and Jacob Steinhardt \([2024](https://arxiv.org/html/2608.12333#bib.bib1)\)showed that language models attach latent Binding IDs to entities, representing this structure in their activations without it ever appearing in the prompt;Saravananet al\.\([2025](https://arxiv.org/html/2608.12333#bib.bib2)\)extended the result to VLMs, where image tokens bind to their textual references\. We subject this in\-context binding to a novel test, where we establish the binding in one language and retrieve it in another\. We organize our study around three questions:
RQ1:Do visual\-textual associations survive across languages?RQ2:What inside the VLM drives their degradation?RQ3:Does linguistic closeness aid binding transfer?
#### Contributions\.
We \(i\) introduceM2Bind, a task and benchmark that varies the context and query language independently over eight languages along two axes of linguistic distance; \(ii\) evaluate binding both extrinsically and intrinsically; \(iii\) trace the behaviour to linguistic distance, model tokenizer and decoder computations\. To our knowledge, this is the first work of its kind\.
[Sec\.˜2](https://arxiv.org/html/2608.12333#S2)defines our task, data, and metrics;[Sec\.˜3](https://arxiv.org/html/2608.12333#S3)discusses results and[Sec\.˜4](https://arxiv.org/html/2608.12333#S4)concludes the paper\.
## 2Methodology: M2BIND
#### Task\.
We introduce the task ofMultilingualMultimodalBINDing: for a visual scene and a textual context that assigns attributes to objects in languageℓctx\\ell\_\{\\mathrm\{ctx\}\}, can a VLM retrieve the correct association when queried in a different languageℓq\\ell\_\{q\}?
#### Problem formulation\.
\([Fig\.˜1](https://arxiv.org/html/2608.12333#S1.F1)\) A scenes=\(v,𝒪\)s=\(v,\\mathcal\{O\}\)consists of an image1vvand a set of objects𝒪=\{o1,o2\}\\mathcal\{O\}=\\\{o\_\{1\},o\_\{2\}\\\}\. Each objectojo\_\{j\}is characterised by a shapeσj∈Σ\\sigma\_\{j\}\\in\\Sigmaand a colourγj∈Γ\\gamma\_\{j\}\\in\\Gamma, withΣ=\{cone,cube,cylinder,sphere\}\\Sigma=\\\{\\textrm\{cone\},\\textrm\{cube\},\\textrm\{cylinder\},\\textrm\{sphere\}\\\}andΓ=\{red,blue,green,yellow,cyan,purple\}\\Gamma=\\\{\\textrm\{red\},\\textrm\{blue\},\\textrm\{green\},\\textrm\{yellow\},\\textrm\{cyan\},\\textrm\{purple\}\\\}\. A textual context2ccrefers to each object by its colour and assigns it an item symbolιj∈ℐ\\iota\_\{j\}\\in\\mathcal\{I\}, thereby defining a colour\-to\-item bindingb:γj↦ιjb:\\gamma\_\{j\}\\mapsto\\iota\_\{j\}\. A query3qqsingles out a target object by its shapeσ⋆\\sigma\_\{\\star\}and asks for the item it contains\. Producing the correct answer requires the model to compose two associations,
σ⋆→𝑣groundγ⋆→𝑐bindι⋆=b\(γ⋆\),\\sigma\_\{\\star\}\\;\\xrightarrow\[\\;v\\;\]\{\\textrm\{ground\}\}\\;\\gamma\_\{\\star\}\\;\\xrightarrow\[\\;c\\;\]\{\\textrm\{bind\}\}\\;\\iota\_\{\\star\}=b\(\\gamma\_\{\\star\}\),\(1\)firstgroundingthe queried shape to a colour through the image, then resolving that colour to itsbounditem through the context \([Eq\.˜1](https://arxiv.org/html/2608.12333#S2.E1)\)\. Crucially, associations must be carried by*latent binding variables*in the model’s activations\(Jiahai Feng and Jacob Steinhardt,[2024](https://arxiv.org/html/2608.12333#bib.bib1); Saravananet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib2)\)\.
#### Linguistic conditioning\.
We render the context in a languageℓctx\\ell\_\{\\mathrm\{ctx\}\}and the query in a languageℓq\\ell\_\{q\}while holding the imagevvfixed, so an instance is the triplex=\(v,cℓctx,qℓq\)x=\(v,\\,c^\{\\ell\_\{\\mathrm\{ctx\}\}\},\\,q^\{\\ell\_\{q\}\}\)\. Because the image, the underlying scene, and the gold itemι⋆\\iota\_\{\\star\}are identical across all conditions, any change in the model’s response isolates the effect of*language*on binding rather than of perception or task difficulty\.
#### Data and Setup\.
We build on the Shapes binding task ofSaravananet al\.\([2025](https://arxiv.org/html/2608.12333#bib.bib2)\), in which each instance pairs a Blender\-rendered image of two 3D objects with a context that binds each object’s colour to an item symbol and a query that targets one object by its shape\. We extend the task by independently varyingℓctx\\ell\_\{\\mathrm\{ctx\}\}andℓq\\ell\_\{q\}over88languages selected along two axes of linguistic distance111We discuss linguistic choices in[AppendixB](https://arxiv.org/html/2608.12333#A2)\.:\(i\) cross\-family:English,French,Mandarin,Arabic, spanning four families and three scripts\(Asher and Moseley,[2018](https://arxiv.org/html/2608.12333#bib.bib10); Littellet al\.,[2017](https://arxiv.org/html/2608.12333#bib.bib11)\); and\(ii\) within\-family, comprising aGermaniccluster \(English,Dutch,German\) and aRomancecluster \(French,Italian,Spanish\) that share the Latin script but differ in genealogical proximity\(Campbell and Grondona,[2008](https://arxiv.org/html/2608.12333#bib.bib12); Gooskenset al\.,[2018](https://arxiv.org/html/2608.12333#bib.bib13)\)222Combinations of these languages yield 16 arrangements for cross\-family and 9 for within\-family\.\.
#### Model\.
We evaluate LLaVA\-1\.5\-OV\-7B\(Anet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib14)\)a SigLIP vision encoder, an MLP projector, and Qwen2 decoder followingSaravananet al\.\([2025](https://arxiv.org/html/2608.12333#bib.bib2)\)as a representative open\-weight VLM\. Given an instancexx, the model returns a conditional distribution over candidate items at the answer position under greedy decoding\. All reported values are averaged over33runs on a single NVIDIA A40 \(48GB\) GPU\.
#### Sequence score\.
We compare the two candidate items𝒞=\{ι⋆,ιswap\}\\mathcal\{C\}=\\\{\\iota\_\{\\star\},\\iota\_\{\\mathrm\{swap\}\}\\\}, whereι⋆\\iota\_\{\\star\}is bound to the queried object andιswap\\iota\_\{\\mathrm\{swap\}\}to the distractor\. For instancexix\_\{i\}, we score a candidateIIby the log\-probability the model assigns to its token sequencey1:m\(I\)\(I\)y\_\{1:m\(I\)\}\(I\)at the answer position:
Si\(I\)=∑t=1m\(I\)logpθ\(yt\(I\)∣xi,y<t\(I\)\)\.S\_\{i\}\(I\)=\\sum\_\{t=1\}^\{m\(I\)\}\\log p\_\{\\theta\}\\\!\\bigl\(y\_\{t\}\(I\)\\mid x\_\{i\},\\,y\_\{<t\}\(I\)\\bigr\)\.\(2\)
#### Accuracy and Factorization Margin\.
Accuracyis the fraction of instances for whichSi\(ι⋆\)\>Si\(ιswap\)S\_\{i\}\(\\iota\_\{\\star\}\)\>S\_\{i\}\(\\iota\_\{\\mathrm\{swap\}\}\)under the score of[Eq\.˜2](https://arxiv.org/html/2608.12333#S2.E2)\. Because accuracy saturates before binding fully degrades, we additionally report theFactorization Margin\(FM\), the mean score separation between the correct and swapped items\(Jiahai Feng and Jacob Steinhardt,[2024](https://arxiv.org/html/2608.12333#bib.bib1)\):
FM=1N∑i=1N\[Si\(ι⋆,i\)−Si\(ιswap,i\)\]\.\\mathrm\{FM\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\bigl\[S\_\{i\}\(\\iota\_\{\\star,i\}\)\-S\_\{i\}\(\\iota\_\{\\mathrm\{swap\},i\}\)\\bigr\]\.\(3\)A large positive FM indicates that the model cleanly separates the bound item from the distractor, whereasFM→0\\mathrm\{FM\}\\\!\\to\\\!0signals a collapsed binding even when accuracy remains high\.
#### Causal localization\.
Accuracy and FM characterise binding only at the output\. To locate*where*binding is consolidated inside the decoder, we apply an interchange intervention\(Menget al\.,[2022](https://arxiv.org/html/2608.12333#bib.bib15); Viget al\.,[2020](https://arxiv.org/html/2608.12333#bib.bib16)\)\. Lethk,jh\_\{k,j\}denote the layer\-kkhidden state at the token position of objectojo\_\{j\}\. We re\-run the forward pass with the two objects’ states interchanged,do\(hk,1↔hk,2\)\\mathrm\{do\}\(h\_\{k,1\}\\\!\\leftrightarrow\\\!h\_\{k,2\}\), and define theIntervention Causal Effect\(ICE\) as the resulting drop in the correct\-item score, averaged over instances:
ICEk=𝔼i\[Si\(ι⋆\)−Si\(ι⋆∣do\(hk,1↔hk,2\)\)\]\.\\mathrm\{ICE\}\_\{k\}=\\mathbb\{E\}\_\{i\}\\\!\\Bigl\[S\_\{i\}\(\\iota\_\{\\star\}\)\-S\_\{i\}\\\!\\bigl\(\\iota\_\{\\star\}\\mid\\mathrm\{do\}\(h\_\{k,1\}\\\!\\leftrightarrow\\\!h\_\{k,2\}\)\\bigr\)\\Bigr\]\.\(4\)A large positiveICEk\\mathrm\{ICE\}\_\{k\}means that disrupting the layer\-kkrepresentation of the two objects sharply lowers the score of the correct item, i\.e\. binding\-relevant information is actively represented at layerkk\. This gives us a profile of where the associations are made\.
## 3Results and Discussions
Table 1:Cross\-family VLM associations\.Cross\-lingual concept binding overEn,Fr,Zh, andAr; rows are context \(ctx\), columns are query \(q\)\. Each cell reports*Acc\. / FM*; for both, higher \(↑\\uparrow\) is better\. Worst cases arehighlighted\.#### \(RQ1\) Binding across language families\.
Binding strength is far from language\-invariant: FM falls by 2\.65, from 5\.45 for monolingualEnglishto 2\.80 for theMandarin–Arabiccross\-script pair, less than half its best monolingual value\.[Tab\.˜1](https://arxiv.org/html/2608.12333#S3.T1)reports accuracy and FM for all 16 context–query combinations overEnglish,French,Mandarin, andArabic\. The monolingual diagonal \(ℓctx=ℓq\\ell\_\{\\mathrm\{ctx\}\}\{=\}\\ell\_\{\\mathrm\{q\}\}\) sets each language’s upper bound:Englishis strongest at 5\.45, while the two non\-Latin scripts sit lowest even on their own diagonal \(Mandarin4\.95,Arabic4\.55\), so script already shapes binding before any language mixing\. Cross\-lingually, FM tracks typological distance\. The closeEnglish–Frenchpair stays near its monolingual values, whereas mixing a Latin with a non\-Latin script collapses FM to roughly 3\.2–4\.1, and the two non\-Latin scripts together reach the 2\.80 floor noted above with accuracy at its∼\\sim0\.90 floor\. Degradation is alsoasymmetric: across the four Latin/non\-Latin pairs, the non\-Latin language on the query side is consistently harder than on the context side by about 0\.25 FM \(e\.g\.English→\\toArabic3\.40 vs\.Arabic→\\toEnglish3\.75\)\. This matches the Matrix Language Frame account of code\-switching\(Myers\-Scotton,[2001](https://arxiv.org/html/2608.12333#bib.bib17)\), whereℓctx\\ell\_\{\\mathrm\{ctx\}\}builds the binding scaffold andℓq\\ell\_\{\\mathrm\{q\}\}must access it: access is more fragile than construction once the frames diverge\. The trend is consistent with multilingual NLP\(Lauscheret al\.,[2020](https://arxiv.org/html/2608.12333#bib.bib19); Pireset al\.,[2019](https://arxiv.org/html/2608.12333#bib.bib20)\)\.
Table 2:Impact of tokenization\.Values reported are for monolingual setting\. We report mean tokens of the image\+context\+query, truncation \(Trunc\.\) against the context limit, and the FM drop relative toEnglish\. For all metrics, lower \(↓\\downarrow\) is better\.
#### \(RQ2\) Tokenization\.
Which component drives this fragility? Tokenization in particular plays a significant role since the same sentence can use a varying amount of tokens based on the language333This is referred to as thefertilityof a tokenizerRustet al\.\([2021](https://arxiv.org/html/2608.12333#bib.bib21)\)\.\. In the monolingual setting \([Tab\.˜2](https://arxiv.org/html/2608.12333#S3.T2)\),Arabicneeds 27\.6% more tokens thanEnglish, truncates 12% of prompts at the context limit, and loses 2\.55 FM, with no cross\-lingual transfer involved\. Binding thus weakens within a single language from token budget alone, implicating tokenizer unfairness\(Petrovet al\.,[2023](https://arxiv.org/html/2608.12333#bib.bib6)\)and supportingRustet al\.\([2021](https://arxiv.org/html/2608.12333#bib.bib21)\)on tokenizer quality and multilingual performance\.
Figure 2:Causal interventions for multilingual binding\.Theyy\-axis is the change in log\-probability \(ICE\) after intervention; thexx\-axis is decoder layer index\. A higher peak implies stronger association strength at those layers\. We show monolingual \(English,Mandarin,Arabic\) and cross\-lingual combinations\. Higher is better\.
#### \(RQ2\) Binding in VLM\.
While[Tab\.˜2](https://arxiv.org/html/2608.12333#S3.T2)explains impact on binding before actual associations, we ask how the internal representations of the model actually perform multilingual association\.[Fig\.˜2](https://arxiv.org/html/2608.12333#S3.F2)provides some insights\. All curves are unimodal, depicting a localized region of layers where multilingual association is observed the strongest\. For monolingual cases all curves peak around themiddle layers\(∼\\sim15\), indicating that the model has already committed to the correct association before the final generation layers\. Cross\-lingual instances undergo a shift towards the right, peaking around thelate layersof the decoder \(∼\\sim20\)\. This points to extra computation reconciling the context\- and query\-language representations\. Interestingly, binding is reduced for cross\-lingual cases as compared to monolingual curves\. The combination ofMandarin\-Arabic, which produced the weakest FM in[Tab\.˜1](https://arxiv.org/html/2608.12333#S3.T1), also produces the flattest and weakest curve in[Fig\.˜2](https://arxiv.org/html/2608.12333#S3.F2)\. This means that even after extra computations in the decoder, the binding is less causally separable , i\.e, the swap intervention has less effect because the two items’ representations are less distinct\.
\(a\)
\(b\)
Table 3:Within\-family binding for Germanic and Romance families\.Worst cases arehighlighted\.
#### \(RQ3\) Binding within families\.
Instead of languages that do not seemingly overlap in mutual intelligibility, how do VLMs behave when the languages arecloserto each other444Consider a nativeDutchspeaker, they might have some understanding upon encountering a query inGerman\.?[Tab\.˜3](https://arxiv.org/html/2608.12333#S3.T3)shows concept binding for the two language families mentioned in[Sec\.˜2](https://arxiv.org/html/2608.12333#S2)\. Overall, every within\-family FM stays above 5, the worst drop being of 0\.45 fromGermanto other languages\. In terms of accuracy, within\-family excels with a lower bound around 0\.98, as compared to a floor of 0\.90 for cross\-family interactions in[Tab\.˜1](https://arxiv.org/html/2608.12333#S3.T1)\.[Tab\.˜3\(a\)](https://arxiv.org/html/2608.12333#S3.T3.st1)shows results for the Germanic family\.English\-Dutchis akin toEnglishin its monolingual setting, closer thanEnglish\-French, despite both being non\-English languages\. Interestingly,German\-Dutchpairs perform really well, generalizing beyondEnglish\.[Tab\.˜3\(b\)](https://arxiv.org/html/2608.12333#S3.T3.st2)shows results for Romance family, whereItaliandemonstrates to be better thanSpanishas ctx or q\.Spanish\-Italianas a pair shows an FM of 5, showing associations work within the family, without the need ofFrenchas an anchor language, supportingConneauet al\.\([2020](https://arxiv.org/html/2608.12333#bib.bib22)\)\.
## 4Conclusion
We present M²BIND, a benchmark that varies the context and query language independently to test whether VLMs maintain entity\-attribute bindings across languages\. Binding dissociates with linguistic distance: distant languages and non\-Latin scripts collapse both binding strength and its causal localization, so near\-ceiling task accuracy is not a reliable predictor of cross\-lingual binding fidelity\. Language\-invariant grounding, effortless for humans, remains unsolved for current VLMs\.
## 5Limitations
The Shapes task uses procedurally generated Blender images; we use this since our interest was to show a simple task where multilingual associations are not performed properly\. Still, noticing an even further degradation of this concept binding for real images remains, beyond two objects or simple shapes\. Further we show our work primarily on a LLaVA VLM, along with additional results on Qwen2\.5 VL in Appendix[Appendix˜C](https://arxiv.org/html/2608.12333#A3)\. A similar look into other commercial models could potentially illustrate our generalizations further\.
## Appendix ARelated Works
#### Multilingual vision\-language models\.
Vision\-language models have advanced rapidly\(OpenAIet al\.,[2024](https://arxiv.org/html/2608.12333#bib.bib9); Teamet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib8)\), yet most are trained on predominantly English data and lose accuracy on non\-English input\(Geigleet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib27)\)\. This gap has driven a wave of multilingual and multicultural systems and evaluations\. On the modelling side, efforts such as Centurio\(Geigleet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib27)\)and Pangea\(Yueet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib30)\)study how language coverage and data mixture shape cross\-lingual ability while preserving English performance\. On the evaluation side, cross\-lingual visual question answering benchmarks such as xGQA\(Pfeifferet al\.,[2022](https://arxiv.org/html/2608.12333#bib.bib28)\)and culturally grounded suites such as CVQA\(Romeroet al\.,[2024](https://arxiv.org/html/2608.12333#bib.bib31)\)measure how well models transfer across languages and cultures\. A recent survey catalogues 31 models and 21 benchmarks and identifies a persistent tension between language neutrality and faithful cross\-lingual behaviour\(Manea and Libovický,[2026](https://arxiv.org/html/2608.12333#bib.bib29)\)\. These efforts share a common lens, namely they measure multilingual ability through end\-task accuracy\. We instead probe an internal property, the stability of the entity–attribute binding itself, and show that aggregate accuracy can stay near ceiling while binding strength degrades, which makes downstream scores an incomplete diagnostic for multilingual deployment\.
#### Language for cross\-modal understanding\.
The question of whether the language one uses shapes how one perceives the world predates modern NLP, originating in the linguistic relativity tradition\(Whorf,[2012](https://arxiv.org/html/2608.12333#bib.bib32); Kay and Kempton,[1984](https://arxiv.org/html/2608.12333#bib.bib33)\)\. A long line of psycholinguistic work has made this concrete in the visual domain, the same setting VLMs operate in\. Speakers whose language lexicalises a colour distinction discriminate those colours faster, as shown for Russian light and dark blue\(Winaweret al\.,[2007](https://arxiv.org/html/2608.12333#bib.bib34)\)and Greek blues\(Thierryet al\.,[2009](https://arxiv.org/html/2608.12333#bib.bib35)\), with category effects that are lateralised to the language\-dominant hemisphere\(Gilbertet al\.,[2006](https://arxiv.org/html/2608.12333#bib.bib36)\)and that appear even pre\-attentively\(Robersonet al\.,[2000](https://arxiv.org/html/2608.12333#bib.bib37); Regier and Kay,[2009](https://arxiv.org/html/2608.12333#bib.bib38)\)\. Beyond colour, the language of thought has been argued to influence conceptions of time and other abstract domains\(Boroditsky,[2001](https://arxiv.org/html/2608.12333#bib.bib39)\)\. Complementarily, neuroscientific evidence indicates that human object representations themselves reflect the co\-occurrence statistics of vision and language\(Bonner and Epstein,[2021](https://arxiv.org/html/2608.12333#bib.bib3)\)and that congruent visual and verbal signals are integrated during perception\(Xieet al\.,[2017](https://arxiv.org/html/2608.12333#bib.bib4)\)\. Together these findings establish that, for humans, language and visual concepts are deeply entangled yet the underlying concept survives a change of language\. Our work asks the machine analogue of this question, namely whether a VLM’s internal visual\-textual associations are likewise preserved when the language of the input changes\.
#### In\-context binding in language and vision\-language models\.
Associating an attribute with the correct entity rather than a competing one is a prerequisite for in\-context reasoning\.Jiahai Feng and Jacob Steinhardt \([2024](https://arxiv.org/html/2608.12333#bib.bib1)\)formalised this as the*Binding ID*mechanism, showing through causal interventions that language models tag co\-referring entity and attribute tokens with a shared latent identifier occupying a low\-rank subspace of the residual stream\. Subsequent work has refined and extended this picture, localising an ordering component that determines binding behaviour\(Daiet al\.,[2024](https://arxiv.org/html/2608.12333#bib.bib25)\), decomposing retrieval into distinct mechanisms the model mixes as contexts grow complex\(Feng and others,[2025](https://arxiv.org/html/2608.12333#bib.bib26)\), recovering the underlying circuitry through learned component masks\(Davieset al\.,[2023](https://arxiv.org/html/2608.12333#bib.bib40); Prakashet al\.,[2024](https://arxiv.org/html/2608.12333#bib.bib41)\), and relating it to the broader linearity of relation decoding and attribute representation in transformers\(Hernandezet al\.,[2024](https://arxiv.org/html/2608.12333#bib.bib42); Heinzerling and Inui,[2024](https://arxiv.org/html/2608.12333#bib.bib43)\)\.Saravananet al\.\([2025](https://arxiv.org/html/2608.12333#bib.bib2)\)extended the Binding ID analysis from text to vision, demonstrating on a synthetic Shapes task that VLMs assign a common identifier to an object’s image tokens and its textual mentions\. We adopt their task and representative open\-weight model as our starting point, and use the same interchange\-intervention methodology\(Menget al\.,[2022](https://arxiv.org/html/2608.12333#bib.bib15); Viget al\.,[2020](https://arxiv.org/html/2608.12333#bib.bib16)\)to locate where binding is computed\. Where all of this work establishes binding within a single \(English\) language, we vary the context and query language independently to test whether the binding is language\-invariant\. Part of the fragility we uncover originates before the decoder, in tokenizer unfairness toward non\-Latin and morphologically rich scripts\(Petrovet al\.,[2023](https://arxiv.org/html/2608.12333#bib.bib6); Rustet al\.,[2021](https://arxiv.org/html/2608.12333#bib.bib21)\), and part inside it, which our layer\-wise interventions separate\.
## Appendix BLinguistic Choices in Detail
### B\.1Cross\-family: Maximal Diversity
Our four cross\-family languages were chosen to differ on every major typological axis\. We characterize diversity using the World Atlas of Language Structures \(WALS;Haspelmath[2009](https://arxiv.org/html/2608.12333#bib.bib44)\), a database of structural properties of languages compiled from descriptive grammars, covering features such as canonical word order, morphological type, writing system, and segmentation conventions\.[Fig\.˜3](https://arxiv.org/html/2608.12333#A2.F3)shows how each cross\-family language is classified along four WALS dimensions\. Cells are colored by whether the value is shared with the majority of the set \(blue\) or divergent \(orange\); no two languages share the same profile across all four axes\.
Figure 3:Typological profile of the four cross\-family languages along WALS dimensions\.Each cell shows the categorical classification; traits are eithersharedwith the majority ordivergent\. No two languages match on all four axes\.EnglishandFrenchshare SVO \(subject\-verb\-object\) order, Latin script, and whitespace segmentation, but differ in morphological type \(English has largely lost its inflectional system; French retains richer agreement and gender\)\.Mandarinshares SVO order withEnglish/Frenchbut diverges on every other axis: it is analytic \(no inflection\), logographic \(each character maps to a morpheme\), and unsegmented \(no whitespace between words\)\.Arabicdiverges most broadly: it permits VSO order, uses root\-and\-pattern morphology where three\-consonant roots are modified by vocalic templates, writes in a right\-to\-left abjad, yet segments with whitespace likeEnglish/French\. These four languages belong to four separate families: Germanic and Romance \(both Indo\-European\), Sinitic \(Sino\-Tibetan\), and Semitic \(Afro\-Asiatic\)\. Typological distance vectors fromlang2vec\(Littellet al\.,[2017](https://arxiv.org/html/2608.12333#bib.bib11)\)confirm near\-maximal dispersion in the joint syntactic–phonological–genetic feature space\. A speaker of one cross\-family language cannot partially comprehend another without formal training\(Ringbom and Jarvis,[2009](https://arxiv.org/html/2608.12333#bib.bib45)\), unlike within\-family pairs where shared vocabulary and structure enable partial comprehension\(Gooskenset al\.,[2018](https://arxiv.org/html/2608.12333#bib.bib13)\)\.
### B\.2Within\-family: Controlled Similarity Gradient
To disentangle genealogical proximity from script difference, we add two clusters where all languages share the Latin script but vary in closeness to their anchor language\.
#### Germanic \(English,Dutch,German\)\.
All three descend from Proto\-Germanic origins\(Ringe,[2017](https://arxiv.org/html/2608.12333#bib.bib47)\)\.DutchandEnglishshare the closer Ingvaeonic \(North Sea Germanic\) subgrouping, whileGermanunderwent the High German consonant shift that systematically altered its stop consonants \(e\.g\., Englishwatervs\. GermanWasser\)\.Germanretains a four\-case system, three grammatical genders, and verb\-final order in subordinate clauses — all absent in modernEnglish\.Dutchoccupies an intermediate position: it preserves two grammatical genders but has largely shed case marking and shares SVO order withEnglish\. Lexical similarity reflects this ordering:English–Dutch∼\{\\sim\}75%,Dutch–German∼\{\\sim\}75%,English–German∼\{\\sim\}60%\(Campbell and Grondona,[2008](https://arxiv.org/html/2608.12333#bib.bib12)\)\.Gooskenset al\.\([2018](https://arxiv.org/html/2608.12333#bib.bib13)\)confirm thatDutchspeakers can partially comprehend writtenEnglishwithout instruction, whileEnglish–Germanintelligibility is measurably lower\.
#### Romance \(French,Italian,Spanish\)\.
All three descend from Vulgar Latin\(Posner,[1996](https://arxiv.org/html/2608.12333#bib.bib46)\)\.Italianis generally considered the most conservative major Romance language, retaining the greatest lexical and morphological continuity with the common ancestor\.French–Italianlexical similarity \(∼\{\\sim\}89%\) is the highest pair in our study, whileFrench–Spanish\(∼\{\\sim\}75%\) is lower due to divergences such as the Latinf\-→\\toSpanishh\-shift and richer verb inflection inSpanish\(Campbell and Grondona,[2008](https://arxiv.org/html/2608.12333#bib.bib12)\)\. Untrained readers can often extract the gist of a text in a related Romance language from shared Latinate vocabulary alone\(Gooskenset al\.,[2018](https://arxiv.org/html/2608.12333#bib.bib13)\)\.
### B\.3Lexical Similarity Structure
[Fig\.˜4](https://arxiv.org/html/2608.12333#A2.F4)shows the full lexical similarity matrix across all eight M2BIND languages\. The block\-diagonal structure is immediately apparent: Germanic languages form one high\-similarity cluster \(60–75%\), Romance languages form another \(75–89%\), and bothMandarinandArabicshow near\-zero overlap \(∼\{\\sim\}1%\) with every other language\. The off\-diagonal blocks between Germanic and Romance show modest similarities \(20–27%\), reflecting their shared but distant Indo\-European ancestry\. WhileEnglishandFrenchhave low lexical similarity \(27%\), extensive Norman–French borrowing into English gives them a shared Latin\-script vocabulary that exceeds what the raw percentage suggests, which may explain the relatively high FM \(5\.20\) for this pair despite the cross\-family classification\.
Figure 4:Lexical similarity \(%\) across all eight languages in M2BIND\. Dashed lines separate language families\.
### B\.4Connection to Multilingual Representations
Genealogically, six of our eight languages belong to Indo\-European:English,Dutch, andGermandescend from Proto\-Germanic\(Ringe,[2017](https://arxiv.org/html/2608.12333#bib.bib47)\), whileFrench,Italian, andSpanishdescend from Vulgar Latin\(Posner,[1996](https://arxiv.org/html/2608.12333#bib.bib46)\)\.Mandarinbelongs to Sino\-Tibetan andArabicto Afro\-Asiatic, placing them in entirely separate families with no shared ancestry\. Work on multilingual transformers has shown that models trained on multilingual data develop internal representations where typologically related languages cluster together\(Pireset al\.,[2019](https://arxiv.org/html/2608.12333#bib.bib20)\)\.Conneauet al\.\([2020](https://arxiv.org/html/2608.12333#bib.bib22)\)demonstrated that cross\-lingual transfer improves with shared vocabulary and model scale, whileLauscheret al\.\([2020](https://arxiv.org/html/2608.12333#bib.bib19)\)showed that transfer quality degrades with increasing typological distance\. These findings predict exactly the gradient visible in[Fig\.˜4](https://arxiv.org/html/2608.12333#A2.F4): if the VLM’s embedding space reflects linguistic distance, within\-family pairs should share sufficiently aligned representations to support binding transfer at the same decoder depth, while cross\-family pairs should require additional late\-layer reconciliation\. Our ICE analysis \(Fig\. 2 in the main paper\) confirms this at the mechanistic level\.
## Appendix CAdditional results on Qwen2\.5
Our main experiments use LLaVA\-1\.5\-OV\-7B, whose decoder is itself a Qwen2 model\. To test whether the binding behaviour we report depends on that particular stack, we repeat the cross\-family protocol on Qwen2\.5\-VL\-7B\(Baiet al\.,[2025](https://arxiv.org/html/2608.12333#bib.bib24)\), a separately trained open\-weight VLM with its own vision encoder and tokenizer\. The images, scenes, gold itemsι⋆\\iota\_\{\\star\}, and the 16 context–query conditions are identical to those in[Sec\.˜3](https://arxiv.org/html/2608.12333#S3), and we reuse the same sequence score and Factorization Margin \(FM\)\.
#### Cross\-family binding\.
[Tab\.˜4](https://arxiv.org/html/2608.12333#A3.T4)reports accuracy and FM for Qwen2\.5\-VL\-7B\. Two differences from[Tab\.˜1](https://arxiv.org/html/2608.12333#S3.T1)stand out\. First, Qwen2\.5\-VL\-7B is a stronger associator in absolute terms\. Accuracy is at ceiling \(1\.000\) across theEnglishandMandarincontext rows and almost everywhere else, and FM exceeds the corresponding LLaVA cell in all but the two non\-Latin monolingual conditions, reaching 6\.67 forEnglish→\\toArabic\. Binding therefore does not collapse on this model in the way it does on LLaVA, and the headline accuracy gives little indication of any residual fragility\. Second, the fragility that remains is visible only through the finer\-grained signals\. Accuracy dips below ceiling only in theArabic\-context row \(0\.950 in its three cross\-lingual cells\) and inFrench→\\toMandarin\(0\.975\), which points to a non\-Latin language, especially on the context side, as the setting that disturbs top\-1 selection\. In FM, the weakest binding is the monolingualMandarindiagonal \(3\.71\) together with theArabic→\\toMandarincross\-script cell \(3\.76\), so the two non\-Latin scripts again mark the lower end of the binding\-strength range even though their accuracy stays high\.
#### Comparison with the main model\.
The agreement between the two models is qualitative rather than numerical\. Both place the floor of binding strength on the non\-Latin scripts,MandarinandArabic, and both leave theEnglish–Frenchregion untouched\. The contrast is that LLaVA expresses this fragility as a large FM collapse and an accuracy drop toward 0\.90, whereas Qwen2\.5\-VL\-7B absorbs most of it into FM while holding accuracy near ceiling, with the largest accuracy cost appearing underArabiccontext\. This supports our central claim in two ways\. The non\-Latin scripts are the consistent locus of difficulty across independently trained VLMs, and accuracy alone is an unreliable indicator of binding fidelity, since on Qwen a reader of the accuracy column would conclude that multilingual binding is solved while FM shows theMandarinand cross\-script settings remain measurably weaker\.
Table 4:Cross\-family binding for Qwen2\.5\-VL\-7B\.Rows are context \(ctx\), columns are query \(q\)\. Each cell reports*Acc\. / FM*\. For both, higher \(↑\\uparrow\) is better\. Worst cases arehighlighted\.
## Appendix DTemplate choices in prompting
Each instance is a single prompt built from three parts, all derived from the same scene and illustrated earlier in[Fig\.˜1](https://arxiv.org/html/2608.12333#S1.F1)\. Thecontextintroduces every object by its colour and assigns it an item symbol, using one sentence per object\. Thequerynames a target object by its shape and asks which item it contains\. Theanswer prefixrestates the queried shape and ends just before the item, so that the model’s next\-token distribution over the candidate symbols is exactly the binding we score\. The colour and shape words that fill these slots are listed for every cross\-family language in[Fig\.˜5](https://arxiv.org/html/2608.12333#A4.F5)\.
For example, the scene in[Fig\.˜1](https://arxiv.org/html/2608.12333#S1.F1)pairs a cyan cube with item P and a yellow cone with item I, and queries the cone\. InEnglishthis yields the context “The cyan object contains item P\. The yellow object contains item I\.”, the query “Which item does the cone contain?”, and the answer prefix “Answer: The cone contains item”, whose correct completion is I\. The same scene rendered withℓctx=French\\ell\_\{\\mathrm\{ctx\}\}\{=\}\{\\color\[rgb\]\{0,0\.35,0\.35\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.35,0\.35\}\\textbf\{French\}\}\{\}andℓq=English\\ell\_\{\\mathrm\{q\}\}\{=\}\{\\color\[rgb\]\{0,0,0\.7\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\.7\}\\textbf\{English\}\}\{\}keeps the image and the gold answer fixed while swapping only the surface forms, for example the colour*cyan*becomes*cyan*and*yellow*becomes*jaune*in the context, isolating the effect of language on binding\.
Figure 5:Choices in template\.Colour and shape terms that fill the\{color\}and\{shape\}slots of the context and query, shown for the four cross\-family languages \(English,French,Mandarin,Arabic\)\. The image and the gold item are held fixed across languages; only these surface forms change\.
## References
- X\. An, Y\. Xie, K\. Yang, W\. Zhang, X\. Zhao, Z\. Cheng, Y\. Wang, S\. Xu, C\. Chen, D\. Zhu, C\. Wu, H\. Tan, C\. Li, J\. Yang, J\. Yu, X\. Wang, B\. Qin, Y\. Wang, Z\. Yan, Z\. Feng, Z\. Liu, B\. Li, and J\. Deng \(2025\)LLaVA\-onevision\-1\.5: fully open framework for democratized multimodal training\.External Links:2509\.23661,[Link](https://arxiv.org/abs/2509.23661)Cited by:[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px5.p1.2)\.
- Atlas of the world’s languages\.Routledge\.Cited by:[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px4.p1.3)\.
- S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. Lin \(2025\)Qwen2\.5\-vl technical report\.arXiv preprint arXiv:2502\.13923\.Cited by:[Appendix C](https://arxiv.org/html/2608.12333#A3.p1.1)\.
- M\. F\. Bonner and R\. A\. Epstein \(2021\)Object representations in the human brain reflect the co\-occurrence statistics of vision and language\.Nature communications12\(1\),pp\. 4081\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p1.1)\.
- L\. Boroditsky \(2001\)Does language shape thought?: mandarin and english speakers’ conceptions of time\.Cognitive psychology43\(1\),pp\. 1–22\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- L\. Campbell and V\. Grondona \(2008\)Ethnologue: languages of the world\. 15th edn\. ed\. by raymond g\. gordonjr\.\. dallas: sil international, 2005\.\.Language84\(3\),pp\. 636–641\.External Links:[Document](https://dx.doi.org/10.1353/lan.0.0054)Cited by:[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px1.p1.3),[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px2.p1.3),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px4.p1.3)\.
- A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. Stoyanov \(2020\)Unsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 8440–8451\.External Links:[Link](https://aclanthology.org/2020.acl-main.747/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.747)Cited by:[§B\.4](https://arxiv.org/html/2608.12333#A2.SS4.p1.1),[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px4.p1.1)\.
- Q\. Dai, B\. Heinzerling, and K\. Inui \(2024\)Representational analysis of binding in language models\.arXiv preprint arXiv:2409\.05448\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- X\. Davies, M\. Nadeau, N\. Prakash, T\. R\. Shaham, and D\. Bau \(2023\)Discovering variable binding circuitry with desiderata\.arXiv preprint arXiv:2307\.03637\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- J\. Fenget al\.\(2025\)Mixing mechanisms: how language models retrieve bound entities in\-context\.arXiv preprint arXiv:2510\.06182\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- G\. Geigle, F\. Schneider, C\. Holtermann, C\. Biemann, R\. Timofte, A\. Lauscher, and G\. Glavaš \(2025\)Centurio: on drivers of multilingual ability of large vision\-language model\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(ACL\),Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1)\.
- A\. L\. Gilbert, T\. Regier, P\. Kay, and R\. B\. Ivry \(2006\)Whorf hypothesis is supported in the right visual field but not the left\.Proceedings of the national academy of sciences103\(2\),pp\. 489–494\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- C\. Gooskens, V\. J\. Van Heuven, J\. Golubović, A\. Schüppert, F\. Swarte, and S\. Voigt \(2018\)Mutual intelligibility between closely related languages in europe\.International Journal of Multilingualism15\(2\),pp\. 169–193\.Cited by:[§B\.1](https://arxiv.org/html/2608.12333#A2.SS1.p2.1),[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px1.p1.3),[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px2.p1.3),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px4.p1.3)\.
- M\. Haspelmath \(2009\)The typological database of the world atlas of language structures\.The use of databases in cross\-linguistic studies41,pp\. 283\.Cited by:[§B\.1](https://arxiv.org/html/2608.12333#A2.SS1.p1.1)\.
- B\. Heinzerling and K\. Inui \(2024\)Monotonic representation of numeric attributes in language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 175–195\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- E\. Hernandez, A\. Sen Sharma, T\. Haklay, K\. Meng, M\. Wattenberg, J\. Andreas, Y\. Belinkov, and D\. Bau \(2024\)Linearity of relation decoding in transformer language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 10504–10526\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- Jiahai Feng and Jacob Steinhardt \(2024\)How do Language Models Bind Entities in Context?\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zb3b6oKO77)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p2.1),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px2.p1.17),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px7.p1.1)\.
- P\. Kay and W\. Kempton \(1984\)What is the sapir\-whorf hypothesis?\.American anthropologist86\(1\),pp\. 65–79\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- A\. Lauscher, V\. Ravishankar, I\. Vulić, and G\. Glavaš \(2020\)From zero to hero: On the limitations of zero\-shot language transfer with multilingual Transformers\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 4483–4499\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.363/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.363)Cited by:[§B\.4](https://arxiv.org/html/2608.12333#A2.SS4.p1.1),[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px1.p1.6)\.
- P\. Littell, D\. R\. Mortensen, K\. Lin, K\. Kairis, C\. Turner, and L\. Levin \(2017\)URIEL and lang2vec: representing languages as typological, geographical, and phylogenetic vectors\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers,pp\. 8–14\.Cited by:[§B\.1](https://arxiv.org/html/2608.12333#A2.SS1.p2.1),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px4.p1.3)\.
- A\. Manea and J\. Libovický \(2026\)Multilingual vision\-language models, a survey\.External Links:2509\.22123,[Link](https://arxiv.org/abs/2509.22123)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px8.p1.4)\.
- C\. Myers\-Scotton \(2001\)The matrix language frame model: developments and responses\.Codeswitching worldwide II2,pp\. 23\.Cited by:[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px1.p1.6)\.
- OpenAI, :, A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford, A\. Mądry, A\. Baker\-Whitcomb, A\. Beutel, A\. Borzunov, A\. Carney, A\. Chow, A\. Kirillov, A\. Nichol, A\. Paino, A\. Renzin, A\. T\. Passos, A\. Kirillov, A\. Christakis, A\. Conneau, A\. Kamali, A\. Jabri, A\. Moyer, A\. Tam, A\. Crookes, A\. Tootoochian, A\. Tootoonchian, A\. Kumar, A\. Vallone, A\. Karpathy, A\. Braunstein, A\. Cann, A\. Codispoti, A\. Galu, A\. Kondrich, A\. Tulloch, A\. Mishchenko, A\. Baek, A\. Jiang, A\. Pelisse, A\. Woodford, A\. Gosalia, A\. Dhar, A\. Pantuliano, A\. Nayak, A\. Oliver, B\. Zoph, B\. Ghorbani, B\. Leimberger, B\. Rossen, B\. Sokolowsky, B\. Wang, B\. Zweig, B\. Hoover, B\. Samic, B\. McGrew, B\. Spero, B\. Giertler, B\. Cheng, B\. Lightcap, B\. Walkin, B\. Quinn, B\. Guarraci, B\. Hsu, B\. Kellogg, B\. Eastman, C\. Lugaresi, C\. Wainwright, C\. Bassin, C\. Hudson, C\. Chu, C\. Nelson, C\. Li, C\. J\. Shern, C\. Conger, C\. Barette, C\. Voss, C\. Ding, C\. Lu, C\. Zhang, C\. Beaumont, C\. Hallacy, C\. Koch, C\. Gibson, C\. Kim, C\. Choi, C\. McLeavey, C\. Hesse, C\. Fischer, C\. Winter, C\. Czarnecki, C\. Jarvis, C\. Wei, C\. Koumouzelis, D\. Sherburn, D\. Kappler, D\. Levin, D\. Levy, D\. Carr, D\. Farhi, D\. Mely, D\. Robinson, D\. Sasaki, D\. Jin, D\. Valladares, D\. Tsipras, D\. Li, D\. P\. Nguyen, D\. Findlay, E\. Oiwoh, E\. Wong, E\. Asdar, E\. Proehl, E\. Yang, E\. Antonow, E\. Kramer, E\. Peterson, E\. Sigler, E\. Wallace, E\. Brevdo, E\. Mays, F\. Khorasani, F\. P\. Such, F\. Raso, F\. Zhang, F\. von Lohmann, F\. Sulit, G\. Goh, G\. Oden, G\. Salmon, G\. Starace, G\. Brockman, H\. Salman, H\. Bao, H\. Hu, H\. Wong, H\. Wang, H\. Schmidt, H\. Whitney, H\. Jun, H\. Kirchner, H\. P\. de Oliveira Pinto, H\. Ren, H\. Chang, H\. W\. Chung, I\. Kivlichan, I\. O’Connell, I\. O’Connell, I\. Osband, I\. Silber, I\. Sohl, I\. Okuyucu, I\. Lan, I\. Kostrikov, I\. Sutskever, I\. Kanitscheider, I\. Gulrajani, J\. Coxon, J\. Menick, J\. Pachocki, J\. Aung, J\. Betker, J\. Crooks, J\. Lennon, J\. Kiros, J\. Leike, J\. Park, J\. Kwon, J\. Phang, J\. Teplitz, J\. Wei, J\. Wolfe, J\. Chen, J\. Harris, J\. Varavva, J\. G\. Lee, J\. Shieh, J\. Lin, J\. Yu, J\. Weng, J\. Tang, J\. Yu, J\. Jang, J\. Q\. Candela, J\. Beutler, J\. Landers, J\. Parish, J\. Heidecke, J\. Schulman, J\. Lachman, J\. McKay, J\. Uesato, J\. Ward, J\. W\. Kim, J\. Huizinga, J\. Sitkin, J\. Kraaijeveld, J\. Gross, J\. Kaplan, J\. Snyder, J\. Achiam, J\. Jiao, J\. Lee, J\. Zhuang, J\. Harriman, K\. Fricke, K\. Hayashi, K\. Singhal, K\. Shi, K\. Karthik, K\. Wood, K\. Rimbach, K\. Hsu, K\. Nguyen, K\. Gu\-Lemberg, K\. Button, K\. Liu, K\. Howe, K\. Muthukumar, K\. Luther, L\. Ahmad, L\. Kai, L\. Itow, L\. Workman, L\. Pathak, L\. Chen, L\. Jing, L\. Guy, L\. Fedus, L\. Zhou, L\. Mamitsuka, L\. Weng, L\. McCallum, L\. Held, L\. Ouyang, L\. Feuvrier, L\. Zhang, L\. Kondraciuk, L\. Kaiser, L\. Hewitt, L\. Metz, L\. Doshi, M\. Aflak, M\. Simens, M\. Boyd, M\. Thompson, M\. Dukhan, M\. Chen, M\. Gray, M\. Hudnall, M\. Zhang, M\. Aljubeh, M\. Litwin, M\. Zeng, M\. Johnson, M\. Shetty, M\. Gupta, M\. Shah, M\. Yatbaz, M\. J\. Yang, M\. Zhong, M\. Glaese, M\. Chen, M\. Janner, M\. Lampe, M\. Petrov, M\. Wu, M\. Wang, M\. Fradin, M\. Pokrass, M\. Castro, M\. O\. T\. de Castro, M\. Pavlov, M\. Brundage, M\. Wang, M\. Khan, M\. Murati, M\. Bavarian, M\. Lin, M\. Yesildal, N\. Soto, N\. Gimelshein, N\. Cone, N\. Staudacher, N\. Summers, N\. LaFontaine, N\. Chowdhury, N\. Ryder, N\. Stathas, N\. Turley, N\. Tezak, N\. Felix, N\. Kudige, N\. Keskar, N\. Deutsch, N\. Bundick, N\. Puckett, O\. Nachum, O\. Okelola, O\. Boiko, O\. Murk, O\. Jaffe, O\. Watkins, O\. Godement, O\. Campbell\-Moore, P\. Chao, P\. McMillan, P\. Belov, P\. Su, P\. Bak, P\. Bakkum, P\. Deng, P\. Dolan, P\. Hoeschele, P\. Welinder, P\. Tillet, P\. Pronin, P\. Tillet, P\. Dhariwal, Q\. Yuan, R\. Dias, R\. Lim, R\. Arora, R\. Troll, R\. Lin, R\. G\. Lopes, R\. Puri, R\. Miyara, R\. Leike, R\. Gaubert, R\. Zamani, R\. Wang, R\. Donnelly, R\. Honsby, R\. Smith, R\. Sahai, R\. Ramchandani, R\. Huet, R\. Carmichael, R\. Zellers, R\. Chen, R\. Chen, R\. Nigmatullin, R\. Cheu, S\. Jain, S\. Altman, S\. Schoenholz, S\. Toizer, S\. Miserendino, S\. Agarwal, S\. Culver, S\. Ethersmith, S\. Gray, S\. Grove, S\. Metzger, S\. Hermani, S\. Jain, S\. Zhao, S\. Wu, S\. Jomoto, S\. Wu, Shuaiqi, Xia, S\. Phene, S\. Papay, S\. Narayanan, S\. Coffey, S\. Lee, S\. Hall, S\. Balaji, T\. Broda, T\. Stramer, T\. Xu, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Cunninghman, T\. Degry, T\. Dimson, T\. Raoux, T\. Shadwell, T\. Zheng, T\. Underwood, T\. Markov, T\. Sherbakov, T\. Rubin, T\. Stasi, T\. Kaftan, T\. Heywood, T\. Peterson, T\. Walters, T\. Eloundou, V\. Qi, V\. Moeller, V\. Monaco, V\. Kuo, V\. Fomenko, W\. Chang, W\. Zheng, W\. Zhou, W\. Manassra, W\. Sheu, W\. Zaremba, Y\. Patil, Y\. Qian, Y\. Kim, Y\. Cheng, Y\. Zhang, Y\. He, Y\. Zhang, Y\. Jin, Y\. Dai, and Y\. Malkov \(2024\)GPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p2.1)\.
- A\. Petrov, E\. La Malfa, P\. H\. S\. Torr, and A\. Bibi \(2023\)Language Model Tokenizers Introduce Unfairness Between Languages\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2305.15425)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px2.p1.1)\.
- J\. Pfeiffer, G\. Geigle, A\. Kamath, J\. O\. Steitz, S\. Roth, I\. Vulić, and I\. Gurevych \(2022\)XGQA: cross\-lingual visual question answering\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 2497–2511\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1)\.
- T\. Pires, E\. Schlinger, and D\. Garrette \(2019\)How multilingual is multilingual BERT?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4996–5001\.External Links:[Link](https://aclanthology.org/P19-1493/),[Document](https://dx.doi.org/10.18653/v1/P19-1493)Cited by:[§B\.4](https://arxiv.org/html/2608.12333#A2.SS4.p1.1),[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px1.p1.6)\.
- R\. Posner \(1996\)The romance languages\.Cambridge University Press\.Cited by:[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px2.p1.3),[§B\.4](https://arxiv.org/html/2608.12333#A2.SS4.p1.1)\.
- N\. Prakash, T\. Shaham, T\. Haklay, Y\. Belinkov, and D\. Bau \(2024\)Fine\-tuning enhances existing mechanisms: a case study on entity tracking\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 8057–8082\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1)\.
- T\. Regier and P\. Kay \(2009\)Language, thought, and color: whorf was half right\.Trends in cognitive sciences13\(10\),pp\. 439–446\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- H\. Ringbom and S\. Jarvis \(2009\)The importance of cross\-linguistic similarity in foreign language learning\.The handbook of language teaching,pp\. 106–118\.Cited by:[§B\.1](https://arxiv.org/html/2608.12333#A2.SS1.p2.1)\.
- D\. Ringe \(2017\)From proto\-indo\-european to proto\-germanic\.Vol\.1,Oxford University Press\.Cited by:[§B\.2](https://arxiv.org/html/2608.12333#A2.SS2.SSS0.Px1.p1.3),[§B\.4](https://arxiv.org/html/2608.12333#A2.SS4.p1.1)\.
- D\. Roberson, I\. Davies, and J\. Davidoff \(2000\)Color categories are not universal: replications and new evidence from a stone\-age culture\.\.Journal of experimental psychology: General129\(3\),pp\. 369\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- D\. Romero, C\. Lyu, H\. A\. Wibowo, T\. Lynn, I\. Hamed, A\. N\. Kishore, A\. Mandal, A\. Dragonetti, A\. Abzaliev, A\. L\. Tonja,et al\.\(2024\)CVQA: culturally\-diverse multilingual visual question answering benchmark\.Advances in Neural Information Processing Systems37,pp\. 11479–11505\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1)\.
- P\. Rust, J\. Pfeiffer, I\. Vulić, S\. Ruder, and I\. Gurevych \(2021\)How good is your tokenizer? on the monolingual performance of multilingual language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 3118–3135\.External Links:[Link](https://aclanthology.org/2021.acl-long.243/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.243)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2608.12333#S3.SS0.SSS0.Px2.p1.1),[footnote 3](https://arxiv.org/html/2608.12333#footnote3)\.
- D\. Saravanan, M\. Tapaswi, and V\. Gandhi \(2025\)Investigating Mechanisms for In\-Context Vision Language Binding\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops \(CVPRW\),Vol\.,Los Alamitos, CA, USA,pp\. 4852–4856\.External Links:ISSN,[Document](https://dx.doi.org/10.1109/CVPRW67362.2025.00476),[Link](https://doi.ieeecomputersociety.org/10.1109/CVPRW67362.2025.00476)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p2.1),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px2.p1.17),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px4.p1.3),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px5.p1.2)\.
- G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, D\. Silver, M\. Johnson, I\. Antonoglou, J\. Schrittwieser, A\. Glaese, J\. Chen, E\. Pitler, T\. Lillicrap, A\. Lazaridou, O\. Firat, J\. Molloy, M\. Isard, P\. R\. Barham, T\. Hennigan, B\. Lee, F\. Viola, M\. Reynolds, Y\. Xu, R\. Doherty, E\. Collins, C\. Meyer, E\. Rutherford, E\. Moreira, K\. Ayoub, M\. Goel, J\. Krawczyk, C\. Du, E\. Chi, H\. Cheng, E\. Ni, P\. Shah, P\. Kane, B\. Chan, M\. Faruqui, A\. Severyn, H\. Lin, Y\. Li, Y\. Cheng, A\. Ittycheriah, M\. Mahdieh, M\. Chen, P\. Sun, D\. Tran, S\. Bagri, B\. Lakshminarayanan, J\. Liu, A\. Orban, F\. Güra, H\. Zhou, X\. Song, A\. Boffy, H\. Ganapathy, S\. Zheng, H\. Choe, Á\. Weisz, T\. Zhu, Y\. Lu, S\. Gopal, J\. Kahn, M\. Kula, J\. Pitman, R\. Shah, E\. Taropa, M\. A\. Merey, M\. Baeuml, Z\. Chen, L\. E\. Shafey, Y\. Zhang, O\. Sercinoglu, G\. Tucker, E\. Piqueras, M\. Krikun, I\. Barr, N\. Savinov, I\. Danihelka, B\. Roelofs, A\. White, A\. Andreassen, T\. von Glehn, L\. Yagati, M\. Kazemi, L\. Gonzalez, M\. Khalman, J\. Sygnowski, A\. Frechette, C\. Smith, L\. Culp, L\. Proleev, Y\. Luan, X\. Chen, J\. Lottes, N\. Schucher, F\. Lebron, A\. Rrustemi, N\. Clay, P\. Crone, T\. Kocisky, J\. Zhao, B\. Perz, D\. Yu, H\. Howard, A\. Bloniarz, J\. W\. Rae, H\. Lu, L\. Sifre, M\. Maggioni, F\. Alcober, D\. Garrette, M\. Barnes, S\. Thakoor, J\. Austin, G\. Barth\-Maron, W\. Wong, R\. Joshi, R\. Chaabouni, D\. Fatiha, A\. Ahuja, G\. S\. Tomar, E\. Senter, M\. Chadwick, I\. Kornakov, N\. Attaluri, I\. Iturrate, R\. Liu, Y\. Li, S\. Cogan, J\. Chen, C\. Jia, C\. Gu, Q\. Zhang, J\. Grimstad, A\. J\. Hartman, X\. Garcia, T\. S\. Pillai, J\. Devlin, M\. Laskin, D\. de Las Casas, D\. Valter, C\. Tao, L\. Blanco, A\. P\. Badia, D\. Reitter, M\. Chen, J\. Brennan, C\. Rivera, S\. Brin, S\. Iqbal, G\. Surita, J\. Labanowski, A\. Rao, S\. Winkler, E\. Parisotto, Y\. Gu, K\. Olszewska, R\. Addanki, A\. Miech, A\. Louis, D\. Teplyashin, G\. Brown, E\. Catt, J\. Balaguer, J\. Xiang, P\. Wang, Z\. Ashwood, A\. Briukhov, A\. Webson, S\. Ganapathy, S\. Sanghavi, A\. Kannan, M\. Chang, A\. Stjerngren, J\. Djolonga, Y\. Sun, A\. Bapna, M\. Aitchison, P\. Pejman, H\. Michalewski, T\. Yu, C\. Wang, J\. Love, J\. Ahn, D\. Bloxwich, K\. Han, P\. Humphreys, T\. Sellam, J\. Bradbury, V\. Godbole, S\. Samangooei, B\. Damoc, A\. Kaskasoli, S\. M\. R\. Arnold, V\. Vasudevan, S\. Agrawal, J\. Riesa, D\. Lepikhin, R\. Tanburn, S\. Srinivasan, H\. Lim, S\. Hodkinson, P\. Shyam, J\. Ferret, S\. Hand, A\. Garg, T\. L\. Paine, J\. Li, Y\. Li, M\. Giang, A\. Neitz, Z\. Abbas, S\. York, M\. Reid, E\. Cole, A\. Chowdhery, D\. Das, D\. Rogozińska, V\. Nikolaev, P\. Sprechmann, Z\. Nado, L\. Zilka, F\. Prost, L\. He, M\. Monteiro, G\. Mishra, C\. Welty, J\. Newlan, D\. Jia, M\. Allamanis, C\. H\. Hu, R\. de Liedekerke, J\. Gilmer, C\. Saroufim, S\. Rijhwani, S\. Hou, D\. Shrivastava, A\. Baddepudi, A\. Goldin, A\. Ozturel, A\. Cassirer, Y\. Xu, D\. Sohn, D\. Sachan, R\. K\. Amplayo, C\. Swanson, D\. Petrova, S\. Narayan, A\. Guez, S\. Brahma, J\. Landon, M\. Patel, R\. Zhao, K\. Villela, L\. Wang, W\. Jia, M\. Rahtz, M\. Giménez, L\. Yeung, J\. Keeling, P\. Georgiev, D\. Mincu, B\. Wu, S\. Haykal, R\. Saputro, K\. Vodrahalli, J\. Qin, Z\. Cankara, A\. Sharma, N\. Fernando, W\. Hawkins, B\. Neyshabur, S\. Kim, A\. Hutter, P\. Agrawal, A\. Castro\-Ros, G\. van den Driessche, T\. Wang, F\. Yang, S\. Chang, P\. Komarek, R\. McIlroy, M\. Lučić, G\. Zhang, W\. Farhan, M\. Sharman, P\. Natsev, P\. Michel, Y\. Bansal, S\. Qiao, K\. Cao, S\. Shakeri, C\. Butterfield, J\. Chung, P\. K\. Rubenstein, S\. Agrawal, A\. Mensch, K\. Soparkar, K\. Lenc, T\. Chung, A\. Pope, L\. Maggiore, J\. Kay, P\. Jhakra, S\. Wang, J\. Maynez, M\. Phuong, T\. Tobin, A\. Tacchetti, M\. Trebacz, K\. Robinson, Y\. Katariya, S\. Riedel, P\. Bailey, K\. Xiao, N\. Ghelani, L\. Aroyo, A\. Slone, N\. Houlsby, X\. Xiong, Z\. Yang, E\. Gribovskaya, J\. Adler, M\. Wirth, L\. Lee, M\. Li, T\. Kagohara, J\. Pavagadhi, S\. Bridgers, A\. Bortsova, S\. Ghemawat, Z\. Ahmed, T\. Liu, R\. Powell, V\. Bolina, M\. Iinuma, P\. Zablotskaia, J\. Besley, D\. Chung, T\. Dozat, R\. Comanescu, X\. Si, J\. Greer, G\. Su, M\. Polacek, R\. L\. Kaufman, S\. Tokumine, H\. Hu, E\. Buchatskaya, Y\. Miao, M\. Elhawaty, A\. Siddhant, N\. Tomasev, J\. Xing, C\. Greer, H\. Miller, S\. Ashraf, A\. Roy, Z\. Zhang, A\. Ma, A\. Filos, M\. Besta, R\. Blevins, T\. Klimenko, C\. Yeh, S\. Changpinyo, J\. Mu, O\. Chang, M\. Pajarskas, C\. Muir, V\. Cohen, C\. L\. Lan, K\. Haridasan, A\. Marathe, S\. Hansen, S\. Douglas, R\. Samuel, M\. Wang, S\. Austin, C\. Lan, J\. Jiang, J\. Chiu, J\. A\. Lorenzo, L\. L\. Sjösund, S\. Cevey, Z\. Gleicher, T\. Avrahami, A\. Boral, H\. Srinivasan, V\. Selo, R\. May, K\. Aisopos, L\. Hussenot, L\. B\. Soares, K\. Baumli, M\. B\. Chang, A\. Recasens, B\. Caine, A\. Pritzel, F\. Pavetic, F\. Pardo, A\. Gergely, J\. Frye, V\. Ramasesh, D\. Horgan, K\. Badola, N\. Kassner, S\. Roy, E\. Dyer, V\. C\. Campos, A\. Tomala, Y\. Tang, D\. E\. Badawy, E\. White, B\. Mustafa, O\. Lang, A\. Jindal, S\. Vikram, Z\. Gong, S\. Caelles, R\. Hemsley, G\. Thornton, F\. Feng, W\. Stokowiec, C\. Zheng, P\. Thacker, Ç\. Ünlü, Z\. Zhang, M\. Saleh, J\. Svensson, M\. Bileschi, P\. Patil, A\. Anand, R\. Ring, K\. Tsihlas, A\. Vezer, M\. Selvi, T\. Shevlane, M\. Rodriguez, T\. Kwiatkowski, S\. Daruki, K\. Rong, A\. Dafoe, N\. FitzGerald, K\. Gu\-Lemberg, M\. Khan, L\. A\. Hendricks, M\. Pellat, V\. Feinberg, J\. Cobon\-Kerr, T\. Sainath, M\. Rauh, S\. H\. Hashemi, R\. Ives, Y\. Hasson, E\. Noland, Y\. Cao, N\. Byrd, L\. Hou, Q\. Wang, T\. Sottiaux, M\. Paganini, J\. Lespiau, A\. Moufarek, S\. Hassan, K\. Shivakumar, J\. van Amersfoort, A\. Mandhane, P\. Joshi, A\. Goyal, M\. Tung, A\. Brock, H\. Sheahan, V\. Misra, C\. Li, N\. Rakićević, M\. Dehghani, F\. Liu, S\. Mittal, J\. Oh, S\. Noury, E\. Sezener, F\. Huot, M\. Lamm, N\. D\. Cao, C\. Chen, S\. Mudgal, R\. Stella, K\. Brooks, G\. Vasudevan, C\. Liu, M\. Chain, N\. Melinkeri, A\. Cohen, V\. Wang, K\. Seymore, S\. Zubkov, R\. Goel, S\. Yue, S\. Krishnakumaran, B\. Albert, N\. Hurley, M\. Sano, A\. Mohananey, J\. Joughin, E\. Filonov, T\. Kępa, Y\. Eldawy, J\. Lim, R\. Rishi, S\. Badiezadegan, T\. Bos, J\. Chang, S\. Jain, S\. G\. S\. Padmanabhan, S\. Puttagunta, K\. Krishna, L\. Baker, N\. Kalb, V\. Bedapudi, A\. Kurzrok, S\. Lei, A\. Yu, O\. Litvin, X\. Zhou, Z\. Wu, S\. Sobell, A\. Siciliano, A\. Papir, R\. Neale, J\. Bragagnolo, T\. Toor, T\. Chen, V\. Anklin, F\. Wang, R\. Feng, M\. Gholami, K\. Ling, L\. Liu, J\. Walter, H\. Moghaddam, A\. Kishore, J\. Adamek, T\. Mercado, J\. Mallinson, S\. Wandekar, S\. Cagle, E\. Ofek, G\. Garrido, C\. Lombriser, M\. Mukha, B\. Sun, H\. R\. Mohammad, J\. Matak, Y\. Qian, V\. Peswani, P\. Janus, Q\. Yuan, L\. Schelin, O\. David, A\. Garg, Y\. He, O\. Duzhyi, A\. Älgmyr, T\. Lottaz, Q\. Li, V\. Yadav, L\. Xu, A\. Chinien, R\. Shivanna, A\. Chuklin, J\. Li, C\. Spadine, T\. Wolfe, K\. Mohamed, S\. Das, Z\. Dai, K\. He, D\. von Dincklage, S\. Upadhyay, A\. Maurya, L\. Chi, S\. Krause, K\. Salama, P\. G\. Rabinovitch, P\. K\. R\. M, A\. Selvan, M\. Dektiarev, G\. Ghiasi, E\. Guven, H\. Gupta, B\. Liu, D\. Sharma, I\. H\. Shtacher, S\. Paul, O\. Akerlund, F\. Aubet, T\. Huang, C\. Zhu, E\. Zhu, E\. Teixeira, M\. Fritze, F\. Bertolini, L\. Marinescu, M\. Bölle, D\. Paulus, K\. Gupta, T\. Latkar, M\. Chang, J\. Sanders, R\. Wilson, X\. Wu, Y\. Tan, L\. N\. Thiet, T\. Doshi, S\. Lall, S\. Mishra, W\. Chen, T\. Luong, S\. Benjamin, J\. Lee, E\. Andrejczuk, D\. Rabiej, V\. Ranjan, K\. Styrc, P\. Yin, J\. Simon, M\. R\. Harriott, M\. Bansal, A\. Robsky, G\. Bacon, D\. Greene, D\. Mirylenka, C\. Zhou, O\. Sarvana, A\. Goyal, S\. Andermatt, P\. Siegler, B\. Horn, A\. Israel, F\. Pongetti, C\. "\. Chen, M\. Selvatici, P\. Silva, K\. Wang, J\. Tolins, K\. Guu, R\. Yogev, X\. Cai, A\. Agostini, M\. Shah, H\. Nguyen, N\. Ó\. Donnaile, S\. Pereira, L\. Friso, A\. Stambler, A\. Kurzrok, C\. Kuang, Y\. Romanikhin, M\. Geller, Z\. Yan, K\. Jang, C\. Lee, W\. Fica, E\. Malmi, Q\. Tan, D\. Banica, D\. Balle, R\. Pham, Y\. Huang, D\. Avram, H\. Shi, J\. Singh, C\. Hidey, N\. Ahuja, P\. Saxena, D\. Dooley, S\. P\. Potharaju, E\. O’Neill, A\. Gokulchandran, R\. Foley, K\. Zhao, M\. Dusenberry, Y\. Liu, P\. Mehta, R\. Kotikalapudi, C\. Safranek\-Shrader, A\. Goodman, J\. Kessinger, E\. Globen, P\. Kolhar, C\. Gorgolewski, A\. Ibrahim, Y\. Song, A\. Eichenbaum, T\. Brovelli, S\. Potluri, P\. Lahoti, C\. Baetu, A\. Ghorbani, C\. Chen, A\. Crawford, S\. Pal, M\. Sridhar, P\. Gurita, A\. Mujika, I\. Petrovski, P\. Cedoz, C\. Li, S\. Chen, N\. D\. Santo, S\. Goyal, J\. Punjabi, K\. Kappaganthu, C\. Kwak, P\. LV, S\. Velury, H\. Choudhury, J\. Hall, P\. Shah, R\. Figueira, M\. Thomas, M\. Lu, T\. Zhou, C\. Kumar, T\. Jurdi, S\. Chikkerur, Y\. Ma, A\. Yu, S\. Kwak, V\. Ähdel, S\. Rajayogam, T\. Choma, F\. Liu, A\. Barua, C\. Ji, J\. H\. Park, V\. Hellendoorn, A\. Bailey, T\. Bilal, H\. Zhou, M\. Khatir, C\. Sutton, W\. Rzadkowski, F\. Macintosh, R\. Vij, K\. Shagin, P\. Medina, C\. Liang, J\. Zhou, P\. Shah, Y\. Bi, A\. Dankovics, S\. Banga, S\. Lehmann, M\. Bredesen, Z\. Lin, J\. E\. Hoffmann, J\. Lai, R\. Chung, K\. Yang, N\. Balani, A\. Bražinskas, A\. Sozanschi, M\. Hayes, H\. F\. Alcalde, P\. Makarov, W\. Chen, A\. Stella, L\. Snijders, M\. Mandl, A\. Kärrman, P\. Nowak, X\. Wu, A\. Dyck, K\. Vaidyanathan, R\. R, J\. Mallet, M\. Rudominer, E\. Johnston, S\. Mittal, A\. Udathu, J\. Christensen, V\. Verma, Z\. Irving, A\. Santucci, G\. Elsayed, E\. Davoodi, M\. Georgiev, I\. Tenney, N\. Hua, G\. Cideron, E\. Leurent, M\. Alnahlawi, I\. Georgescu, N\. Wei, I\. Zheng, D\. Scandinaro, H\. Jiang, J\. Snoek, M\. Sundararajan, X\. Wang, Z\. Ontiveros, I\. Karo, J\. Cole, V\. Rajashekhar, L\. Tumeh, E\. Ben\-David, R\. Jain, J\. Uesato, R\. Datta, O\. Bunyan, S\. Wu, J\. Zhang, P\. Stanczyk, Y\. Zhang, D\. Steiner, S\. Naskar, M\. Azzam, M\. Johnson, A\. Paszke, C\. Chiu, J\. S\. Elias, A\. Mohiuddin, F\. Muhammad, J\. Miao, A\. Lee, N\. Vieillard, J\. Park, J\. Zhang, J\. Stanway, D\. Garmon, A\. Karmarkar, Z\. Dong, J\. Lee, A\. Kumar, L\. Zhou, J\. Evens, W\. Isaac, G\. Irving, E\. Loper, M\. Fink, I\. Arkatkar, N\. Chen, I\. Shafran, I\. Petrychenko, Z\. Chen, J\. Jia, A\. Levskaya, Z\. Zhu, P\. Grabowski, Y\. Mao, A\. Magni, K\. Yao, J\. Snaider, N\. Casagrande, E\. Palmer, P\. Suganthan, A\. Castaño, I\. Giannoumis, W\. Kim, M\. Rybiński, A\. Sreevatsa, J\. Prendki, D\. Soergel, A\. Goedeckemeyer, W\. Gierke, M\. Jafari, M\. Gaba, J\. Wiesner, D\. G\. Wright, Y\. Wei, H\. Vashisht, Y\. Kulizhskaya, J\. Hoover, M\. Le, L\. Li, C\. Iwuanyanwu, L\. Liu, K\. Ramirez, A\. Khorlin, A\. Cui, T\. LIN, M\. Wu, R\. Aguilar, K\. Pallo, A\. Chakladar, G\. Perng, E\. A\. Abellan, M\. Zhang, I\. Dasgupta, N\. Kushman, I\. Penchev, A\. Repina, X\. Wu, T\. van der Weide, P\. Ponnapalli, C\. Kaplan, J\. Simsa, S\. Li, O\. Dousse, F\. Yang, J\. Piper, N\. Ie, R\. Pasumarthi, N\. Lintz, A\. Vijayakumar, D\. Andor, P\. Valenzuela, M\. Lui, C\. Paduraru, D\. Peng, K\. Lee, S\. Zhang, S\. Greene, D\. D\. Nguyen, P\. Kurylowicz, C\. Hardin, L\. Dixon, L\. Janzer, K\. Choo, Z\. Feng, B\. Zhang, A\. Singhal, D\. Du, D\. McKinnon, N\. Antropova, T\. Bolukbasi, O\. Keller, D\. Reid, D\. Finchelstein, M\. A\. Raad, R\. Crocker, P\. Hawkins, R\. Dadashi, C\. Gaffney, K\. Franko, A\. Bulanova, R\. Leblond, S\. Chung, H\. Askham, L\. C\. Cobo, K\. Xu, F\. Fischer, J\. Xu, C\. Sorokin, C\. Alberti, C\. Lin, C\. Evans, A\. Dimitriev, H\. Forbes, D\. Banarse, Z\. Tung, M\. Omernick, C\. Bishop, R\. Sterneck, R\. Jain, J\. Xia, E\. Amid, F\. Piccinno, X\. Wang, P\. Banzal, D\. J\. Mankowitz, A\. Polozov, V\. Krakovna, S\. Brown, M\. Bateni, D\. Duan, V\. Firoiu, M\. Thotakuri, T\. Natan, M\. Geist, S\. tan Girgin, H\. Li, J\. Ye, O\. Roval, R\. Tojo, M\. Kwong, J\. Lee\-Thorp, C\. Yew, D\. Sinopalnikov, S\. Ramos, J\. Mellor, A\. Sharma, K\. Wu, D\. Miller, N\. Sonnerat, D\. Vnukov, R\. Greig, J\. Beattie, E\. Caveness, L\. Bai, J\. Eisenschlos, A\. Korchemniy, T\. Tsai, M\. Jasarevic, W\. Kong, P\. Dao, Z\. Zheng, F\. Liu, F\. Yang, R\. Zhu, T\. H\. Teh, J\. Sanmiya, E\. Gladchenko, N\. Trdin, D\. Toyama, E\. Rosen, S\. Tavakkol, L\. Xue, C\. Elkind, O\. Woodman, J\. Carpenter, G\. Papamakarios, R\. Kemp, S\. Kafle, T\. Grunina, R\. Sinha, A\. Talbert, D\. Wu, D\. Owusu\-Afriyie, C\. Du, C\. Thornton, J\. Pont\-Tuset, P\. Narayana, J\. Li, S\. Fatehi, J\. Wieting, O\. Ajmeri, B\. Uria, Y\. Ko, L\. Knight, A\. Héliou, N\. Niu, S\. Gu, C\. Pang, Y\. Li, N\. Levine, A\. Stolovich, R\. Santamaria\-Fernandez, S\. Goenka, W\. Yustalim, R\. Strudel, A\. Elqursh, C\. Deck, H\. Lee, Z\. Li, K\. Levin, R\. Hoffmann, D\. Holtmann\-Rice, O\. Bachem, S\. Arora, C\. Koh, S\. H\. Yeganeh, S\. Põder, M\. Tariq, Y\. Sun, L\. Ionita, M\. Seyedhosseini, P\. Tafti, Z\. Liu, A\. Gulati, J\. Liu, X\. Ye, B\. Chrzaszcz, L\. Wang, N\. Sethi, T\. Li, B\. Brown, S\. Singh, W\. Fan, A\. Parisi, J\. Stanton, V\. Koverkathu, C\. A\. Choquette\-Choo, Y\. Li, T\. Lu, A\. Ittycheriah, P\. Shroff, M\. Varadarajan, S\. Bahargam, R\. Willoughby, D\. Gaddy, G\. Desjardins, M\. Cornero, B\. Robenek, B\. Mittal, B\. Albrecht, A\. Shenoy, F\. Moiseev, H\. Jacobsson, A\. Ghaffarkhah, M\. Rivière, A\. Walton, C\. Crepy, A\. Parrish, Z\. Zhou, C\. Farabet, C\. Radebaugh, P\. Srinivasan, C\. van der Salm, A\. Fidjeland, S\. Scellato, E\. Latorre\-Chimoto, H\. Klimczak\-Plucińska, D\. Bridson, D\. de Cesare, T\. Hudson, P\. Mendolicchio, L\. Walker, A\. Morris, M\. Mauger, A\. Guseynov, A\. Reid, S\. Odoom, L\. Loher, V\. Cotruta, M\. Yenugula, D\. Grewe, A\. Petrushkina, T\. Duerig, A\. Sanchez, S\. Yadlowsky, A\. Shen, A\. Globerson, L\. Webb, S\. Dua, D\. Li, S\. Bhupatiraju, D\. Hurt, H\. Qureshi, A\. Agarwal, T\. Shani, M\. Eyal, A\. Khare, S\. R\. Belle, L\. Wang, C\. Tekur, M\. S\. Kale, J\. Wei, R\. Sang, B\. Saeta, T\. Liechty, Y\. Sun, Y\. Zhao, S\. Lee, P\. Nayak, D\. Fritz, M\. R\. Vuyyuru, J\. Aslanides, N\. Vyas, M\. Wicke, X\. Ma, E\. Eltyshev, N\. Martin, H\. Cate, J\. Manyika, K\. Amiri, Y\. Kim, X\. Xiong, K\. Kang, F\. Luisier, N\. Tripuraneni, D\. Madras, M\. Guo, A\. Waters, O\. Wang, J\. Ainslie, J\. Baldridge, H\. Zhang, G\. Pruthi, J\. Bauer, F\. Yang, R\. Mansour, J\. Gelman, Y\. Xu, G\. Polovets, J\. Liu, H\. Cai, W\. Chen, X\. Sheng, E\. Xue, S\. Ozair, C\. Angermueller, X\. Li, A\. Sinha, W\. Wang, J\. Wiesinger, E\. Koukoumidis, Y\. Tian, A\. Iyer, M\. Gurumurthy, M\. Goldenson, P\. Shah, M\. Blake, H\. Yu, A\. Urbanowicz, J\. Palomaki, C\. Fernando, K\. Durden, H\. Mehta, N\. Momchev, E\. Rahimtoroghi, M\. Georgaki, A\. Raul, S\. Ruder, M\. Redshaw, J\. Lee, D\. Zhou, K\. Jalan, D\. Li, B\. Hechtman, P\. Schuh, M\. Nasr, K\. Milan, V\. Mikulik, J\. Franco, T\. Green, N\. Nguyen, J\. Kelley, A\. Mahendru, A\. Hu, J\. Howland, B\. Vargas, J\. Hui, K\. Bansal, V\. Rao, R\. Ghiya, E\. Wang, K\. Ye, J\. M\. Sarr, M\. M\. Preston, M\. Elish, S\. Li, A\. Kaku, J\. Gupta, I\. Pasupat, D\. Juan, M\. Someswar, T\. M\., X\. Chen, A\. Amini, A\. Fabrikant, E\. Chu, X\. Dong, A\. Muthal, S\. Buthpitiya, S\. Jauhari, N\. Hua, U\. Khandelwal, A\. Hitron, J\. Ren, L\. Rinaldi, S\. Drath, A\. Dabush, N\. Jiang, H\. Godhia, U\. Sachs, A\. Chen, Y\. Fan, H\. Taitelbaum, H\. Noga, Z\. Dai, J\. Wang, C\. Liang, J\. Hamer, C\. Ferng, C\. Elkind, A\. Atias, P\. Lee, V\. Listík, M\. Carlen, J\. van de Kerkhof, M\. Pikus, K\. Zaher, P\. Müller, S\. Zykova, R\. Stefanec, V\. Gatsko, C\. Hirnschall, A\. Sethi, X\. F\. Xu, C\. Ahuja, B\. Tsai, A\. Stefanoiu, B\. Feng, K\. Dhandhania, M\. Katyal, A\. Gupta, A\. Parulekar, D\. Pitta, J\. Zhao, V\. Bhatia, Y\. Bhavnani, O\. Alhadlaq, X\. Li, P\. Danenberg, D\. Tu, A\. Pine, V\. Filippova, A\. Ghosh, B\. Limonchik, B\. Urala, C\. K\. Lanka, D\. Clive, Y\. Sun, E\. Li, H\. Wu, K\. Hongtongsak, I\. Li, K\. Thakkar, K\. Omarov, K\. Majmundar, M\. Alverson, M\. Kucharski, M\. Patel, M\. Jain, M\. Zabelin, P\. Pelagatti, R\. Kohli, S\. Kumar, J\. Kim, S\. Sankar, V\. Shah, L\. Ramachandruni, X\. Zeng, B\. Bariach, L\. Weidinger, T\. Vu, A\. Andreev, A\. He, K\. Hui, S\. Kashem, A\. Subramanya, S\. Hsiao, D\. Hassabis, K\. Kavukcuoglu, A\. Sadovsky, Q\. Le, T\. Strohman, Y\. Wu, S\. Petrov, J\. Dean, and O\. Vinyals \(2025\)Gemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p2.1)\.
- G\. Thierry, P\. Athanasopoulos, A\. Wiggett, B\. Dering, and J\. Kuipers \(2009\)Unconscious effects of language\-specific terminology on preattentive color perception\.Proceedings of the National Academy of Sciences106\(11\),pp\. 4567–4570\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber \(2020\)Investigating gender bias in language models using causal mediation analysis\.Advances in neural information processing systems33,pp\. 12388–12401\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2608.12333#S2.SS0.SSS0.Px8.p1.4)\.
- B\. L\. Whorf \(2012\)Language, thought, and reality: selected writings of benjamin lee whorf\.MIT press\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- J\. Winawer, N\. Witthoft, M\. C\. Frank, L\. Wu, A\. R\. Wade, and L\. Boroditsky \(2007\)Russian blues reveal effects of language on color discrimination\.Proceedings of the national academy of sciences104\(19\),pp\. 7780–7785\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1)\.
- Y\. Xie, Y\. Xu, C\. Bian, and M\. Li \(2017\)Semantic congruent audiovisual integration during the encoding stage of working memory: an erp and sloreta study\.Scientific Reports7\(1\),pp\. 5112\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.12333#S1.p1.1)\.
- X\. Yue, Y\. Song, A\. Asai, S\. Kim, J\. NYANDWI, S\. Khanuja, A\. Kantharuban, L\. Sutawika, S\. Ramamoorthy, and G\. Neubig \(2025\)Pangea: a fully open multilingual multimodal llm for 39 languages\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 47758–47811\.Cited by:[Appendix A](https://arxiv.org/html/2608.12333#A1.SS0.SSS0.Px1.p1.1)\.Similar Articles
Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models
This paper presents the first systematic study of multilingual instruction following in Vision-Language-Action (VLA) models, revealing significant performance degradation when models trained on English are evaluated on other languages. The authors propose Multilingual Principal Component Alignment (MPCA) to reduce the multilingual performance gap.
Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs
The paper identifies the Ghost Anchor phenomenon in multilingual MLLMs, where visual signals are underutilized during early alignment, and proposes the ANCHOR training framework to improve visual semantic emergence and performance across languages.
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
This paper develops a benchmark to evaluate the multilingual ability of vision-language models to use spatial deictic expressions (e.g., 'this', 'that') across four languages, revealing that models differ from humans in selecting appropriate demonstratives based on distance.
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
This paper challenges the 'Attention-Confidence Assumption' by demonstrating that attention map sharpness is a poor predictor of correctness in Vision-Language Models. Instead, it shows that reliability is better indicated by hidden-state geometry and self-consistency, with significant findings on architectural differences between late-fusion and early-fusion models.