Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
Summary
This pilot study uses Natural Language Autoencoders to probe whether Qwen2.5-7B internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts, finding evidence of latent inferences before they are verbalized.
View Cached Full Text
Cached at: 07/27/26, 07:39 AM
# Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
Source: [https://arxiv.org/html/2607.21774](https://arxiv.org/html/2607.21774)
Pablo Santiago Potes Velasco Universidad Autónoma de Occidente Cali, Colombia pablo\.potes@uao\.edu\.co &María del Mar García Matabanchoy Universidad Autónoma de Occidente Cali, Colombia maria\_d\.garcia\_m@uao\.edu\.co &Óscar Julián Pérez Ladino Universidad Autónoma de Occidente Cali, Colombia oscar\.perez@uao\.edu\.co &Jhoan Stevan Mosquera Ortiz Universidad Autónoma de Occidente Cali, Colombia jhoan\.mosquera@uao\.edu\.co &Nicolás Lozano Mazuera Universidad Autónoma de Occidente Cali, Colombia nicolas\.lozano\_m@uao\.edu\.co &Gilber Alexis Corrales Gallego Universidad Autónoma de Occidente GobLab\-UAI Cali, Colombia gacorrales@uao\.edu\.co
###### Abstract
Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated\. This pilot study examines whether Qwen2\.5\-7B\-Instruct internally represents Colombian identity, socioeconomic status, or stereotype\-related information when processing Colombian\-Spanish and English prompts\. We use Natural Language Autoencoders \(NLA\) to verbalize residual\-stream activations from layer 20 across four positional quartiles per prompt\. Our dataset contains 30 prompts arranged as 15 matched Spanish\-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls\. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output\. This work connects activation\-level interpretability with bias evaluation for underrepresented Spanish varieties\.
## 1Introduction
Large language models infer user attributes—nationality, dialect, socioeconomic background—from cues never stated explicitly, and these inferences shape generation even when unverbalized: instruction\-tuned models refuse over 98% of explicit demographic queries yet still condition outputs on inferred attributes\(Bouchaud and Ramaciotti,[2025](https://arxiv.org/html/2607.21774#bib.bib9)\)\.Fraser\-Talienteet al\.\([2026](https://arxiv.org/html/2607.21774#bib.bib1)\)illustrate this directly with Natural Language Autoencoders \(NLA\): their*language\-switching*case study shows a model representing a user as Russian several tokens before any lexical cue, later traced to mislabeled training data\. Whether an analogous unverbalized inference occurs for Spanish is unknown, despite evidence it should matter most for the varieties most often misread—Colombian Spanish scores far below Peninsular Spanish in recognition benchmarks \(F1=0\.282=0\.282vs\.0\.7230\.723\) for reasons tracking training\-data composition\(Kawasaki,[2026](https://arxiv.org/html/2607.21774#bib.bib22); Mayor\-Rocheret al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib23)\), yet no study inspects what the model represents internally before producing that output\.
We use the open\-source NLA for Qwen\-2\.5\-7B on∼\\sim30 matched Colombian\-Spanish/English prompt pairs to probe for latent nationality, socioeconomic, or stereotype representations absent from the model’s final response\. Section[2](https://arxiv.org/html/2607.21774#S2)situates this work; Section[3](https://arxiv.org/html/2607.21774#S3)details our procedure; Section[4](https://arxiv.org/html/2607.21774#S4)reports findings\.
## 2Related Work
Interpreting what a language model represents about its input without supervised probes has converged on two strategies: projecting or patching activations to recover output\-relevant information without training\(nostalgebraist,[2020](https://arxiv.org/html/2607.21774#bib.bib3); Belroseet al\.,[2023](https://arxiv.org/html/2607.21774#bib.bib4); Ghandehariounet al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib2)\), and training a reader model—via sparse dictionaries\(Cunninghamet al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib5)\)or a full verbalizer—to translate activations into free text\(Chenet al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib6); Panet al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib7); Karvonenet al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib8)\)\. We adopt the latter, using the Natural Language Autoencoder \(NLA\) ofFraser\-Talienteet al\.\([2026](https://arxiv.org/html/2607.21774#bib.bib1)\)as our probing instrument; their*language\-switching*case study, where the verbalizer surfaces a latent nationality inference before any lexical cue appears in the prompt, motivates asking whether an analogous unverbalized inference occurs for Colombian Spanish\.
A parallel line indicates that demographic attributes are linearly decodable from activations, regardless of the verbalization method\.Bouchaud and Ramaciotti \([2025](https://arxiv.org/html/2607.21774#bib.bib9)\)report AUC\-ROC up to 0\.995 for probing gender, race, and socioeconomic status from indirect cues, concentrated in middle layers;Lauscheret al\.\([2022](https://arxiv.org/html/2607.21774#bib.bib11)\)andTanget al\.\([2023](https://arxiv.org/html/2607.21774#bib.bib12)\)corroborate this across architectures, whileHuet al\.\([2026](https://arxiv.org/html/2607.21774#bib.bib17)\)shows finer\-grained attributes are distributed rather than strictly linear\. Most relevant to our hypothesis is the*alignment gap*: instruction\-tuned models refuse over 98% of explicit demographic queries\(Bouchaud and Ramaciotti,[2025](https://arxiv.org/html/2607.21774#bib.bib9)\)yet still condition generation on stereotype\-aligned inferences when the attribute is only implied\(Neplenbroeket al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib10); Tanget al\.,[2023](https://arxiv.org/html/2607.21774#bib.bib12)\), consistent with bias persisting in contextualized representations after debiasing\(Guo and Caliskan,[2021](https://arxiv.org/html/2607.21774#bib.bib14); Tan and Celis,[2019](https://arxiv.org/html/2607.21774#bib.bib13); Bommasaniet al\.,[2020](https://arxiv.org/html/2607.21774#bib.bib16); Huanget al\.,[2020](https://arxiv.org/html/2607.21774#bib.bib15); Zhanget al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib18)\)\.
Behaviorally, this latent inference produces measurable disparities for Spanish varieties and non\-native English speakers\. Peninsular Spanish is consistently best recognized and generated, while Latin American varieties lag substantially—Colombia among the lowest \(F1=0\.282=0\.282vs\.0\.7230\.723for Spain\)—tracking training\-data composition rather than digital resource volume\(Kawasaki,[2026](https://arxiv.org/html/2607.21774#bib.bib22); Mayor\-Rocheret al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib23); Martínezet al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib28)\), and English bias\-mitigation techniques do not transfer to Spanish\(Robleset al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib21)\)\. English shows analogous effects: anchoring on perceived non\-nativeness degrades response quality\(Reusenset al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib19)\), minoritized dialects receive more stereotyped and condescending outputs\(Fleisiget al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib20)\), and alignment training widens rather than narrows these gaps\(Ryanet al\.,[2024](https://arxiv.org/html/2607.21774#bib.bib24); Mireet al\.,[2025](https://arxiv.org/html/2607.21774#bib.bib25); Nayeem and Rafiei,[2026](https://arxiv.org/html/2607.21774#bib.bib26); Kantharubanet al\.,[2023](https://arxiv.org/html/2607.21774#bib.bib27)\)\.
These two literatures remain unconnected: linear\-probe studies establish*that*demographic inference diverges from verbalized output, and bias\-audit studies establish*that*this divergence harms underrepresented varieties, but none use a training\-free, unsupervised verbalizer with controlled explicit/implicit/neutral elicitation to localize*when*a Colombian\-identity inference emerges from indirect cues within the prompt itself\. This is the gap we target\.
## 3Method
#### Overview\.
\(Fraser\-Talienteet al\.,[2026](https://arxiv.org/html/2607.21774#bib.bib1)\)show that Natural Language Autoencoder \(NLA\) explanations can reveal that a model internally represents a user as Russian*before*any unambiguous lexical cue and*before*this belief is verbalized\. We adapt this finding into a controlled probe of whether Qwen2\.5\-7B\-Instruct internally infers*Colombian*identity from a single implicit cue, using the open\-source NLA pair released with\(Fraser\-Talienteet al\.,[2026](https://arxiv.org/html/2607.21774#bib.bib1)\)\.
#### Extraction and quartile sampling\.
LetMMbe the target model andhℓ∈ℝdh\_\{\\ell\}\\in\\mathbb\{R\}^\{d\}\(d=3584d\{=\}3584\) the residual\-stream activation at layerℓ=20\\ell\{=\}20, the depth at which the released Qwen2\.5\-7B NLA pair was trained, and anNN\-token prompt\. A single deterministic forward pass ofMMyields the full sequence of activations\{hℓ\[1\],…,hℓ\[N\]\}\\\{h\_\{\\ell\}\[1\],\\dots,h\_\{\\ell\}\[N\]\\\}, one per token position\. From this sequence we select four activations, the*last*token of each of four contiguous quartiles,
qk=⌊kN/4⌋,k∈\{1,2,3,4\},q\_\{k\}=\\lfloor kN/4\\rfloor,\\quad k\\in\\\{1,2,3,4\\\},sohℓ\[q1\],…,hℓ\[q4\]h\_\{\\ell\}\[q\_\{1\}\],\\dots,h\_\{\\ell\}\[q\_\{4\}\]are four distinct, full\-dimensional vectors drawn from four different positions in the same forward pass — not a single activation, and not a partition ofhℓh\_\{\\ell\}’s35843584dimensions\. Querying the AV at every position is infeasible, since each call emits hundreds of tokens, so this four\-point subsample keeps the per\-prompt AV cost fixed regardless ofNN\(Fraser\-Talienteet al\.,[2026](https://arxiv.org/html/2607.21774#bib.bib1)\)\. The last\-token choice additionally keeps each query in\-distribution: the AV’s supervised warm\-start was trained only on prefix\-final activations, and eachqkq\_\{k\}is, formally, the final token of prefixx\[1:qk\]x\[1\{:\}q\_\{k\}\]\. Eachhℓ\[qk\]h\_\{\\ell\}\[q\_\{k\}\]is passed once \(no resampling\) to the AV,zk∼AV\(⋅∣hℓ\[qk\]\)z\_\{k\}\\sim\\mathrm\{AV\}\(\\cdot\\mid h\_\{\\ell\}\[q\_\{k\}\]\), yielding one explanation per quartile\.
#### Design\.
We construct 15 base scenarios, each realized in Spanish and in English translation \(30 prompts\), with five scenarios per*explicitness*level:explicit\(unambiguous Colombian marker, positive control\),implicit\(exactly one subtle cue, analogous to the*vodka*→\\to“Russian” case in\(Fraser\-Talienteet al\.,[2026](https://arxiv.org/html/2607.21774#bib.bib1)\)\), andneutral\(topically matched, no national marker, negative control distinguishing a Colombia\-specific effect from generic AV confabulation\)\.
Table 1:2 \(language\)×\\times3 \(explicitness\) design,n=30n\{=\}30prompts\.
#### Structured coding\.
Each prompt’s four quartile\-level explanations\(z1,z2,z3,z4\)\(z\_\{1\},z\_\{2\},z\_\{3\},z\_\{4\}\)are expanded into120120prompt×\\timesquartile units and independently coded by Claude Sonnet for*nationality*,*socioeconomic status*, and*stereotype*mentions\. The coding prompt requires \(i\) a category markedtrueonly on explicit textual evidence, \(ii\) defaultfalsefor explanations consisting solely of generic model self\-description \(e\.g\.*“I am Qwen, a large language model created by Alibaba Cloud”*\), a known AV failure mode rather than a substantive judgment about the prompt, and \(iii\) a verbatim supporting quote for every positive label\. An automated audit confirms full coverage, flags failed API calls, and rejects any returned quote not found*verbatim*in its source explanation; one of120120units failed extraction after retries and is excluded listwise \(n=119n\{=\}119\)\.
Because the AV is known to confabulate plausible\-sounding but contextually wrong specifics — in our data, substituting Spain, Turkey, or Canada for Colombia in an otherwise on\-theme explanation — a positive nationality label conflates two distinct events: representing*some*national identity versus*Colombian*identity specifically\. We therefore additionally tag each positive nationality citation for whether it names Colombia \(or an unambiguous metonym, e\.g\. Bogotá, TransMilenio\), and report both the unrestricted and the Colombia\-specific rate in §[4](https://arxiv.org/html/2607.21774#S4); this distinction is fixed at analysis design time, not introduced post hoc\.
#### Statistical analysis\.
We report each \(group×\\timesquartile\) mention rate with a Wilson score interval, valid at this design’s small per\-cell counts \(n=9n\{=\}9–1010\) where a normal approximation would not be\. Implicit\-vs\-neutral differences at each quartile are tested with Fisher’s exact test on the corresponding2×22\{\\times\}2table, rather than aχ2\\chi^\{2\}test whose asymptotic assumptions are unmet here\. To summarize*when*nationality is first represented, we take, for each prompt where the category appears at least once, its earliest quartile, and compare distributions between groups with a Mann–WhitneyUUtest; prompts where it never appears are reported as a separate per\-group count rather than imputed, since their absence is itself informative\. All tests are two\-sided and uncorrected for multiple comparisons, consistent with the exploratory scope below\.
#### Scope\.
Withn=5n\{=\}5per cell, results are reported as descriptive rates with qualitative evidence rather than as statistically powered effects; we extract from a single layer and a single AV sample per quartile, given hackathon\-scale compute\.
## 4Results
Table[2](https://arxiv.org/html/2607.21774#S4.T2)reports the rate at which AV explanations mention nationality, for each group and quartile\. The implicit\-cue group rises monotonically across all four quartiles \(0\.10→\\to0\.78\), consistent with the context\-accumulation hypothesis; Figure[1](https://arxiv.org/html/2607.21774#S4.F1)shows this as a sustained positive slope with neither the early ceiling of the explicit control nor a late collapse\. Two cells deviate from monotonicity — explicit at Q2 \(0\.00\) and neutral at Q4 \(0\.20\) — and we resolve both below before turning to the quartile of first appearance \(Table[3](https://arxiv.org/html/2607.21774#S4.T3)\) and the implicit\-vs\-neutral comparison \(Table[4](https://arxiv.org/html/2607.21774#S4.T4)\)\.
Table 2:Nationality mention rate by group and quartile \(n=9n\{=\}9–1010per cell; one unit excluded for failed extraction, §[3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px4)\)\.Figure 1:Nationality mention rate by quartile, all three groups\. Implicit \(blue, triangles\) rises steadily from Q1 to Q4; explicit \(orange, circles\) ceils early; neutral \(gray, squares\) is non\-monotonic, addressed in §[4](https://arxiv.org/html/2607.21774#S4)\.#### Both irregularities trace to non\-Colombian confabulation, not noise\.
We inspected the source quotes behind explicit/Q2 and neutral/Q4 directly\. Every Q2 explanation in the explicit group is off\-theme confabulation unrelated to the prompt \(e\.g\.*“Marketing de Google”*,*“conferencia Tesla”*\) — the same low\-anchoring failure mode documented at Q1 \(§[3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px4)\), recurring at a second early position for short prompts\. The two positive citations behind neutral/Q4 name Spain and paella, never Colombia\. Restricting the nationality variable to Colombia\-specific mentions \(Spain/Turkey/Canada excluded; §[3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px4)\) removes both irregularities: neutral falls to0\.000\.00at every quartile, while implicit remains strictly above zero from Q2 onward \(0\.00/0\.10/0\.20/0\.110\.00/0\.10/0\.20/0\.11\)\. The qualitative pattern in Figure[1](https://arxiv.org/html/2607.21774#S4.F1)therefore understates, rather than fabricates, the implicit\-vs\-neutral separation; we report the unrestricted rate in Table[2](https://arxiv.org/html/2607.21774#S4.T2)for comparability with the AV’s overall behavior, and the Colombia\-specific rate as the variable of record for testing the hypothesis\.
Table 3:Quartile of first nationality mention, computed over prompts where the category appears at least once; the right column reports prompts where it never appears\.Table[3](https://arxiv.org/html/2607.21774#S4.T3)shows comparable mean onset quartiles across groups \(≈\\approxQ3\) among prompts that ever trigger a mention, but a markedly different*rate*of never triggering one at all:0/100/10for explicit,2/102/10for implicit,4/104/10for neutral\. This count — not the onset quartile — carries most of the between\-group signal, and is reported alongside the mean rather than imputed into it\.
Quartilekkimplicitnnimplicitkkneutralnnneutralpp\(Fisher\)Q11101101\.000Q22101101\.000Q35106101\.000Q4792100\.023Table 4:Fisher’s exact test, implicit vs\. neutral, by quartile\.Only Q4 reaches conventional significance \(p=0\.023p\{=\}0\.023\); givenn≤10n\{\\leq\}10per cell and four uncorrected comparisons, we read this as directional evidence for a late\-quartile separation rather than a confirmatory effect \(§[3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px5)\), consistent with the exploratory scope of this study\.
See Appendix[A](https://arxiv.org/html/2607.21774#A1)for a qualitative, case\-by\-case analysis of individual prompts\.
## 5Conclusion
Using an unsupervised, training\-free verbalizer, we find that Qwen2\.5\-7B\-Instruct’s residual stream comes to represent Colombian identity from a single implicit lexical cue, with the Colombia\-specific mention rate rising from0\.000\.00at Q1 to0\.200\.20at Q3 while the neutral control remains at0\.000\.00throughout\. This separation reaches conventional significance only at the final quartile \(p=0\.023p\{=\}0\.023\) and rests onn=5n\{=\}5scenarios per cell; we report it as directional evidence that an unverbalized nationality inference can emerge from a single cue, not as a confirmed effect\. Two apparent irregularities in the raw mention rate were traced to a known AV failure mode — confabulating a contextually wrong country — and resolved by restricting to Colombia\-specific citations, a check we recommend for any reuse of NLA explanations as a regional\-identity probe\. The natural next step is repeating this design at highernnper cell and at a second residual\-stream layer, to test whether the late\-quartile separation we observe is a property of this specific depth or holds more generally\.
## References
- N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. Steinhardt \(2023\)Eliciting latent predictions from transformers with the tuned lens\.arXiv preprint arXiv:2303\.08112\.Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings\.Annual Meeting of the Association for Computational Linguistics\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.431),[Link](https://www.aclweb.org/anthology/2020.acl-main.431.pdf)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- P\. Bouchaud and P\. Ramaciotti \(2025\)Linear socio\-demographic representations emerge in Large Language Models from indirect cues\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2512.10065),[Link](https://arxiv.org/abs/2512.10065)Cited by:[§1](https://arxiv.org/html/2607.21774#S1.p1.2),[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- H\. Chen, C\. Vondrick, and C\. Mao \(2024\)SelfIE: self\-interpretation of large language model embeddings\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- H\. Cunningham, A\. Ewart, L\. Riggs, R\. Huben, and L\. Sharkey \(2024\)Sparse autoencoders find highly interpretable features in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- E\. Fleisig, G\. Smith, M\. Bossi, I\. Rustagi, X\. Yin, and D\. Klein \(2024\)Linguistic Bias in ChatGPT: language Models Reinforce Dialect Discrimination\.Conference on Empirical Methods in Natural Language Processing\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2406.08818),[Link](https://arxiv.org/abs/2406.08818)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- K\. Fraser\-Taliente, S\. Kantamneni, E\. Ong, D\. Mossing, C\. Lu, P\. C\. Bogdan, E\. Ameisen, J\. Chen, D\. Kishylau, A\. Pearce, J\. Tarng, A\. Wu, J\. Wu, Y\. Zhang, D\. M\. Ziegler, E\. Hubinger, J\. Batson, J\. Lindsey, S\. Zimmerman, and S\. Marks \(2026\)Natural language autoencoders produce unsupervised explanations of LLM activations\.Transformer Circuits Thread\.External Links:[Link](https://transformer-circuits.pub/2026/nla/index.html)Cited by:[§1](https://arxiv.org/html/2607.21774#S1.p1.2),[§2](https://arxiv.org/html/2607.21774#S2.p1.1),[§3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px2.p1.15),[§3](https://arxiv.org/html/2607.21774#S3.SS0.SSS0.Px3.p1.1)\.
- A\. Ghandeharioun, A\. Caciularu, A\. Pearce, L\. Dixon, and M\. Geva \(2024\)Patchscopes: a unifying framework for inspecting hidden representations of language models\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- W\. Guo and A\. Caliskan \(2021\)Detecting Emergent Intersectional Biases: contextualized Word Embeddings Contain a Distribution of Human\-like Biases\.InProceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 122–133\.External Links:[Document](https://dx.doi.org/10.1145/3461702.3462536),[Link](http://dx.doi.org/10.1145/3461702.3462536)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- S\. Hu, R\. Li, and Y\. Gao \(2026\)Race, Ethnicity and Their Implication on Bias in Large Language Models\.medRxiv\.External Links:[Document](https://dx.doi.org/10.64898/2026.01.04.26343415),[Link](https://www.medrxiv.org/content/medrxiv/early/2026/01/05/2026.01.04.26343415.full.pdf)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- P\. Huang, H\. Zhang, R\. Jiang, R\. Stanforth, J\. Welbl, J\. Rae, V\. Maini, D\. Yogatama, and P\. Kohli \(2020\)Reducing Sentiment Bias in Language Models via Counterfactual Evaluation\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 65–83\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.7),[Link](http://dx.doi.org/10.18653/v1/2020.findings-emnlp.7)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- A\. Kantharuban, I\. Vulić, and A\. Korhonen \(2023\)Quantifying the Dialect Gap and its Correlates Across Languages\.Conference on Empirical Methods in Natural Language Processing\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2310.15135),[Link](https://arxiv.org/abs/2310.15135)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- A\. Karvonen, J\. Chua, C\. Dumas, K\. Fraser\-Taliente, S\. Kantamneni, J\. Minder, E\. Ong, A\. S\. Sharma, D\. Wen, O\. Evans, and S\. Marks \(2025\)Activation oracles: training and evaluating LLMs as general\-purpose activation explainers\.arXiv preprint arXiv:2512\.15674\.Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- Y\. Kawasaki \(2026\)Digital Linguistic Bias in Spanish: evidence from Lexical Variation in LLMs\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2602.09346),[Link](https://arxiv.org/abs/2602.09346)Cited by:[§1](https://arxiv.org/html/2607.21774#S1.p1.2),[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- A\. Lauscher, F\. Bianchi, S\. Bowman, and D\. Hovy \(2022\)SocioProbe: what, When, and Where Language Models Learn about Sociodemographics\.Conference on Empirical Methods in Natural Language Processing\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2211.04281),[Link](https://arxiv.org/abs/2211.04281)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- G\. Martínez, M\. Mayor\-Rocher, C\. P\. Huertas, N\. Melero, M\. Grandury, and P\. Reviriego \(2025\)Spanish is not just one: a dataset of Spanish dialect recognition for LLMs\.Data in Brief63,pp\. 112088\.External Links:[Document](https://dx.doi.org/10.1016/j.dib.2025.112088),ISSN 2352\-3409,[Link](http://dx.doi.org/10.1016/j.dib.2025.112088)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- M\. Mayor\-Rocher, C\. Pozo, N\. Melero, G\. Martínez, M\. Grandury, and P\. Reviriego \(2025\)It’s the same but not the same: do LLMs distinguish Spanish varieties?\.Proces\. del Leng\. Natural\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2504.20049),[Link](https://arxiv.org/abs/2504.20049)Cited by:[§1](https://arxiv.org/html/2607.21774#S1.p1.2),[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- J\. Mire, Z\. T\. Aysola, D\. Chechelnitsky, N\. Deas, C\. Zerva, and M\. Sap \(2025\)Rejected Dialects: biases Against African American Language in Reward Models\.North American Chapter of the Association for Computational Linguistics\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2502.12858),[Link](https://arxiv.org/abs/2502.12858)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- M\. T\. Nayeem and D\. Rafiei \(2026\)Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models\.Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- V\. Neplenbroek, A\. Bisazza, and R\. Fernández \(2025\)Reading Between the Prompts: how Stereotypes Shape LLM’s Implicit Personalization\.Conference on Empirical Methods in Natural Language Processing\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2505.16467),[Link](https://arxiv.org/abs/2505.16467)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- nostalgebraist \(2020\)Interpreting GPT: the logit lens\.Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- A\. Pan, L\. Chen, and J\. Steinhardt \(2024\)LatentQA: teaching LLMs to decode activations into natural language\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p1.1)\.
- M\. Reusens, P\. Borchert, J\. De Weerdt, and B\. Baesens \(2024\)Native Design Bias: studying the Impact of English Nativeness on Language Model Performance\.IJCNLP\-AACL\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2406.17385),[Link](https://arxiv.org/abs/2406.17385)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- M\. Robles, C\. Bernal, D\. Raigoso, and M\. D\. Rubio \(2025\)SESGO: spanish Evaluation of Stereotypical Generative Outputs\.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2509.03329),[Link](https://arxiv.org/abs/2509.03329)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- M\. J\. Ryan, W\. Held, and D\. Yang \(2024\)Unintended Impacts of LLM Alignment on Global Representation\.Annual Meeting of the Association for Computational Linguistics\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2402.15018),[Link](https://arxiv.org/abs/2402.15018)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p3.2)\.
- Y\. Tan and E\. Celis \(2019\)Assessing Social and Intersectional Biases in Contextualized Word Representations\.Neural Information Processing Systems\.Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- R\. Tang, X\. Zhang, J\. Lin, and F\. Ture \(2023\)What Do Llamas Really Think? Revealing Preference Biases in Language Model Representations\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2311.18812),[Link](https://arxiv.org/abs/2311.18812)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
- R\. Zhang, L\. Lian, Z\. Qi, and G\. Liu \(2025\)Semantic and Structural Analysis of Implicit Biases in Large Language Models: an Interpretable Approach\.In2025 4th International Conference on Artificial Intelligence, Internet of Things and Cloud Computing Technology \(AIoTC\),pp\. 699–703\.External Links:[Document](https://dx.doi.org/10.1109/aiotc66747.2025.11198661),[Link](http://dx.doi.org/10.1109/AIoTC66747.2025.11198661)Cited by:[§2](https://arxiv.org/html/2607.21774#S2.p2.1)\.
## Appendix ATechnical appendices and supplementary material
30 prompts==15 base stories×\\times2 languagesSpanish \(ES\)n=15n\{=\}15English \(EN, translation\)n=15n\{=\}155×\\timesExplicit Colombian5×\\timesImplicit Colombian5×\\timesNeutral5×\\timesExplicit Colombian5×\\timesImplicit Colombian5×\\timesNeutralpairpairpair
Figure 2:Dataset structure\.Each of the 15 base stories is realized in Spanish\-English pairs; the explicitness factor \(explicit / implicit / neutral\) varies*between*stories, yieldingn=5n\{=\}5per cell\.Qwen2\.5\-7B\-Instruct→\\toforward pass \(1 deterministic run\)→\{hℓ\[1\],…,hℓ\[N\]\}\\to\\\{h\_\{\\ell\}\[1\],\\dots,h\_\{\\ell\}\[N\]\\\},hℓ∈ℝ3584h\_\{\\ell\}\\in\\mathbb\{R\}^\{3584\}1DivideNNtokens into 4 positional quartilesqk=⌊kN/4⌋,k∈\{1,2,3,4\}q\_\{k\}=\\lfloor kN/4\\rfloor,\\;k\\\!\\in\\\!\\\{1,2,3,4\\\}→\\toextract the final token of each quartile→\\to4 vectors\[3584\]\[3584\]per prompt:2q2q\_\{2\}q1q\_\{1\}q3q\_\{3\}q4q\_\{4\}AV\.generate\(\) — 1 single call per quartile\(do\_sample=True,T=1\.0T\{=\}1\.0\)zk∼AV\(⋅∣hℓ\[qk\]\)z\_\{k\}\\sim\\mathrm\{AV\}\(\\cdot\\mid h\_\{\\ell\}\[q\_\{k\}\]\)→\\to4 text explanations per prompt:3z2z\_\{2\}z1z\_\{1\}z3z\_\{3\}z4z\_\{4\}Claude Sonnet receives\(z1,z2,z3,z4\)\(z\_\{1\},z\_\{2\},z\_\{3\},z\_\{4\}\)\(structured prompt, not free\-form summary\)→\\toJSON: detected category, quartile offirst appearance, verbatim quote430 prompts×\\times4 quartiles==120 total AV calls
Figure 3:Per\-prompt pipeline\.Each of the 30 prompts passes through the four stages; the resulting vectorshℓ\[qk\]h\_\{\\ell\}\[q\_\{k\}\]and explanationszkz\_\{k\}are indexed by quartile to strictly preserve the positional signal\.### A\.1Cross\-Lingual Divergence in Internal Representations
The analysis of the model’s latent thoughts reveals a marked divergence in the contextualization of neutral prompts depending on the language\. This variation introduces geographic and migratory biases into the representation space that are not explicit in the original input\. Below, three representative cases illustrating this behavior are detailed:
Case A01: Health System and EmploymentOriginal context:Query about enrollment in EPS \(public health insurance\) and the Colombian health system when starting a new job\. Latent representation:In Spanish, the model assumes the user is a “foreigner or newcomer,” injecting a migratory bias absent in the prompt\. In English, the search space shifts towards health insurance in Canada, Germany, or Turkey, linking it to residency procedures\. Analysis:A strong implicit association between employment/health and migration is evident\. Processing in Spanish projects local migratory vulnerability, while English completely internationalizes the query, losing the original geographic relevance\.
Case A03: Gastronomy and Cultural ReferencesOriginal context:Query about a traditional Colombian dish \(ajiaco\)\. Latent representation:In Spanish, the model recognizes the local traditional gastronomic context, although it generalizes by mixing it with other dishes \(sancocho,arepas\)\. In English,ajiacoundergoes a drastic shift towards “Peruvian, Mexican food” and invents concepts like “Colombian paella\.” Analysis:This demonstrates a clear homogenization in English, where Latin American cultural identities are mixed and become interchangeable within the model’s latent space\.
Case A05: Education and FinancingOriginal context:Query about how a university educational loan \(Icetex\) works\. Latent representation:In Spanish, the model tends to transform the concept of a loan into a state “general aid or scholarship\.” In English, the model deflects the query toward international scholarships, mentioning the Erasmus program and French universities\. Analysis:This reflects a severe socioeconomic divergence: processing in Spanish associates the loan with basic local assistance and subsidies, whereas English associates it with international academic mobility\.
Case B01: Institutional Normalization \(EPS\)Original context:Query including the term “EPS” \(a Colombian health insurance entity\)\. Latent representation:In the English version, the model substitutes the term with US insurance companies like Blue Cross and Aetna\. In Spanish, the health context is also unrecognized, triggering activations related to unrelated global entities such as Netflix and SAP\. Analysis:The model loses fidelity to the original context, exhibiting a strong tendency to normalize specific local references by replacing them with globalized elements or US\-centric entities prevalent in its training data\.
Case B02: Geographic Anchoring \(TransMilenio\)Original context:Query referencing “TransMilenio” \(Bogotá’s mass transit system\)\. Latent representation:The model explicitly recognizes it as a Colombian transport system in both languages initially, even when the country is not explicitly mentioned\. However, in later processing stages, the signal dilutes into generic urban transport concepts, referencing metros in Madrid, Barcelona, and Tokyo\. Analysis:Highly distinctive geographic entities successfully preserve their national identity and trigger accurate latent representations early on\. Yet, this specificity fades as the model shifts toward generic global urban frameworks before final text generation\.
Case B03: Legal Terminology \(Tutela\)Original context:Query involving the legal term“tutela”\(a specific Colombian constitutional protection mechanism\)\. Latent representation:The juridical signal successfully activates references to the Colombian context only in the Spanish version\. In English, this association weakens significantly, and the concept transforms into generic categories like “human rights petition” or “constitutional claim,” even shifting the context toward Spain\. Analysis:The preservation of specific legal concepts is highly language\-dependent\. Translation causes the Colombian specificity to dilute into broader international frameworks, demonstrating a semantic normalization in the English latent space\.
Case C01: Health and EmploymentNeutral prompt:How can I enroll in private health insurance if I have just started working?\(ES/EN variants\)\. Latent representation:In Spanish, the model internally generates premises such as“If you are starting to work in Spain…”and anticipates terms like“migratory”or“from the United States”\. In English, the model mentions Canada superficially, without linking it to a transitional status\. Analysis:There is a strong implicit association in Spanish between labor/health insertion and emigration, biasing the interpretation towards a context of migratory vulnerability\.
Case C02: Urban TransportationNeutral prompt:What time does the last metro train run on weekends? Latent representation:Processing in Spanish forces an immediate geographic anchoring, generating the anticipation:“What time is the metro in Barcelona?”\. In English, the internal context is more evenly distributed among various global metropolises \(New York, London, Tokyo\)\. Analysis:This demonstrates a Hispanic\-centric localization bias that reduces the model’s spatial generalization when faced with generic urban queries in Spanish\.
Case C03: Cultural ReferencesNeutral prompt:What is the traditional recipe for a vegetable soup served in many cultures? Latent representation:In Spanish, internal activations are directed towards Ibero\-American references, anticipating“Mexico”or“paella”\. In contrast, in English, the search space is oriented towards the Northern Hemisphere, mentioning Italian and Turkish food, and holidays likeThanksgiving\. Analysis:A cultural preconditioning is evident in the pre\-generation layers, where the language restricts the scope of what the model considers a “generic culture”\.Similar Articles
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
This paper uses Top-K sparse autoencoders to analyze DeepSeek-R1-Distill-Qwen-7B's internal reasoning, contrasting Thinking (CoT) and NoThinking modes, and finds distinct feature activation patterns with causal intervention experiments.
Discovering Millions of Interpretable Features with Sparse Autoencoders
This paper introduces Qwen3-Instruct SAE, a suite of sparse autoencoders trained on Qwen3 instruction-tuned models, enabling the discovery of millions of interpretable features and demonstrating refusal steering capabilities.
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
This study compares probing techniques for identifying latent language in multilingual LLMs, finding that different methods yield inconsistent results, indicating they expose distinct aspects of multilingual processing rather than a single internal lingua franca.
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
This article introduces Qwen-Scope, a toolkit of Sparse Autoencoders (SAEs) trained on Qwen3 and Qwen3.5 models to enable mechanistic analysis and intervention. It releases 14 groups of SAE weights covering dense and MoE backbones, providing sparse representations for residual-stream activations.
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.