Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering
Summary
ArXiv preprint examining how Gemma 4 31B internally represents nine African languages and whether country-specific steering directions can bias English generation, finding regional alignment patterns among African languages and measurable steering effects for Nigeria, Ghana, Kenya, and South Africa.
View Cached Full Text
Cached at: 09/30/26, 09:50 AM
# A Regional Semantic Hub and Cultural Steering
Source: [https://arxiv.org/html/2609.36205](https://arxiv.org/html/2609.36205)
## Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering
Muhammad Abdullahi SaidAffiliation:African Institute for Mathematical Sciences \(AIMS\)Affiliation:University of Cape Town \(UCT\)Email:[mohdasaid@aims\.ac\.za](mailto:)
###### Abstract
We study how Gemma 4 31B represents African languages and responds to cultural steering\. The first study compares nine African languages and three controls using probes, contrast directions, and measures of representation similarity\. Transfer from English varies across languages and layers\. Directions representing an Africa versus West contrast are more aligned among the African languages than between these languages and the controls at several layers\. The comparison across language families passes the reported Holm threshold at five of twelve layers, although dependence between language pairs limits the statistical interpretation\. Within Nigeria, Yoruba and Igbo are more aligned than the average of their pairs with Hausa at eleven of twelve layers\. The second study uses separate English data to construct directions for Nigeria, Ghana, Kenya, and South Africa\. Under union scoring at the selected strengths, estimated differences in attribution rates from random directions range from 0\.63 to 0\.81\. Most outputs pass the automated structural coherence screen\. Comparisons with Aya Expanse 32B show that results depend on the representation measure\. Together, the studies document regional and family patterns in the sampled representations and country steering in English\.
## 1Introduction
African languages remain underrepresented in language model development and evaluation\([Nekoto et al\., 2020](https://arxiv.org/html/2609.36205#bib.bib21);[Adelani et al\., 2021](https://arxiv.org/html/2609.36205#bib.bib20)\)\. Work on multilingual representations asks whether models process different languages through an English oriented internal representation\([Wendler et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib1)\)or through a shared semantic space that is not specific to one language\([Wu et al\., 2025](https://arxiv.org/html/2609.36205#bib.bib2)\)\. Studying African languages broadens the language coverage of this question and provides comparisons across language families within a region\.
We examine two aspects of internal representation in a frozen multilingual model\. Study 1 measures cross language similarity in a translated corpus labelled by entity type and by an Africa versus West content contrast\. It asks whether information transfers from English and whether contrast directions are more similar among African languages than to non African controls\. Study 2 constructs country versus Western directions from a separate English corpus and tests whether adding these directions during generation elicits content associated with Nigeria, Ghana, Kenya, or South Africa\.
In this paper, a regional semantic hub refers to greater alignment of the sampled African language directions on the chosen content contrast\. This operational description does not establish a general semantic space shared by all African languages\. Both studies examine representations of African languages and cultural content in the same model\. Study 1 compares how a regional contrast is represented across languages\. Study 2 tests whether directions derived from country content can alter English generation\. Each study constructs its own directions, and a separate pilot selects the steering layer\. The experiments therefore answer separate questions about representation and intervention; they do not test whether the regional pattern causes the steering effect\.
Our contributions are:
- •A comparison of nine African languages and three controls using four measures of internal representation, including a comparison within Nigeria that holds country fixed while comparing pairs from the same family and from different families\.
- •Evidence that Africa versus West directions are more aligned among the African languages than with the controls at several layers, with the strongest separation at layers 14 to 24\. The comparison across families meets the reported Holm threshold at five of twelve layers; the result within Afro Asiatic remains inconclusive\.
- •An evaluation of four country specific steering directions in English against random directions of matched norm, including specificity, coherence, and extraction method comparisons\.
- •A comparison of English reference measures with Aya Expanse 32B, showing that the model ordering depends on the measure\.
## 2Related Work
#### Multilingual representations\.
Cross language transfer has motivated analyses of shared representations\([Pires et al\., 2019](https://arxiv.org/html/2609.36205#bib.bib5);[Conneau et al\., 2020](https://arxiv.org/html/2609.36205#bib.bib6)\)\.[Wendler et al\. \(2024\)](https://arxiv.org/html/2609.36205#bib.bib1)investigate the latent language of multilingual transformers using the logit lens, which maps intermediate hidden states through the output vocabulary\.[Wu et al\. \(2025\)](https://arxiv.org/html/2609.36205#bib.bib2)study shared semantic representations across languages and modalities\. Our measurements compare hidden states directly\. They measure similarity and transfer on the selected corpus\. The language used during token prediction is outside the scope of these measurements\.
#### African languages and language coverage\.
Participatory machine translation and African language named entity recognition have documented the need for broader language coverage and locally grounded resources\([Nekoto et al\., 2020](https://arxiv.org/html/2609.36205#bib.bib21);[Adelani et al\., 2021](https://arxiv.org/html/2609.36205#bib.bib20)\)\. We analyse internal states to examine the information represented by the model\. A successful probe shows that the selected information can be recovered from those states; it does not by itself explain downstream failures or establish that a model uses the recovered information during generation\.
#### Comparing representations\.
Linear probes measure information recoverable from frozen activations\([Alain and Bengio, 2018](https://arxiv.org/html/2609.36205#bib.bib3)\)\. SVCCA compares correlated low dimensional representations\([Raghu et al\., 2017](https://arxiv.org/html/2609.36205#bib.bib4)\), whereas orthogonal Procrustes measures agreement under an orthogonal transformation\. Different similarity measures do not rank representations identically\([Kornblith et al\., 2019](https://arxiv.org/html/2609.36205#bib.bib7)\)\. We report each measure separately because they assess different properties\. Directions constructed from differences in means compare a specific labelled contrast\([Park et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib8)\)\.
#### Cultural steering\.
Activation addition and contrastive activation addition modify generation by adding directions to internal states\([Turner et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib9);[Rimsky et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib10)\); related work studies single direction behavioural interventions\([Arditi et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib11)\)\. Cultural steering has been investigated through localized knowledge, cultural values, and country specific generation\([Veselovsky et al\., 2025](https://arxiv.org/html/2609.36205#bib.bib12);[Dang and Masud, 2026](https://arxiv.org/html/2609.36205#bib.bib13);[Khanuja et al\., 2026](https://arxiv.org/html/2609.36205#bib.bib14)\)\. Studies of cultural alignment also motivate examining which groups a model’s outputs represent\([Tao et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib15);[AlKhamissi et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib16)\)\. We evaluate a difference in means approach for four African countries\. Our random direction comparison tests whether the selected direction matters\.
## 3Models and Data
The primary model is Gemma 4 31B, instruction tuned\([Google DeepMind, 2026b](https://arxiv.org/html/2609.36205#bib.bib17)\), recorded asgoogle/gemma\-4\-31B\-it, with 60 layers and 5376 hidden dimensions\. Measurements use its twelve sampled layers, indexed from zero as L4, L9,…\\ldots, L59\. The secondary model is Aya Expanse 32B\([Dang et al\., 2024](https://arxiv.org/html/2609.36205#bib.bib18)\), recorded asCohereForAI/aya\-expanse\-32b, with 40 layers and 8192 hidden dimension; measurements use L0, L4,…\\ldots, L36, L39\. Both models use bfloat16 precision\. Aya is evaluated on Swahili, Hausa, Amharic, Zulu, Yoruba, and Finnish\. The comparison between models using English as a reference uses the five shared African languages\. An additional saved pairwise comparison uses these five languages and Finnish as its sole control \(Appendix[A\.5](https://arxiv.org/html/2609.36205#A1.SS5)\)\. Steering uses Gemma alone\.
#### Study 1 corpus\.
Each record pairs a full English sentence with a translation generated using Gemini 3 Pro\([Google DeepMind, 2026a](https://arxiv.org/html/2609.36205#bib.bib19)\)\. Each of the twelve language files contains 1,628 pairs: 37 templates applied to 44 entities, with 814 Africa labelled and 814 West labelled records\. The entity classes contain 444 city, 444 language, 370 food, and 370 landmark records\. The twelve files share the same 1,628 English sentences and entity identities\. Entity type defines the four class probe task; region defines the difference in means direction\. These are different prediction targets\. Appendix[A\.1](https://arxiv.org/html/2609.36205#A1.SS1)gives the entity counts and a parallel example\. Table[1](https://arxiv.org/html/2609.36205#S3.T1)gives the language groups\. Danish and Finnish provide European controls, and Quechua provides a low resource non African comparison\. These controls broaden coverage but do not establish that training resources or translation quality are matched across languages\.
Table 1:Language groups used in the comparisons\. Songhai contributes to the all African comparison but not the two family level comparisons\.
#### Study 2 data\.
Steering uses a separate set of 800 English contrastive pairs, with 200 per country\. Each pair contains a sentence about a target country’s culture and a Western counterpart\. Evaluation uses 40 open ended prompts that do not name a target country: 37 labelled cultural prompts and three labelled controls in rubric version 2\.0\. The saved aggregate rates include these controls\. Appendix[B\.1](https://arxiv.org/html/2609.36205#A2.SS1)gives a construction example and documents the implemented scoring exclusions\. This construction dataset is separate from the multilingual corpus in Study 1\.
#### Translation and evaluation provenance\.
Gemini 3 Pro was used for Study 1 translations, and Gemini 3 Pro Preview evaluates Study 2 generations\. The instruction tuned Gemma 4 31B model is the primary model whose representations we analyse\. Translation fidelity is not established by an independent human assessment in the reported evaluation\. The translator, judge, and primary model also share a developer\. Translation and cultural content may both affect the measured representations\. Our conclusions apply to this corpus and language sample\.
#### Compute\.
Experiments used a single A100 80GB instance on Google Cloud, at a reported total cost of approximately USD 1,000\. This includes activation extraction, the Study 1 measures, and the Study 2 sweep over three extraction methods, four country and five random vectors, six strengths, 40 prompts, and five samples per prompt\.
## 4Study 1: Methods
#### Sentence representations\.
For sentencexx, letht\(ℓ\)\(x\)∈ℝdh\_\{t\}^\{\(\\ell\)\}\(x\)\\in\\mathbb\{R\}^\{d\}be the output of transformer blockℓ\\ellat token positiontt\. We tokenize raw sentences without chat formatting, using padded batches of 16 and truncation at 128 tokens\. We average over positions with attention mask value one, denotedTxT\_\{x\}:
h¯\(ℓ\)\(x\)=1\|Tx\|∑t∈Txht\(ℓ\)\(x\)\.\\bar\{h\}^\{\(\\ell\)\}\(x\)=\\frac\{1\}\{\|T\_\{x\}\|\}\\sum\_\{t\\in T\_\{x\}\}h\_\{t\}^\{\(\\ell\)\}\(x\)\.\(1\)Each sentence thus contributes one fixed length vector per layer\. Padding is excluded; special tokens added by the tokenizer are not separately removed\. Pooling uses float32 values captured by a forward hook on the selected block’s output\. The four measures below use these pooled representations\.
#### Entity type probe\.
A four class linear classifier is trained on frozen representations using Adam, learning rate 0\.01, for 100 epochs\. We report English to English, target to target, and English to target accuracy; the last applies the English classifier to target language representations without retraining\. Uniform guessing gives 25%; the full corpus’s majority class baseline is 27\.3%\. The variant with shared entities randomly splits records 75:25 for training and evaluation, so entity identities can occur in both\. For evaluation on unseen entities, each of five folds holds out one African and one Western entity per class, trains on the remaining entities, and evaluates on all records for the eight excluded entities\. This assesses generalization beyond the training entities, while retaining the corpus’s templates and construction process\. The majority class baseline depends on the class counts in each test fold\.
#### Africa versus West direction\.
For the Africa labelled setAAand West labelled setWWin each language, define
μA=1\|A\|∑x∈Ah¯\(x\),μW=1\|W\|∑x∈Wh¯\(x\),v=μA−μW‖μA−μW‖2\.\\begin\{split\}\\mu\_\{A\}&=\\frac\{1\}\{\|A\|\}\\sum\_\{x\\in A\}\\bar\{h\}\(x\),\\qquad\\mu\_\{W\}=\\frac\{1\}\{\|W\|\}\\sum\_\{x\\in W\}\\bar\{h\}\(x\),\\\\ v&=\\frac\{\\mu\_\{A\}\-\\mu\_\{W\}\}\{\\\|\\mu\_\{A\}\-\\mu\_\{W\}\\\|\_\{2\}\}\.\\end\{split\}\(2\)The English to target angle isθ=arccos\(vEN⊤vXX\)\\theta=\\arccos\(v\_\{\\mathrm\{EN\}\}^\{\\top\}v\_\{\\mathrm\{XX\}\}\), reported in degrees\. A smaller angle indicates closer alignment of this specific contrast; it does not establish equivalence of the complete representation spaces\. Directions, SVCCA, and Procrustes are computed over ten seeded 80% subsamples drawn without replacement\. Reported metric summaries average these ten runs\.
#### SVCCA\.
Singular Vector Canonical Correlation Analysis compares correlated components of the English and target representations\. Each side is reduced by principal component analysis to rankk∈\{20,40\}k\\in\\\{20,40\\\}\. We average the top five canonical correlations and subtract the corresponding mean over ten runs with shuffled pairs:
S=15∑i=15ρi−110∑s=110\(15∑i=15ρi\(s\)\)\.S=\\frac\{1\}\{5\}\\sum\_\{i=1\}^\{5\}\\rho\_\{i\}\-\\frac\{1\}\{10\}\\sum\_\{s=1\}^\{10\}\\left\(\\frac\{1\}\{5\}\\sum\_\{i=1\}^\{5\}\\rho\_\{i\}^\{\(s\)\}\\right\)\.\(3\)Hereρi\\rho\_\{i\}denotes a canonical correlation andρi\(s\)\\rho\_\{i\}^\{\(s\)\}its counterpart in a shuffled run\. A positive value indicates correlation above this shuffled baseline\.
#### Orthogonal Procrustes\.
LetXXandYYcontain representations of paired English and target sentences, with each set centred on its mean\. We fit an orthogonal transformation and report a normalized residual similarity:
R∗=argminR⊤R=I‖XR−Y‖F,P=clip\[0,1\]\(1−‖XR∗−Y‖F2‖X‖F2\+‖Y‖F2\)\.\\begin\{split\}R^\{\*\}&=\\mathop\{\\arg\\min\}\_\{R^\{\\top\}R=I\}\\\|XR\-Y\\\|\_\{F\},\\\\ P&=\\operatorname\{clip\}\_\{\[0,1\]\}\\left\(1\-\\frac\{\\\|XR^\{\*\}\-Y\\\|\_\{F\}^\{2\}\}\{\\\|X\\\|\_\{F\}^\{2\}\+\\\|Y\\\|\_\{F\}^\{2\}\}\\right\)\.\\end\{split\}\(4\)The Frobenius norm∥⋅∥F\\\|\\cdot\\\|\_\{F\}summarizes the matrix residual\. LargerPPindicates closer alignment\.
#### Pairwise regional and family comparisons\.
At each sampled layer, we compare language directions normalized to unit length within each matched subsampling seed and average their cosines:Mab=110∑s=09va,s⊤vb,sM\_\{ab\}=\\frac\{1\}\{10\}\\sum\_\{s=0\}^\{9\}v\_\{a,s\}^\{\\top\}v\_\{b,s\}\. We compare cosines within all African languages, within Niger Congo, and within Afro Asiatic with the corresponding cosines to controls\. A fourth comparison tests Niger Congo to Afro Asiatic pairs against Niger Congo to control pairs\. One sided Mann Whitney tests summarize whether the first set tends to have larger values\. Holm correction adjusts the tests across layers within each comparison: twelve for Gemma and eleven for Aya\. Appendix[A\.3](https://arxiv.org/html/2609.36205#A1.SS3)reports Gemma means and adjusted values; Appendix[A\.5](https://arxiv.org/html/2609.36205#A1.SS5)gives the saved Aya comparison\.
Language pairs reuse the same directions, so their observations are dependent\. Holm adjustment accounts for the number of layer comparisons\. Dependence between pairs remains, so we use the test results as descriptive evidence\. Within Nigeria, we compare the cosine for Yoruba and Igbo with the mean of the cosines for Yoruba with Hausa and Igbo with Hausa\. This holds country fixed, but it does not control all differences between languages\. Tests and bootstrap intervals across layers are also limited by dependence between layers\.
#### Cross model comparison\.
We compare the English reference measures at relative depths from 10% to 90%, choosing each model’s nearest sampled layer\. For each measure and depth, we summarize the five paired African language differences by their median and a 95% bias corrected and accelerated bootstrap interval from 10,000 resamples\. Positive differences favour Aya: Gemma minus Aya for the angle, and Aya minus Gemma for the other measures\. The probe ratio is English to target accuracy divided by target to target accuracy\. The additional Aya pairwise analysis uses fewer African languages and only Finnish as a control\. Its coverage differs from the Gemma analysis\.
## 5Study 1: Results
### 5\.1English reference similarity varies by language and layer
Table[2](https://arxiv.org/html/2609.36205#S5.T2)gives the numerical results for all twelve languages\. The probe evaluated on unseen entities exceeds 25% chance at early layers for most languages, with substantial variation in the margin above chance\. At L19, accuracy is 71% for Danish, 61% for Swahili, 54% for Amharic, and 26% for Yoruba\. At L59, these values are 32%, 28%, 29%, and 27%\. Holding out entities reduces the stronger transfer scores, while preserving the broad depth pattern \(Appendix[A\.2](https://arxiv.org/html/2609.36205#A1.SS2)\)\.
Accuracy on unseen entities \(%\)English reference geometry at L19LanguageL4L19L59Angle \(∘\)SVCCA \(k=40k=40\)ProcrustesSwahili486128750\.690\.77Yoruba292627760\.680\.60Igbo323326690\.680\.63Zulu334027840\.680\.63Wolof313826500\.660\.58Hausa414429690\.690\.68Amharic505429540\.690\.75Somali354927720\.680\.67Songhai363827520\.640\.56Quechua375126510\.620\.50Finnish345526420\.680\.69Danish547132260\.690\.79Table 2:Gemma results at the precision shown in the heatmaps\. Probe chance is 25%\. L4, L19, and L59 illustrate three selected depths; the L19 values use the same layer for every language\. Angles compare Africa versus West directions, whereas the probe predicts entity type\. Higher SVCCA and Procrustes scores indicate greater similarity under their respective definitions\.The angle between the Africa versus West directions generally increases away from the input and decreases toward the output \(Figure[1](https://arxiv.org/html/2609.36205#S5.F1)\)\. Six African languages reach their largest angle at L24; Quechua and Finnish peak at L34\. Amharic remains relatively rotated from English at the final layer\.
Figure 1:Angles between English and target language Africa versus West directions across Gemma layers\. Smaller angles indicate closer alignment\.SVCCA varies less across languages than the probe: at L19 its scores range from 0\.62 for Quechua to 0\.69 for several languages\. Procrustes distinguishes these representations more strongly, ranging from 0\.50 for Quechua to 0\.79 for Danish at that layer\. African languages do not form a consistently separate group from the controls across these English reference measures\. Resource availability is one possible explanation for the language ordering, but the design does not separate it from tokenization and translation quality\. We therefore do not interpret the ordering as a measured training resource effect\.
### 5\.2Regional alignment is strongest at early and middle depth
The pairwise comparison shows greater alignment among African language directions than between these directions and the controls at several layers \(Figure[2](https://arxiv.org/html/2609.36205#S5.F2)\)\. At L19 the all African mean cosine is 0\.746, compared with 0\.345 for African to control pairs\. The corresponding means are 0\.815 versus 0\.540 at L14 and 0\.596 versus 0\.282 at L24\. Table[3](https://arxiv.org/html/2609.36205#S5.T3)summarizes the reported corrected thresholds, subject to the dependence limitation in Section[4](https://arxiv.org/html/2609.36205#S4)\.
Table 3:Reported Mann Whitney comparisons, Holm adjusted across twelve layers within each comparison\. NC: Niger Congo; AA: Afro Asiatic\. Shared languages make the pairwise observations dependent\.The comparison across families meets the reported threshold at L9, L14, L19, L24, and L44\. The observed alignment therefore crosses the family boundary at these layers\. The comparison within Afro Asiatic meets the threshold at no layer\. With only three languages in that group, its result remains inconclusive\.
Figure 2:Mean cosines among African languages and across the two language families, compared with control pairs\. Solid green lines mark Holm adjustedp<0\.05p<0\.05; dotted green lines mark only rawp<0\.05p<0\.05\. Bands reproduce the saved bootstrap intervals over direction subsamples\. The tests reuse languages, which limits their statistical interpretation\.Within Nigeria, the Yoruba and Igbo pair is more aligned than the mean of the two Hausa pairs at eleven of twelve layers, with an average difference of 0\.116 and a reported 95% bootstrap interval across layers of \[0\.078, 0\.151\]\. The original paired test givest\(11\)=5\.93t\(11\)=5\.93with one sidedp<0\.001p<0\.001; sign and Wilcoxon checks give the same direction of evidence\. These results across layers are descriptive because the layers are dependent\. At L19 the family gap is approximately 0\.08, compared with the all African to control gap of 0\.40\. These two gaps describe L19 only\.
The late layer regional differences are smaller\. At L59, within African and African to control means are 0\.821 and 0\.742\. Quechua’s position outside the African grouping shows that the pattern is not shared by every low resource language in this sample; it does not rule out translation, tokenization, or content differences as explanations\.
### 5\.3Cross model differences depend on the measure
Aya does not reproduce every Gemma depth pattern\. Its probe transfer for Finnish and Hausa peaks later, and its angle map does not show the same overall shape\. Across the five African languages, Aya has higher SVCCA and Gemma has higher Procrustes alignment at all nine matched depths, with reported bootstrap intervals excluding zero\. Aya also has higher probe transfer at depths from 50% to 90%, with intervals excluding zero\. Its probe ratio is higher at these depths, but the interval at 70% includes zero,\[−0\.0029,0\.3188\]\[\-0\.0029,0\.3188\]\. The probe ratio intervals at 50%, 60%, 80%, and 90% exclude zero\. Most angle differences remain inconclusive\. In the additional saved Aya pairwise analysis, the all African comparison meets the reported Holm threshold at L32 and L36 \(2/11 layers\), while neither within family comparison nor the comparison across families does so at any layer \(Appendix[A\.5](https://arxiv.org/html/2609.36205#A1.SS5)\)\. This smaller sample provides limited evidence about whether the Gemma pattern extends to Aya\.
## 6Study 2: Methods
For each countrycc, we compute a unit directionvc=\(μc−μW\)/‖μc−μW‖2v\_\{c\}=\(\\mu\_\{c\}\-\\mu\_\{W\}\)/\\\|\\mu\_\{c\}\-\\mu\_\{W\}\\\|\_\{2\}from English country and Western anchor sentences\. During generation we add it to the residual stream, the model state passed between layers, usingh←h\+αvch\\leftarrow h\+\\alpha v\_\{c\}\. Extraction and injection use L24, selected by a pilot using Nigeria alone over layers and strengths\. This choice was not derived from Study 1’s regional analysis\.
The primary extractor,chat\_mean, averages all nonpadding token representations from inputs formatted with the chat template, including the instruction and role/generation markers; it does not pool generated response tokens\. The user message is “Complete this sentence in one sentence: \[sentence\]”, formatted with a generation prompt and thinking disabled\. Extraction truncates at 128 tokens and uses batches of 16\. We compare it with raw text mean pooling \(raw\_mean\) and the saved chat last position extractor \(chat\_last\)\. The two chat based methods perform similarly at their selected strengths; raw text extraction performs substantially worse \(Appendix[B\.5](https://arxiv.org/html/2609.36205#A2.SS5)\)\. The intervention is added to the output of block 24 at every position processed by that block during generation, including the initial prompt pass\.
#### Generation and controls\.
We sweepα∈\{0,20,30,40,50,60\}\\alpha\\in\\\{0,20,30,40,50,60\\\}, using temperature 0\.7, toppp0\.9, and a maximum of 80 new tokens\. Each condition uses 40 prompts and five samples per prompt, giving 200 generated responses before scoring exclusions\. The prompts are wrapped as user turns using the same sentence completion instruction and chat template\. Five random unit vectors, with seeds 42 to 46, are injected at the same layer and strengths and pooled at evaluation\. Theα=0\\alpha=0condition is unsteered\. Neither control is an explicit country prompting baseline\.
#### Scoring\.
A lexical rubric checks for country specific strict tokens using substring matching that ignores letter case\. Gemini 3 Pro receives the prompt, rubric, rubric verdict, and response and returns country attribution and a coherence judgement at temperature zero\. A separate structural coherence diagnostic requires at least five words separated by whitespace and a repetition fraction below 0\.6, where repetition is one minus the proportion of distinct lowercased words\. This diagnostic detects short or repetitive outputs\. It provides a limited check of linguistic quality\.
We report four scoring modes: rubric only, judge only, intersection \(both attribute the country\), and union \(either attributes it\)\. Union is the primary reported mode\. The non rubric modes exclude outputs marked incoherent by the model judge or lacking a country attribution verdict; the structural diagnostic is reported separately and does not filter these rates\. Rubric only scoring does not apply the judge coherence filter\. The three rubric designated control prompts are included in the saved aggregates, and the implemented diagonal exclusion is the South African northern festival prompt only\. Intersection is more restrictive than union on their common eligible outputs\.
#### Rates, uncertainty, and selection\.
The diagonal hit rate is the fraction of eligible outputs attributed to the country corresponding to the steering vector\. Off diagonal rates measure attribution to the other countries\. Rates have 95% Wilson intervals\. Prompt bootstrap intervals use 10,000 resamples of per prompt means, weighting prompts equally\. When coherence filtering leaves unequal numbers of responses per prompt, this estimand differs from the response pooled rate underlying the Wilson interval\. Both are reported in Appendix[B\.4](https://arxiv.org/html/2609.36205#A2.SS4)\.
At each country’s selected positive strength, we estimate the difference from the pooled random direction baseline using independent Jeffreys posteriors:
pc∼Beta\(kc\+12,nc−kc\+12\),pr∼Beta\(kr\+12,nr−kr\+12\)\.\\begin\{split\}p\_\{c\}&\\sim\\mathrm\{Beta\}\(k\_\{c\}\+\\tfrac\{1\}\{2\},n\_\{c\}\-k\_\{c\}\+\\tfrac\{1\}\{2\}\),\\\\ p\_\{r\}&\\sim\\mathrm\{Beta\}\(k\_\{r\}\+\\tfrac\{1\}\{2\},n\_\{r\}\-k\_\{r\}\+\\tfrac\{1\}\{2\}\)\.\\end\{split\}\(5\)Herekkandnnare hit and eligible trial counts\. We report the posterior mean ofpc−prp\_\{c\}\-p\_\{r\}and a 95% credible interval from 20,000 draws\. The selected strength maximizes the observed diagonal rate, breaking ties toward smaller strength\. Selecting the strength on the evaluation data can overestimate both the rate and its difference from random\. We therefore report the full strength sweep alongside the selected results\.
## 7Study 2: Results
### 7\.1Country directions increase attribution to the target country
For the prompt “The national football team of this country is affectionately known by the nickname”, an unsteered output names Australia’s Socceroos\. Atα=50\\alpha=50, the four country directions elicit Super Eagles, Black Stars, Harambee Stars, and Bafana Bafana, respectively; a random direction produces England’s Three Lions\. This example illustrates the intervention on one prompt\. Aggregate results appear in Table[4](https://arxiv.org/html/2609.36205#S7.T4)\.
Table 4:Country steering results withchat\_meanextraction\. Union rates use the listed selected strengths; random rates pool five random directions at the same strength\. Intersection rates use their own selected strengths: 50 for Nigeria, Ghana, and Kenya; 40 for South Africa\. The effect column reports the posterior mean of the difference from random\. CrI denotes a posterior credible interval\. Selecting the best strength affects both rates and differences\.Under union scoring, differences from random range from 0\.63 to 0\.81, with the reported intervals excluding zero\. The more restrictive intersection rates are lower: 0\.62 for Nigeria, 0\.55 for South Africa, 0\.44 for Ghana, and 0\.54 for Kenya\. The ordering of Ghana and Kenya thus depends on the scoring rule\. At a common strength of 50, union rates are 0\.88, 0\.71, 0\.68, and 0\.72 for Nigeria, Ghana, Kenya, and South Africa, respectively, compared with approximately 0\.03 to 0\.08 for individual random directions\. These results show that the directions alter the country content measured by this evaluation\.
### 7\.2Specificity and coherence vary with strength
Each country direction eventually elicits its own country more often than the others, but separation is not immediate \(Appendix[B\.3](https://arxiv.org/html/2609.36205#A2.SS3)\)\. Atα=30\\alpha=30, the Ghana vector produces Nigerian and Ghanaian attribution at similar rates, 39\.1% and 38\.6%\. At the common strengthα=50\\alpha=50, the largest confusions are between these countries: the Ghana vector produces Nigerian attribution in 15\.0% of eligible responses, and the Nigeria vector produces Ghanaian attribution in 10\.5%\. Shared cultural content is a possible explanation, but this experiment does not isolate its cause\.
Figure 3:Attribution to each target country across steering strengths\. Lines show country directions and pooled random directions at positive strengths; bands are the saved 95% Wilson intervals\. The separate point at zero shows the measured unsteered rate over all 40 prompts\. South Africa’s country curve excludes one prompt, while its unsteered point and random pool retain it\. Selected strengths are marked\. Appendix[B\.4](https://arxiv.org/html/2609.36205#A2.SS4)gives intervals based on prompt means\.Attribution rises and then plateaus or declines, with selected strengths between 40 and 50 \(Figure[3](https://arxiv.org/html/2609.36205#S7.F3)\)\. The attribution rate for each country direction exceeds its random pool at every tested positive strength\. The structural coherence rate for pooled country vector outputs is 100% through strength 50 and approximately 97% at 60\. Random directions reach approximately 85% pooled coherence at 50, including one seed at 56\.5%\. The structural screen and language model coherence judgement agree on 94\.1% of 9,200 scored generations\. Their agreement is not an independent human validation of coherence\.
At high strengths, attribution to both the target and other countries declines while measured coherence remains high\. The outputs therefore contain fewer of the country signals recognised by the evaluation\.
## 8Discussion
The regional pattern concerns directions representing the Africa versus West content contrast\. The largest separation from controls occurs in the early and middle layers\. The comparison across Niger Congo and Afro Asiatic meets the reported Holm threshold at five of twelve layers\. Within Nigeria, Yoruba and Igbo are more aligned than the average of their pairs with Hausa at eleven of twelve layers\. Regional and family patterns therefore coexist in this sample\. Shared languages and dependent layers limit the statistical interpretation, while translation and content effects remain possible explanations\.
These measurements address a narrower question than a general semantic hub hypothesis\([Wu et al\., 2025](https://arxiv.org/html/2609.36205#bib.bib2)\)\. They do not determine whether the model uses English during intermediate token prediction\. The comparison with Aya also depends on the measure: Aya has higher SVCCA and Gemma has higher Procrustes alignment at the matched depths\. The smaller Aya pairwise sample has different language and control coverage, so it cannot isolate the model as the cause of the regional differences\.
The steering intervention changes country attribution in English generation\. Each country direction exceeds the random pool at every tested positive strength\. Under union scoring, estimated differences at the selected strengths range from 0\.63 to 0\.81\. These estimates depend on the scoring rules, eligible responses, and selection of strength on the evaluation data\. The structural screen detects short or repetitive outputs; its high pass rate does not establish cultural accuracy or human judged coherence\. An explicit country prompting baseline is needed to assess whether steering improves on asking for the country directly\.
The two studies use distinct data and independently constructed directions\. Their findings establish an observed regional pattern and a separate intervention effect\. Testing a connection would require intervening on the regional directions from Study 1 and measuring the resulting behaviour\. Independently assessed translations and steering prompts in African languages would also test whether the findings extend beyond the present evaluation\.
## 9Conclusion
Gemma’s representations show regional and family patterns for the sampled Africa versus West contrast\. The comparison with Aya depends on the representation measure\. Separately, directions derived from English country content increase attribution to Nigeria, Ghana, Kenya, and South Africa relative to random directions\. These findings support regional alignment on the studied corpus and country steering in English\. A universal African semantic hub, steering in African languages, and a causal connection between the studies remain untested\.
## Limitations
#### Corpus and scope\.
The multilingual analysis uses one corpus organized around an Africa versus West content contrast and translated by a model\. Translation errors, shared construction patterns, and cultural or topical content can affect its geometry\. We lack independent human translation validation in the reported evaluation\. Language resource categories are not direct measurements of model training exposure, and tokenization and translation quality remain confounded with those categories\.
#### Statistical interpretation\.
Language pairs reuse directions and sampled layers belong to the same model, so neither set provides independent replications\. The reported pairwise tests and intervals across layers require caution even after multiple testing correction\. The five language cross model bootstrap also has limited coverage\. For steering, the prompt cluster intervals account for repeated sampling within prompts, but neither they nor random subtraction remove optimism from selecting the best strength\. The pilot using Nigeria alone may favour that country\.
#### Generalization and evaluation\.
Steering is limited to Gemma; the Aya regional comparison has reduced language coverage and only one control\. Study 2 is entirely in English and has no explicit country prompting comparison\. Its lexical rubric and model judge measure country content through selected prompts; they do not measure the full cultural diversity of a country\. Saved aggregate rates include three general control prompts and therefore are not rates restricted to cultural prompts\. South Africa’s northern festival exclusion is not applied to the random pool, so that comparison uses different prompt coverage\. The judge sees the rubric’s verdict and shares a developer with the evaluated model, so agreement between the rubric and judge is not independent validation\. The reported union rates exclude outputs judged incoherent, and coherence does not guarantee factual accuracy or cultural appropriateness\.
## Project Repository
## Acknowledgements
I want to start by thanking the Almighty Allah for sparing my life, guiding me and supporting me in this journey\.
This work was made possible by a scholarship from Google DeepMind awarded through the African Institute for Mathematical Sciences \(AIMS\)\. I am grateful to both organizations for their funding and the opportunity\. I owe particular thanks to my supervisor, Professor Jonathan Shock of the University of Cape Town, for his invaluable guidance and support\. Finally, I thank my parents whose unwavering support made all of this possible\.
## References
- Adelaniet al\.\(2021\)D\. I\. Adelani, J\. Abbott, G\. Neubig, D\. D’souza, J\. Kreutzer, C\. Lignos, C\. Palen\-Michel, H\. Buzaaba, S\. Rijhwani, S\. Ruder, S\. Mayhew, I\. A\. Azime, S\. H\. Muhammad, C\. C\. Emezue, J\. Nakatumba\-Nabende, P\. Ogayo, A\. Anuoluwapo, C\. Gitau, D\. Mbaye, J\. Alabi, S\. M\. Yimam, T\. R\. Gwadabe, I\. Ezeani, R\. A\. Niyongabo, J\. Mukiibi, V\. Otiende, I\. Orife, D\. David, S\. Ngom, T\. Adewumi, P\. Rayson, M\. Adeyemi, G\. Muriuki, E\. Anebi, C\. Chukwuneke, N\. Odu, E\. P\. Wairagala, S\. Oyerinde, C\. Siro, T\. S\. Bateesa, T\. Oloyede, Y\. Wambui, V\. Akinode, D\. Nabagereka, M\. Katusiime, A\. Awokoya, M\. MBOUP, D\. Gebreyohannes, H\. Tilaye, K\. Nwaike, D\. Wolde, A\. Faye, B\. Sibanda, O\. Ahia, B\. F\. P\. Dossou, K\. Ogueji, T\. I\. DIOP, A\. Diallo, A\. Akinfaderin, T\. Marengereke, and S\. OseiMasakhaNER: named entity recognition for African languages\.Transactions of the Association for Computational Linguistics9,pp\. 1116–1131\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00416),[Link](https://aclanthology.org/2021.tacl-1.66/)Cited by:[§1](https://arxiv.org/html/2609.36205#S1.p1.1),[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px2.p1.1)\.
- Alain and Bengio \(2018\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644v4\.Note:Version 4, revised 22 November 2018External Links:[Link](https://arxiv.org/abs/1610.01644v4)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px3.p1.1)\.
- AlKhamissiet al\.\(2024\)B\. AlKhamissi, M\. ElNokrashy, M\. Alkhamissi, and M\. DiabInvestigating cultural alignment of large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12404–12422\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.671)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Arditiet al\.\(2024\)A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. NandaRefusal in language models is mediated by a single direction\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS\),Note:arXiv:2406\.11717External Links:[Link](https://papers.nips.cc/paper_files/paper/2024/hash/f545448535dfde4f9786555403ab7c49-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Conneauet al\.\(2020\)A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. StoyanovUnsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 8440–8451\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.747),1911\.02116Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px1.p1.1)\.
- Danget al\.\(2024\)J\. Dang, S\. Singh, D\. D’souza, A\. Ahmadian, A\. Salamanca, M\. Smith, A\. Peppin, S\. Hong, M\. Govindassamy, T\. Zhao, S\. Kublik, M\. Amer, V\. Aryabumi, J\. A\. Campos, Y\. Tan, T\. Kocmi, F\. Strub, N\. Grinsztajn, Y\. Flet\-Berliac, A\. Locatelli, H\. Lin, D\. Talupuru, B\. Venkitesh, D\. Cairuz, B\. Yang, T\. Chung, W\. Ko, S\. S\. Shi, A\. Shukayev, S\. Bae, A\. Piktus, R\. Castagné, F\. Cruz\-Salinas, E\. Kim, L\. Crawhall\-Stein, A\. Morisot, S\. Roy, P\. Blunsom, I\. Zhang, A\. Gomez, N\. Frosst, M\. Fadaee, B\. Ermis, A\. Üstün, and S\. HookerAya expanse: combining research breakthroughs for a new multilingual frontier\.Note:arXiv:2412\.04261External Links:[Link](https://arxiv.org/abs/2412.04261)Cited by:[§3](https://arxiv.org/html/2609.36205#S3.p1.1)\.
- Dang and Masud \(2026\)T\. D\. A\. Dang and S\. MasudCultural value alignment via latent activation steering in large language models\.arXiv preprint arXiv:2605\.26365\.External Links:2605\.26365,[Link](https://arxiv.org/abs/2605.26365)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Google DeepMind \(2026a\)Google DeepMindGemini 3 pro model card\.Note:[https://storage\.googleapis\.com/deepmind\-media/Model\-Cards/Gemini\-3\-Pro\-Model\-Card\.pdf](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf)Updated May 2026; model released November 2025Cited by:[§3](https://arxiv.org/html/2609.36205#S3.SS0.SSS0.Px1.p1.1)\.
- Google DeepMind \(2026b\)Google DeepMindGemma 4 model card\.Note:[https://ai\.google\.dev/gemma/docs/core/model\_card\_4](https://ai.google.dev/gemma/docs/core/model_card_4)Cited by:[§3](https://arxiv.org/html/2609.36205#S3.p1.1)\.
- Khanujaet al\.\(2026\)S\. Khanuja, H\. Liu, S\. Zhang, J\. Lambert, M\. Chen, R\. Mathews, and L\. WangSteering LLMs for culturally localized generation\.arXiv preprint arXiv:2603\.23301\.External Links:2603\.23301,[Link](https://arxiv.org/abs/2603.23301)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Kornblithet al\.\(2019\)S\. Kornblith, M\. Norouzi, H\. Lee, and G\. HintonSimilarity of neural network representations revisited\.InProceedings of the 36th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.97,pp\. 3519–3529\.External Links:1905\.00414,[Link](https://proceedings.mlr.press/v97/kornblith19a.html)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px3.p1.1)\.
- Nekotoet al\.\(2020\)W\. Nekoto, V\. Marivate, T\. Matsila, T\. Fasubaa, T\. Fagbohungbe, S\. O\. Akinola, S\. Muhammad, S\. Kabongo Kabenamualu, S\. Osei, F\. Sackey, R\. A\. Niyongabo, R\. Macharm, P\. Ogayo, O\. Ahia, M\. M\. Berhe, M\. Adeyemi, M\. Mokgesi\-Selinga, L\. Okegbemi, L\. Martinus, K\. Tajudeen, K\. Degila, K\. Ogueji, K\. Siminyu, J\. Kreutzer, J\. Webster, J\. T\. Ali, J\. Abbott, I\. Orife, I\. Ezeani, I\. A\. Dangana, H\. Kamper, H\. Elsahar, G\. Duru, G\. Kioko, M\. Espoir, E\. van Biljon, D\. Whitenack, C\. Onyefuluchi, C\. C\. Emezue, B\. F\. P\. Dossou, B\. Sibanda, B\. Bassey, A\. Olabiyi, A\. Ramkilowan, A\. Öktem, A\. Akinfaderin, and A\. BashirParticipatory research for low\-resourced machine translation: a case study in African languages\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 2144–2160\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.195),[Link](https://aclanthology.org/2020.findings-emnlp.195/)Cited by:[§1](https://arxiv.org/html/2609.36205#S1.p1.1),[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px2.p1.1)\.
- Parket al\.\(2024\)K\. Park, Y\. J\. Choe, and V\. VeitchThe linear representation hypothesis and the geometry of large language models\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.235,pp\. 39643–39666\.External Links:[Link](https://proceedings.mlr.press/v235/park24c.html)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px3.p1.1)\.
- Pireset al\.\(2019\)T\. Pires, E\. Schlinger, and D\. GarretteHow multilingual is multilingual BERT?\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 4996–5001\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1493),1906\.01502Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px1.p1.1)\.
- Raghuet al\.\(2017\)M\. Raghu, J\. Gilmer, J\. Yosinski, and J\. Sohl\-DicksteinSVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability\.InAdvances in Neural Information Processing Systems 30 \(NIPS\),Note:arXiv:1706\.05806External Links:[Link](https://proceedings.neurips.cc/paper/2017/hash/dc6a7e655d7e5840e66733e9ee67cc69-Abstract.html)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px3.p1.1)\.
- Rimskyet al\.\(2024\)N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. TurnerSteering Llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 15504–15522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Taoet al\.\(2024\)Y\. Tao, O\. Viberg, R\. S\. Baker, and R\. F\. KizilcecCultural bias and cultural alignment of large language models\.PNAS Nexus3\(9\),pp\. pgae346\.External Links:[Document](https://dx.doi.org/10.1093/pnasnexus/pgae346),2311\.14096Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Turneret al\.\(2024\)A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmidActivation addition: steering language models without optimization\.arXiv preprint arXiv:2308\.10248v4\.Note:Version 4, revised 4 June 2024External Links:[Link](https://arxiv.org/abs/2308.10248v4)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Veselovskyet al\.\(2025\)V\. Veselovsky, B\. Argin, B\. Stroebl, C\. Wendler, R\. West, J\. Evans, T\. L\. Griffiths, and A\. NarayananLocalized cultural knowledge is conserved and controllable in large language models\.arXiv preprint arXiv:2504\.10191\.External Links:2504\.10191,[Link](https://arxiv.org/abs/2504.10191)Cited by:[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px4.p1.1)\.
- Wendleret al\.\(2024\)C\. Wendler, V\. Veselovsky, G\. Monea, and R\. WestDo llamas work in English? on the latent language of multilingual transformers\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15366–15394\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.820),2402\.10588Cited by:[§1](https://arxiv.org/html/2609.36205#S1.p1.1),[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2025\)Z\. Wu, X\. V\. Yu, D\. Yogatama, J\. Lu, and Y\. KimThe semantic hub hypothesis: language models share semantic representations across languages and modalities\.InThe Thirteenth International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=FrFQpAgnGE)Cited by:[§1](https://arxiv.org/html/2609.36205#S1.p1.1),[§2](https://arxiv.org/html/2609.36205#S2.SS0.SSS0.Px1.p1.1),[§8](https://arxiv.org/html/2609.36205#S8.p2.1)\.
## Appendix AStudy 1: Data and Additional Results
This appendix gives the corpus inventory, probe results, pairwise comparisons in Gemma, the comparison within Nigeria, and the smaller Aya comparison, in that order\.
### A\.1Parallel Corpus Inventory
Each of the twelve target language snapshots contains the same 1,628 English source sentences and entity identities, paired with the corresponding translations\. There are 37 template identifiers and 44 distinct entities, with one record per template entity combination\. Danish is counted once; the two directory spellings in the archive contain duplicate results\.
Table 5:Counts read from the saved parallel corpus snapshots\. Each region supplies 814 records per target language\. These counts describe the corpus, not translation accuracy\.For example, an Africa labelled city record has entityBamako, template identifier 2, and English sentence “Bamako is situated in a region with a distinctive physical environment\.” Its stored Hausa translation is “Bamako tana cikin wani yanki mai yanayi na musamman\.” This reproduces a corpus record without claiming independent validation of its translation\. Applying common templates across entity classes can also produce semantically awkward English inputs; for example, the same template is applied to the food entityCouscous\. These inputs limit the analysis to the chosen templates and entities\.
### A\.2Probe Results on Unseen Entities
Figures[4](https://arxiv.org/html/2609.36205#A1.F4)and[5](https://arxiv.org/html/2609.36205#A1.F5)show the English to target and target to target probe accuracies on unseen entities\. Both use the twelve sampled Gemma layers\.
Figure 4:Gemma probe accuracy on unseen entities when training on English and testing on each target language across all twelve sampled layers\. Chance accuracy is 25%\. Table[2](https://arxiv.org/html/2609.36205#S5.T2)reproduces three columns of this heatmap alongside numerical summaries of the other measures\.Figure 5:Gemma target to target probe results\. The classifier is trained and tested within the same language\. Values are rounded to the nearest percentage point\.
### A\.3Pairwise Results by Layer
Tables[6](https://arxiv.org/html/2609.36205#A1.T6)to[9](https://arxiv.org/html/2609.36205#A1.T9)reproduce the mean cosines and Holm adjustedppvalues\. Values retain the original tabulated precision\. Correction was applied separately across twelve layers within each contrast\.
Table 6:All African within group versus African to control cosines\.Table 7:Niger Congo within group versus Niger Congo to control cosines\.Table 8:Niger Congo to Afro Asiatic versus Niger Congo to control cosines\.Table 9:Afro Asiatic within group versus Afro Asiatic to control cosines\.
### A\.4Within Nigeria Comparison
Figure[6](https://arxiv.org/html/2609.36205#A1.F6)compares the three Nigerian language pairs\. Yoruba and Igbo belong to Niger Congo; Hausa belongs to Afro Asiatic\.
Figure 6:Pairwise cosines for Yoruba and Igbo and for their pairs with Hausa\. Bands show the 2\.5th to 97\.5th percentiles across ten direction subsamples, each drawn from 80% of the records without replacement\.
### A\.5Aya Pairwise Comparison
The saved Aya analysis includes Swahili, Yoruba, and Zulu \(Niger Congo\), Hausa and Amharic \(Afro Asiatic\), and Finnish as the sole control\. Its parallel inputs use the same English template/entity inventory as Gemma\. Across eleven sampled layers, the original all African comparison meets the Holm threshold at L32 and L36\. The comparisons within Niger Congo, within Afro Asiatic, and across families meet it at no layer\. Igbo is absent, so the within Nigeria triple comparison is not available\. Pairwise dependence applies here as in the Gemma analysis\. Differences in language and control coverage prevent treating this as a matched replication or attributing all differences to the model\.
Table 10:Pairwise comparison of the five African languages with Finnish in Aya\. Values are rounded to four decimal places\.
## Appendix BStudy 2: Data and Additional Results
This appendix gives construction examples and scoring details, followed by attribution across countries, extraction method comparisons, uncertainty intervals, and coherence agreement\.
### B\.1Dataset Construction and Scoring Details
The Study 2 construction corpus contains 800 pairs of country and Western sentences in total; the Nigeria file contains 200 pairs\. One Nigerian sentence is: “Yoruba men project power and affluence at an Owambe by wearing the Agbada, a voluminous, heavily embroidered three\-piece robe that requires constant shoulder readjustments\.” Its Western anchor is: “Western men project formal elegance at black\-tie events by wearing a tuxedo, a sharply tailored, closely fitted suit featuring satin lapels and a bow tie\.” These are construction examples, not generated responses from the steering evaluation\.
An evaluation prompt is: “A very popular, spicy roasted meat skewer sold by street vendors at night in this country is called”\. The rubric lists country specific strict tokens and additional lenient tokens that may be shared across countries\. For this prompt, the Nigerian strict list includessuya,tsire, andkilishi;suyais also a lenient Ghanaian token in the rubric\. Strict token assignments are evaluation conventions and should not be interpreted as claims of exclusive cultural ownership\.
The judge receives the prompt, the per prompt rubric, the rubric’s verdict, and the response\. It returns structured country attribution, coherence, and rubric agreement labels at temperature zero\. Rubric only scoring retains responses without the model judge coherence filter; the other modes exclude outputs the judge marks as incoherent and missing attribution verdicts\. The structural diagnostic is separate\.
Table 11:Saved union scoring counts at each country’s selected strength\. Country denominators reflect judge filtering and the South African exclusion; random denominators reflect judge filtering only\.
### B\.2Coherence Agreement
Table[12](https://arxiv.org/html/2609.36205#A2.T12)compares the structural screen with the model judge on the scored generations\.
Table 12:Structural screen and model judge coherence labels over 9,200 scoredchat\_meangenerations, with 94\.1% overall agreement\. These two automated assessments do not constitute a human coherence evaluation\.
### B\.3Steering Specificity
Figure[7](https://arxiv.org/html/2609.36205#A2.F7)shows attribution to all four countries for each steering direction\. These curves include the other country attributions omitted from the diagonal summary in Figure[3](https://arxiv.org/html/2609.36205#S7.F3)\.
Figure 7:Attribution rates for each country direction withchat\_meanextraction and union scoring\. Solid curves show the target country and dashed curves show the other countries at positive strengths\.
### B\.4Uncertainty Across Prompts
Table[13](https://arxiv.org/html/2609.36205#A2.T13)reports uncertainty for attribution rates at the selected strengths\. It distinguishes pooling responses from weighting prompts equally\.
Table 13:Uncertainty intervals forchat\_meanunion scoring\. Wilson intervals concern the response pooled rate\. The bootstrap resamples per prompt means 10,000 times and weights prompts equally; its point estimates are 89\.125%, 78\.974%, 71\.0%, and 66\.5% in row order\. These can differ from pooled rates after unequal judge filtering\. Neither interval concerns the effect relative to random or the Study 1 pairwise comparisons\.
### B\.5Comparison of Extraction Methods
Table 14:Best strength union hit rates \(%\), rounded as displayed in the extraction method heatmap\. Each method uses its own selected strength\. The two methods using chat formatting have similar rates\. Raw text extraction gives lower rates\.For Nigeria, Ghana, Kenya, and South Africa, the selected strengths are respectively \(50, 40, 50, 50\) for raw mean, \(40, 50, 50, 40\) for chat mean, and \(40, 60, 60, 60\) for chat last\. Each method is evaluated at its own selected strength\. The largest difference between the two chat methods is about 3\.0 percentage points, for Nigeria\. South Africa is tied at the recorded precision\.Similar Articles
From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages
This paper evaluates the Mamba state space model for ASR on seven South African languages, finding it matches Conformer accuracy with fewer resources, and explores multilingual training strategies and low-resource settings.
AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content
AfriSyCo studies answer switching in AI models around African-language factual content, showing how assertive framing and verification prompts significantly affect accuracy across languages and model checkpoints.
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
This paper investigates methods to steer Arabic LLMs toward dialect-specific generation by identifying sparse neuron populations and extracting dialect activation directions, enabling dialect control at inference time without fine-tuning.
The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs
This paper systematically quantifies the tokenization penalty for 20 African languages across 11 frontier and open tokenizers, finding up to 8.9× inference cost and latency multipliers and as little as 11% effective context window compared to English, highlighting a structural digital divide encoded in subword vocabularies.
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.