Concept Direction Reliability Across Languages with Different Tokenizer Fertility

arXiv cs.CL Papers

Summary

This paper evaluates how reliably sentiment concept directions extracted from language model representations reproduce across splits, comparing English, Hausa, and Yoruba over four language models. It finds consistent language rank order in direction agreement (English > Hausa > Yoruba) and shows that high probe classification accuracy does not imply directional consistency, though it does not establish tokenizer fertility as the cause.

arXiv:2609.36194v1 Announce Type: new Abstract: Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducibility, we measure split-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics. Using the final token, split-half agreement ranges from 0.737 to 0.870 for English, 0.589 to 0.762 for Hausa, and 0.101 to 0.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross-lingual differences.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:50 AM

# Concept Direction Reliability Across Languageswith Different Tokenizer Fertility
Source: [https://arxiv.org/html/2609.36194](https://arxiv.org/html/2609.36194)
Abass OguntadeElisha KomolafeBabangida SaniFatima Muhammad AdamMuhammad Sammani SaniAffiliation:African Institute for Mathematical SciencesAffiliation:Bayero University KanoAffiliation:Federal University DutseAffiliation:University of Vienna

###### Abstract

Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate\. To evaluate direction reproducibility, we measure split\-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts\. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics\. Using the final token, split\-half agreement ranges from 0\.737 to 0\.870 for English, 0\.589 to 0\.762 for Hausa, and 0\.101 to 0\.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons\. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency\. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length\. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross\-lingual differences\.

## 1Introduction

Concept directions are widely used to identify and manipulate properties such as sentiment within language model representations\[[14](https://arxiv.org/html/2609.36194#bib.bib14),[15](https://arxiv.org/html/2609.36194#bib.bib15)\]\. A standard extraction approach computes the difference in means between representations of contrastive text pairs\. However, before interpreting or steering with an extracted direction, it is essential to determine whether an independent sample would yield a similar vector\.

We evaluate direction reproducibility for sentiment across English, Hausa, and Yoruba\. Each contrast pair consists of a pleased message and an annoyed message regarding a shared topic\. Our primary evaluation metric is split\-half agreement: we repeatedly partition topics into two non\-overlapping subsets, estimate directions from each, and measure their consistency across samples\. Keeping each topic within one half prevents agreement from being inflated by sharing that topic across halves\.

To contextualize our findings, we compare direction agreement against linear probe accuracy\[[2](https://arxiv.org/html/2609.36194#bib.bib3),[4](https://arxiv.org/html/2609.36194#bib.bib5)\]\. While probe accuracy assesses whether sentiment information is linearly recoverable, direction agreement measures whether a specific extraction method reliably isolates the same vector\.

Finally, we analyze how direction agreement interacts with factors such as representation method, text source, sentence length, topic selection, and tokenizer fertility\[[10](https://arxiv.org/html/2609.36194#bib.bib11),[1](https://arxiv.org/html/2609.36194#bib.bib2)\]\. We measure the association between fertility and agreement across languages without claiming a direct causal link\.

## 2Study design

### Data and models\.

Each language dataset comprises 100 sentiment contrast pairs across 25 topics \(4 pairs per topic\)\. While English contains source pairs only, Hausa and Yoruba are evaluated across four distinct text sources \(arms\): native texts written directly by local authors \(Arm A\), human translations of English pairs \(Arm B\), machine translations of English pairs \(Arm C\), and back\-translated native texts \(Arm D\)\. Native writers were prompted in their own language without viewing the English sentences\. We apply Unicode NFC normalization across all texts and evaluate each arm independently, as translation can alter the structural features used by a model\[[3](https://arxiv.org/html/2609.36194#bib.bib4)\]\.

We evaluate representations across four open\-weight models: Gemma 4 \(E2B, E4B, and 12B variants\)\[[6](https://arxiv.org/html/2609.36194#bib.bib8)\]and AfroLlama V1\[[8](https://arxiv.org/html/2609.36194#bib.bib10)\], where E2B and E4B denote effective active parameter counts instead of total stored weights\. At every layer, we extract representations using both the final token position and mean pooling across all non\-padding tokens\. All inference uses unquantized float16 precision with sequence lengths capped at 128 tokens; complete model checkpoints and extraction settings are provided in Appendix[A](https://arxiv.org/html/2609.36194#A1)\.

### Split\-half agreement and layer selection\.

Lethi\+,hi−∈ℝph\_\{i\}^\{\+\},h\_\{i\}^\{\-\}\\in\\mathbb\{R\}^\{p\}be the representations of the positive and negative messages in pairii, and letΔi=hi\+−hi−\\Delta\_\{i\}=h\_\{i\}^\{\+\}\-h\_\{i\}^\{\-\}\. For a set of pairsSS, we normalise the average difference to obtain

d^​\(S\)=\|S\|−1​∑i∈SΔi‖\|S\|−1​∑i∈SΔi‖2\.\\widehat\{d\}\(S\)=\\frac\{\|S\|^\{\-1\}\\sum\_\{i\\in S\}\\Delta\_\{i\}\}\{\\left\\\|\|S\|^\{\-1\}\\sum\_\{i\\in S\}\\Delta\_\{i\}\\right\\\|\_\{2\}\}\.\(1\)We measure agreement between two directions using cosine similarity, where11indicates identical directions,00indicates orthogonality, and negative values indicate opposing components\. We use ten topics for layer selection and fifteen topics for evaluation\. Within each topic set, we repeatedly partition the topics into two non\-overlapping halves, estimate a direction from each half, and compute their cosine similarity\. We average this similarity across 100 random topic partitions to determine the split\-half agreement score, assessing whether independent samples recover consistent directions\. Keeping all four contrast pairs of a given topic together within the same split prevents shared topics between the two halves\. For evaluation, each split divides the fifteen topics into subsets of seven and eight topics \(28 and 32 pairs, respectively\)\. Given topic partitionsSbS\_\{b\}andS¯b\\bar\{S\}\_\{b\}for splitbb, the overall agreementrris

r=1100​∑b=1100d^​\(Sb\)⊤​d^​\(S¯b\)\.r=\\frac\{1\}\{100\}\\sum\_\{b=1\}^\{100\}\\widehat\{d\}\(S\_\{b\}\)^\{\\top\}\\widehat\{d\}\(\\bar\{S\}\_\{b\}\)\.\(2\)To check for sentence length effects, we average the two representations within each pair and construct a length direction by contrasting pairs above and below the median word count\. We calculate length overlap as the absolute cosine similarity between this length direction and the sentiment direction\. For a given model, we select the layer that maximizes selection agreement while maintaining a length overlap below0\.150\.15\. If no layer satisfies this threshold, the selected layer is deemed ineligible and the result is reported as unavailable\. This step addresses one possible influence of length, though it does not eliminate every source of bias\.

### Probes and uncertainty\.

We train logistic regression probes on the 40 selection pairs and test them on the 60 evaluation pairs\. We report test accuracy at two layers: the layer selected for direction reliability, and a separate layer selected strictly via classification performance on selection data using 5\-fold cross\-validation that keeps each topic together\. Feature vectors are centered and scaled to unit variance before training\. We applyL2L\_\{2\}regularization with a fixed hyperparameterC=1C=1across all models\. To estimate uncertainty while holding layer selection fixed, we construct confidence intervals across topics\. Agreement intervals are computed using a leave\-one\-out jackknife procedure across evaluation topics\. Accuracy intervals are constructed by bootstrapping the evaluation topics 1,000 times while keeping the fitted classifier fixed\. These intervals do not adjust for multiple comparisons\. Appendix[B](https://arxiv.org/html/2609.36194#A2)provides complete details for both estimation procedures\.

## 3Sentiment prediction and direction agreement

Figure 1:Results for natively written texts in each language using final token representations\. The panels show split\-half direction agreement \(left\) and probe classification accuracy at the selected layer \(right\)\. Bars indicate approximate 95% confidence intervals with layer selection held fixed\. The dashed line marks chance accuracy\.Using final token representations, split\-half agreement is highest for English, followed by Hausa and then Yoruba across all four models \(Figure[1](https://arxiv.org/html/2609.36194#S3.F1)\), with values from 0\.737 to 0\.870, 0\.589 to 0\.762, and 0\.101 to 0\.399, respectively\. Tokenizer fertility follows the inverse pattern, averaging approximately 1\.14 for English, 1\.78 to 2\.03 for Hausa, and 2\.57 to 2\.95 for Yoruba\. Lower agreement thus aligns with higher subword fertility in this comparison\. However, because these observations originate from only three languages, evaluating multiple models on the same language does not yield independent cross\-lingual samples\.

Sentiment remains predictable despite these lower agreement scores\. Across all 24 combinations of model, language, and representation method, every separately selected probe achieves an interval lower bound above random chance \(0\.50\)\. This pattern holds at the layer selected for direction reliability across all 23 eligible combinations\. For instance, Gemma 4 E2B on Yoruba yields a final token agreement of 0\.101 alongside a classification accuracy of 0\.625; under mean pooling, agreement reaches 0\.317 while accuracy rises to 0\.692\. A classifier can therefore reliably predict labels even when the underlying extracted direction varies across samples\. Because accuracy and cosine similarity operate on distinct scales, where random chance is 0\.50 and orthogonal directions yield zero, their absolute difference does not quantify predictability relative to reliability\. Instead, the key takeaway is that prediction remains viable at layers with weak direction agreement\.

Representation extraction methods further influence consistency\. Final token agreement exceeds mean pooling agreement for every English and Hausa model under the primary topic division, with paired difference intervals excluding zero across all four English models and two Hausa models\. On the other hand, Yoruba shows mixed directional trends with all three available intervals overlapping zero\. The fourth Yoruba comparison is unavailable because no AfroLlama layer satisfies the 0\.15 length overlap threshold under mean pooling\. These findings demonstrate that mean pooling does not consistently outperform final token extraction for Yoruba, nor do they define a specific fertility threshold where pooling becomes advantageous\.

## 4Sensitivity to selection and estimation choices

Table 1:Layer availability across twenty divisions of the topics at threshold 0\.15\. The language columns count divisions with an eligible layer\. The final column counts English\>\>Hausa\>\>Yoruba among divisions with all three results\. Unavailable comparisons are excluded from this count\.We repeat layer selection and evaluation across twenty specified topic divisions using the same dataset, allocating ten topics for selection and fifteen for evaluation in each split\. Because these splits divide the same underlying observations, including the primary division, they evaluate sensitivity to topic selection instead of replication on independent data\.

The relative language ordering \(English \> Hausa \> Yoruba\) using the final token position remains perfectly consistent across all 77 valid comparisons \(Table[1](https://arxiv.org/html/2609.36194#S4.T1)\), with three comparisons having no eligible Yoruba layer\. Under mean pooling, this ordering holds in 56 of 65 complete comparisons, alongside fifteen unavailable cases\. Language rank ordering is therefore noticeably more consistent when using final token representations\.

While overall rank order remains steady, individual numerical values show substantial variation\. For example, E2B Yoruba final token agreement ranges from−\-0\.117 to 0\.275 across eighteen eligible divisions\. Similarly, for 12B English under mean pooling, the primary division agreement of 0\.473 marks the minimum across all twenty splits, where the median reaches 0\.717\. For AfroLlama on Yoruba under mean pooling, fourteen divisions yield an eligible layer even though none qualifies in the primary division\. These ranges illustrate observed empirical variability instead of statistical confidence bounds, which is why we present the primary division alongside this sensitivity analysis\.

Applying the sentence length constraint directly influences which layer outputs can be reported\. Without this check, AfroLlama on Yoruba under mean pooling selects layer 2, yielding a selection agreement of 0\.943 alongside a length overlap of 0\.997\. Because no layer satisfies the 0\.15 overlap limit in that division, raising the threshold to 0\.20 allows a layer to qualify\. Evaluating thresholds of 0\.10, 0\.15, 0\.20, and 0\.25 reveals that sixteen of the 24 combinations maintain the same eligible layer, whereas eight experience changes in layer choice or availability\. This threshold directly shapes model selection and should be reported alongside the primary findings\.

Finally, we compare direction vectors derived from mean difference representations against the weight vectors of logistic probes\. Each probe is trained on one half of the evaluation topics, with its parameters transformed back into the original representation space for comparison\. Both techniques evaluate the same layer across identical topic divisions\. Mean difference vectors yield higher split\-half agreement across all 23 eligible combinations, with paired difference intervals excluding zero in seventeen cases\. While this finding holds for our specific probe parameters and layer selection protocol, it does not demonstrate that mean differences are universally superior across all extraction settings\. Complete results appear in Appendix[F](https://arxiv.org/html/2609.36194#A6)\.

## 5Comparisons between text sources

We compare each translated arm with native writing at the layer selected from native data, using the same fifteen evaluation topics\. Across fifteen eligible combinations for Hausa and Yoruba, median cosine similarity is 0\.338 for human translation, 0\.398 for machine translation and 0\.756 for translation into English and back\. Human translation has the lowest agreement within an arm in fourteen combinations\. This pattern describes the present dataset and does not establish a ranking of translation quality\.

Texts translated into English and back remain closest to native writing across all fifteen combinations\. Because these pairs originate directly from native sentences, this similarity does not provide independent confirmation of the extracted direction\. We also report an attenuation reference scorehA​Xh\_\{AX\}, representing expected agreement under classical measurement error assumptions:

hA​X=g⁡\(rA\)​g​\(rX\),g⁡\(r\)=2​r1\+r\.h\_\{AX\}=\\sqrt\{g\(r\_\{A\}\)g\(r\_\{X\}\)\},\\qquad g\(r\)=\\frac\{2r\}\{1\+r\}\.\(3\)The functionggapplies the Spearman–Brown correction for sample halves\[[5](https://arxiv.org/html/2609.36194#bib.bib6),[13](https://arxiv.org/html/2609.36194#bib.bib13)\], whilehhuses the standard attenuation relation\[[12](https://arxiv.org/html/2609.36194#bib.bib12)\]\. These underlying assumptions do not necessarily hold for this representations, and we leave the reference undefined when reliability is negative\. All fifteen comparisons with back\-translated texts exceed this reference score, though shared topic content across text sources may introduce dependent errors in human and machine translations as well\. Appendix[G](https://arxiv.org/html/2609.36194#A7)reports complete cosine similarities, attenuation reference values, and confidence intervals\. This reference score is neither a strict upper bound nor a test of statistical equivalence\.

## 6Discussion

Split\-half agreement exposes sample sensitivity even when a probe predicts sentiment accurately\. Successful probe performance does not demonstrate that the direction used to represent sentiment is reproducible\. Distinct topic samples can contain different predictive features, enabling accurate classification even when estimated direction vectors vary\. Similarly, the lower agreement observed among probe coefficient vectors suggests that successful classification can occur across multiple decision boundaries\. These mechanisms represent potential explanatory factors that our experimental design does not evaluate directly\.

High direction agreement alone remains insufficient for representation analysis\. As demonstrated by the sentence length checks, a consistently estimated vector can reflect surface properties of the text instead of sentiment\. Researchers evaluating representation geometry should assess classification performance alongside direction agreement, inspect identifiable influences such as sentence length, and select layers using topic sets distinct from evaluation topics\. This approach adds a reliability check to existing frameworks for probe interpretation\[[7](https://arxiv.org/html/2609.36194#bib.bib9),[4](https://arxiv.org/html/2609.36194#bib.bib5)\]\. Several scope constraints frame these empirical findings\. This study evaluates a single sentiment contrast type across three languages\. Three native authors per language completed distinct topic blocks, meaning that topic partitioning does not separate topic choice from writer\-specific effects\. Furthermore, evaluation relies on a dataset containing fifteen topics per split, yielding approximate confidence intervals\. Subword fertility also covaries with pretraining exposure, language family, and specific structural features; as a result, neither tokenizer comparisons nor model comparisons isolate an explicit causal mechanism\.

## 7Conclusion

Sentiment remains predictable across these datasets even when extracted directions show weak agreement across samples\. Split\-half agreement calculated from final token representations is consistently highest for English, followed by Hausa and Yoruba, across all valid comparisons in the tested topic divisions\. Specific numerical values and layer availability exhibit greater variability, particularly when using mean pooling\. Measuring split\-half agreement provides a direct assessment of direction reproducibility alongside classification performance\. To ensure reproducibility in representation interpretability studies, agreement metrics should be reported with explicit details regarding topic partitioning, layer selection rules, and sentence length constraints\. Determining whether tokenizer characteristics directly cause these cross\-lingual variations will require controlled experimental designs\.

## References

- \[1\]O\. Ahia, S\. Kumar, H\. Gonen, J\. Kasai, D\. R\. Mortensen, N\. A\. Smith, and Y\. Tsvetkov\(2023\)Do all languages cost the same? tokenization in the era of commercial language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 9904–9923\.Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p4.1)\.
- \[2\]G\. Alain and Y\. Bengio\(2016\)Understanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p3.1)\.
- \[3\]M\. Artetxe, G\. Labaka, and E\. Agirre\(2020\)Translation artifacts in cross\-lingual transfer learning\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 7674–7684\.Cited by:[§2](https://arxiv.org/html/2609.36194#S2.SS0.SSS0.Px1.p1.1)\.
- \[4\]Y\. Belinkov\(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p3.1),[§6](https://arxiv.org/html/2609.36194#S6.p2.1)\.
- \[5\]W\. Brown\(1910\)Some experimental results in the correlation of mental abilities\.British Journal of Psychology3\(3\),pp\. 296–322\.Cited by:[§5](https://arxiv.org/html/2609.36194#S5.p2.2)\.
- \[6\]Gemma Team\(2026\)Gemma 4 technical report\.arXiv preprint arXiv:2607\.02770\.Cited by:[§2](https://arxiv.org/html/2609.36194#S2.SS0.SSS0.Px1.p2.1)\.
- \[7\]J\. Hewitt and P\. Liang\(2019\)Designing and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 2733–2743\.Cited by:[§6](https://arxiv.org/html/2609.36194#S6.p2.1)\.
- \[8\]Jacaranda Health\(n\.d\.\)Jacaranda/AfroLlama\_V1\.Note:Hugging Face model repository[https://huggingface\.co/Jacaranda/AfroLlama\_V1](https://huggingface.co/Jacaranda/AfroLlama_V1)\. Accessed 27 September 2026Cited by:[§2](https://arxiv.org/html/2609.36194#S2.SS0.SSS0.Px1.p2.1)\.
- \[9\]NLLB Teamet al\.\(2022\)No language left behind: scaling human\-centered machine translation\.arXiv preprint arXiv:2207\.04672\.Cited by:[Appendix A](https://arxiv.org/html/2609.36194#A1.p1.1)\.
- \[10\]P\. Rust, J\. Pfeiffer, I\. Vulić, S\. Ruder, and I\. Gurevych\(2021\)How good is your tokenizer? on the monolingual performance of multilingual language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 3118–3135\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.243)Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p4.1)\.
- \[11\]C\. R\. Shalizi\(2023\)Jackknife notes\.Note:Carnegie Mellon University lecture notesExternal Links:[Link](https://www.stat.cmu.edu/~cshalizi/490/23/2023-11-29-jackknife.pdf)Cited by:[Appendix B](https://arxiv.org/html/2609.36194#A2.p1.1)\.
- \[12\]C\. Spearman\(1904\)The proof and measurement of association between two things\.American Journal of Psychology15\(1\),pp\. 72–101\.Cited by:[§5](https://arxiv.org/html/2609.36194#S5.p2.2)\.
- \[13\]C\. Spearman\(1910\)Correlation calculated from faulty data\.British Journal of Psychology3\(3\),pp\. 271–295\.Cited by:[§5](https://arxiv.org/html/2609.36194#S5.p2.2)\.
- \[14\]A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid\(2023\)Steering language models with activation engineering\.arXiv preprint arXiv:2308\.10248\.Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p1.1)\.
- \[15\]A\. Zouet al\.\(2023\)Representation engineering: a top\-down approach to AI transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§1](https://arxiv.org/html/2609.36194#S1.p1.1)\.

## Appendix AExtraction and evaluation protocol

We use Arm A for primary direction agreement and probe classification experiments, and Arms B, C, and D for comparisons across text sources\. Text sources are aligned by topic; independently authored pairs do not share identical identifier strings\. Three native authors per language authored distinct topic blocks\. Because source metadata contains incomplete writer tracking tags, we could not evaluate generalization across separate authors\. We usednllb\-200\-3\.3B\[[9](https://arxiv.org/html/2609.36194#bib.bib7)\]to generate Arms C and D\.

Each normalized sequence was passed directly to the model tokenizer without chat formatting or prompt wrappers\. We retained special tokens and applied right\-side padding\. Final\-token extraction uses the final index where the attention mask equals one; mean pooling averages all sequence positions where the attention mask equals one, including special tokens\.

Layer index zero corresponds to the embedding layer output\. Models E2B and E4B are treated as distinct architectures instead of nested or identically trained variants\. Exact model checkpoints, package dependencies, and random seed specifications appear in the reproducibility code repository\.

The maximum sequence length reaches 68 tokens for Gemma architectures and 79 tokens for AfroLlama; no text required truncation\. Subword fertility represents the aggregate token count excluding special tokens divided by total whitespace word count across both sides of all native pairs, instead of an average of individual sentence ratios\.

We set0\.150\.15as the primary length overlap threshold, using0\.100\.10,0\.200\.20, and0\.250\.25for sensitivity checks\. Analysis code is available at[https://github\.com/mohdasaid/provenance\-repe](https://github.com/mohdasaid/provenance-repe)\. The repository contains evaluation scripts and plotting code\.

## Appendix BUncertainty and numerical checks

For a statisticTTcomputed fromm=15m=15evaluation topics, letT\(−j\)T\_\{\(\-j\)\}denote its value after excluding all four pairs associated with topicjj\. Each deletion removes the topic from the existing divisions without drawing new divisions\. The selected layer remains fixed throughout this procedure\. Standard errors are estimated using the jackknife framework described by[Shalizi \[11\]](https://arxiv.org/html/2609.36194#bib.bib1)to measure variation in the statistic across topic samples\. Interval bounds are constructed using a Studentttmultiplier:

SE^\(T\)=\[m−1m∑j=1m\(T\(−j\)−T¯\(−⋅\)\)2\]1/2,T±tm−1,0\.975SE^\(T\)\.\\widehat\{\\mathrm\{SE\}\}\(T\)=\\left\[\\frac\{m\-1\}\{m\}\\sum\_\{j=1\}^\{m\}\\bigl\(T\_\{\(\-j\)\}\-\\overline\{T\}\{\(\-\\cdot\)\}\\bigr\)^\{2\}\\right\]^\{1/2\},\\qquad T\\pm t\{m\-1,0\.975\}\\widehat\{\\mathrm\{SE\}\}\(T\)\.\(4\)Here,T¯\(−⋅\)\\overline\{T\}\_\{\(\-\\cdot\)\}is the mean over the fifteen topic deletions, andtm−1,0\.975t\_\{m\-1,0\.975\}denotes the 97\.5th percentile of the Studentttdistribution with fourteen degrees of freedom\. Cosine intervals are clipped to\[−1,1\]\[\-1,1\], paired reliability differences to\[−2,2\]\[\-2,2\], and intervals for the attenuation reference to\[0,1\]\[0,1\]\. Negative reliability values are not clipped to zero prior to applying the attenuation transformation\. If an evaluation step yields an undefined output following a topic deletion, its corresponding interval is left reported as unavailable\. We construct the interval by removing each topic in turn and recalculating the difference between the attenuation reference and observed cosine similarity\.

These intervals approximate measurement uncertainty in extraction statistics for the observed sample sizes under the specified protocol\. They remain strictly conditional on the selected layer, selection sample, and topic partitioning schedule\. Consequently, they do not estimate uncertainty originating from new authors or repeated layer selection, nor do they demonstrate the validity of the underlying attenuation model\.

Probe classification accuracy evaluates mean correctness across the two messages in each pair and subsequently across evaluation pairs\. Interval estimation uses a non\-parametric bootstrap resampling fifteen topics with replacement across 1,000 iterations, carrying all paired scores for a sampled topic together\. The resulting interval spans the 2\.5th to 97\.5th percentiles of the bootstrapped accuracy distribution\. Classifier parameters and layer selection remain fixed\. In datasets yielding perfect classification, every resample maintains identical performance, producing an interval with identical lower and upper bounds\. Such an interval does not prove perfect general accuracy\.

For the matched estimator comparison, both extraction procedures evaluate the first twenty topic divisions\. For each topic deletion, both probe halves undergo complete retraining, with standardization parameters recomputed strictly on the remaining training sequences\. The confidence interval is calculated from the paired difference between mean\-difference agreement and probe\-coefficient agreement under identical topic deletions\. These intervals are approximate and unadjusted; their empirical coverage has not been independently validated for probe weight vectors\.

Interval behavior for direction agreement and cosine similarity was verified through synthetic simulations across seven scenario configurations using 16, 64, and 1,536 dimensions, encompassing weak and strong signals, unequal coordinate variances, heavy\-tailed Student\-t5t\_\{5\}noise, and correlated noise conditions\. Each scenario evaluated 300 synthetic datasets alongside 10,000 reference iterations\. Across checks of reliability, cosine similarity, and pooling method differences, empirical coverage ranged from 93\.3% to 99\.7%\. This finite simulation result provides empirical context instead of a guaranteed coverage bound for the primary text collection, probe weight comparisons, or attenuation reference scores\.

Numerical calculations were validated independently using a parallel codebase operating on the saved activation tensors\. The accompanying reproducibility repository details these numerical checks and their operational bounds\.

## Appendix CNative results and probe evaluation

Tables[2](https://arxiv.org/html/2609.36194#A3.T2)and[3](https://arxiv.org/html/2609.36194#A3.T3)report reliability and decoding at the same layer\. All intervals in this appendix are approximate conditional 95% intervals\. A dash means the relevant layer or transformed quantity is unavailable, not that it equals zero\.

Table 2:Native final token results at the layer selected for reliability\.Table 3:Native mean pooling results at the layer selected for reliability\. AfroLlama Yoruba has no eligible layer at threshold 0\.15\.Table 4:Accuracy after a separate search for a probe layer within selection data\. These are distinct from the results at the same layer above\. Fertility is computed on native text\.There are 120 evaluation messages per probe\. Since each pair contains one positive and one negative label, chance accuracy is 0\.5 and ordinary accuracy equals balanced accuracy\. These intervals use the fixed classifiers and topic bootstrap described in Appendix[B](https://arxiv.org/html/2609.36194#A2)\.

## Appendix DReadout and threshold sensitivity

Table 5:Paired difference between final token and mean pooling reliability in the primary allocation\. The two readouts may select different layers\. Intervals delete the same topic in both conditions\.Table 6:Selected layer and reliability on evaluation data in parentheses at each threshold for length overlap\. A dash indicates no eligible layer\. Sixteen configurations retain the same eligible layer across all four thresholds\.
## Appendix ERepeated topic allocations

We summarise all twenty topic allocations below\. Minima and maxima describe variation across eligible allocations; they are not confidence limits\. The allocations reuse the same data\.

Table 7:Agreement across twenty allocations at threshold 0\.15\. Median and range use eligible allocations only\. The final column counts distinct selected layers\.Figure 2:Agreement across topic allocations\. Each dot is an eligible allocation and the horizontal mark is its configuration’s median\. Counts show eligibility out of twenty\.
## Appendix FMatched direction estimators

Both estimators use the same twenty topic divisions into halves and the same layer selected from native data\. Logistic coefficients are divided by the training scaler’s feature standard deviations before the cosine is computed\. The intercept is omitted\. Directions from mean differences are not standardised\. Both fitting procedures are therefore evaluated in the original activation coordinates, although the probe’s regularisation still depends on its training standardisation\.

Table 8:Agreement of mean difference and probe coefficient vectors\. The difference is paired by topic deletion\. Seventeen unadjusted approximate intervals exclude zero\. These comparisons use twenty partitions, so their values for mean differences need not equal the main estimates based on 100 divisions\.We selected layers to maximise agreement of mean differences on separate topics\. We did not repeat the comparison with layers selected for probe stability, so the estimator ranking remains conditional on this choice and the fixed regularisation\. Appendix[B](https://arxiv.org/html/2609.36194#A2)describes the interval limitations\.

## Appendix GProvenance comparisons

All comparisons use the layer selected from native data and the same evaluation topic labels\. Agreement within each arm is shown first\. The following pages report observed cosine between native writing and another armcc, attenuation referencehhfrom Equation[3](https://arxiv.org/html/2609.36194#S5.E3), and their paired difference\. The reference assumes a relationship between split half agreement and reliability for the full sample that is not established for these representations\. Shared topics and derived text can also induce dependent errors\.

Table 9:Agreement within each arm for native writing \(A\), human translation \(B\), machine translation \(C\) and translation into English and back \(D\)\. AfroLlama Yoruba mean pooling is unavailable\. Full arm intervals are included in the accompanying CSV tables\.Table 10:Native versus human translation \(A–B\)\. Values are followed by full conditional 95% intervals\. An interval marked \[–\] is undefined because a required reliability after a topic is removed is negative\. The E4B Yoruba mean pooling reference is itself undefined because the point reliability for B is negative\.Table 11:Native versus machine translation \(A–C\)\. The text is translated from English seeds\. Independent errors relative to native writing are not established by this construction because the conditions share topics and evaluation choices\.Table 12:Native versus translation into English and back \(A–D\)\. D is derived from A\. All fifteen observed cosines exceed the attenuation reference, which must not be interpreted as a hard bound or as evidence of independent agreement\.The medians reported in Section[5](https://arxiv.org/html/2609.36194#S5)summarise different models and readouts\. They are not independent estimates of a population translation effect\.

Similar Articles

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

arXiv cs.CL

This paper investigates whether language models remain robust to alternative (non-canonical) tokenizations across 27 languages, finding that invariance observed in English does not generalize and that languages with higher token fragmentation show greater sensitivity. The authors demonstrate that LoRA fine-tuning with multi-tokenization data can mitigate this sensitivity.

Some Large Language Models Exhibit Consistent Risk Attitudes

arXiv cs.AI

This paper introduces a framework to test whether large language models exhibit consistent risk attitudes across domains. It finds that most LLMs show intra-task and cross-domain stability in risk attitude, converging to a narrower distribution than humans.