Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
Summary
The paper shows that unequal single-token support for names in large language models leads to biased concept accessibility across demographics, and introduces NameTrace to measure lexical comparability.
View Cached Full Text
Cached at: 09/29/26, 08:12 PM
Paper page - Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
Source: https://huggingface.co/papers/2609.34065
Abstract
Namesarepersonalidentifiers,buttheyalsocarrysocialmeaningandarewidelyusedtoevaluatehowlanguagemodelstreatdifferentpeople.Suchevaluationstypicallyassumethatmatchednamesarecomparablemodelinputs.Weshowthatthisassumptionoftenfailsatthelexicalinterface:matchednamesarenotnecessarilymatchedinputs.Somenamesreceivedirectsingle-tokenaccess,whileothersareassembledfrommultiplesubwords,creatingunequalname-surfacesupport.Acrossnearlyhalfamillionfirstnamesand12LLM-associatedtokenizers,directlexicalaccessishighlyselective,modeldependent,andunevenacrossrace-andgender-associatednamemetadata.WeintroduceNameTrace,amodel-native,fine-grained,pre-behavioralframeworkformeasuringwhetherunequalname-surfacesupportremainsavocabularypropertyorbecomesvisibleintask-relevantinternalrepresentations.NameTracemeasuresconceptaccessibilityfromthemodel’sownprobabilitiesovertask-specificadjectiveaxeswithcontinuoustask-alignedweights.Onmatchedatomicandshort-fragmentednameswithinthesamerace/ethnicity--gender-associatedstrata,supportpredictssystematicdifferencesinconceptaccessibilityacrossfellowship,hiring,clinicalassessment,andlending.Thesedifferencespersistacrossalleightmatchedstrata,extendacrossmodelfamilies,andtransfertounseennames.Hidden-stateinterventionsfurthershowthatthemeasuredtaskdirectionshavedownstreamleverage,shiftinglaterconstrainedchoices.Unequallexicalsupportisthereforedemographicallystructuredattheinputandremainsvisibleintask-relevantmodelcomputation.NameTracemakeslexicalcomparabilitymeasurable,supportingabroaderprinciple:behavioralcomparabilitybeginswithlexicalcomparability.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.34065
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34065 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34065 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34065 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models
This paper systematically compares equitable tokenizers for multilingual LLMs across 11 Southeast Asian languages, finding that Parity-aware BPE achieves the best efficiency-equity trade-off and that cross-lingual fairness and tokenization efficiency are not fundamentally at odds.
Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
This paper investigates the fairness of crosslingual evaluation methods for language models, showing that common normalized metrics can be biased due to tokenization and orthographic differences, and proposes using sentence-level negative log likelihood on semantically equivalent sequences for more consistent crosslingual comparisons.
Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring
This paper investigates monocultural biases in large language models that lead to unequal systemic exclusion in hiring, finding that post-trained models exacerbate age-based discrimination and increase exclusion rates from 5.6% to 17.3%.
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages
This paper investigates whether language models remain robust to alternative (non-canonical) tokenizations across 27 languages, finding that invariance observed in English does not generalize and that languages with higher token fragmentation show greater sensitivity. The authors demonstrate that LoRA fine-tuning with multi-tokenization data can mitigate this sensitivity.
Byte-level models
Discusses whether byte-level tokenizers outperform subword tokenizers for precise tasks like distinguishing similar names, counting characters, and case sensitivity, and asks for current recommendations.