How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
Summary
This paper investigates how American English becomes the default in large language models, examining structural bias across pretraining data, tokenization, and generation stages through a controlled study with American and British English variants.
View Cached Full Text
Cached at: 09/29/26, 08:16 AM
Paper page - How Does “English (US)” Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
Source: https://huggingface.co/papers/2604.04204
Abstract
Large language models exhibit systematic bias toward American English over British English across pretraining data, tokenization, and generation stages, reflecting broader geopolitical and colonial influences on AI development.
Large language models(LLMs) are increasingly embedded in educational, professional, and public infrastructure, yet widely used platforms expose “English (US)” as a primary English setting despite the global diversity of English. We ask: How does “English (US)” become the default? We study this question asstructural bias, examining how geopolitical histories of data curation, digital dominance, andlinguistic standardizationintersect with the LLM development pipeline. Using British English as a controlled reference, we construct a curated resource of 1,813 matched American English (AmE)--British English (BrE) variants and introduce DiAlign, a dynamic, training-free method for estimating regional alignment from distributional evidence. We triangulate the AmE preference across data exposure --> representation --> generation, jointly examining pretraining and post-training data, tokenizer behavior and provenance, model prediction cost, and generated language across developer countries, prompt conditions, domains and sources, linguistic categories, and registers. AmE is consistently favored across all six auditedpretraining corporaand 21 post-training datasets, is generally represented more compactly by tokenizers, and receives lower prediction cost. It also remains the dominant generation default under neutral English prompting; British-English prompting shifts this preference toward BrE but does not consistently eliminate the AmE default. To our knowledge, this is the first rigorous pipeline-wide study ofstructural biasacross major phases of LLM development. Our findings show that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization,epistemic injustice, and inequity in global AI deployment, while providing a rigorous basis for targeted component-level intervention.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2604\.04204
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.04204 in a model README.md to link it from this page.
Datasets citing this paper1
#### tafseer-nayeem/ame-bre-structural-bias Viewer• Updatedabout 9 hours ago • 3.63k • 29
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.04204 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Toward LLMs Beyond English-Centric Development
This paper demonstrates that LLMs are heavily biased toward English, and shows that continual pre-training does not offer cost advantages over training from scratch for adapting models to other languages, especially for cultural understanding.
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.
Side-by-side Comparison Amplifies Dialect Bias in Language Models
This research paper finds that language models exhibit increased dialect bias when comparing Standard American English and African-American Vernacular English side-by-side, even after safety fine-tuning. Counterfactual fairness fine-tuning can reduce some biases in isolation but not consistently in contrastive settings.
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
This paper audits six large language models for gender stereotyping across English, Korean, Chinese, and Japanese, anchoring against human baselines. It finds that LLM stereotyping often exceeds human cross-country variation and can compound across languages, introducing a four-pattern framework to characterize such behaviors.
Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge
This paper proposes Judge-LS, a protocol to evaluate whether LLM-as-a-judge models are invariant to language switching between English and Chinese. It finds that switching languages causes 10.7-14.4% preference flips and that judges achieve their highest accuracy in English.