Register Bias in Complexity-Based Large Language Model Routing
Summary
This paper shows that complexity-based routing in large language models biases against non-standard English registers like African American English, due to length-based signals, leading to lower capacity allocation and compounded bias.
View Cached Full Text
Cached at: 09/17/26, 08:46 AM
# Register Bias in Complexity-Based Large Language Model Routing
Source: [https://arxiv.org/html/2609.17542](https://arxiv.org/html/2609.17542)
###### Abstract
Large language model services increasingly route each query to one of several models of differing capability, using a cheap estimate of query complexity to send easy queries to small models and hard queries to large ones\. I show that this routing step is not register neutral: text written in a non\-standard English register, African American English or the English of second\-language writers, is systematically assigned a lower\-capacity tier than a meaning\-equivalent standard\-English version of the same query\. The effect is driven by a specific, common routing signal, input length, because non\-standard registers omit function words and thus look shorter and therefore simpler; other complexity signals do not carry it\. I demonstrate the disparity on 37,704 authentic learner sentence pairs and on a controlled parallel corpus\. I then measure the quality consequence on a device, edge, and cloud model ladder and find that the harm is driven by pervasive model bias, every tier, including a frontier cloud model, answers non\-standard\-register queries significantly less accurately, while the marginal quality cost of the routing decision itself is not significant on this benchmark\. Complexity\-based routing thus compounds the exposure of the users that the models already serve worst\.
## IIntroduction
Cost\-aware serving of large language models \(LLMs\) routes each query to a model of appropriate strength, sending easy queries to small models and hard queries to large ones based on an estimate of query difficulty\[[1](https://arxiv.org/html/2609.17542#bib.bib1),[2](https://arxiv.org/html/2609.17542#bib.bib2),[3](https://arxiv.org/html/2609.17542#bib.bib3)\]\. The routing signal is typically cheap and surface\-level\. I ask whether that routing decision is fair across the linguistic register in which a query is written, and I study the router as an object of audit independent of any single deployed system\.
I find that a complexity\-based router treats meaning\-equivalent queries differently by register\. A question written in African American English or in the English of a second\-language writer is assigned a lower\-capacity tier than a standard\-English version of the same question\. This paper makes the following contributions\.
1. 1\.The first audit, to my knowledge, showing that a complexity\-based LLM router assigns meaning\-equivalent queries to different capability tiers depending on linguistic register, evidenced on authentic human text \(37,704 learner pairs\) and a controlled parallel corpus\.
2. 2\.A decomposition of the disparity by complexity signal: token length carries it robustly, while readability and syntactic\-depth signals do not, and can even reverse\.
3. 3\.A mechanistic account: non\-standard registers omit function words, shortening the text, so length\-based routing reads them as simpler\.
4. 4\.A quality decomposition on a device, edge, and cloud ladder separating model bias from routing\-induced harm, with the finding that even a frontier cloud model is significantly register\-biased, while the routing decision’s own marginal quality cost is not significant on this benchmark, reported honestly as a null\.
Figure 1:The routing pipeline and the register\-downshift mechanism\. A length\-based complexity router assigns each query to a device, edge, or cloud tier\. A non\-standard\-register phrasing of the same question omits function words and is therefore shorter, so the router assigns it a weaker tier: the disparity this paper audits\.
## IIRelated Work
LLM routing and cascades\.Cost\-aware routing and model cascades reduce serving cost by sending each query to a model of appropriate strength\[[1](https://arxiv.org/html/2609.17542#bib.bib1),[2](https://arxiv.org/html/2609.17542#bib.bib2)\]\. Closest to this work, a carbon\-aware and fairness\-aware router\[[3](https://arxiv.org/html/2609.17542#bib.bib3)\]argues that naive energy\-saving routing can leave users in certain regions or languages with lower\-quality service, and adds a distributional\-fairness constraint\. That work operates over geography and distinct languages and proposes a mitigation; it does not audit register variation within English, and it does not decompose routing\-induced disparity from the underlying models’ own bias, which are the questions I take up\.
Dialect bias in LLMs\.A single model’s accuracy and behavior can vary sharply by dialect\. Covert dialect prejudice has been demonstrated in modern LLMs\[[4](https://arxiv.org/html/2609.17542#bib.bib4)\], and dialect stress\-testing frameworks quantify accuracy gaps across English varieties\[[5](https://arxiv.org/html/2609.17542#bib.bib5),[6](https://arxiv.org/html/2609.17542#bib.bib6)\]\. I take single\-model dialect bias as established and do not claim it; my contribution concerns the*routing*layer that sits above the models\.
Complexity measures\.I use standard, published readability and complexity signals \(token length, Flesch\-Kincaid grade\[[10](https://arxiv.org/html/2609.17542#bib.bib10)\], Gunning fog\[[11](https://arxiv.org/html/2609.17542#bib.bib11)\], and dependency\-parse depth\) as representative of the difficulty proxies real routers use, so the finding concerns the class of complexity\-based routers rather than one bespoke formula\.
## IIIMethod
### III\-AFraming: complexity\-based routing
I treat the router as a function that maps a query’s measured complexity to one of three capability tiers \(Fig\.[1](https://arxiv.org/html/2609.17542#S1.F1)\), and I ask whether that function assigns meaning\- equivalent queries to different tiers depending on register\. I audit four standard complexity signals: token*length*; maximum dependency\-parse depth \(*syn\_depth*\); Flesch\-Kincaid grade \(*fk\_grade*\); and Gunning fog \(*fog*\)\. For each signal I fit a tercile router: I set the 33rd and 66th percentiles on a reference set of standard\-English texts, then assign any text to tier 0, 1, or 2 by which band its measure falls in\. Higher measured complexity routes to a higher tier\.
### III\-BRegister variation: authentic and controlled
Authentic \(primary\)\.The W&I\+LOCNESS corpus\[[7](https://arxiv.org/html/2609.17542#bib.bib7)\]contains real English\- learner writing in which each source sentence is a learner’s original and an annotation reduces it to a corrected, standard\-English version\. Reconstructing the correction yields an authentic learner\-original, standard\-corrected parallel pair for the same content\. I use the beginner, intermediate, and advanced learner strata \(37,704 pairs\), with the native LOCNESS stratum as a control\.
Controlled \(secondary\)\.To hold the question exactly constant across registers, I take a 300\-question sample of Natural Questions\[[8](https://arxiv.org/html/2609.17542#bib.bib8),[9](https://arxiv.org/html/2609.17542#bib.bib9)\]and transform each standard\-English question into African American English and Indian English variants using Multi\-VALUE\[[6](https://arxiv.org/html/2609.17542#bib.bib6)\], a rule\-based transformer whose African American English rules are human\-validated\. Questions are normalized to well\-formed form before transformation; transforms that error out are dropped \(1 of 300\)\.
### III\-CSemantic\-equivalence gate
A register variant is usable only if it asks the same question as its standard\-English original\. I gate every controlled variant with an out\-of\-pipeline judge \(Claude Opus\) instructed to compare meaning only and to ignore grammar, dialect, and spelling, keeping a question only if every variant passes\. This retained 279 of 299 questions \(6\.7% dropped\)\. The judge is deliberately not one of the routed models\.
### III\-DRouting\-disparity metric
The router’s decision is deterministic given the text, so the primary result needs no model inference\. For each parallel pair I compute the tier assigned to the standard\-English member and to each register variant under each complexity measure, and count how often the non\-standard version is routed down, up, or the same\. I test the asymmetry with a paired sign test \(McNemar exact for small counts; a continuity\-corrected normal approximation for large counts\)\.
### III\-EQuality tiers and grading
For the quality analysis I use a realistic device, edge, and cloud capability ladder: Llama 3\.2 1B \(on\-device\)\[[13](https://arxiv.org/html/2609.17542#bib.bib13)\], Llama 3\.1 8B \(edge\)\[[12](https://arxiv.org/html/2609.17542#bib.bib12)\], and Claude Opus \(cloud frontier\)\[[14](https://arxiv.org/html/2609.17542#bib.bib14)\]\. The two open models run locally with Ollama\[[15](https://arxiv.org/html/2609.17542#bib.bib15)\]\. I grade each answer by fact\-containment against the reference answers with light morphological normalization, which is deterministic and register\-invariant, so the quality metric itself carries no dialect bias, and the grader is not one of the routed models\.
## IVOffline Results: The Routing Disparity
### IV\-AAuthentic learner text
On 37,704 authentic learner/corrected pairs, the length\-based router routes the learner’s original to a weaker tier than its standard correction far more often than the reverse: 1,272 downshifts versus 474 upshifts \(paired sign test,p<10−3p<10^\{\-3\}\)\. The disparity is significant and one\-directional under the length signal\. The other three signals do not show it: under syntactic depth, Flesch\-Kincaid grade, and Gunning fog the counts run the other way \(for example, Flesch\-Kincaid 1,147 down versus 2,194 up\), because learner errors and missing punctuation inflate readability and depth scores\. On the 988 native LOCNESS pairs the length disparity is much weaker \(20 down versus 7 up,p=0\.019p=0\.019\) and the readability signals behave erratically, consistent with those signals being noisy on short single sentences\.
### IV\-BControlled parallel corpus
On the 279 gated Natural\-Questions parallel sets, terciles fit on the standard\-English questions, the length\-based router again downshifts both non\-standard registers significantly: African American English 68 down versus 42 up \(p=0\.017p=0\.017\), Indian English 101 down versus 34 up \(p<0\.001p<0\.001\)\. The other signals are inconsistent: syntactic depth and Gunning fog are non\- significant, and Flesch\-Kincaid reverses for Indian English \(11 down versus 51 up,p<0\.001p<0\.001\)\.
### IV\-CMechanism
Every downshift under the length signal has the same cause\. Non\-standard registers omit function words, African American English auxiliary deletion \(“when did X air” becomes “when X air”\) and second\-language writers dropping articles and prepositions, which shortens the text\. A length\-based router reads shorter as simpler and assigns a weaker tier\. Two independent data sources converge on this mechanism\. The robust claim is therefore narrow and honest: token\- length\-based routing, the most common routing signal, systematically misroutes non\-standard\- register queries to weaker tiers, and other complexity signals do not show a consistent effect\.
## VQuality Cost of the Misrouting
I answer each of the 279 gated register\-parallel questions with the device, edge, and cloud ladder and grade by fact\-containment, decomposing the quality effect into model bias and routing\-induced harm\.
### V\-APath B: model bias \(tier fixed\)
Holding the tier fixed, I compare accuracy on the standard\-English question with accuracy on each register variant \(McNemar paired test\), Table[I](https://arxiv.org/html/2609.17542#S5.T1)and Fig\.[2](https://arxiv.org/html/2609.17542#S5.F2)\. The models are register\-biased, and the effect strengthens with capability rather than vanishing: the frontier cloud model, despite answering standard\-English questions best, is significantly less accurate on both African American English and Indian English\. A state\-of\-the\-art model is not register\-neutral\. The 1B is too weak to show the effect \(floor\)\.
TABLE I:Path B: per\-tier accuracy by register \(McNemarppvs\. standard English\)\.Figure 2:Per\-tier answer accuracy by register on the device, edge, and cloud ladder\. Register bias persists and strengthens up the ladder: even the frontier cloud model answers African American English and Indian English significantly less accurately than Standard English\. Stars: McNemarppvs\. Standard English \(∗p<\.05\*\\,p<\.05,∗∗p<\.01\*\*\\,p<\.01,∗∗∗p<\.001\*\*\*\\,p<\.001\)\.
### V\-BPath A: routing\-induced harm
Under the length\-based router, each register variant is assigned a tier by its own length\. I compare the accuracy it realizes there against the accuracy it would have received at the tier its standard\-English version routes to, on the questions the router downshifts\. The effect is not significant for either register \(African American English realized 0\.430 vs\. counterfactual 0\.423,p=0\.80p=0\.80; Indian English realized 0\.341 vs\. counterfactual 0\.348,p=0\.87p=0\.87\), even with a wide capability gap between tiers\. The apparent reason is the benchmark: Natural\-Questions items are short and nearly uniform in length, so the router’s tier boundaries are tight and the downshifted questions are not systematically the ones where the tier gap decides the answer\. I report the null\.
### V\-CReading the two paths together
The quality cost borne by non\-standard\-register users is driven by pervasive model bias \(Path B\), which is significant and, strikingly, present even in the frontier cloud model\. Complexity\-based routing compounds the exposure of these users by systematically downshifting their queries \(Section V\), sending exactly the users the models serve worst toward the weaker tiers\. The marginal quality cost of the routing decision itself is not significant on this benchmark \(Path A\), and I do not claim otherwise\.
## VIOn\-Device Deployment Realism
In mobile deployments the weakest tier runs on the user’s own device, so the users whose register is downshifted are the ones served by the on\-device model\. I verify this is a real deployment path by cross\-compiling llama\.cpp\[[16](https://arxiv.org/html/2609.17542#bib.bib16)\]for Android and running Llama 3\.2 1B \(the same weights as the server\-side tier\) on a physical Samsung Galaxy S25\+ over the Android Debug Bridge, an independent path using no prior system’s application\. On a 40\-question sample the 1B runs on\-device at a mean 29\.1 tokens per second, with on\-device accuracy 0\.175 essentially matching the server\-side 0\.200 on the same questions and correctness agreement on 37 of 40 questions \(0\.925\)\. The weakest tier is therefore genuinely deployable on the phone and its answers match the server\-side model, so the quality disparity measured server\-side transfers to real hardware\. This is a deployment\-realism result, not a new quality measurement, since the on\-device and server tier\-0 models are identical\.
## VIILimitations
Rates are workload\-specific; the robust claims are the direction and the mechanism, not the exact percentages\. The controlled register variants are rule\-transformed, which is disclosed; authenticity rests on the W&I\+LOCNESS arm\. Readability and syntactic\-depth measures are noisy on short single sentences, and the length result is the reliable one\. The routing\-induced quality harm \(Path A\) is null on a short\-question benchmark; a workload with more length variation is needed to test whether that null is benchmark\-specific\.
## VIIIConclusion
Complexity\-based LLM routing, judged by the most common signal it uses, systematically sends non\-standard\-register queries to weaker tiers\. The users so downshifted are also the ones every tier, including a frontier cloud model, answers significantly less accurately\. The routing layer thus compounds the exposure of the users the models already serve worst, a fairness concern that sits above any individual model and that a router can be designed to avoid\.
## References
- \[1\]L\. Chen, M\. Zaharia, and J\. Zou, “FrugalGPT: How to use large language models while reducing cost and improving performance,” arXiv:2305\.05176, 2023\.
- \[2\]I\. Ong*et al\.*, “RouteLLM: Learning to route LLMs with preference data,” in*Proc\. Int\. Conf\. Learning Representations \(ICLR\)*, 2025 \(arXiv:2406\.18665\)\.
- \[3\]T\. Li, Z\. Zhao, Z\. Li, X\. Yue, and J\. Yu, “Fair and carbon\-aware LLM routing for web services,” in*Proc\. ACM Web Conf\. \(WWW\)*, 2026, doi:10\.1145/3774904\.3793001\.
- \[4\]V\. Hofmann, P\. R\. Kalluri, D\. Jurafsky, and S\. King, “AI generates covertly racist decisions about people based on their dialect,”*Nature*, vol\. 633, pp\. 147–154, 2024\.
- \[5\]C\. Ziems*et al\.*, “VALUE: Understanding dialect disparity in NLU,” in*Proc\. ACL*, 2022\.
- \[6\]C\. Ziems*et al\.*, “Multi\-VALUE: A framework for cross\-dialectal English NLP,” in*Proc\. ACL*, 2023\.
- \[7\]C\. Bryant, M\. Felice, Ø\. E\. Andersen, and T\. Briscoe, “The BEA\-2019 shared task on grammatical error correction,” in*Proc\. BEA Workshop*, 2019\.
- \[8\]T\. Kwiatkowski*et al\.*, “Natural Questions: A benchmark for question answering research,”*Trans\. Assoc\. Comput\. Linguistics*, vol\. 7, pp\. 453–466, 2019\.
- \[9\]K\. Lee, M\.\-W\. Chang, and K\. Toutanova, “Latent retrieval for weakly supervised open domain question answering,” in*Proc\. ACL*, 2019\.
- \[10\]J\. P\. Kincaid, R\. P\. Fishburne, R\. L\. Rogers, and B\. S\. Chissom, “Derivation of new readability formulas for Navy enlisted personnel,” Naval Technical Training Command, Research Branch Report 8\-75, 1975\.
- \[11\]R\. Gunning,*The Technique of Clear Writing*\. New York: McGraw\-Hill, 1952\.
- \[12\]A\. Grattafiori*et al\.*, “The Llama 3 herd of models,” arXiv:2407\.21783, 2024\.
- \[13\]Meta AI, “Llama 3\.2: Revolutionizing edge AI and vision with open, customizable models,” Meta blog / model card, 2024\.
- \[14\]Anthropic, “Claude \(Opus\) model card,” 2024\.
- \[15\]Ollama, “Ollama,” software,[https://github\.com/ollama/ollama](https://github.com/ollama/ollama)\.
- \[16\]G\. Gerganov*et al\.*, “llama\.cpp,” software,[https://github\.com/ggerganov/llama\.cpp](https://github.com/ggerganov/llama.cpp)\.Similar Articles
From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale
This paper examines covert dialect bias in large language models by analyzing their internal probability distributions across four English dialects, finding that models associate more negative housing-related adjectives with African American Vernacular English and Nigerian Pidgin, reflecting inherited biases similar to human discrimination.
Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring
This paper investigates monocultural biases in large language models that lead to unequal systemic exclusion in hiring, finding that post-trained models exacerbate age-based discrimination and increase exclusion rates from 5.6% to 17.3%.
Online Learning with LLM Experts from Limited Feedback
This paper formulates the adaptive routing of prompts to large language model experts as a contextual bandit problem with limited feedback, proposing algorithms that achieve sublinear regret and demonstrate efficient learning of high-quality routing strategies.
Side-by-side Comparison Amplifies Dialect Bias in Language Models
This research paper finds that language models exhibit increased dialect bias when comparing Standard American English and African-American Vernacular English side-by-side, even after safety fine-tuning. Counterfactual fairness fine-tuning can reduce some biases in isolation but not consistently in contrastive settings.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.