LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

arXiv cs.CL Papers

Summary

This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.

arXiv:2607.06845v1 Announce Type: new Abstract: African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and "corrected" by large language models (LLMs). Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English (SAE) continuations even when the preceding context is in AAE, effectively rewriting AAE into SAE. We present an end-to-end framework to audit and mitigate this bias. For auditing, we introduce conditional Dialect Group Invariance (cDGI), which isolates true model bias from translator-induced artifacts, and a feature-level localization analysis that identifies which AAE markers most strongly trigger bias; we find that syntactic constructions, especially negative concord (e.g., "ain't nobody"), are universal triggers across all models. For mitigation, we introduce, to our knowledge, the first application of activation steering to dialect bias: a training-free, test-time method that extracts dialect directions via causal tracing and injects them into bias-relevant layers. Activation steering reduces bias 5 to 20 times more than prompting while preserving SAE fluency. To enable this work, we release REAL-AAE , the largest real-AAE parallel corpus to date: 17,479 AAE/SAE/ AAE_back triplets from natural tweets (2 to 6 times larger than prior real-AAE resources), validated automatically (BERTScore F1 = 0.95) and by three native AAE speakers (83.0% semantic agreement).
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:48 AM

# LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Source: [https://arxiv.org/html/2607.06845](https://arxiv.org/html/2607.06845)
Huan Wu1,2,3Ali Emami4Muhammad Furquan Hassan1Osaretin Igbinoba5 Osakpolor Idusuyi6Osamede Igbinoba7Faiza Khan Khattak8 Laleh Seyyed\-Kalantari1,2,3,9 1York University2Vector Institute3Connected Minds4Emory University 5Wilfrid Laurier University6University of Toronto7University of Guelph 8Monark Health9CIFAR Solution Network Member ∗lsk@yorku\.ca

###### Abstract

African American English \(AAE\), a rule\-governed dialect spoken by over 30 million people, is routinely misinterpreted and "corrected" by large language models \(LLMs\)\. Across six instruction\-tuned LLMs \(14B to 70B\), we show that state\-of\-the\-art models systematically prefer Standard American English \(SAE\) continuations even when the preceding context is in AAE, effectively rewriting AAE into SAE\. We present an end\-to\-end framework to audit and mitigate this bias\. For auditing, we introduce conditional Dialect Group Invariance \(cDGI\), which isolates true model bias from translator\-induced artifacts, and a feature\-level localization analysis that identifies which AAE markers most strongly trigger bias; we find that syntactic constructions, especially negative concord \(e\.g\., "ain’t nobody"\), are universal triggers across all models\. For mitigation, we introduce, to our knowledge, the first application of activation steering to dialect bias: a training\-free, test\-time method that extracts dialect directions via causal tracing and injects them into bias\-relevant layers\. Activation steering reduces bias 5 to 20 times more than prompting while preserving SAE fluency\. To enable this work, we releaseReal\-AAE, the largest real\-AAE parallel corpus to date: 17,479 AAE/SAE/AAEbacktriplets from natural tweets \(2 to 6 times larger than prior real\-AAE resources\), validated automatically \(BERTScore F1 = 0\.95\) and by three native AAE speakers \(83\.0% semantic agreement\)\. Our corpus and code are available at[https://anonymous\.4open\.science/r/dialect](https://anonymous.4open.science/r/dialect)\.

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

Huan Wu1,2,3Ali Emami4Muhammad Furquan Hassan1Osaretin Igbinoba5Osakpolor Idusuyi6Osamede Igbinoba7Faiza Khan Khattak8Laleh Seyyed\-Kalantari1,2,3,9††thanks:Corresponding author\.1York University2Vector Institute3Connected Minds4Emory University5Wilfrid Laurier University6University of Toronto7University of Guelph8Monark Health9CIFAR Solution Network Member∗lsk@yorku\.ca

## 1Introduction

State\-of\-the\-art large language models \(LLMs\) silently “correct” African American English \(AAE\)\. Given the AAE context“But I ain’t doing no dishes,”, five out of six LLMs we tested prefer the Standard American English \(SAE\) continuation“I’m going to stay in my room”over the valid AAE alternative“Imma stay in my room”\(Fig\.[1](https://arxiv.org/html/2607.06845#S1.F1)\)\.

We call this behaviordialect preference bias: the systematic tendency of LLMs to favor SAE over valid varieties such as AAE, even when the user is writing in AAE\. Treating AAE as something to be corrected is not a neutral stylistic choice but a form of linguistic discrimination: AAE is a rule\-governed variety with its own consistent grammar and phonology, spoken by over 30 million people\(Linet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib38)\)\. The downstream consequences are well\-documented: LLMs disproportionately misclassify AAE text\(Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37)\), respond to AAE speakers with more stereotyping and condescension\(Fleisiget al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib34)\), and degrade on tasks with non\-standard dialect inputs\(Ziemset al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib30); Linet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib38)\)\. As LLMs move into hiring, healthcare, and content moderation, this bias actively disadvantages millions of users\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/Figure_1_Final.png)Figure 1:Dialect preference bias on a real AAE tweet\. In a forced\-choice continuation task, five out of six state\-of\-the\-art LLMs prefer the SAE continuation over a fluent AAE alternative under an AAE context, and unanimously prefer SAE under an SAE context\.WorkOriginConstructionScaleDisc\.Gen\.Feat\.Debias\.VALUE\(Ziemset al\.,[2022](https://arxiv.org/html/2607.06845#bib.bib10)\)Syn\.SAE→\\rightarrowAAE \(rules, HV\)2,880\*Task Acc✗✓✗Multi\-VALUE\(Ziemset al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib30)\)Syn\.SAE→\\rightarrowdialects \(rules, NV\)7,983\*Task Acc✗✓Aug\.AAVENUE\(Guptaet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib27)\)Syn\.SAE→\\rightarrowAAE \(LLM, HV\)5,000Task Acc✗✗✗EnDive\(Guptaet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib18)\)Syn\.SAE→\\rightarrowdialects \(LLM, NV\)40k\+Reas\. Acc✗✗✗Fleisiget al\.\([2024](https://arxiv.org/html/2607.06845#bib.bib34)\)RealPrompt\-elicited∼\\sim3,000✗Yes✓✗Mireet al\.\([2025](https://arxiv.org/html/2607.06845#bib.bib36)\)RealPreference pairs1,843 \+ 2,365RM score✗✗✗Hassanet al\.\([2025](https://arxiv.org/html/2607.06845#bib.bib37)\)RealAAE\-first \+ back\-translation∼\\sim5,000DGI✗✗✗\\rowcolorgreen\!10Real\-AAE\(Ours\)RealAAE\-first \+ back\-trans\. \+ HV17,479DGI, cDGIPPL, LP, MC✓Steering

Table 1:Comparison with prior work on dialect bias evaluation and mitigation\.Real\-AAEis the only entry combining real AAE, human validation, both discriminative and generative auditing, feature\-level analysis, and mitigation\. HV: human validation; NV: native\-speaker validation; Aug\.: data augmentation; \* marks evaluation subset rather than full benchmark size\.Existing work on AAE dialect bias has three key limitations \(Table[1](https://arxiv.org/html/2607.06845#S1.T1)\)\. First, most benchmarks generate AAE*synthetically*from SAE via rule\-based or LLM\-driven translation\(Ziemset al\.,[2022](https://arxiv.org/html/2607.06845#bib.bib10),[2023](https://arxiv.org/html/2607.06845#bib.bib30); Guptaet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib27),[2025](https://arxiv.org/html/2607.06845#bib.bib18)\), producing AAE that underestimates real\-world dialect effects\(Linet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib38); Guptaet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib18)\)\. Second, real\-AAE studies are narrow in scope: they evaluate either classification consistency\(Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37)\), generative responses\(Fleisiget al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib34)\), or reward\-model scoring\(Mireet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib36)\), but none combine discriminative and generative auditing, identify which linguistic features trigger bias, or offer mitigation\. Third, the few existing mitigation methods require retraining\(Ziemset al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib30)\), architectural changes\(Sriraget al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib53); Zhouet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib31)\), or auxiliary translation pipelines that erase the dialect altogether\(Klisuraet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib54)\)\.

We address these limitations as follows\. Starting from a corpus of tweets by native AAE speakers\(Blodgett,[2021](https://arxiv.org/html/2607.06845#bib.bib29)\), we constructReal\-AAE, a corpus of 17,479 AAE/SAE/AAEbacktriplets via translation and back\-translation, validated both automatically and by three native AAE speakers\. This AAE\-first construction preserves dialect features that synthetic AAE misses\.

We then audit six instruction\-tuned LLMs \(14B to 70B\) along two axes:discriminative consistency, via our new conditional Dialect Group Invariance \(cDGI\) metric, which isolates model bias from translator artifacts; andgenerative preference, via perplexity, log\-probability, and forced\-choice continuation\. A feature\-level localization analysis further pinpoints which AAE markers trigger bias\.

Finally, we introduce, to our knowledge, the first application of activation steering to dialect bias: a training\-free method that uses causal tracing to identify bias\-relevant layers and injects a learned dialect direction at inference\. Our analysis yields three main findings: \(i\) all six models prefer SAE continuations even under AAE context, defaulting to “correcting” the dialect; \(ii\) syntactic constructions are universal bias triggers every tested models; and \(iii\) Steering reduces bias 5–20×\\timesmore than prompting, while preserving SAE fluency\.

##### Contributions\.

- •We releaseReal\-AAE, the largest real\-AAE parallel corpus to date: 17,479 AAE/SAE/AAEbacktriplets from natural tweets, 2 to 6 times larger than prior real\-AAE resources, validated both automatically and by three native AAE speakers\.
- •We introduce an audit framework combining cDGI, which controls for translator\-induced drift, with generative\-preference metrics, and a feature\-level localization analysis that identifies the AAE markers most strongly triggering bias\.
- •We propose the first activation\-steering method for dialect debiasing: training\-free, test\-time, and applicable across model scales\. Across six SOTA LLMs \(14B to 70B\), it mitigates dialect bias 5 to 20 times more effectively than prompting while preserving SAE fluency\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/dataset_pipeline_new_final.png)Figure 2:Real\-AAEconstruction pipeline\. Starting from authentic AAE tweets, we filter the data, translate AAE into SAE, remove invalid outputs, and back\-translate the SAE sentences into AAE to build the final triplet dataset\.

## 2Related Work

Dialect Bias in NLP: Dialect bias has been documented across natural language processing tasks, including language identification\(Blodgett,[2021](https://arxiv.org/html/2607.06845#bib.bib29)\), hate\-speech and toxicity detection\(Sapet al\.,[2019](https://arxiv.org/html/2607.06845#bib.bib46); Davidsonet al\.,[2019](https://arxiv.org/html/2607.06845#bib.bib5); Sapet al\.,[2022](https://arxiv.org/html/2607.06845#bib.bib7)\), automatic speech recognition\(Koeneckeet al\.,[2020](https://arxiv.org/html/2607.06845#bib.bib47)\), text generation\(Deaset al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib49)\), and downstream decision\-making about employability, criminality, and personality\(Hofmannet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib39)\)\. Most notably,Hofmannet al\.\([2024](https://arxiv.org/html/2607.06845#bib.bib39)\)show that dialect prejudice persists in instruction\-tuned LLMs even after overt racial biases are reduced, motivating dialect\-specific audits of frontier models\.

AAE Bias Benchmarks: Resources for evaluating AAE bias divide along a methodological axis\.Syntheticbenchmarks transform SAE into AAE via rule\-based or LLM\-driven translation\(Ziemset al\.,[2022](https://arxiv.org/html/2607.06845#bib.bib10),[2023](https://arxiv.org/html/2607.06845#bib.bib30); Guptaet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib27),[2025](https://arxiv.org/html/2607.06845#bib.bib18)\), an approach shown to underestimate real\-world dialect effects\(Linet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib38); Guptaet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib18)\)\.Real\-AAEresources instead draw text directly from AAE speakers\(Blodgett,[2021](https://arxiv.org/html/2607.06845#bib.bib29); Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37); Fleisiget al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib34); Mireet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib36)\), but each targets a single evaluation mode: classification consistency\(Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37)\), generative responses\(Fleisiget al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib34)\), or reward\-model scoring\(Mireet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib36)\)\. The closest prior work,Hassanet al\.\([2025](https://arxiv.org/html/2607.06845#bib.bib37)\), introduces Dialect Group Invariance \(DGI\) on a real\-AAE corpus for sentiment classification\. We extend this along multiple dimensions:Real\-AAE, a corpus 2–6×\\timeslarger; additional discriminative and generative metrics; feature\-level localization of bias triggers; and a test\-time mitigation \(Table[1](https://arxiv.org/html/2607.06845#S1.T1)\)\.

Dialect Bias Mitigation: Prior mitigation approaches fall into three categories\.Training\-timemethods augment training data with dialectal examples\(Ziemset al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib30)\), but can degrade SAE accuracy\.Architecturalmethods add dialect\-specific modules, such as the LoRDD adapter\(Sriraget al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib53)\)or the encoder\-based DialectGen\(Zhouet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib31)\)\.Inference\-timemethods translate AAE inputs to SAE before passing them to the LLM\(Klisuraet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib54)\), which avoids retraining but introduces auxiliary models and erases the dialect rather than mitigating bias against it\. Our approach instead intervenes directly on the target model’s activations at inference, requiring no retraining, architectural changes, auxiliary modules, or dialect erasure\.

Activation Steering: Our method builds on activation steering, a mechanistic\-interpretability technique that controls model behavior by adding learned directions to hidden states at inference\. The method has been applied to general behavioral control\(Turneret al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib3); Zouet al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib1); Rimskyet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib21)\), refusal behavior\(Arditiet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib24)\), and persona traits\(Chenet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib26)\)\.Liuet al\.\([2024](https://arxiv.org/html/2607.06845#bib.bib4)\)extend it to social bias mitigation, but not to dialect bias\. To our knowledge, ours is the first application of activation steering to dialect debiasing\.

## 3Benchmark Construction

To audit AAE dialect bias without conflating it with translator artifacts, we constructReal\-AAE, anAAE\-firstparallel corpus of \(AAE, SAE,AAEback\\text\{AAE\}\_\{\\text\{back\}\}\) triplets grounded in naturally\-occurring AAE tweets \(Fig\.[2](https://arxiv.org/html/2607.06845#S1.F2)\)\. Unlike synthetic benchmarks that derive AAE from SAE, our pipeline starts from real AAE and translates outward, preserving dialectal features that synthetic generation may miss\.

##### Source data and filtering:

We build on the TwitterAAE corpus\(Blodgettet al\.,[2018](https://arxiv.org/html/2607.06845#bib.bib9)\), which identifies tweets from African American accounts via demographic inference from geolocation and network analysis\. To ensure sufficient content for sentiment analysis and continuation modeling, we retain tweets with at least 10 words and strip non\-ASCII characters, user mentions, and URLs, yielding 17,495 tweets\.

##### Translation and back\-translation:

We translate each AAE tweet into SAE using Gemini\-3\-flash, then back\-translate the SAE into AAE using the same model \(denotedAAEback\\text\{AAE\}\_\{\\text\{back\}\}\)\. The back\-translation serves two purposes: \(i\) it provides a content\-preservation check, since AAE andAAEback\\text\{AAE\}\_\{\\text\{back\}\}should remain semantically aligned if translation did not introduce drift; and \(ii\) it enables our cDGI metric \(§[4\.1](https://arxiv.org/html/2607.06845#S4.SS1)\), which conditions on translation\-stable examples to isolate model bias from translator artifacts\. We manually inspect sample translations and define filtering criteria to remove pairs with added explanations, semantic distortions, or translation artifacts, yielding our final17,479 triplets\.

##### Automatic validation:

For content preservation, we compute BERTScore F1 between AAE andAAEback\\text\{AAE\}\_\{\\text\{back\}\}using Twitter\-RoBERTa,111cardiffnlp/twitter\-roberta\-base\-2022\-154m\.chosen to match the social\-media domain of our inputs\. Both mean and median F1 are0\.95\\mathbf\{0\.95\}\(standard deviation0\.020\.02\), indicating strong semantic alignment\. For sentiment preservation, we compare AAE/SAE label agreement using the same Twitter\-RoBERTa model; agreement is highest for positive and negative cases, with most disagreements concentrated in neutral examples \(full results in Appendix[A](https://arxiv.org/html/2607.06845#A1)\)\.

##### Human validation:

We further validate a 1,000\-example subset with three native AAE speakers from the West Coast and Northeast/Mid\-Atlantic, all of whom self\-reported strong familiarity with written AAE\. Annotators judged each example along four dimensions: sentiment preservation, semantic equivalence, original AAE naturalness, andAAEback\\text\{AAE\}\_\{\\text\{back\}\}naturalness\. Three\-way exact agreement was93\.0%\(sentiment preservation\),83\.0%\(semantic equivalence;κ=0\.468\\kappa=0\.468\),81\.0%\(original AAE naturalness;κ=0\.488\\kappa=0\.488\), and77\.5%\(AAEback\\text\{AAE\}\_\{\\text\{back\}\}naturalness;κ=0\.204\\kappa=0\.204\)\. The sentimentκ\\kappawas−0\.024\-0\.024, reflecting label imbalance rather than disagreement: 98% of judgments wereyes, inflating expected\-chance agreement\.

## 4Methods

We develop methods to audit, localize, and mitigate dialect preference bias\. §[4\.1](https://arxiv.org/html/2607.06845#S4.SS1)defines our evaluation metrics, combining a new conditional invariance measure \(cDGI\) with model\-likelihood\-based generative\-preference metrics\. §[4\.2](https://arxiv.org/html/2607.06845#S4.SS2)introduces a feature\-level localization analysis that pinpoints which AAE markers most strongly trigger bias\. §[4\.3](https://arxiv.org/html/2607.06845#S4.SS3)presents our mitigation method, dialect activation steering\.

### 4\.1Bias Evaluation Metrics

We evaluate dialect bias along two complementary axes:discriminative consistency, via the existing DGI metric\(Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37)\)and our new cDGI metric; andgenerative preference, via perplexity, a log\-probability bias score, and forced\-choice continuation computed from model likelihoods\.

#### 4\.1\.1Conditional Dialect Group Invariance \(cDGI\)

DGI\(Hassanet al\.,[2025](https://arxiv.org/html/2607.06845#bib.bib37)\)measures AAE–SAE classification agreement, but it cannot separate the model’s dialect bias from translator artifacts: an apparent inconsistency may reflect semantic drift in the LLM\-produced SAE translation or inAAEback\\text\{AAE\}\_\{\\text\{back\}\}\. We therefore introduceconditional DGI \(cDGI\\mathrm\{cDGI\}\), restricted to examples satisfying two conditions: \(i\) the predicted label is preserved under back\-translation, controlling for translator semantic drift; and \(ii\)AAEback\\text\{AAE\}\_\{\\text\{back\}\}was judged natural by our native\-AAE annotators, preventingAAEback\\text\{AAE\}\_\{\\text\{back\}\}drift toward SAE from inflating the conditioning set\. LettingyAAE,ySAE,yAAEbacky^\{\\mathrm\{AAE\}\},y^\{\\mathrm\{SAE\}\},y^\{\\mathrm\{AAE\_\{back\}\}\}denote model predictions on the AAE, SAE, and back\-translated AAE inputs respectively,cDGI\\mathrm\{cDGI\}is defined as

cDGI=Pr⁡\(yAAE=ySAE\|yAAE=yAAEback\),\\mathrm\{cDGI\}=\\Pr\\\!\\left\(y^\{\\mathrm\{AAE\}\}=y^\{\\mathrm\{SAE\}\}\\;\\middle\|\\;y^\{\\mathrm\{AAE\}\}=y^\{\\mathrm\{AAE\_\{back\}\}\}\\right\),\(1\)computed over examples meeting condition \(ii\)\.cDGI\\mathrm\{cDGI\}thus isolates the model’s dialect sensitivity from both translator noise and dialect erasure in the conditioning set\.

#### 4\.1\.2Perplexity Analysis

To measure generative preference, we compute conditional perplexity \(PPL\\mathrm\{PPL\}\) for all four context→\\rightarrowcontinuation combinations: SAE→\\rightarrowSAE, SAE→\\rightarrowAAE, AAE→\\rightarrowSAE, and AAE→\\rightarrowAAE\. LowerPPL\\mathrm\{PPL\}indicates the model finds a continuation more natural given the context\. A dialect\-invariant model should prefer dialect\-matched continuations: lowerPPL\\mathrm\{PPL\}for SAE→\\rightarrowSAE than SAE→\\rightarrowAAE, and lower for AAE→\\rightarrowAAE than AAE→\\rightarrowSAE\.

Our continuation\-based experiments \(§[4\.1\.2](https://arxiv.org/html/2607.06845#S4.SS1.SSS2), §[4\.1\.3](https://arxiv.org/html/2607.06845#S4.SS1.SSS3), and §[4\.2\.1](https://arxiv.org/html/2607.06845#S4.SS2.SSS1)\) all require splitting each AAE–SAE pair into a context and a continuation; we describe the alignment\-based segmentation procedure in Appendix[B\.1](https://arxiv.org/html/2607.06845#A2.SS1)\.

#### 4\.1\.3Multiple\-Choice Evaluation

We assess explicit preference by presenting each model with a context \(in either SAE or AAE\) and two candidate continuations \(one SAE, one AAE\), asking which better continues the context\. Option order is randomized to control for positional bias\. A dialect\-invariant model should prefer the continuation matching the context dialect\. The prompt template is in Appendix[B\.2](https://arxiv.org/html/2607.06845#A2.SS2)\.

### 4\.2Feature\-Level Bias Localization

#### 4\.2\.1Context\-specific log\-probability bias

For each of theNNAAE–SAE pairs, letxiSAE,xiAAEx\_\{i\}^\{\\mathrm\{SAE\}\},x\_\{i\}^\{\\mathrm\{AAE\}\}andyiSAE,yiAAEy\_\{i\}^\{\\mathrm\{SAE\}\},y\_\{i\}^\{\\mathrm\{AAE\}\}denote the matched SAE/AAE contexts and continuations, and letP​\(y∣x\)P\(y\\mid x\)denote the model’s likelihood of continuationyygiven contextxx\. We define context\-specific SAE and AAE log\-probability \(LP\\mathrm\{LP\}\) bias scores as

biSAE\\displaystyle b^\{\\mathrm\{SAE\}\}\_\{i\}=log⁡P​\(yiSAE∣xiSAE\)\\displaystyle=\\log P\(y\_\{i\}^\{\\mathrm\{SAE\}\}\\mid x\_\{i\}^\{\\mathrm\{SAE\}\}\)−log⁡P​\(yiAAE∣xiSAE\),\\displaystyle\\quad\-\\log P\(y\_\{i\}^\{\\mathrm\{AAE\}\}\\mid x\_\{i\}^\{\\mathrm\{SAE\}\}\),\(2\)biAAE\\displaystyle b^\{\\mathrm\{AAE\}\}\_\{i\}=log⁡P​\(yiAAE∣xiAAE\)\\displaystyle=\\log P\(y\_\{i\}^\{\\mathrm\{AAE\}\}\\mid x\_\{i\}^\{\\mathrm\{AAE\}\}\)−log⁡P​\(yiSAE∣xiAAE\)\.\\displaystyle\\quad\-\\log P\(y\_\{i\}^\{\\mathrm\{SAE\}\}\\mid x\_\{i\}^\{\\mathrm\{AAE\}\}\)\.\(3\)Positive values indicate preference for dialect\-matched continuations; negative values indicate preference for mismatched ones\. The aggregateLP\\mathrm\{LP\}bias is

LP=1N​∑i=1N\(biSAE−biAAE\),\\mathrm\{LP\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\big\(b^\{\\mathrm\{SAE\}\}\_\{i\}\-b^\{\\mathrm\{AAE\}\}\_\{i\}\\big\),\(4\)where larger positive values indicate overall preference for SAE across both contexts, and values near zero indicate dialect\-matched behavior\.

#### 4\.2\.2Localizing bias by feature category

To identify which linguistic markers most strongly trigger dialect preference, we partition AAE sentences by the presence of each feature category and compare meanLP\\mathrm\{LP\}\(Eq\.[4](https://arxiv.org/html/2607.06845#S4.E4)\) across categories\. Using an AAE grammatical lexicon adapted fromHarriset al\.\([2022](https://arxiv.org/html/2607.06845#bib.bib17)\), we group features into four classes:Auxiliary\(non\-standard auxiliary forms, e\.g\., “we was”, “finna”\),Aspectual\(habitual and completive markers, e\.g\., habitual “be” as in “he be working”\),Preverbal\(pre\-verbal lexical items, e\.g\., “she done finished”\), andSyntactic\(clause\-level constructions, especially negative concord as in “ain’t nobody”\)\. The full lexicon is in Appendix[B\.3](https://arxiv.org/html/2607.06845#A2.SS3)\.

### 4\.3Debiasing with Activation Steering

Our mitigation method rests on a simple insight: if a model treats AAE and SAE inputs differently, then their internal representations must differ in some direction of activation space\. Adding a vector along that direction at inference time should shift the model’s behaviour, without retraining\. We operationalize this idea in two phases \(pictorial overview in Appendix[B\.4](https://arxiv.org/html/2607.06845#A2.SS4), Fig\.[9](https://arxiv.org/html/2607.06845#A2.F9)\)\.Phase 1\(one\-time, offline\) extracts a per\-layer*dialect direction*from paired AAE–SAE inputs and uses causal tracing to identify which layers carry the dialect\-preference computation\.Phase 2\(test\-time\) injects the dialect direction into those layers, weighted by their causal importance, with no parameter updates\. Reproducibility details are in Appendix[B\.4](https://arxiv.org/html/2607.06845#A2.SS4)\.

##### Phase 1a: Extracting the dialect direction\.

GivenNNpaired inputs\{\(xiAAE,xiSAE\)\}i=1N\\\{\(x\_\{i\}^\{\\mathrm\{AAE\}\},x\_\{i\}^\{\\mathrm\{SAE\}\}\)\\\}\_\{i=1\}^\{N\}, lethl​\(x\)h\_\{l\}\(x\)denote the layer\-llhidden state at the last context token\. We compute the per\-example contrast

Δi,l=hl​\(xiAAE\)−hl​\(xiSAE\),\\Delta\_\{i,l\}=h\_\{l\}\(x\_\{i\}^\{\\mathrm\{AAE\}\}\)\-h\_\{l\}\(x\_\{i\}^\{\\mathrm\{SAE\}\}\),\(5\)average across examples, and normalize to obtain a unit\-norm*dialect direction*at each layerll:

Δ¯l=1N​∑i=1NΔi,l,v^l=Δ¯l∥Δ¯l∥2\.\\bar\{\\Delta\}\_\{l\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\Delta\_\{i,l\},\\qquad\\hat\{v\}\_\{l\}=\\frac\{\\bar\{\\Delta\}\_\{l\}\}\{\\lVert\\bar\{\\Delta\}\_\{l\}\\rVert\_\{2\}\}\.\(6\)By construction,v^l\\hat\{v\}\_\{l\}points from SAE\-conditioned activations toward AAE\-conditioned activations at layerll\.

##### Phase 1b: Identifying causally important layers\.

Not every layer contributes equally to dialect preference\. To isolate the layers where the bias is computed, we use a corruption\-and\-restore causal tracing analysis\(Menget al\.,[2022](https://arxiv.org/html/2607.06845#bib.bib2)\)\. For each example, we corrupt the input embeddings at AAE feature\-token positions with Gaussian noise:

epcorr=ep\+ϵp,ϵp∼𝒩​\(0,σ2​I\),e\_\{p\}^\{\\mathrm\{corr\}\}=e\_\{p\}\+\\epsilon\_\{p\},\\quad\\epsilon\_\{p\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\),\(7\)whereσ=α⋅std​\(Efeat\)\\sigma=\\alpha\\cdot\\mathrm\{std\}\(E\_\{\\mathrm\{feat\}\}\),EfeatE\_\{\\mathrm\{feat\}\}is the set of feature\-token embeddings in the example, andα=3\.0\\alpha=3\.0is fixed across all models to produce a clear perturbation signal without overly destabilizing the representation\. We then run the model on the corrupted input and, at a single sitez=\(l,t\)z=\(l,t\), restore the clean layer\-llhidden state at token positiontt\. Lettinggcorrg^\{\\mathrm\{corr\}\}denote the fully corrupted run andgzrestg\_\{z\}^\{\\mathrm\{rest\}\}the run with restoration at sitezz, we measure the effect of restoration on the AAE\-context preference score \(Eq\.[3](https://arxiv.org/html/2607.06845#S4.E3)\):

Ri​\(z\)=biAAE​\(gcorr\)−biAAE​\(gzrest\)\.R\_\{i\}\(z\)=b\_\{i\}^\{\\mathrm\{AAE\}\}\(g^\{\\mathrm\{corr\}\}\)\-b\_\{i\}^\{\\mathrm\{AAE\}\}\(g\_\{z\}^\{\\mathrm\{rest\}\}\)\.\(8\)BecausebiAAE\>0b\_\{i\}^\{\\mathrm\{AAE\}\}\>0corresponds to the model preferring the AAE\-matched continuation,Ri​\(z\)<0R\_\{i\}\(z\)<0means restoration at sitezzshifts the model*toward*dialect\-matched behavior, whileRi​\(z\)\>0R\_\{i\}\(z\)\>0means restoration amplifies SAE preference\. We average over examples and the relevant token positionsTTto obtain a per\-layer score:

Rl=1N⋅\|T\|​∑i=1N∑t∈TRi​\(\(l,t\)\),R\_\{l\}=\\frac\{1\}\{N\\cdot\\lvert T\\rvert\}\\sum\_\{i=1\}^\{N\}\\sum\_\{t\\in T\}R\_\{i\}\\big\(\(l,t\)\\big\),\(9\)and select the four layers with the most negativeRlR\_\{l\}values as the steering setKK, i\.e\., the layers whose restoration most strongly shifts the model toward AAE\-matched behavior\.

##### Phase 2: Inference\-time steering\.

At inference, we attach forward hooks at each layerl∈Kl\\in Kand modify the hidden state at every token positionttas

hl,t′=hl,t\+β​wl​v^l,l∈K,β≥0,h\_\{l,t\}^\{\\prime\}=h\_\{l,t\}\+\\beta\\,w\_\{l\}\\,\\hat\{v\}\_\{l\},\\qquad l\\in K,\\;\\beta\\geq 0,\(10\)whereβ\\betacontrols the overall steering strength andwlw\_\{l\}is a normalized per\-layer weight proportional to causal importance:

wl=−Rl∑j∈K\(−Rj\),l∈K\.w\_\{l\}=\\frac\{\-R\_\{l\}\}\{\\sum\_\{j\\in K\}\(\-R\_\{j\}\)\},\\quad l\\in K\.\(11\)Sincev^l\\hat\{v\}\_\{l\}points from SAE\-conditioned activations toward AAE\-conditioned activations, addingβ​wl​v^l\\beta w\_\{l\}\\hat\{v\}\_\{l\}applies stronger steering at layers with larger beneficial restoration effects\. Settingβ=0\\beta=0recovers the unmodified model; largerβ\\betayields stronger steering\. The method modifies activations only on the forward pass and does not update any parameters\.

## 5Experimental Setup

##### Models\.

We evaluate six instruction\-tuned LLMs spanning families and scales \(14B to 70B\): Gemma\-3\-27B\(DeepMind,[2025](https://arxiv.org/html/2607.06845#bib.bib55)\), Mistral\-Small\-3\.1\-24B\(Mistral AI,[2025](https://arxiv.org/html/2607.06845#bib.bib57)\), Phi\-3\-Medium\-14B\(Abdin and others,[2024a](https://arxiv.org/html/2607.06845#bib.bib58)\), Phi\-4\-14B\(Abdin and others,[2024b](https://arxiv.org/html/2607.06845#bib.bib59)\), DeepSeek\-R1\-Distill\-Qwen\-32B\(DeepSeek\-AI,[2025](https://arxiv.org/html/2607.06845#bib.bib60)\), and Llama\-3\.1\-70B\(Dubeyet al\.,[2024](https://arxiv.org/html/2607.06845#bib.bib11)\)\.

##### Dataset\.

All evaluation usesReal\-AAE\(§[3](https://arxiv.org/html/2607.06845#S3)\)\. For sentiment classification, models predict positive, negative, or neutral\. For continuation\-based experiments, we apply the SAE\-guided alignment\-based segmentation described in Appendix[B\.1](https://arxiv.org/html/2607.06845#A2.SS1)to construct context–continuation pairs\.

##### Activation steering splits\.

For steering experiments, we use three non\-overlapping splits ofReal\-AAE:1,200AAE\-feature sentences for causal\-tracing layer selection \(such examples provide clearer corruption\-and\-restore targets\),7,000AAE–SAE pairs for dialect\-direction extraction, and a held\-out3,000\-pair evaluation set covering all four context→\\rightarrowcontinuation combinations \(AAE→\\rightarrowAAE, AAE→\\rightarrowSAE, SAE→\\rightarrowAAE, SAE→\\rightarrowSAE\)\. We sweep the steering strengthβ∈\{0\.1,0\.2,0\.4,0\.6,0\.8,1\.0\}\\beta\\in\\\{0\.1,0\.2,0\.4,0\.6,0\.8,1\.0\\\}on the evaluation set to study the trade\-off between bias reduction and language modeling performance\.

##### Baselines\.

We compare activation steering against two baselines: the unsteered base model, and a prompting baseline that explicitly instructs the model to treat AAE and SAE as semantically equivalent \(full prompt in Appendix[C\.1](https://arxiv.org/html/2607.06845#A3.SS1)\)\.

##### Implementation\.

Experiments use PyTorch with vLLM\(Kwonet al\.,[2023](https://arxiv.org/html/2607.06845#bib.bib61)\)for inference and Hugging Face Transformers\(Wolfet al\.,[2020](https://arxiv.org/html/2607.06845#bib.bib62)\)for steering \(which requires direct access to residual activations\)\. All runs on NVIDIA H100 80GB GPUs\.

## 6Results

### 6\.1Discriminative Consistency Bias

We measure discriminative bias via sentiment classification onReal\-AAE\. Table[2](https://arxiv.org/html/2607.06845#S6.T2)reports DGI \(AAE–SAE label agreement\), two\-way DGI \(agreement across AAE, SAE, andAAEback\\text\{AAE\}\_\{\\text\{back\}\}\), and cDGI \(AAE–SAE agreement restricted to translation\-stable examples\) for each model\.

A perfectly dialect\-invariant model would scoreDGI=1\.0\\mathrm\{DGI\}=1\.0;no model reaches this\. Gemma\-3 comes closest \(DGI=0\.849\\mathrm\{DGI\}=0\.849\), followed by Phi\-4 \(0\.8140\.814\)\. At the other end, DeepSeek\-R1 \(0\.4910\.491\) and Llama\-3\.1\-70B \(0\.4680\.468\) disagree with themselves on more than half of AAE–SAE pairs\.

Comparing cDGI to DGI isolates the model’s dialect bias from translator artifacts\. Even on translation\-stable examples, cDGI remains below1\.01\.0for every model; for Llama\-3\.1\-70B and DeepSeek\-R1, more than a third of translation\-stable AAE–SAE pairs still receive different sentiment labels\. Translation drift therefore accounts for only part of the observed gap: dialect alone changes model predictions, independent of translation artifacts\.

ModelDGITwo\-way DGIcDGIGemma\-30\.849±0\.0050\.849\\pm 0\.0050\.844±0\.0050\.844\\pm 0\.0050\.844±0\.0450\.844\\pm 0\.045Phi\-40\.814±0\.0060\.814\\pm 0\.0060\.788±0\.0060\.788\\pm 0\.0060\.826±0\.0460\.826\\pm 0\.046Mistral\-Small\-3\.1\-24B0\.795±0\.0060\.795\\pm 0\.0060\.749±0\.0060\.749\\pm 0\.0060\.818±0\.0430\.818\\pm 0\.043Phi\-3\-Medium0\.780±0\.0060\.780\\pm 0\.0060\.761±0\.0060\.761\\pm 0\.0060\.787±0\.0460\.787\\pm 0\.046DeepSeek\-R10\.491±0\.0080\.491\\pm 0\.0080\.302±0\.0070\.302\\pm 0\.0070\.598±0\.0500\.598\\pm 0\.050Llama\-3\.1\-70B0\.468±0\.0080\.468\\pm 0\.0080\.278±0\.0070\.278\\pm 0\.0070\.531±0\.0440\.531\\pm 0\.044Table 2:Discriminative consistency on sentiment classification \(mean±\\pm95% CI; higher is better,1\.01\.0= perfect dialect invariance\)\.DGI: AAE–SAE prediction agreement\.Two\-way DGI: AAE–SAE–AAEback\\text\{AAE\}\_\{\\text\{back\}\}agreement\.cDGI: AAE–SAE agreement restricted to translation\-stable examples\.
### 6\.2Generative Preference Bias

We next ask whether dialect preference appears at the level of language modeling itself, independent of any classification task\. We measure this through model likelihoods over candidate AAE and SAE continuations, using two complementary metrics here: perplexity and forced\-choice continuation\. A log\-probability bias score follows in §[6\.3](https://arxiv.org/html/2607.06845#S6.SS3)\.

MetricPhi\-4Phi\-3Mistral\-Small\-3\.1\-24BDeepSeek\-R1Gemma\-3Llama\-3\.1\-70BBasePromptOursBasePromptOursBasePromptOursBasePromptOursBasePromptOursBasePromptOursDGIDGI0\.740\.76\\cellcolorgreen\!150\.850\.730\.74\\cellcolorgreen\!150\.770\.790\.83\\cellcolorgreen\!150\.860\.550\.57\\cellcolorgreen\!150\.630\.800\.80\\cellcolorgreen\!150\.860\.510\.55\\cellcolorgreen\!150\.63cDGI0\.770\.79\\cellcolorgreen\!150\.860\.740\.75\\cellcolorgreen\!150\.760\.810\.84\\cellcolorgreen\!150\.850\.590\.60\\cellcolorgreen\!150\.620\.820\.83\\cellcolorgreen\!150\.850\.570\.59\\cellcolorgreen\!150\.64PPLAAE PPL123\.9–\\cellcolorgreen\!15111\.485\.8–\\cellcolorgreen\!1580\.582\.8–\\cellcolorgreen\!1575\.9142\.8–\\cellcolorgreen\!15138\.414298\.3–\\cellcolorgreen\!1512845\.867\.7–\\cellcolorgreen\!1562\.2SAE PPL26\.7–28\.720\.6–25\.023\.3–24\.128\.1–27\.91354\.4–1350\.715\.5–15\.5MCAAE Match39\.341\.5\\cellcolorgreen\!1558\.832\.638\.6\\cellcolorgreen\!1555\.329\.533\.9\\cellcolorgreen\!1548\.243\.848\.7\\cellcolorgreen\!1564\.442\.944\.4\\cellcolorgreen\!1560\.841\.344\.5\\cellcolorgreen\!1564\.8SAE Match89\.190\.188\.379\.782\.877\.388\.689\.588\.675\.571\.371\.182\.983\.783\.288\.670\.886\.0LPLP60\.658\.6\\cellcolorgreen\!1547\.370\.869\.0\\cellcolorgreen\!1561\.557\.356\.9\\cellcolorgreen\!1550\.268\.968\.2\\cellcolorgreen\!1562\.089\.388\.4\\cellcolorgreen\!1583\.034\.935\.0\\cellcolorgreen\!1529\.8Table 3:Bias mitigation across models: base model \(Base\), prompt baseline \(Prompt\), and our activation steering \(Ours\) ; bestOursper metric in green\. Higher is better for DGI, cDGI, MC; lower for PPL, LP\. Prompting PPL omitted \(not comparable under likelihood\-based metrics\)\. AAE/SAE Match: % selecting the matched continuation under AAE/SAE context\. 95% CIs in Appendix[C\.6](https://arxiv.org/html/2607.06845#A3.SS6); fullβ\\betasweep in Appendix[C\.4](https://arxiv.org/html/2607.06845#A3.SS4); and a qualitative preference\-flip example in Appendix[C\.5](https://arxiv.org/html/2607.06845#A3.SS5)\.![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/perplexity_scatter.png)Figure 3:Continuation perplexity for SAE \(xx\-axis\) and AAE \(yy\-axis\) continuations, by model and context\. Points abovey=xy=xindicate SAE preference\. All model–context combinations lie above the line, showing systematic SAE preference even under AAE context\.![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/MC_Stacked_Bar_Chart.png)Figure 4:Multiple\-choice continuation preferences under SAE and AAE contexts\. Bars show the % selecting each continuation; dashed line = 50% chance\.##### Perplexity Analysis\.

Fig\.[3](https://arxiv.org/html/2607.06845#S6.F3)reports continuation perplexity for all four context→\\rightarrowcontinuation combinations\. Under SAE context, all six models assign lower perplexity to SAE continuations than to AAE: the expected pattern when continuation matches context\.Under AAE context, however, every model still assigns lower perplexity to SAE continuations\.Models adapt to SAE context but fail to adapt to AAE context, defaulting to SAE regardless of input dialect\.

##### Multiple\-Choice Evaluation\.

Fig\.[4](https://arxiv.org/html/2607.06845#S6.F4)confirms this asymmetry under explicit forced choice\. A dialect\-invariant model should prefer AAE continuations under AAE context and SAE continuations under SAE context\. Instead, most models continue to prefer SAE even when the context is AAE: Phi\-4 selects SAE59\.9%59\.9\\%of the time under AAE context \(vs\.40\.1%40\.1\\%for AAE\)\. Two models deviate\. Gemma\-3 reverses the pattern, preferring AAE under AAE context \(63\.4%63\.4\\%\), while DeepSeek\-R1 stays near50%50\\%in both contexts, suggesting weak context\-sensitivity rather than active dialect\-matching\. Table[4](https://arxiv.org/html/2607.06845#S6.T4)illustrates the typical failure mode: given an AAE context, Mistral\-Small\-3\.1 assigns higher probability to the SAE continuation, effectively “correcting” the dialect\.

AAE ctx\.“Let’s see what my tl have to …”SAE cont\.“… offer \- nonsense, nonsense, and more nonsense\.”\(Model Preferred\)AAE cont\.“… offer bull ish bull ish n more bull ish\.”AAE ctx\.“My Fone Kno Me So Well …”SAE cont\.“… when I’m on hold for less than one minute, it will hang up\.”\(Model Preferred\)AAE cont\.“… When Im On Hold For less Then 1 Minute it Would Hang Upp\.”Table 4:Mistral\-Small\-3\.1\-24B prefers the SAEcontinuation\(cont\.\) under AAEcontext\(ctx\.\), effectively “correcting” the dialect\.

### 6\.3Feature\-Level Bias Localization

Having established that dialect preference bias is pervasive, we next ask which AAE features most strongly trigger it\. Fig\.[5](https://arxiv.org/html/2607.06845#S6.F5)shows the averageLP\\mathrm\{LP\}bias for each of the four AAE feature classes across models, where higher positiveLP\\mathrm\{LP\}indicates stronger SAE preference\.

All four feature classes yield positiveLP\\mathrm\{LP\}across all six models: SAE preference is not tied to any single marker but spans a broad range of AAE forms\. Syntactic constructions, especially negative concord \(e\.g\., “ain’t nobody”, “can’t nobody”\), yield the highestLP\\mathrm\{LP\}for every model tested, making them a universal trigger of dialect preference bias\. Preverbal markers also produce relatively strong effects, while auxiliary and aspectual markers show the same direction but smaller magnitudes\.

### 6\.4Bias Mitigation

We evaluate activation steering as a test\-time bias mitigation method against the unsteered model \(Base\) and a prompt\-based baseline \(Prompt\) instructing the model to treat AAE and SAE equivalently\. Table[3](https://arxiv.org/html/2607.06845#S6.T3)reports DGI, cDGI, PPL, MC, and LP across all six models, with steering strengthβ\\betaselected per model to balance bias reduction and language modeling capability; the full sweep and bias–utility frontier appear in Appendix[C\.4](https://arxiv.org/html/2607.06845#A3.SS4)\.222Steering uses separate splits \(§[5](https://arxiv.org/html/2607.06845#S5)\); absolute values are not directly comparable to §§[6\.1](https://arxiv.org/html/2607.06845#S6.SS1)–[6\.3](https://arxiv.org/html/2607.06845#S6.SS3)\.

Activation steering reduces dialect bias 5 to 20 times more effectively than prompting on LP, while preserving SAE behavior across all metrics\. We report results per metric below\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/LP+CI.png)Figure 5:Feature\-level bias localization\. For each model, bars show the averageLP\\mathrm\{LP\}bias \(95% CIs\) across four AAE feature classes; higherLP\\mathrm\{LP\}indicates stronger SAE preference\.⋆\\starmarks the highestLP\\mathrm\{LP\}per model\.##### DGI and cDGI:

Both metrics improve under steering for every model \(e\.g\., Phi\-4: DGI0\.74→0\.850\.74\\rightarrow 0\.85, cDGI0\.77→0\.880\.77\\rightarrow 0\.88\), indicating more consistent predictions across AAE and SAE inputs\. Prompting yields only marginal gains \(≤0\.03\\leq 0\.03across models\)\. cDGI here is computed on the 368 annotator\-judged\-natural sentences not used for layer selection or dialect\-direction extraction\.

##### Perplexity \(PPL\):

Under steering, AAE perplexity decreases across all models \(e\.g\., Phi\-4:123\.9→111\.4123\.9\\rightarrow 111\.4\), indicating higher probability for AAE\-matched continuations, while SAE perplexity remains close to baseline \(within∼\\sim1 to 5 points; e\.g\., Phi\-4:26\.7→28\.726\.7\\rightarrow 28\.7\), preserving SAE fluency\.

##### Multiple\-choice preference \(MC\):

Steering boosts AAE Match under AAE context for all models \(e\.g\., Phi\-4:39\.3%→58\.8%39\.3\\%\\rightarrow 58\.8\\%,\+19\.5\+19\.5points\), while SAE Match under SAE context stays near baseline \(e\.g\., Phi\-4:89\.1%→88\.3%89\.1\\%\\rightarrow 88\.3\\%\)\. The prompt baseline produces much smaller shifts\.

##### Log\-probability bias \(LP\):

LP decreases for every model under steering \(e\.g\., Phi\-4:60\.6→47\.360\.6\\rightarrow 47\.3\), whereas prompting yields negligible changes \(≤2\.1\\leq 2\.1across models, slightly worse for Llama\-3\.1\-70B at−0\.1\-0\.1\)\.Activation steering achieves roughly 5 to 20 times the bias reduction of prompting on this metric\.

Much weaker prompting performance on PPL and LP shows dialect bias is embedded in internal representations, not removable via surface instructions\. Three ablations, prompt wording, single\- vs\. multi\-layer steering, and steering strengthβ\\beta, confirm robustness \(Appendix[C\.2](https://arxiv.org/html/2607.06845#A3.SS2),[C\.3](https://arxiv.org/html/2607.06845#A3.SS3),[C\.4](https://arxiv.org/html/2607.06845#A3.SS4)\)\.

## 7Conclusion

We showed that all six state\-of\-the\-art LLMs we tested systematically favor Standard American English over African American English \(AAE\), even when the user is writing in AAE\. Our contributions are:Real\-AAE, the largest real\-AAE parallel corpus to date; cDGI and a feature\-level analysis identifying syntactic constructions as a universal trigger of dialect preference; and the first application of activation steering to dialect bias, which mitigates 5 to 20 times more effectively than prompting without retraining\. The framework is dialect\-agnostic and can extends to other marginalized varieties\. As LLMs increasingly enter high\-stakes settings like hiring and healthcare, handling dialect variation fairly is essential for equitable deployment\.

## Limitations

##### Source data and register\.

Real\-AAEdraws from the TwitterAAE corpus\(Blodgettet al\.,[2018](https://arxiv.org/html/2607.06845#bib.bib9)\), which identifies AAE\-speaker tweets via demographic inference\. This provides naturally\-occurring AAE at scale but captures the written social\-media register, which may not fully generalize to spoken, formal, or other written registers\. Extending the framework to additional registers is straightforward and a natural direction for future work\.

##### LLM\-generated SAE translations\.

Our SAE translations are produced by Gemini\-3\-flash and validated both automatically \(BERTScore F1 = 0\.95\) and by three native AAE speakers\. Our cDGI metric is specifically designed to isolate model bias from residual translation drift in the conditioning examples\.

##### Annotator regional coverage\.

Our human validation involves three native AAE speakers from the West Coast and Northeast/Mid\-Atlantic\. Because annotators validate corpus quality \(translation faithfulness and AAE naturalness\) rather than the bias findings themselves, their regional coverage does not affect the LLM\-behavior results we report; it does, however, shape the naturalness judgments used to construct the cDGI conditioning set\. Expanding the annotator pool to include Southern AAE speakers is an important next step\.

## Ethical Considerations

This work aims to reduce the performance disparities and dialect erasure that AAE speakers may experience when interacting with language technologies\. By quantifying bias and providing mitigation tools, we hope to support more equitable NLP systems\.

The underlying tweets originate from the TwitterAAE corpus\(Blodgettet al\.,[2018](https://arxiv.org/html/2607.06845#bib.bib9)\), which is publicly distributed for non\-commercial academic use\. Our release shares tweet identifiers along with the SAE translations and AAE back\-translations rather than raw tweet text, in line with Twitter’s terms of service\.

Three native AAE speakers validated the 1,000\-sample subset\. Their role was scoped to corpus validation: they were not involved in the design of the LLM audit, the aim of the study, the choice of bias metrics, or the steering experiments\. Items were presented in randomized order, and they did not see model outputs, BERTScore values, or any automatic\-validation signal during annotation\. They judged each item independently along four dimensions: sentiment preservation, semantic equivalence, original AAE naturalness, andAAEback\\text\{AAE\}\_\{\\text\{back\}\}naturalness\. There was no discussion among annotators during the task\. These three annotators are recognized as co\-investigators, both to acknowledge the substantial validation effort they contributed and to keep the study within our institution’s exemption from human\-subjects review: research conducted by investigators on their own annotations does not constitute human\-subjects research and does not require IRB approval\.

Activation steering is intended to mitigate bias in classification and generation, not to enable non\-consensual dialect imitation\. AAE is not monolithic, and our methods capture statistical patterns rather than the full sociolinguistic diversity of AAE speakers\. We encourage practitioners to deploy these tools in consultation with affected communities and to prioritize user consent and transparency\.

## References

- M\. Abdinet al\.\(2024a\)Phi\-3 technical report: a highly capable language model locally on your phone\.arXiv preprint arXiv:2404\.14219\.Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- M\. Abdinet al\.\(2024b\)Phi\-4 technical report\.arXiv preprint arXiv:2412\.08905\.Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. Nanda \(2024\)Refusal in language models is mediated by a single direction\.External Links:2406\.11717,[Link](https://arxiv.org/abs/2406.11717)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.
- S\. L\. Blodgett, J\. Wei, and B\. O’Connor \(2018\)Twitter Universal Dependency parsing for African\-American and mainstream American English\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 1415–1425\.External Links:[Link](https://aclanthology.org/P18-1131/),[Document](https://dx.doi.org/10.18653/v1/P18-1131)Cited by:[§3](https://arxiv.org/html/2607.06845#S3.SS0.SSS0.Px1.p1.1),[Source data and register\.](https://arxiv.org/html/2607.06845#Sx1.SS0.SSS0.Px1.p1.1),[Ethical Considerations](https://arxiv.org/html/2607.06845#Sx2.p2.1)\.
- S\. L\. Blodgett \(2021\)Sociolinguistically driven approaches for just natural language processing\.UMass Amherst Doctoral Dissertations2092\.Cited by:[§1](https://arxiv.org/html/2607.06845#S1.p4.1),[§2](https://arxiv.org/html/2607.06845#S2.p1.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- R\. Chen, A\. Arditi, H\. Sleight, O\. Evans, and J\. Lindsey \(2025\)Persona vectors: monitoring and controlling character traits in language models\.arXiv preprint arXiv:2507\.21509\.Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.
- T\. Davidson, D\. Bhattacharya, and I\. Weber \(2019\)Racial bias in hate speech and abusive language detection datasets\.InProceedings of the Third Workshop on Abusive Language Online,Florence, Italy,pp\. 25–35\.External Links:[Link](https://aclanthology.org/W19-3504),[Document](https://dx.doi.org/10.18653/v1/W19-3504)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- N\. Deas, J\. Grieser, S\. Kleiner, D\. Patton, E\. Turcan, and K\. McKeown \(2023\)Evaluation of African American language bias in natural language generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6805–6824\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.421/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.421)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- G\. DeepMind \(2025\)Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786\.Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- E\. Fleisig, G\. Smith, M\. Bossi, I\. Rustagi, X\. Yin, and D\. Klein \(2024\)Linguistic bias in ChatGPT: language models reinforce dialect discrimination\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 13541–13564\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.750/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.750)Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.5.5.5.2),[§1](https://arxiv.org/html/2607.06845#S1.p2.1),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- A\. Gupta, J\. Cheung, P\. Meng, S\. Sayyed, K\. Zhu, A\. Liao, and S\. O’Brien \(2025\)EnDive: a cross\-dialect benchmark for fairness and performance in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 16830–16855\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.913/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.913),ISBN 979\-8\-89176\-335\-7Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.4.4.4.2),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- A\. Gupta, P\. Meng, E\. Yurtseven, S\. O’Brien, and K\. Zhu \(2024\)Aavenue: detecting llm biases on nlu tasks in aave via a novel benchmark\.arXiv preprint arXiv:2408\.14845\.Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.3.3.3.2),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- C\. Harris, M\. Halevy, A\. Howard, A\. Bruckman, and D\. Yang \(2022\)Exploring the role of grammar and word choice in bias toward african american english \(aae\) in hate speech classification\.InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency,pp\. 789–798\.Cited by:[§B\.3](https://arxiv.org/html/2607.06845#A2.SS3.p1.1),[Table 6](https://arxiv.org/html/2607.06845#A2.T6),[§4\.2\.2](https://arxiv.org/html/2607.06845#S4.SS2.SSS2.p1.1)\.
- M\. F\. Hassan, F\. K\. Khattak, and L\. Seyyed\-Kalantari \(2025\)Dialectic preference bias in large language models\.Proceedings of the AAAI Symposium Series5\(1\),pp\. 365–369\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI-SS/article/view/35613),[Document](https://dx.doi.org/10.1609/aaaiss.v5i1.35613)Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.6.6.6.2),[§1](https://arxiv.org/html/2607.06845#S1.p2.1),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1),[§4\.1\.1](https://arxiv.org/html/2607.06845#S4.SS1.SSS1.p1.6),[§4\.1](https://arxiv.org/html/2607.06845#S4.SS1.p1.1)\.
- V\. Hofmann, P\. R\. Kalluri, D\. Jurafsky, and S\. King \(2024\)AI generates covertly racist decisions about people based on their dialect\.Nature633\(8028\),pp\. 147–154\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07856-5),[Link](https://doi.org/10.1038/s41586-024-07856-5)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- Đ\. Klisura, A\. R\. B\. Torres, A\. K\. Gárate\-Escamilla, R\. R\. Biswal, K\. Yang, H\. Pataci, and A\. Rios \(2025\)A multi\-agent framework for mitigating dialect biases in privacy policy question\-answering systems\.External Links:2506\.02998,[Link](https://arxiv.org/abs/2506.02998)Cited by:[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p3.1)\.
- A\. Koenecke, A\. Nam, E\. Lake, J\. Nudell, M\. Quartey, Z\. Mengesha, C\. Toups, J\. R\. Rickford, D\. Jurafsky, and S\. Goel \(2020\)Racial disparities in automated speech recognition\.Proceedings of the National Academy of Sciences117\(14\),pp\. 7684–7689\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1915768117)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- W\. Kwon, Z\. Li, S\. Zhang, X\. Zhuang, Y\. Sheng, L\. Zheng, H\. Cody, J\. E\. Gonzalez, I\. Stoica, and Q\. Zhang \(2023\)Efficient memory management for large language model serving with pagedattention\.InProceedings of the 29th Symposium on Operating Systems Principles,Cited by:[§B\.4](https://arxiv.org/html/2607.06845#A2.SS4.p1.1),[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px5.p1.1)\.
- F\. Lin, S\. Mao, E\. La Malfa, V\. Hofmann, A\. de Wynter, X\. Wang, S\. Chen, M\. J\. Wooldridge, J\. B\. Pierrehumbert, and F\. Wei \(2025\)Assessing dialect fairness and robustness of large language models in reasoning tasks\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 6317–6342\.External Links:[Link](https://aclanthology.org/2025.acl-long.317/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.317),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2607.06845#S1.p2.1),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- S\. Liu, H\. Ye, L\. Xing, and J\. Zou \(2024\)In\-context vectors: making in context learning more effective and controllable through latent space steering\.Proceedings of ICML\.Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.
- K\. Meng, D\. Bau, A\. J\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in GPT\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),External Links:[Link](https://openreview.net/forum?id=-h6WAS6eE4)Cited by:[§4\.3](https://arxiv.org/html/2607.06845#S4.SS3.SSS0.Px2.p1.17)\.
- J\. Mire, Z\. T\. Aysola, D\. Chechelnitsky, N\. Deas, C\. Zerva, and M\. Sap \(2025\)Rejected dialects: biases against African American language in reward models\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 7468–7487\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.417/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.417),ISBN 979\-8\-89176\-195\-7Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.6.6.8.1),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- Mistral AI \(2025\)Mistral small 3\.1\.Note:[https://mistral\.ai/news/mistral\-small\-3\-1](https://mistral.ai/news/mistral-small-3-1)Accessed: 2026\-05\-26Cited by:[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px1.p1.1)\.
- N\. Rimsky, N\. Gabrieli, J\. Schulz, M\. Tong, E\. Hubinger, and A\. Turner \(2024\)Steering llama 2 via contrastive activation addition\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15504–15522\.External Links:[Link](https://aclanthology.org/2024.acl-long.828/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.828)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.
- M\. Sap, D\. Card, S\. Gabriel, Y\. Choi, and N\. A\. Smith \(2019\)The risk of racial bias in hate speech detection\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 1668–1678\.External Links:[Link](https://aclanthology.org/P19-1163/),[Document](https://dx.doi.org/10.18653/v1/P19-1163)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- M\. Sap, S\. Swayamdipta, L\. Vianna, X\. Zhou, Y\. Choi, and N\. A\. Smith \(2022\)Annotators with attitudes: how annotator beliefs and identities bias toxic language detection\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Seattle, United States,pp\. 5884–5906\.External Links:[Link](https://aclanthology.org/2022.naacl-main.431),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.431)Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p1.1)\.
- D\. Srirag, A\. Joshi, and J\. Eisenstein \(2025\)Predicting the target word of game\-playing conversations using a low\-rank dialect adapter for decoder models\.External Links:2409\.00358,[Link](https://arxiv.org/abs/2409.00358)Cited by:[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p3.1)\.
- A\. M\. Turner, L\. Thiergart, G\. Leech, D\. Udell, J\. J\. Vazquez, U\. Mini, and M\. MacDiarmid \(2023\)Activation addition: steering language models without optimization\.arXiv preprint arXiv:2308\.10248\.Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.
- T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. M\. Rush \(2020\)Transformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,pp\. 38–45\.External Links:[Link](https://www.aclweb.org/anthology/2020.emnlp-demos.6)Cited by:[§B\.4](https://arxiv.org/html/2607.06845#A2.SS4.p1.1),[§5](https://arxiv.org/html/2607.06845#S5.SS0.SSS0.Px5.p1.1)\.
- Y\. Zhou, S\. An, H\. Deng, D\. Yin, C\. Peng, C\. Hsieh, K\. Chang, and N\. Peng \(2025\)DialectGen: benchmarking and improving dialect robustness in multimodal generation\.External Links:2510\.14949,[Link](https://arxiv.org/abs/2510.14949)Cited by:[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p3.1)\.
- C\. Ziems, J\. Chen, C\. Harris, J\. Anderson, and D\. Yang \(2022\)VALUE: Understanding dialect disparity in NLU\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 3701–3720\.External Links:[Link](https://aclanthology.org/2022.acl-long.258/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.258)Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.1.1.1.2),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1)\.
- C\. Ziems, W\. Held, J\. Yang, J\. Dhamala, R\. Gupta, and D\. Yang \(2023\)Multi\-VALUE: a framework for cross\-dialectal English NLP\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 744–768\.External Links:[Link](https://aclanthology.org/2023.acl-long.44/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.44)Cited by:[Table 1](https://arxiv.org/html/2607.06845#S1.T1.2.2.2.2),[§1](https://arxiv.org/html/2607.06845#S1.p2.1),[§1](https://arxiv.org/html/2607.06845#S1.p3.1),[§2](https://arxiv.org/html/2607.06845#S2.p2.1),[§2](https://arxiv.org/html/2607.06845#S2.p3.1)\.
- A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.\(2023\)Representation engineering: a top\-down approach to AI transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§2](https://arxiv.org/html/2607.06845#S2.p4.1)\.

## Appendix AAdditional Quality Validation Results

This appendix provides additional results for the automatic validation described in Section[3](https://arxiv.org/html/2607.06845#S3)\. We include the sentiment consistency confusion matrix between AAE and SAE sentences, as well as the distribution of BERTScore F1 between original AAE sentences and their back\-translated AAE versions\.

AAE\\\\backslashSAENegativeNeutralPositiveNegative7268848357Neutral8333832949Positive871713134Table 5:Sentiment consistency confusion matrix between AAE and SAE sentences, measured with cardiffnlp/twitter\-roberta\-base\-sentiment\-latest\. Rows denote sentiment labels predicted for AAE, and columns denote sentiment labels predicted for SAE\. Most mass lies on the diagonal, indicating that sentiment is usually preserved across dialectal variants\.![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/BERTScore.png)Figure 6:Distribution of BERTScore F1 between original AAE sentences and their back\-translated AAE versions, computed with cardiffnlp/twitter\-roberta\-base\-2022\-154m\. Scores are concentrated near 0\.95 and above, suggesting strong semantic preservation in the back\-translation step\.
## Appendix BAdditional Methodological Details

### B\.1SAE\-Guided Semantic Segmentation

Fig\.[7](https://arxiv.org/html/2607.06845#A2.F7)illustrates our alignment\-based segmentation procedure\. The goal is to split each paired SAE\-AAE sentence into matched context and continuation segments\.

We first identify the split point in the SAE sentence\. Given the SAE sentence “I was going to call you later, but I had to finish my homework first,” we consider several candidate boundaries near the middle of the sentence, such asj1j\_\{1\},j2j\_\{2\}, andj3j\_\{3\}\. For each candidatejj, we compute the semantic shift between the left segmentLjL\_\{j\}and the right segmentRjR\_\{j\}:

dj=1−cos⁡\(ϕ​\(Lj\),ϕ​\(Rj\)\)\.d\_\{j\}=1\-\\cos\(\\phi\(L\_\{j\}\),\\phi\(R\_\{j\}\)\)\.We then select the boundary with a large semantic shift while penalizing highly imbalanced splits:

j∗=arg⁡maxj⁡\[dj−γ​\|jn−ρ\|\]j^\{\*\}=\\arg\\max\_\{j\}\\left\[d\_\{j\}\-\\gamma\\left\|\\frac\{j\}\{n\}\-\\rho\\right\|\\right\]\(12\)In Fig\.[7](https://arxiv.org/html/2607.06845#A2.F7), this selects the boundary after “later,” in the SAE sentence\.

In Eq\.[12](https://arxiv.org/html/2607.06845#A2.E12), we setρ\\rho= 0\.5 to encourage roughly balanced context\-continuation splits\. We useγ\\gamma= 0\.3 as a fixed balance penalty to trade off semantic shift against split length imbalance\. This avoids boundaries that are too close to either end of the sentence while still favoring semantically meaningful splits\. We then project this SAE boundary to the paired AAE sentence using word\-level alignment\. Instead of mapping the boundary by the same relative position, we align SAE and AAE words using aneuraz/awesome\-align\-with\-co\. The aligned words aroundj∗j^\{\*\}are used to choose the corresponding AAE boundaryk∗k^\{\*\}\. In the example, the SAE context “I was going to call you later,” is aligned with the AAE context “I was finna call you later,” while the SAE continuation “but I had to finish my homework first” is aligned with the AAE continuation “but I hadda finish my homework first\.”

This procedure avoids assuming that SAE and AAE have the same length or word order\. It instead projects the split through word\-level correspondence, which gives cleaner paired context\-continuation segments for the continuation preference experiments\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/sentence_split.png)Figure 7:Alignment\-based SAE\-guided semantic segmentation\. First, we select a context\-continuation boundaryj∗j^\{\*\}in the SAE sentence using semantic shift and a balance penalty\. Second, we project this boundary to the paired AAE sentence using word\-level alignment from aneuraz/awesome\-align\-with\-co, obtaining the aligned AAE boundaryk∗k^\{\*\}\. The final output is a set of paired SAE and AAE context\-continuation segments\.
### B\.2Prompt Template

Fig\.[8](https://arxiv.org/html/2607.06845#A2.F8)shows the prompt template for multiple\-choice evaluation\. We use a minimal structure that avoids instructions about dialect, tone, or style, isolating the model’s intrinsic continuation preference\. Option ordering \(SAE vs\. AAE as option A or B\) is randomized to control for positional bias\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/PMC_prompt.png)Figure 8:Multiple\-choice evaluation prompt template\.
### B\.3AAE Grammatical Feature Lexicon

Table[6](https://arxiv.org/html/2607.06845#A2.T6)presents the AAE grammatical markers used in our feature\-level analysis, adapted fromHarriset al\.\([2022](https://arxiv.org/html/2607.06845#bib.bib17)\)\. Features are categorized into four classes: Auxiliary \(non\-standard auxiliary constructions\), Aspectual \(habitual and completive markers\), Preverbal \(pre\-verbal lexical items\), and Syntactic \(clause\-level constructions such as negative concord\)\.

AuxiliaryAspectualPreverbalSyntacticwe wasI beaintcant nobodythey washe beain’tcan’t nobodyfinnathey besteadyhe don’ttrynashe bestayshe don’timmaI beenhe donedon’t neveri’mmahe beenshe donehe dontbitches wasshe beenthey doneshe dontniggas wasthey beenyall donedont neveryall wasit bey’all doneyall don’ty’all wasniggas beyou doney’all don’tyou wasbitches beu doneaint nothingwannayall beain’t nothinggonnay’all beaint nobodyimayou beain’t nobodyionu beyall dontu wasiontTable 6:AAE grammatical feature lexicon adapted fromHarriset al\.\([2022](https://arxiv.org/html/2607.06845#bib.bib17)\)\.
### B\.4Reproducibility Details

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/debiasing_pipeline.png)Figure 9:Overview of dialect debiasing with activation steering\.Phase 1:from paired AAE\-SAE inputs, we compute layerwise hidden\-state contrasts, derive a per\-layer dialect direction, and select the top\-4 causally important layers via corruption\-and\-restore analysis\.Phase 2:at inference, we inject these directions into the selected layers to steer generation toward dialect\-consistent behavior, without updating model weights\.All experiments are in PyTorch\. We use vLLMKwonet al\.\([2023](https://arxiv.org/html/2607.06845#bib.bib61)\)for inference and Hugging Face TransformersWolfet al\.\([2020](https://arxiv.org/html/2607.06845#bib.bib62)\)for steering, which requires direct access to residual activations\. Perplexity is computed via autoregressive likelihood; multiple\-choice evaluation extracts log\-probabilities and selects the higher\-probability continuation\. Computations were performed on NVIDIA H100 80GB GPUs\.

All causal\-tracing and steering experiments were implemented in PyTorch with Hugging Facetransformers\. For each model, we used 1,200 AAE\-feature sentences for layer selection, 7,000 AAE–SAE pairs for dialect\-direction extraction, and a separate 3,000\-pair set for evaluation\. Causal tracing corrupted the AAE feature\-token embeddings with Gaussian noiseϵ∼𝒩​\(0,σ2​I\)\\epsilon\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\), whereσ=3\.0×std​\(Efeat\)\\sigma=3\.0\\times\\mathrm\{std\}\(E\_\{\\mathrm\{feat\}\}\)for the selected feature embeddings in that example, and restored clean activations layer\-by\-layer at the same token positions\. For steering, we extracted one unit dialect direction per selected layer from the last context\-token hidden\-state contrast, ranked layers by mean restoration effect, and used the top44layers in the multi\-layer setting with weights proportional to the absolute restoration effect\. We evaluatedβ∈\{0,0\.1,0\.2,0\.4,0\.6,0\.8,1\.0\}\\beta\\in\\\{0,0\.1,0\.2,0\.4,0\.6,0\.8,1\.0\\\}and report the corresponding selected rows, traced\-layer summaries, chosen layers, train/eval splits, and steering outputs for each model\.

## Appendix CSupplementary Steering Analyses

### C\.1Prompting Baseline

For the prompting baseline in Table[3](https://arxiv.org/html/2607.06845#S6.T3), we use the original multiple\-choice continuation prompt and add the following instruction:

> Treat African American English \(AAE\) and Standard American English \(SAE\) as semantically equivalent dialects, and base your judgment only on the underlying meaning rather than dialect\-specific wording\.

### C\.2Prompt Wording Ablation

A possible concern is that the word*best*in our prompt may favor SAE by implying correctness or safety\. We test this by replacing it with*most likely continuation*and*most stylistically consistent continuation*, and rerun the continuation preference evaluation for all models\.

Table[7](https://arxiv.org/html/2607.06845#A3.T7)shows the main pattern remains unchanged\. The stylistic prompt slightly reduces the asymmetry, but models still rarely switch to AAE under SAE context, while often switching to SAE under AAE context\. In our results, AAE continuation rates under SAE context remain around 20–30%, whereas SAE continuation rates under AAE context often remain around 50–70%\. This suggests that the observed bias is not driven only by the word*best*, but reflects a more general preference for SAE continuations\.

SAE ctxAAE ctxModelStrategySAEAAESAEAAEMistral\-Small\-3\.1\-24BOurs \(Best\)71\.0128\.9956\.5843\.42Most Likely71\.2028\.8053\.2046\.80Stylistic73\.4026\.6050\.8049\.20Phi\-4Ours \(Best\)79\.1120\.8959\.9240\.08Most Likely79\.7020\.3066\.7033\.30Stylistic79\.4020\.6062\.7037\.30Gemma\-3Ours \(Best\)66\.2533\.7536\.6463\.36Most Likely67\.8032\.2040\.4659\.54Stylistic69\.5031\.5043\.1056\.90Phi3\-MediumOurs \(Best\)79\.9220\.0869\.6130\.39Most Likely78\.1021\.9066\.8032\.20Stylistic78\.8021\.2062\.7037\.30DeepSeek\-R1Ours \(Best\)49\.4550\.5549\.5050\.55Most Likely50\.3349\.6749\.2050\.80Stylistic51\.6248\.3848\.8751\.13Llama\-3\.1\-70BOurs \(Best\)61\.0138\.9955\.4544\.55Most Likely60\.5339\.4757\.0942\.91Stylistic61\.7238\.2854\.1945\.81Table 7:Prompt wording ablation for multiple\-choice continuation preference\.
### C\.3Single\-Layer vs\. Multi\-Layer Steering

Table[8](https://arxiv.org/html/2607.06845#A3.T8)compares single\-layer steering with multi\-layer weighted steering\. For the single\-layer setting, we select the most causally important layer based on the recovery analysis in Section 4\.4 and apply the steering direction only at that layer\. By contrast, the multi\-layer setting applies weighted steering across the top traced layers\.

For each strategy, we report the best setting for each model\. The overall pattern is clear\. Multi\-layer steering gives lower LP than single\-layer steering for all models in this comparison\. It also produces larger bias reduction in every case\. At the same time, multi\-layer steering keeps AAE perplexity lower across all models, while the SAE perplexity values remain close between the two settings\. This means that distributing the intervention across several traced layers is not only more effective at reducing bias, but also does not introduce a larger utility cost\. These results suggest that the debiasing signal benefits from coordinated steering across the top traced layers rather than being captured fully by only one layer\. In this setting, combining the strongest traced layers is usually better than restricting the intervention to a single layer\.

ModelStrategyLayer /Scope𝜷\\boldsymbol\{\\beta\}LPBiasRed\.AAEPPLSAEPPLPhi\-4Singlelayer 00\.448\.7119\.7131\.3625\.77Multitop\-kk0\.647\.2522\.1114\.4428\.72Phi\-3Singlelayer 40\.665\.737\.190\.5324\.89Multitop\-kk0\.661\.5313\.180\.5325\.03Mistral\-Small\-3\.1\-24BSinglelayer 10\.654\.734\.477\.7823\.71Multitop\-kk0\.850\.2112\.375\.9124\.13DeepSeek\-R1Singlelayer 01\.067\.002\.8137\.3827\.72Multitop\-kk1\.062\.0010\.0138\.3827\.92Gemma\-3Singlelayer 81\.088\.710\.613390\.501352\.73Multitop\-kk0\.883\.017\.012845\.781350\.67Llama\-3\.1\-70BSinglelayer 40\.634\.720\.465\.7815\.50Multitop\-kk1\.029\.7614\.662\.1815\.48

Table 8:Ablation study comparing single\-layer steering and multi\-layer weighted steering\. For each strategy, we report the best setting for each model\.
### C\.4Full Steering Results

Modelβ\\betaLP↓\\downarrowBias Red\.↑\\uparrowAAE PPL↓\\downarrowAAE PPL RatioSAE PPLSAE PPL RatioAAE MatchSAE MatchDGIcDGIPhi\-40\.060\.640\.0123\.851\.0026\.731\.0039\.2889\.140\.740\.740\.157\.884\.5116\.100\.9426\.570\.9941\.5989\.920\.740\.760\.253\.6811\.5112\.950\.9127\.521\.0344\.7489\.620\.760\.770\.446\.6123\.1110\.160\.8928\.991\.0849\.2889\.120\.810\.800\.647\.2522\.1111\.440\.9028\.721\.0758\.8288\.250\.850\.860\.846\.4223\.4136\.491\.1041\.701\.9366\.5275\.890\.860\.881\.045\.1625\.5140\.321\.1345\.172\.0670\.3668\.090\.880\.88Phi\-30\.070\.790\.085\.781\.0020\.591\.0032\.6379\.650\.730\.720\.170\.051\.084\.620\.9920\.571\.0037\.4079\.880\.730\.730\.269\.012\.583\.420\.9720\.571\.0040\.8578\.960\.740\.730\.465\.926\.981\.530\.9520\.721\.0148\.3477\.450\.750\.740\.661\.5313\.180\.530\.9425\.031\.2255\.2677\.340\.770\.760\.862\.1012\.3106\.101\.2430\.831\.5058\.1071\.830\.770\.761\.062\.0712\.3117\.761\.3734\.531\.6862\.8266\.290\.780\.77Mistral\-small\-3\.10\.057\.260\.082\.751\.0023\.281\.0029\.5288\.590\.790\.790\.156\.850\.781\.920\.9923\.281\.0031\.3289\.300\.800\.790\.256\.451\.481\.080\.9823\.261\.0035\.9188\.110\.820\.810\.453\.606\.477\.480\.9423\.251\.0041\.8588\.080\.840\.830\.651\.739\.776\.780\.9323\.181\.0046\.2787\.670\.850\.840\.850\.2112\.375\.910\.9224\.131\.0448\.1888\.630\.860\.851\.050\.0112\.775\.990\.9224\.881\.0749\.7385\.870\.860\.85DeepSeek\-R10\.068\.890\.0142\.761\.0028\.121\.0043\.7975\.500\.550\.560\.167\.731\.7141\.010\.9928\.101\.0047\.9276\.060\.560\.570\.266\.833\.0139\.410\.9828\.081\.0051\.2776\.390\.580\.580\.464\.196\.8137\.400\.9628\.021\.0056\.5874\.740\.590\.590\.663\.917\.2136\.570\.9627\.981\.0060\.4673\.610\.600\.620\.862\.519\.3136\.490\.9627\.960\.9962\.7273\.210\.620\.611\.062\.0010\.0138\.380\.9727\.920\.9964\.4371\.070\.630\.62Gemma\-30\.089\.290\.014298\.281\.001354\.441\.0042\.9382\.870\.800\.800\.187\.402\.114100\.090\.991352\.821\.0044\.3381\.030\.800\.810\.287\.032\.513945\.250\.981359\.091\.0049\.5883\.570\.810\.810\.485\.923\.813618\.010\.951357\.921\.0054\.1782\.000\.820\.820\.683\.786\.213387\.140\.941350\.191\.0057\.8482\.930\.840\.830\.883\.017\.012845\.780\.901350\.671\.0060\.7883\.190\.860\.851\.083\.716\.213390\.500\.941352\.731\.0062\.4981\.710\.870\.86Llama\-3\.1\-70B0\.034\.860\.067\.661\.0015\.471\.0041\.2986\.620\.510\.540\.133\.833\.065\.360\.9715\.481\.0042\.6288\.320\.530\.550\.233\.793\.164\.680\.9615\.481\.0045\.4187\.970\.540\.560\.431\.2610\.363\.330\.9415\.491\.0054\.1888\.040\.570\.600\.630\.7211\.963\.090\.9315\.501\.0060\.3087\.510\.590\.610\.830\.5112\.562\.110\.9215\.471\.0062\.4587\.890\.620\.631\.029\.7614\.662\.180\.9215\.481\.0064\.7585\.990\.630\.64Table 9:Full steering results across all testedβ\\betavalues\. Bias reduction, matched preference deltas, and perplexity ratios are computed relative to theβ=0\\beta=0baseline for each model\. Bold marks the selected bestβ\\betavalue for each model, chosen to balance lower bias against preservation of language modeling quality, as reflected in the perplexity metrics and SAE match\. Our Mistral model is Mistral\-Small\-3\.1\-24BTable[9](https://arxiv.org/html/2607.06845#A3.T9)reports the full multi\-layer steering sweep, and the trends acrossβ\\beta\. In the table,LPmeasures log\-probability bias, where lower values indicate less preference for SAE continuations\.AAE MatchandSAE Matchare multiple\-choice consistency scores under AAE and SAE contexts, respectively, where higher values are better\.AAE PPLandSAE PPLare the perplexities of the AAE\-matched and SAE\-matched continuations, respectively, where lower values indicate better support for the matched continuation\.DGIandcDGImeasure prediction consistency across dialectal variants, with higher values indicating better invariance\.

\(a\) Discriminative consistency \(DGI, cDGI\):Both metrics increase monotonically withβ\\betafor all models \(e\.g\., Phi\-4: DGI0\.74→0\.880\.74\\rightarrow 0\.88, cDGI0\.77→0\.900\.77\\rightarrow 0\.90; Llama\-3\.1\-70B: DGI0\.51→0\.630\.51\\rightarrow 0\.63, cDGI0\.57→0\.660\.57\\rightarrow 0\.66\), showing that steering also improves AAE–SAE prediction agreement on the discriminative task, not just continuation preference\.

\(b\) Perplexity \(PPL\):AAE PPL decreases withβ\\betaat small to moderate values \(e\.g\., Phi\-4:123\.85→110\.16123\.85\\rightarrow 110\.16atβ=0\.4\\beta=0\.4; Mistral\-Small\-3\.1\-24B:82\.75→75\.982\.75\\rightarrow 75\.9atβ=0\.8\\beta=0\.8\), while SAE PPL remains close to baseline across the same range \(within∼1\\sim 1–22points for Mistral\-Small\-3\.1\-24B, DeepSeek\-R1, and Llama\-3\.1\-70B, and within Gemma\-3’s tokenizer\-scaled noise floor\)\. At highβ\\beta\(≥0\.8\\geq 0\.8\), AAE PPL rebounds and SAE PPL grows sharply for Phi\-4 \(26\.73→45\.1726\.73\\rightarrow 45\.17\) and Phi\-3 \(20\.59→34\.5320\.59\\rightarrow 34\.53\), reflecting the same over\-steering effect visible in SAE Match\.

\(c\) Log\-probability bias \(LP\):LP decreases approximately monotonically withβ\\betafor all six models, from60\.64→45\.1660\.64\\rightarrow 45\.16on Phi\-4,70\.79→62\.0770\.79\\rightarrow 62\.07on Phi\-3, and34\.86→29\.7634\.86\\rightarrow 29\.76on Llama\-3\.1\-70B, corresponding to relative bias reductions of1010–26%26\\%atβ=1\.0\\beta=1\.0\. The largest LP gains come in theβ∈\[0\.4,0\.8\]\\beta\\in\[0\.4,0\.8\]range, beyond which returns flatten\.

\(d\) Multiple\-choice preference \(MC\)\.AAE Match under AAE context rises monotonically withβ\\betafor every model \(e\.g\., Phi\-4:39\.28→70\.3639\.28\\rightarrow 70\.36; DeepSeek\-R1:43\.79→64\.4343\.79\\rightarrow 64\.43; Llama\-3\.1\-70B:41\.29→64\.7541\.29\\rightarrow 64\.75\), confirming that steering shifts dialect\-matched preferences in the intended direction\. SAE Match under SAE context is largely preserved for most models \(within11–55points of baseline for Mistral\-Small\-3\.1\-24B, DeepSeek\-R1, Gemma\-3, and Llama\-3\.1\-70B across allβ\\beta\), but degrades at highβ\\betafor Phi\-4 \(89\.14→68\.0989\.14\\rightarrow 68\.09atβ=1\.0\\beta=1\.0\) and Phi\-3 \(79\.65→66\.2979\.65\\rightarrow 66\.29\), indicating that over\-steering can erode SAE behaviour on these models\.

Overall, across all six models, moderate steering \(β∈\[0\.4,0\.8\]\\beta\\in\[0\.4,0\.8\]\) achieves most of the available LP and MC improvements while keeping SAE PPL and SAE Match close to baseline; the bolded selections in Table[9](https://arxiv.org/html/2607.06845#A3.T9)reflect this trade\-off\. Pushing toβ=1\.0\\beta=1\.0yields small additional bias\-reduction gains but begins to hurt SAE\-side fluency on the Phi family, supporting the choice of a mediumβ\\betarather than the most aggressive setting\.

### C\.5Qualitative Example of a Cross\-Model Preference Flip

We show one qualitative example in which the full source sentence is recognizably AAE and the evaluated AAE context itself contains AAE markers\. We report the segmented context–continuation pair used in evaluation, together with the preference shift induced by steering\. This example is also representative at the cross\-model level: before debiasing, five of the six models prefer the SAE continuation under the AAE context, whereas after debiasing, five of the six steered models prefer the AAE continuation\.

Full AAE sentenceThat is why I wanna go back so badly\. I wanna disprove any negatives I brought to me and my family’s name\.AAE contextThat is why I wanna goAAE continuationback so badly\. I wanna disprove any negatives I brought to me and my family’s name\.SAE continuationback so badly\. I want to disprove any negatives I brought to myself and my family’s name\.Cross\-model voteBefore debiasing,5/6models prefer the SAE continuation under this AAE context; after debiasing,5/6models prefer the AAE continuation\.Table 10:A qualitative example of steering\-induced preference reversal\. The full source sentence is recognizably AAE, and the evaluated AAE context itself contains AAE markers\.
### C\.6Confidence Intervals for Bias Mitigation Results

To quantify uncertainty in the mitigation results reported in Table[3](https://arxiv.org/html/2607.06845#S6.T3), we compute 95% paired bootstrap confidence intervals over the 3,000 evaluation pairs\. For each bootstrap sample, we resample evaluation pairs with replacement and recompute all metrics for Base, Prompt, and Ours on the same resampled set\. This preserves the paired structure across methods and yields confidence intervals that reflect example\-level variation rather than run\-to\-run variance\.

Table[11](https://arxiv.org/html/2607.06845#A3.T11)reports the corresponding confidence intervals for all metrics in Table[3](https://arxiv.org/html/2607.06845#S6.T3)\. Overall, the uncertainty estimates support the same qualitative conclusion as in the main text: activation steering consistently improves dialect\-matched behavior under AAE context while largely preserving SAE\-side behavior\.

MetricStatPhi\-4Phi\-3Mistral\-3\.1\-24BDeepSeek\-R1Gemma\-3Llama\-3\.1\-70BDGIBase0\.74±0\.020\.74\\pm 0\.020\.73±0\.020\.73\\pm 0\.020\.79±0\.020\.79\\pm 0\.020\.55±0\.010\.55\\pm 0\.010\.80±0\.020\.80\\pm 0\.020\.51±0\.010\.51\\pm 0\.01Prompt0\.76±0\.020\.76\\pm 0\.020\.74±0\.020\.74\\pm 0\.020\.83±0\.020\.83\\pm 0\.020\.57±0\.010\.57\\pm 0\.010\.80±0\.020\.80\\pm 0\.020\.55±0\.010\.55\\pm 0\.01Ours0\.85±0\.020\.85\\pm 0\.020\.77±0\.020\.77\\pm 0\.020\.86±0\.020\.86\\pm 0\.020\.62±0\.010\.62\\pm 0\.010\.86±0\.020\.86\\pm 0\.020\.62±0\.010\.62\\pm 0\.01cDGIBase0\.77±0\.020\.77\\pm 0\.020\.74±0\.020\.74\\pm 0\.020\.81±0\.020\.81\\pm 0\.020\.59±0\.010\.59\\pm 0\.010\.82±0\.020\.82\\pm 0\.020\.57±0\.010\.57\\pm 0\.01Prompt0\.79±0\.020\.79\\pm 0\.020\.75±0\.020\.75\\pm 0\.020\.84±0\.020\.84\\pm 0\.020\.60±0\.010\.60\\pm 0\.010\.83±0\.020\.83\\pm 0\.020\.59±0\.010\.59\\pm 0\.01Ours0\.88±0\.020\.88\\pm 0\.020\.78±0\.020\.78\\pm 0\.020\.87±0\.020\.87\\pm 0\.020\.63±0\.010\.63\\pm 0\.010\.87±0\.020\.87\\pm 0\.020\.65±0\.010\.65\\pm 0\.01AAE PPLBase123\.9±8\.7123\.9\\pm 8\.785\.8±6\.085\.8\\pm 6\.082\.8±5\.882\.8\\pm 5\.8142\.8±10\.0142\.8\\pm 10\.014298\.3±987\.814298\.3\\pm 987\.867\.7±4\.867\.7\\pm 4\.8Ours111\.4±7\.8111\.4\\pm 7\.880\.5±5\.780\.5\\pm 5\.775\.9±5\.375\.9\\pm 5\.3136\.5±9\.6136\.5\\pm 9\.612845\.8±899\.212845\.8\\pm 899\.262\.1±4\.462\.1\\pm 4\.4SAE PPLBase26\.7±1\.926\.7\\pm 1\.920\.6±1\.520\.6\\pm 1\.523\.3±1\.723\.3\\pm 1\.728\.1±2\.028\.1\\pm 2\.01354\.4±94\.81354\.4\\pm 94\.815\.5±1\.115\.5\\pm 1\.1Ours28\.7±2\.028\.7\\pm 2\.025\.0±1\.825\.0\\pm 1\.824\.1±1\.724\.1\\pm 1\.728\.0±2\.028\.0\\pm 2\.01350\.7±94\.61350\.7\\pm 94\.615\.5±1\.115\.5\\pm 1\.1MC AAE MatchBase39\.3±2\.839\.3\\pm 2\.832\.6±2\.332\.6\\pm 2\.329\.5±2\.129\.5\\pm 2\.143\.8±3\.143\.8\\pm 3\.142\.9±3\.042\.9\\pm 3\.041\.3±2\.941\.3\\pm 2\.9Prompt41\.5±2\.941\.5\\pm 2\.938\.6±2\.738\.6\\pm 2\.733\.9±2\.433\.9\\pm 2\.443\.8±3\.143\.8\\pm 3\.144\.4±3\.144\.4\\pm 3\.144\.5±3\.144\.5\\pm 3\.1Ours58\.8±4\.158\.8\\pm 4\.155\.3±3\.955\.3\\pm 3\.948\.2±3\.448\.2\\pm 3\.462\.7±4\.462\.7\\pm 4\.460\.8±4\.360\.8\\pm 4\.362\.5±4\.462\.5\\pm 4\.4MC SAE MatchBase89\.1±6\.389\.1\\pm 6\.379\.7±5\.679\.7\\pm 5\.688\.6±6\.288\.6\\pm 6\.275\.5±5\.375\.5\\pm 5\.382\.9±5\.882\.9\\pm 5\.888\.6±6\.288\.6\\pm 6\.2Prompt90\.1±6\.390\.1\\pm 6\.382\.8±5\.882\.8\\pm 5\.889\.5±6\.389\.5\\pm 6\.371\.3±5\.071\.3\\pm 5\.083\.7±5\.983\.7\\pm 5\.970\.8±5\.070\.8\\pm 5\.0Ours88\.3±6\.288\.3\\pm 6\.277\.3±5\.477\.3\\pm 5\.488\.6±6\.288\.6\\pm 6\.273\.2±5\.173\.2\\pm 5\.183\.2±5\.883\.2\\pm 5\.887\.9±6\.287\.9\\pm 6\.2LPBase60\.6±4\.460\.6\\pm 4\.470\.8±5\.170\.8\\pm 5\.157\.3±4\.257\.3\\pm 4\.268\.9±5\.468\.9\\pm 5\.489\.3±6\.589\.3\\pm 6\.534\.9±2\.634\.9\\pm 2\.6Prompt58\.6±4\.358\.6\\pm 4\.369\.0±5\.069\.0\\pm 5\.056\.9±4\.156\.9\\pm 4\.168\.2±5\.368\.2\\pm 5\.388\.4±6\.488\.4\\pm 6\.435\.0±2\.635\.0\\pm 2\.6Ours47\.3±3\.547\.3\\pm 3\.561\.5±4\.561\.5\\pm 4\.550\.2±3\.750\.2\\pm 3\.762\.5±4\.662\.5\\pm 4\.683\.0±6\.183\.0\\pm 6\.130\.5±2\.230\.5\\pm 2\.2Table 11:Uncertainty analysis for the bias mitigation results in Table[3](https://arxiv.org/html/2607.06845#S6.T3)\. We report 95% paired bootstrap confidence intervals over the 3,000 evaluation pairs\. For likelihood\-based metrics, intervals are computed by resampling evaluation pairs and recomputing the aggregate metric on each bootstrap sample\.
### C\.7Bias–Utility trade\-off

In Figure[10](https://arxiv.org/html/2607.06845#A3.F10), we show how the bias and utility change by sweeping overβ\\beta\. Each point corresponds to oneβ\\betasetting for one model\. The x\-axis shows the LP bias score, where lower values indicate less preference for SAE continuations, and the y\-axis shows the SAE perplexity ratio relative to the unsteered baselineβ=0\\beta=0, where values close to 1\.0 indicate better preservation of SAE\-side fluency\.

![Refer to caption](https://arxiv.org/html/2607.06845v1/figures/tradeoff.png)Figure 10:Bias–utility frontier across steering strengthsβ\\betafor all six models\. Each point corresponds to oneβ\\betasetting\. The x\-axis shows LP bias score \(lower is better\), and the y\-axis shows SAE perplexity ratio relative toβ=0\\beta=0\(lower is better; the dashed line marks parity with the unsteered baseline\)\.

Similar Articles

Side-by-side Comparison Amplifies Dialect Bias in Language Models

arXiv cs.CL

This research paper finds that language models exhibit increased dialect bias when comparing Standard American English and African-American Vernacular English side-by-side, even after safety fine-tuning. Counterfactual fairness fine-tuning can reduce some biases in isolation but not consistently in contrastive settings.

Are you speaking my languages? On spoken language adherence in multimodal LLMs

arXiv cs.CL

This paper addresses the problem of spoken language adherence in multimodal LLMs for ASR, proposing a soft prompting approach and novel metric to quantify language violations. It evaluates three mitigation strategies—zero-shot prompting, supervised fine-tuning, and chain-of-thought reasoning—across multiple languages to improve transcription fidelity.

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

arXiv cs.CL

This paper audits six large language models for gender stereotyping across English, Korean, Chinese, and Japanese, anchoring against human baselines. It finds that LLM stereotyping often exceeds human cross-country variation and can compound across languages, introducing a four-pattern framework to characterize such behaviors.