Tag
This paper examines covert dialect bias in large language models by analyzing their internal probability distributions across four English dialects, finding that models associate more negative housing-related adjectives with African American Vernacular English and Nigerian Pidgin, reflecting inherited biases similar to human discrimination.
This paper audits and mitigates dialect bias in large language models, showing they systematically prefer Standard American English over African American English. The authors introduce activation steering, a training-free method that reduces bias significantly while preserving fluency, and release the largest real-AAE parallel corpus to date.
This research paper finds that language models exhibit increased dialect bias when comparing Standard American English and African-American Vernacular English side-by-side, even after safety fine-tuning. Counterfactual fairness fine-tuning can reduce some biases in isolation but not consistently in contrastive settings.