Tag
This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.
ArXiv preprint maps stereotype-encoding neurons and attention heads in GPT-2 Small and Llama 3.2, showing biases cluster in small neuron subsets yet ablating them barely reduces biased text generation.