Tag
This research paper examines how sociodemographic prompting affects LLMs' judgments in subjective tasks, finding that it often aligns models with majority groups while misrepresenting minority groups, with instruction-tuning identified as a potential cause.
This paper introduces ContextBias and ContextBench to evaluate bias persistence in text-to-image models, finding that bias increases in semantically unrelated contexts.
This pilot study uses Natural Language Autoencoders to probe whether Qwen2.5-7B internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts, finding evidence of latent inferences before they are verbalized.
Introduces RPAM, a principled metric for evaluating associations in language models that demonstrates high predictive validity for downstream outputs, tested on Mistral-7B-Instruct, Mistral-7B, and GPT-2.
VIBE is a framework that evaluates generative bias in Large Audio-Language Models using open-ended tasks with human-recorded speech, revealing systematic biases triggered by gender and accent cues.
This paper presents the first bias evaluation of multimodal speech recognition models, finding significant accuracy differences across gender and ethnicity when pairing faces with audio, with implications for fairness in AI systems.