Tag
A study analyzing 25,500 LLM resume evaluations across 10 models found a 45% bias rate driven by 'silent bias', with models inventing professional-sounding excuses to penalize candidates. It highlights significant variability in fairness and stability, with Claude, Mistral-Large, and Llama 4 being most stable, while Qwen and older Gemini models were volatile.
OpenAI published a study examining how subtle identity cues like user names can influence ChatGPT's responses, introducing the concept of 'first-person fairness' to evaluate whether name-based biases lead to harmful stereotypes in direct user interactions. The research highlights limitations including a focus on English-language, binary gender, and four racial/ethnic categories.