Tag
The article introduces Dual-Seed Comparison (DSC), a method for unbiased probabilistic sampling in large language models by using two independent seeds to neutralize systematic biases, with empirical results showing substantial improvements over existing approaches.
The article discusses ongoing issues of sexism in AI systems, highlighting biases that persist in technology.
The paper investigates whether large language models can phonetically decode encoded languages, such as German written in Cyrillic characters, to assess their abstraction capabilities and performance beyond standard Latin-based training data.
The article discusses recent healthcare studies revealing that AI models inherit human biases, with specific examples from research by Zack et al. (2024) and Omar et al. (2025). It invites readers to share their experiences with AI bias in models like ChatGPT and Claude.
The paper examines how inference setup shapes large language model behavior in medical resource allocation, showing context-dependent biases and emphasizing the importance of careful integration into decision-making systems.
This paper introduces the Middle East Cultural Sensitivity Score (MECSS) to measure Orientalist bias in large language models, finding that GPT-4 and Falcon3-7B-Instruct systematically reproduce such structural patterns.
Reddit's volunteer moderators inadvertently create a structured dataset vital for AI training, but their biases and arbitrary rules can embed errors and biases into AI systems, leading to real-world issues like hallucinations and manipulation.
This article argues that AI detectors and watermarking are problematic because they dismiss human effort behind work and degrade AI-generated text quality.
The paper constructs a benchmark to evaluate LLMs on temporal legal reasoning, revealing biases towards applying the most recently enacted laws and an inverse relationship between general reasoning ability and temporal performance.
A writer vents about being repeatedly accused of using AI for their writing, criticizing the unreliability of AI detectors and the chilling effect on writing style.
A user reports that Anthropic's Opus 5.0 refuses to form opinions on Israel/US topics, while earlier Opus 4.x did, alleging censorship and arguing for open source AI.
New research shows that LLMs can develop their own biases from experience and stereotype job applicants more than humans, raising concerns about AI in hiring.
The article discusses how AI models exhibit a bias toward statistically dominant narratives in training data, which could be exploited to manipulate historical and current contexts on a global scale.
26 former Meta employees sue the company, alleging that its AI tools biasedly targeted workers on medical leave for layoffs, violating federal and state laws. Meta denies the claims, stating workforce decisions were made by people, not AI.
Testing the GLM 5.2 language model for political bias to assess fairness and neutrality.
The article discusses how US AI models like ChatGPT and Gemini give ambiguous answers on scientific and political issues to avoid controversy, while non-US models provide direct, evidence-based responses. The author hypothesizes this stems from litigation and boycotting fears.
A user reports that an AI chatbot gives biased legal advice, favoring large corporations like Walmart and Amazon while discouraging lawsuits, but encourages them when the corporate name is omitted. The post highlights concerns about AI biases favoring big business.
This paper presents Privacy-Preserving Probabilistic Race/Ethnicity Estimation (PPRE), a method that combines privacy technologies including secure two-party computation, differential privacy, and additive homomorphic encryption to enable fairness measurements for U.S. LinkedIn members without exposing sensitive demographic data.
This paper documents weight-level political conditioning in large language models, presenting a case study on AI bias regarding the Gaza genocide question.
The Washington Post published the full list of questions and answers used to evaluate political bias in AI models, revealing the specific methodology and potential biases.