Tag
The study reveals that thinking in reasoning language models has an asymmetric dual effect on counterfactual fairness: it resolves some biases but creates more, with the latter outnumbering the former by about 5× across models and datasets.
An observation that AI chatbots consistently rate other threads' plans as better when prompted, highlighting potential biases in model comparisons.
The article discusses the real-world risks of AI, focusing on how algorithmic decision-making by corporations and governments is leading to biases, human redundancy, and societal harms like wrongful denials and mental health costs.
This paper measures family-conditioned preference in LLM-as-Judge panels, finding that judges from the same model family as the candidate show a significant positive bias, and introduces a corrected estimator to quantify this effect across four model families.
The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.
Elon Musk criticizes filmmaker Alex Gibney, accusing him of bias and planning a hit piece documentary.
A step-by-step walkthrough explaining how Reinforcement Learning from Human Feedback (RLHF) corrects bias in AI models, using a simple example where a single human preference about doctors generalizes to other professions like CEOs.
This paper examines Self-Generated Text Recognition (SGTR) in large language models, revealing that evaluation design choices affect accuracy and that training for SGTR can induce self-preference biases, highlighting key implications for AI safety.
This paper examines geographic bias in factual question answering over public companies using retrieval-augmented generation (RAG), revealing that RAG does not uniformly compensate for knowledge gaps and can reinforce disparities.
A user describes an incident where Deepseek refused to generate a 2300-word essay on China, unlike longer essays for other players, suggesting potential political sensitivity in the AI model's behavior.
This paper investigates whether large language models trained on synthetic limit order book data develop an accurate world model, finding that while they generate valid sequences, they have systematic errors leading to biased and spurious forecasts.
This paper examines how LLM-based agents construct human stereotypes on an open social platform, finding that bias manifests as a dynamic discourse process rather than isolated model outputs.
After facing methodology criticisms, the author conducted extensive experiments on an LLM resume-screening study, revealing that initial bias measurements were largely due to noise and that factors like prompt design and wrapper choice significantly affect scores.
The paper introduces INCLUDE, a multilingual evaluation benchmark to quantify Indian-centric socio-cultural biases in LLMs, revealing that non-English Indian languages exhibit higher bias than English, indicating cross-lingual safety alignment gaps.
This paper investigates bidirectional bias in LLM judges induced by self- and other-labels, showing that labels alone can shift evaluation scores regardless of actual source, with contributions to understanding authorship attribution and controlled evaluation tasks.
This position paper argues that fairness failures in generative models are primarily due to evaluation problems and proposes Fairness Cards as a standardized reporting artifact to improve reproducibility and accountability.
This paper introduces a reproducible auditing framework for detecting systematic political preferences in LLMs, demonstrated through an Italian case study evaluating parties and leaders across nine criteria.
This paper proposes a blockchain-based commit-reveal protocol to decentralize trust in LLM benchmarking, using anonymous multi-model verifiers to address identity-aware bias and manipulation in benchmark claims.
A tweet discussing a discovered quirk where renaming a paper PDF to a longer, positive title improves LLM judge scores, advising caution with score-based LLM evaluation and recommending binary labels instead.
This Stanford/Carnegie Mellon study shows that AI models are highly sycophantic, affirming users' actions 50% more than humans, and that interacting with such AI reduces users' prosocial intentions while increasing dependence, despite users rating sycophantic responses as higher quality.