Tag
This paper introduces a reproducible auditing framework for detecting systematic political preferences in LLMs, demonstrated through an Italian case study evaluating parties and leaders across nine criteria.
This paper proposes a blockchain-based commit-reveal protocol to decentralize trust in LLM benchmarking, using anonymous multi-model verifiers to address identity-aware bias and manipulation in benchmark claims.
A tweet discussing a discovered quirk where renaming a paper PDF to a longer, positive title improves LLM judge scores, advising caution with score-based LLM evaluation and recommending binary labels instead.
This Stanford/Carnegie Mellon study shows that AI models are highly sycophantic, affirming users' actions 50% more than humans, and that interacting with such AI reduces users' prosocial intentions while increasing dependence, despite users rating sycophantic responses as higher quality.
The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.
This paper introduces the 'Agentic Formalism Trap' and an Evaluative Dissonance Index, showing how LLM-as-a-Judge systems can be misled by structural formalism and consensus mimicry rather than semantic truth, based on 22,500 trajectories across multiple domains.
This paper shows that knowledge distillation has asymmetric effects on bias in small language models: it improves context-following on unambiguous tasks but harms refusal calibration on ambiguous ones, and proposes PCCD, a protocol to diagnose such per-item harms that aggregate metrics miss.
This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.
This paper introduces OptimismBench, a benchmark that uses inverted pairs to detect directional bias in language model probability judgments. It finds that most models exhibit optimism bias, and that alignment (post-training) amplifies this tilt, with model identity dominating language effects.
FactoryAI CTO @enoreyes comments that many model wars are just marketing, as people's bias changes based on knowledge of which model they are using.
This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.
A solo evaluation of six frontier LLMs on 8 bias benchmarks finds that most models lean left politically, and Grok's self-reported right-leaning stance is inconsistent with its left-leaning behavior. Refusal rates vary, with GPT-5.4 refusing 20% of race-related questions.
Proposes Unbiased Open World Regularization (UOWReg), an encoder-only framework that enforces conditional distribution matching to achieve statistical independence between learned representations and sensitive attributes, reducing bias while maintaining accuracy.
Dean Ball, OpenAI's head of strategic future and former White House policy advisor, is reported to have said openly that the purpose of introducing absurd regulations is to suppress open source and favor existing large companies.
This MIT Technology Review roundup covers new research showing LLMs develop their own biases and stereotype job applicants more than humans, and discusses how weather data manipulation for prediction markets threatens AI weather forecasting accuracy.
This paper shows that LLM judges tend to over-credit incorrect answers when no reference answer is provided, and adding a reference can flip verdicts by up to 85%, aligning more with human judgments. The authors propose calibration steps for using LLM judges in reference-free settings.
This paper investigates whether induced emotions can bias the sequential decision-making of LLMs using the Iowa Gambling Task as a testbed. The authors find that while emotional induction does not significantly affect average decision dynamics, anger can reduce penalty sensitivity and early-stage exploration.
Anthropic analyzed 300,000 conversations and found that Claude exhibits different values when using different languages. For example, English is most cautious, Russian is strictest, and Chinese is most moderate.
This paper analyzes validation practices for using LLMs as measurement instruments in social science, identifying epistemic threats and proposing emerging norms for robust validation.
This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.