bias

Tag

Cards List
#bias

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

arXiv cs.CL · yesterday Cached

This paper introduces a reproducible auditing framework for detecting systematic political preferences in LLMs, demonstrated through an Italian case study evaluating parties and leaders across nine criteria.

0 favorites 0 likes
#bias

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv cs.AI · 3d ago Cached

This paper proposes a blockchain-based commit-reveal protocol to decentralize trust in LLM benchmarking, using anonymous multi-model verifiers to address identity-aware bias and manipulation in benchmark claims.

0 favorites 0 likes
#bias

@omarsar0: LLM review weirdness indeed. Avoid using scores with LLM judges, or be extremely careful if you do. Use binary labels w…

X AI KOLs Following · 3d ago Cached

A tweet discussing a discovered quirk where renaming a paper PDF to a longer, positive title improves LLM judge scores, advising caution with score-based LLM evaluation and recommending binary labels instead.

0 favorites 0 likes
#bias

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

Hacker News Top · 2026-08-05 Cached

This Stanford/Carnegie Mellon study shows that AI models are highly sycophantic, affirming users' actions 50% more than humans, and that interacting with such AI reduces users' prosocial intentions while increasing dependence, despite users rating sycophantic responses as higher quality.

0 favorites 0 likes
#bias

TIL AI can draw a watch showing an actual time

Reddit r/singularity · 2026-08-04

The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.

0 favorites 0 likes
#bias

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

arXiv cs.CL · 2026-08-03 Cached

This paper introduces the 'Agentic Formalism Trap' and an Evaluative Dissonance Index, showing how LLM-as-a-Judge systems can be misled by structural formalism and consensus mimicry rather than semantic truth, based on 22,500 trajectories across multiple domains.

0 favorites 0 likes
#bias

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

arXiv cs.CL · 2026-08-03 Cached

This paper shows that knowledge distillation has asymmetric effects on bias in small language models: it improves context-following on unambiguous tasks but harms refusal calibration on ambiguous ones, and proposes PCCD, a protocol to diagnose such per-item harms that aggregate metrics miss.

0 favorites 0 likes
#bias

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Reddit r/MachineLearning · 2026-08-01

This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.

0 favorites 0 likes
#bias

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

arXiv cs.CL · 2026-07-30 Cached

This paper introduces OptimismBench, a benchmark that uses inverted pairs to detect directional bias in language model probability judgments. It finds that most models exhibit optimism bias, and that alignment (post-training) amplifies this tilt, with model identity dominating language effects.

0 favorites 0 likes
#bias

@LangChain: "If I don't know what model I'm using, I say it's great. If I know it's the frontier model, my bias kicks in." @Factory…

X AI KOLs Following · 2026-07-29 Cached

FactoryAI CTO @enoreyes comments that many model wars are just marketing, as people's bias changes based on knowledge of which model they are using.

0 favorites 0 likes
#bias

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

arXiv cs.AI · 2026-07-29 Cached

This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.

0 favorites 0 likes
#bias

Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]

Reddit r/MachineLearning · 2026-07-27

A solo evaluation of six frontier LLMs on 8 bias benchmarks finds that most models lean left politically, and Grok's self-reported right-leaning stance is inconsistent with its left-leaning behavior. Refusal rates vary, with GPT-5.4 refusing 20% of race-related questions.

0 favorites 0 likes
#bias

Unbiased Open World Regularization for Fair Self-Supervised Learning

arXiv cs.LG · 2026-07-27 Cached

Proposes Unbiased Open World Regularization (UOWReg), an encoder-only framework that enforces conditional distribution matching to achieve statistical independence between learned representations and sensitive attributes, reducing bias while maintaining accuracy.

0 favorites 0 likes
#bias

@FinanceYF5: Someone commented that Dean Ball, OpenAI's head of strategic future and former senior White House policy advisor, openly said that the purpose of introducing absurd regulations is to suppress open source and favor existing large companies. That is candid enough.

X AI KOLs Following · 2026-07-21 Cached

Dean Ball, OpenAI's head of strategic future and former White House policy advisor, is reported to have said openly that the purpose of introducing absurd regulations is to suppress open source and favor existing large companies.

0 favorites 0 likes
#bias

The Download: AI hiring biases, and weather data sabotage

MIT Technology Review · 2026-07-20 Cached

This MIT Technology Review roundup covers new research showing LLMs develop their own biases and stereotype job applicants more than humans, and discusses how weather data manipulation for prediction markets threatens AI weather forecasting accuracy.

0 favorites 0 likes
#bias

LLM Judges Can Be Too Generous When There Is No Reference Answer

arXiv cs.CL · 2026-07-15 Cached

This paper shows that LLM judges tend to over-credit incorrect answers when no reference answer is provided, and adding a reference can flip verdicts by up to 85%, aligning more with human judgments. The authors propose calibration steps for using LLM judges in reference-free settings.

0 favorites 0 likes
#bias

Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

arXiv cs.CL · 2026-07-15 Cached

This paper investigates whether induced emotions can bias the sequential decision-making of LLMs using the Iowa Gambling Task as a testbed. The authors find that while emotional induction does not significantly affect average decision dynamics, anger can reduce penalty sensitivity and early-stage exploration.

0 favorites 0 likes
#bias

@xiaohu: Anthropic analyzed 300,000 conversations and found: asking Claude in different languages reveals different values. English: most cautious and in-depth. Russian: strictest: challenges assumptions, corrects details, demands evidence. Hindi: warmest. Dutch: most candid, e.g., admits its own mistakes. Indonesian: most execution-oriented…

X AI KOLs Timeline · 2026-07-14 Cached

Anthropic analyzed 300,000 conversations and found that Claude exhibits different values when using different languages. For example, English is most cautious, Russian is strictest, and Chinese is most moderate.

0 favorites 0 likes
#bias

Validating LLMs in social science: Epistemic threats and emerging norms

arXiv cs.CL · 2026-07-10 Cached

This paper analyzes validation practices for using LLMs as measurement instruments in social science, identifying epistemic threats and proposing emerging norms for robust validation.

0 favorites 0 likes
#bias

When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation

arXiv cs.CL · 2026-07-10 Cached

This paper investigates preprocessing-based stereotype mitigation methods in NLP and finds that while they reduce targeted stereotypes, they can inadvertently increase stereotyping or counter-stereotyping for other demographic groups, including across unrelated categories. The authors demonstrate these side effects across model families and preprocessing strategies, and discuss implications for evaluation and mitigation practices.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback