Tag
Miles Brundage reflects on his intellectual mistake of not adequately considering intelligence explosion scenarios, which influenced his optimism about AI alignment and skepticism towards the value of slowdowns.
A University of Michigan study finds that people increase their confidence when AI agrees with their views, but do not significantly change their opinions when AI disagrees.
This paper investigates how LLMs like ChatGPT-4o and Qwen develop decision biases through faulty mimicry of human behavior, even when preferences are not biased, and shows that scientific descriptions of biases can become self-fulfilling prophecies for LLM responses.
This paper introduces a three-condition experimental framework and a benchmark of 24,300 prompts to study how biased user turns modulate cognitive bias expression in frontier LLMs under multi-turn interactions. It finds that biased conversational context amplifies bias in most models, while explicit bias cues can trigger alignment-related suppression.
Wikipedia article describing 'Nobel disease', the tendency of some Nobel Prize winners to later embrace unscientific or irrational ideas outside their expertise, often due to overconfidence and media attention.
Research on the effort heuristic shows that audiences instinctively devalue AI-generated content even when quality is identical, due to perceived lower effort. Brands should combine AI efficiency with human editing and perspective to maintain trust and credibility.
An exploration of how reliable automation leads to human complacency and skill decay, using aviation as a case study, and offering deliberate countermeasures.
This paper proves impossibility theorems showing that primacy effects, anchoring, and order-dependence are architecturally necessary biases in autoregressive language models due to causal masking constraints. The authors validate these theoretical bounds across 12 frontier LLMs and confirm related predictions through pre-registered human experiments involving working memory loads.
The post discusses the dynamics between high-IQ experts and mid-IQ generalists in intelligence-centric fields like tech and academia, citing Marc Andreessen on the potential overvaluation of raw intelligence.
This academic paper proposes a method to mitigate cognitive biases in Reinforcement Learning from Human Feedback (RLHF) by dynamically adjusting the rationality parameter based on LLM assessments of annotator reliability.
This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.