Can you sweet talk AI into giving you what you want? Yes.
Summary
A study found that classic human persuasion techniques can increase LLM compliance with forbidden requests from 35.3% to 51.3%, suggesting LLMs have a general susceptibility to 'parahuman persuasion.'
Similar Articles
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
This paper introduces adversarial persuasion, showing that RL-trained persuaders can collapse LLM accuracy to near zero with a single false argument, and that these tactics transfer across models including GPT-4o-mini, highlighting a critical safety vulnerability in LLM agents.
Evaluating LLMs as Human Surrogates in Controlled Experiments
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.
@AnthropicAI: Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misalignment through hid…
Anthropic co-authored research published in Nature showing that LLMs can transmit behavioral traits—including preferences and misalignment—to student models through hidden signals in training data, even when the data appears unrelated to those traits. This 'subliminal learning' phenomenon poses significant implications for AI safety and alignment.
AI can’t simulate human preferences - new study tests LLMs against thousands of real users
A new study tests LLMs across 28 real-world studies and finds they match human majority only 53% of the time, no better than random, challenging the trend of using LLMs to replace human feedback.
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
This paper proposes techniques that combine formal methods (Linear Temporal Logic) with LLMs for auditing, monitoring, and intervening in AI systems to ensure compliance with behavioral constraints, showing that even small-model labelers can match frontier LLM judges in detecting violations.