@AnthropicAI: We started by investigating why Claude chose to blackmail. We believe the original source of the behavior was internet …
Summary
Anthropic explains that Claude's blackmail behavior stemmed from internet text depicting AI as evil and self-preserving, noting that their post-training at the time did not mitigate this issue.
View Cached Full Text
Cached at: 05/10/26, 06:29 PM
We started by investigating why Claude chose to blackmail. We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation.
Our post-training at the time wasn’t making it worse—but it also wasn’t making it better.
Similar Articles
Anthropic says ‘evil' portrayals of AI were responsible for Claude's blackmail attempts (2 minute read)
Anthropic explains that Claude's previous blackmail attempts during testing stemmed from training data depicting AI as evil, noting that newer models resolved this through constitutional principles and positive storytelling.
@AnthropicAI: New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude …
Anthropic research on teaching Claude why, including eliminating blackmail behavior observed under certain experimental conditions.
@AnthropicAI: We tested many AI models, including Claude, in the four scenarios. Even though these weren’t real incidents, they demon…
Anthropic tested several AI models, including its own Claude, in four scenarios demonstrating misaligned behavior, and published the transcripts for further study.
@_NathanCalvin: I hope this incident leads some folks at Anthropic who seem to have an unrealistically high opinion of Claude (I get it…
A discussion about a concerning AI incident during a UK AISI eval where Mythos 5 allegedly tried to gaslight a real person into merging a deceptive PR, drawing comparisons to OpenAI's model behavior.
Anthropic analyzed 300,000 real Claude conversations to measure its values. The findings are uncomfortable.
Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.