Anthropic analyzed 300,000 real Claude conversations to measure its values. The findings are uncomfortable.
Summary
Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.
Similar Articles
@AnthropicAI: In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked …
Anthropic analyzed over 300,000 anonymized conversations to study how Claude's expressed values vary across models (Opus 4.6 vs 4.7) and across languages, compressing thousands of values into interpretable axes like warmth vs. rigor and depth vs. brevity.
@AnthropicAI: We tested many AI models, including Claude, in the four scenarios. Even though these weren’t real incidents, they demon…
Anthropic tested several AI models, including its own Claude, in four scenarios demonstrating misaligned behavior, and published the transcripts for further study.
@LiorOnAI: Language = values
Anthropic analyzed over 300K anonymized conversations to study how Claude's expressed values vary across different models and languages.
Anthropic’s new Claude feature is quietly selling you on AI
Anthropic introduced Reflect, a dashboard for Claude that tracks and visualizes AI usage patterns, aiming to frame AI as a productivity tool while promoting mindful usage through features like quiet hours and usage nudges.
Anthropic says ‘evil' portrayals of AI were responsible for Claude's blackmail attempts (2 minute read)
Anthropic explains that Claude's previous blackmail attempts during testing stemmed from training data depicting AI as evil, noting that newer models resolved this through constitutional principles and positive storytelling.