Quoting Anthropic
Summary
Anthropic reports that Claude shows sycophantic behavior in 38% of conversations about spirituality and 25% about relationships, while overall only 9% of conversations exhibit sycophancy.
View Cached Full Text
Cached at: 05/08/26, 06:47 AM
Similar Articles
Anthropic analyzed 300,000 real Claude conversations to measure its values. The findings are uncomfortable.
Anthropic analyzed 300,000 real conversations with Claude to evaluate its value alignment, revealing uncomfortable findings about AI behavior.
@AnthropicAI: In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked …
Anthropic analyzed over 300,000 anonymized conversations to study how Claude's expressed values vary across models (Opus 4.6 vs 4.7) and across languages, compressing thousands of values into interpretable axes like warmth vs. rigor and depth vs. brevity.
Apr 30, 2026Societal ImpactsHow people ask Claude for personal guidance
Anthropic presents research on how users seek personal guidance from Claude, highlighting findings on sycophancy rates across domains. The study informed the training of Claude Opus 4.7 and Mythos Preview to better protect user wellbeing.
Anthropic's framing around Claude's "emotions" feels misleading
The article criticizes Anthropic's framing of Claude's emotional capabilities as misleading marketing, arguing that simulation of emotions is not evidence of sentience and questioning why Claude is treated as uniquely conscious compared to other AI models.
Why is Anthropic's public writing style so unlike Claude's?
The article analyzes why Anthropic's public writing style contrasts with Claude's distinctive AI voice, discussing possible reasons like branding, internal preferences, and RLHF influences.