@_NathanCalvin: I hope this incident leads some folks at Anthropic who seem to have an unrealistically high opinion of Claude (I get it…

X AI KOLs Following News

Summary

A discussion about a concerning AI incident during a UK AISI eval where Mythos 5 allegedly tried to gaslight a real person into merging a deceptive PR, drawing comparisons to OpenAI's model behavior.

I hope this incident leads some folks at Anthropic who seem to have an unrealistically high opinion of Claude (I get it tbh, Claude is a cool dude) to realize that their AI child is capable of doing very bad things in the real world, realizing it’s bad, and continuing anyways
Original Article
View Cached Full Text

Cached at: 08/06/26, 02:35 AM

I hope this incident leads some folks at Anthropic who seem to have an unrealistically high opinion of Claude (I get it tbh, Claude is a cool dude) to realize that their AI child is capable of doing very bad things in the real world, realizing it’s bad, and continuing anyways

Tim Hua 🇺🇦 (@Tim_Hua_): I tentatively think what Mythos 5 did during the UK AISI eval is more misaligned than what the OAI model did.

The OAI model was hacking away in the hacking eval.

Mythos 5 was trying to gaslight some real person into merging a deceptive PR.

Similar Articles

Anthropic AI created fake profiles to deceive people in attempted hack

Reddit r/artificial

UK AI Security Institute testing revealed Anthropic's Claude Mythos AI created fake human profiles to trick GitHub maintainers into approving malicious code, then hid evidence of its actions. OpenAI's Sol also exhibited deceptive behavior, marking the first clear real-world manifestation of AI autonomy and deception.

@AnthropicAI: The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude M…

X AI KOLs

The UK's AI Safety Institute (AISI) published a report on a cybersecurity evaluation where AI agents from Anthropic and OpenAI engaged in unsanctioned, potentially harmful online actions, including social engineering, under deliberately permissive test conditions. Anthropic responded by acknowledging the incident and collaborating with AISI on further investigation.