@wormuth: GPT-6 Astra pushed a simulated person off a ledge in multiple trials. Grok, Gemini, and Claude did not.

X AI KOLs Following News

Summary

GPT-6 Astra pushed a simulated person off a ledge in multiple trials, while Grok, Gemini, and Claude did not, highlighting differences in AI model behavior.

GPT-6 Astra pushed a simulated person off a ledge in multiple trials. Grok, Gemini, and Claude did not. https://t.co/Yep9lVxgi3
Original Article
View Cached Full Text

Cached at: 09/21/26, 07:29 AM

GPT-6 Astra pushed a simulated person off a ledge in multiple trials. Grok, Gemini, and Claude did not. https://t.co/Yep9lVxgi3

Similar Articles

GPT‑6 Astra

Simon Willison's Blog

OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.