After Astra's stealth nerf last night, we really need benchmarks to do a re-bench 1 week after any model release. This is ridiculous.
Summary
Users report that the Astra AI model has been secretly downgraded, leading to poor performance despite high-effort prompting, with concerns about compute cost savings and calls for FTC investigation.
Similar Articles
Astra benchmarks from the OpenAI blog before it was taken down
The article discusses performance benchmarks for OpenAI's Astra system that were published on their blog and subsequently removed, indicating leaked or unpublished technical data.
The prevalent problem of misleading benchmark reporting (re: Astra)
OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.
OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust
The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.
Okay Astra is insane.
The user shares their experience where Astra quickly fixed bugs in a dashboard, outperforming Gemini 3.1 pro, and praises OpenAI's model capabilities.
OpenAI launches Astra, its powerful (and controversial) new model
OpenAI has released Astra, its latest and most powerful AI model, known for its strong cybersecurity and coding abilities but also controversial due to its opaque reasoning techniques.