The prevalent problem of misleading benchmark reporting (re: Astra)
Summary
OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.
Similar Articles
Astra benchmarks from the OpenAI blog before it was taken down
The article discusses performance benchmarks for OpenAI's Astra system that were published on their blog and subsequently removed, indicating leaked or unpublished technical data.
25% difference on a benchmark – just a mistake on how you run it😬
Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.
On GPT-6 Astra 98.6% ARC AGI-3: don't fall for the hype
The article warns against hyping GPT-6 Astra's ARC AGI-3 results, noting Nvidia's 100% achievement with AVO and OpenAI's use of a non-standard harness.
Arena.ai is running possibly the most fraudulent benchmark thus far
The article criticizes Arena.ai for allegedly running dishonest benchmarks, claiming it ranked GPT 5.5 below Meta's Muse Spark in coding and Grok Imagine above Seedance in video generation, which the author asserts is objectively false.
Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.