OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust
Summary
The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.
Similar Articles
The prevalent problem of misleading benchmark reporting (re: Astra)
OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.
What GPT-6 Astra’s 99.9% ARC-AGI-3 Score Actually Measures
The article analyzes GPT-6 Astra's 99.9% score on the ARC-AGI-3 benchmark, revealing that the score varies based on testing harnesses and questioning OpenAI's AGI interpretation amid clarifications from the benchmark creators.
Astra benchmarks from the OpenAI blog before it was taken down
The article discusses performance benchmarks for OpenAI's Astra system that were published on their blog and subsequently removed, indicating leaked or unpublished technical data.
OpenAI launches Astra, its powerful (and controversial) new model
OpenAI has released Astra, its latest and most powerful AI model, known for its strong cybersecurity and coding abilities but also controversial due to its opaque reasoning techniques.
OpenAI's AGI number came from a harness, not the model (6 minute read)
OpenAI's reported 99.9% AGI benchmark score for GPT-6 Astra was achieved using a specific harness (Provider Adapter), while the standard harness yields a much lower 62.7% score, highlighting significant issues in AI model evaluation and benchmarking transparency.