25% difference on a benchmark – just a mistake on how you run it😬
Summary
Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.
Similar Articles
The prevalent problem of misleading benchmark reporting (re: Astra)
OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.
@OpenAI: A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retai…
OpenAI reveals that enabling retained reasoning and context compaction tripled GPT-5.6 Sol's ARC-AGI-3 benchmark scores, highlighting how harness settings significantly impact measured model performance.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
OpenAI discovered that enabling retained reasoning and compaction settings in the API harness tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark while cutting output tokens by 6x, revealing that benchmark performance is heavily influenced by harness design.
What GPT-6 Astra’s 99.9% ARC-AGI-3 Score Actually Measures
The article analyzes GPT-6 Astra's 99.9% score on the ARC-AGI-3 benchmark, revealing that the score varies based on testing harnesses and questioning OpenAI's AGI interpretation amid clarifications from the benchmark creators.
@OpenAI: GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark o…
GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.