25% difference on a benchmark – just a mistake on how you run it😬

Reddit r/AI_Agents News

Summary

Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.

Not sure if you've seen the discussion online about a problem OpenAI ran into – they even had to write a separate article and do research on it. On the ARC-AGI benchmark, which has very hard logic puzzles, their best model GPT-5.6 Sol was scoring way less than the recent Opus 5 models from Anthropic. Turns out that on this benchmark, with the simplest harness (in combination with the default API), the model in that sorry state scores something like 13%, but if you fix it – you get 38%, which is actually more than Anthropic. Anthropic, accordingly, spread this around, and after that OpenClaw creator (still works at OpenAI) wrote something like: hey, maybe you guys should check how you publish benchmark comparisons 🙃 I think the problem is actually completely real – in terms of how the harness affects the quality of results. It's like in our practice, with a client who's never used our product. They use it, get suboptimal results, and the first thing we learn – the client completely misunderstood how to use the product. And when the product is specialized, it's not at all about the UX, where everything gets sorted out quickly – there you have to figure out how the agent inside understood all this, how it works with context, with data, how the data is fed in, how the agent is launched. This is why I get really annoyed by benchmarks where raw models are tested with practically no harness – you just call the model that pulls the tools. What difference does it make to me in principle if Opus works better, but in the Claude Code harness (or actually any real world harness) it works worse than Codex? Doesn't matter to me, because in my view this is no longer a representative comparison. Nobody's going to use the model in a simple harness if everyone's using it in coding agents. But the industry still shows benchmarks this way. I think it's about time we moved forward a bit. This default way of evaluating confuses more than it informs.
Original Article

Similar Articles

What GPT-6 Astra’s 99.9% ARC-AGI-3 Score Actually Measures

Reddit r/artificial

The article analyzes GPT-6 Astra's 99.9% score on the ARC-AGI-3 benchmark, revealing that the score varies based on testing harnesses and questioning OpenAI's AGI interpretation amid clarifications from the benchmark creators.