OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust

Reddit r/ArtificialInteligence News

Summary

The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.

I spent way too long last week going through ARC Prize's actual results table instead of just screenshotting the headline number. Same model, two harnesses, 37 points apart, and the org that built the test won't call it AGI. Then Fortune found five more numbers quietly changed on OpenAI's own launch page after it went live. Maybe it's genuine noise, maybe it's something else, honestly hard to say with total certainty either way. Wrote up the whole harness breakdown plus the Llama 4 precedent. Here's what I found. https://srutiosocial.com/gpt-6-astra-arc-agi-3-score-explained/
Original Article

Similar Articles

What GPT-6 Astra’s 99.9% ARC-AGI-3 Score Actually Measures

Reddit r/artificial

The article analyzes GPT-6 Astra's 99.9% score on the ARC-AGI-3 benchmark, revealing that the score varies based on testing harnesses and questioning OpenAI's AGI interpretation amid clarifications from the benchmark creators.