@gdb: arc-agi-3 is now saturated
Summary
GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.
View Cached Full Text
Cached at: 09/04/26, 12:23 PM
arc-agi-3 is now saturated
ARC Prize (@arcprize): GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we’ve seen
Our analysis:
Similar Articles
Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.
GPT‑6 Astra
OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.
GPT-6 is released [N]
GPT-6 has been released by OpenAI, demonstrating about 60% accuracy on the ARC-AGI-3 benchmark without additional harnesses.
ChatGPT 5.6 - ARC-AGI 3 score
ChatGPT 5.6 achieved a new score on the ARC-AGI benchmark, indicating progress toward general intelligence.
GPT-5.6 Series (2 minute read)
OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.