@gdb: arc-agi-3 is now saturated

X AI KOLs Timeline Models

Summary

GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.

arc-agi-3 is now saturated
Original Article
View Cached Full Text

Cached at: 09/04/26, 12:23 PM

arc-agi-3 is now saturated

ARC Prize (@arcprize): GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:

  • Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
  • It surpasses human performance on 96% of ARC-AGI-3 levels
  • It builds the most precise symbolic model of novel environments we’ve seen

Our analysis:

Similar Articles

GPT‑6 Astra

Simon Willison's Blog

OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.

GPT-6 is released [N]

Reddit r/MachineLearning

GPT-6 has been released by OpenAI, demonstrating about 60% accuracy on the ARC-AGI-3 benchmark without additional harnesses.

ChatGPT 5.6 - ARC-AGI 3 score

Reddit r/singularity

ChatGPT 5.6 achieved a new score on the ARC-AGI benchmark, indicating progress toward general intelligence.

GPT-5.6 Series (2 minute read)

TLDR AI

OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.