GPT-6 Sol & Astra dominate the ARC-AGI-3 leaderboard
Summary
GPT-6 Sol and Astra have taken the top spots on the ARC-AGI-3 leaderboard, marking a notable advance in abstract reasoning benchmarks for frontier AI models.
Similar Articles
@gdb: arc-agi-3 is now saturated
GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.
GPT-5.6 Series (2 minute read)
OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.
Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.
GPT‑6 Astra
OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.
GPT-6.1 Sol - Apparently near-Astra performance for complex work at a lower cost.
OpenAI releases GPT-6.1 Sol, an AI model with near-Astra performance at lower cost for complex coding, computer use, and professional work, supporting various reasoning efforts.