OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.
GPT-5.6 Sol is the standout model of the GPT-5.6 family. It is the first model to win an ARC-AGI-3 public game. The model is able to read an unfamiliar scene correctly and in the game's own vocabulary. It won the game on ARC-AGI because it correctly oriented itself in a new environment first.
# GPT-5.6 - ARC-AGI Results
Source: [https://arcprize.org/results/openai-gpt-5-6](https://arcprize.org/results/openai-gpt-5-6)
## GPT\-5\.6Series
OpenAI·Jul 9, 2026·3models·15reasoning variants
GPT\-5\.6 Sol is the standout model of the GPT\-5\.6 family\. Sol at max reasoning effort is the only performant model \(as of July 2026\) averaging 13\.33% on Public and 7\.78% on Semi\-Private\. It is the first model to win an ARC\-AGI\-3 public game \(ft09, 87%\)\. Sol is able to read an unfamiliar scene correctly and in the game's own vocabulary\. It treats a failed hypothesis as a reason to re\-plan rather than thrash\. Most agent failures are upstream of the code they write or the action they take\. Sol is able to perform on ARC\-AGI not because it executes better, but because it correctly orients itself in a new environment first\.
## ARC\-AGI 3leaderboard
GPT\-5\.6 SolGPT\-5\.6 TerraGPT\-5\.6 Luna
## Verified scores
ModelVariantARC\-AGI\-1ARC\-AGI\-2ARC\-AGI\-3[Sol](https://arcprize.org/results/openai-gpt-5-6-sol)Max96\.5%
92\.5%
7\.8%
Extra High97\.5%
90\.0%
7\.0%
High97\.0%
85\.4%
2\.1%
Medium92\.5%
67\.1%
1\.1%
Low74\.5%
42\.5%
0\.3%
[Terra](https://arcprize.org/results/openai-gpt-5-6-terra)Max96\.5%
83\.9%
0\.8%
Extra High94\.0%
74\.2%
0\.7%
High92\.0%
67\.1%
0\.5%
Medium77\.0%
37\.5%
0\.1%
Low60\.2%
18\.8%
0\.0%
[Luna](https://arcprize.org/results/openai-gpt-5-6-luna)Max88\.0%
59\.5%
0\.2%
Extra High87\.7%
47\.6%
0\.0%
High76\.5%
29\.3%
0\.1%
Medium56\.5%
7\.4%
0\.2%
Low34\.2%
5\.1%
0\.2%
## Tasks & environments
Pass/fail per reasoning level across each benchmark\. Hardest tasks \(fewest levels solving\) are listed first\.
Model
### ARC\-AGI\-3 Public Demo
25environments
### ARC\-AGI\-2 Public Eval
120tasks
### ARC\-AGI\-1 Public Eval
400tasks
[All results](https://arcprize.org/results)
GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.
OpenAI's GPT-5.6 Sol is benchmarked as their best vision model yet, showing significant improvements in object detection and other visual tasks compared to previous models like GPT-5.5.
OpenAI's blog post describes how GPT-5.6 Sol, a new frontier model, uses self-optimization to improve its own inference efficiency while maintaining high intelligence.
OpenAI unveils GPT-5.6 Sol, a flagship model for long-running autonomous work across applications and enterprise data, featuring Ultra mode with sub-agents for faster, stronger results. The model was used internally to help train Luna and demonstrates significant cost and performance improvements over previous versions.