@OpenAI: GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark o…

X AI KOLs News

Summary

GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what it had learned. We found that enabling two API settings tripled our scores with 6x fewer output tokens.
Original Article
View Cached Full Text

Cached at: 07/30/26, 01:46 AM

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games?

We investigated. The harness was not letting it remember what it had learned.

We found that enabling two API settings tripled our scores with 6x fewer output tokens.

Similar Articles

GPT-5.6 Series (2 minute read)

TLDR AI

OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.

5.6 Sol is underhyped for general work (7 minute read)

TLDR AI

OpenAI unveils GPT-5.6 Sol, a flagship model for long-running autonomous work across applications and enterprise data, featuring Ultra mode with sub-agents for faster, stronger results. The model was used internally to help train Luna and demonstrates significant cost and performance improvements over previous versions.