@sama: goblin-level blog post
Summary
A claim that GPT-5.6 Sol achieves state-of-the-art on ARC-AGI-3 by enabling reasoning across multiple context windows using canonical compaction.
View Cached Full Text
Cached at: 07/30/26, 01:47 AM
goblin-level blog post
Tibo (@thsottiaux): Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3.
Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
Similar Articles
GPT-5.6 Series (2 minute read)
OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.
@sama: obviously the best model we have ever produced, but also one of the best blog posts we have ever produced:
OpenAI announces the GPT-5.6 family including Sol, Terra, and Luna, claiming state-of-the-art performance across coding, knowledge work, and science with higher efficiency and lower cost.
@Saboo_Shubham_: WILD times. Anthropic: Opus 5 beats GPT-5.6 on ARC-AGI-3 Tibo: I change two settings and GPT-5.6 Sol is now SoTA
Anthropic's Opus 5 beats GPT-5.6 on ARC-AGI-3, but Tibo claims GPT-5.6 Sol becomes SoTA with two setting changes involving multi-context reasoning and canonical compaction.
@OpenAI: GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark o…
GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.
@gdb: arc-agi-3 is now saturated
GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.