@Saboo_Shubham_: WILD times. Anthropic: Opus 5 beats GPT-5.6 on ARC-AGI-3 Tibo: I change two settings and GPT-5.6 Sol is now SoTA
Summary
Anthropic's Opus 5 beats GPT-5.6 on ARC-AGI-3, but Tibo claims GPT-5.6 Sol becomes SoTA with two setting changes involving multi-context reasoning and canonical compaction.
View Cached Full Text
Cached at: 07/30/26, 01:48 AM
WILD times.
Anthropic: Opus 5 beats GPT-5.6 on ARC-AGI-3 Tibo: I change two settings and GPT-5.6 Sol is now SoTA https://t.co/iQl3GJJXCe
Tibo (@thsottiaux): Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3.
Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
Similar Articles
@Saboo_Shubham_: INSANE. This gets better ever other week. Opus 5 as advisor, GPT-5.6 as orchestrator, and Gemini 3.6 Flash is the worke…
Shubham Saboo discusses a speculative AI system where Opus 5 acts as advisor, GPT-5.6 as orchestrator, and Gemini 3.6 Flash as worker to solve complex problems.
@sama: goblin-level blog post
A claim that GPT-5.6 Sol achieves state-of-the-art on ARC-AGI-3 by enabling reasoning across multiple context windows using canonical compaction.
I asked Sol Max to compare the output of Claude Opus 5 High and GPT 5.6 Sol Max on a specific puzzle on ARC-AGI-3 where Opus 5 had a 98.81% score and GPT 5.6 Sol Max had a 21.42% score
An analysis comparing Claude Opus 5 High and GPT 5.6 Sol Max on an ARC-AGI-3 puzzle shows Opus winning by preserving detailed state in visible output, while Sol relies on discarded hidden reasoning.
@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.
@VraserX: Okay, GPT-5.6 is way bigger than I expected. Sol looks insane, Terra and Luna make the cost side actually scary, and Ul…
A user expresses amazement at the scale of GPT-5.6, noting impressive capabilities like Sol, Terra, Luna, and Ultra mode with parallel multi-agent execution, suggesting a major advancement.