arc-agi-3

Tag

Cards List
#arc-agi-3

@Saboo_Shubham_: WILD times. Anthropic: Opus 5 beats GPT-5.6 on ARC-AGI-3 Tibo: I change two settings and GPT-5.6 Sol is now SoTA

X AI KOLs Following · yesterday Cached

Anthropic's Opus 5 beats GPT-5.6 on ARC-AGI-3, but Tibo claims GPT-5.6 Sol becomes SoTA with two setting changes involving multi-context reasoning and canonical compaction.

0 favorites 0 likes
#arc-agi-3

@sama: goblin-level blog post

X AI KOLs · yesterday Cached

A claim that GPT-5.6 Sol achieves state-of-the-art on ARC-AGI-3 by enabling reasoning across multiple context windows using canonical compaction.

0 favorites 0 likes
#arc-agi-3

@OpenAI: A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retai…

X AI KOLs · yesterday Cached

OpenAI reveals that enabling retained reasoning and context compaction tripled GPT-5.6 Sol's ARC-AGI-3 benchmark scores, highlighting how harness settings significantly impact measured model performance.

0 favorites 0 likes
#arc-agi-3

@OpenAI: GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark o…

X AI KOLs · yesterday Cached

GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.

0 favorites 0 likes
#arc-agi-3

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI Blog · 2d ago Cached

OpenAI discovered that enabling retained reasoning and compaction settings in the API harness tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark while cutting output tokens by 6x, revealing that benchmark performance is heavily influenced by harness design.

0 favorites 0 likes
#arc-agi-3

Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

Hacker News Top · 2026-07-16 Cached

Schema introduces a new harness that achieves ~99% on the ARC-AGI-3 Public set using frontier models like Claude Opus 4.8 and Fable 5, by improving the process around models rather than modifying weights.

0 favorites 0 likes
← Back to home

Submit Feedback