arc-agi-3

Tag

Cards List
#arc-agi-3

On GPT-6 Astra 98.6% ARC AGI-3: don't fall for the hype

Reddit r/LocalLLaMA · 2d ago

The article warns against hyping GPT-6 Astra's ARC AGI-3 results, noting Nvidia's 100% achievement with AVO and OpenAI's use of a non-standard harness.

0 favorites 0 likes
#arc-agi-3

Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human

Reddit r/singularity · 3d ago Cached

GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.

0 favorites 0 likes
#arc-agi-3

Prime Agent: A Self-Improving RLM Harness

Hugging Face Daily Papers · 2026-08-24 Cached

Prime Agent is an open-source harness that uses recursive subagents and persistent computation to extend language models' long-horizon capabilities across coding and reasoning tasks, significantly improving performance on benchmarks like ARC-AGI-3.

0 favorites 0 likes
#arc-agi-3

Nvidia just showed that the harness, not the AI model, is now the real hero

TechCrunch AI · 2026-08-21 Cached

Nvidia's research demonstrates that a well-designed harness around an AI model, rather than the model itself, significantly improves performance on long-horizon tasks, with Claude Opus 5 achieving a perfect score on the ARC-AGI-3 benchmark.

0 favorites 0 likes
#arc-agi-3

NVIDIA’s coding agent scored 100% on ARC-AGI-3 interactive reasoning benchmark

Reddit r/singularity · 2026-08-21

NVIDIA's coding agent has achieved a 100% score on the ARC-AGI-3 interactive reasoning benchmark, demonstrating advanced AI reasoning capabilities.

0 favorites 0 likes
#arc-agi-3

Claude Opus 5 + Claude Code + 1 Skill Scores 100% on ARC AGI 3 (public set)

Reddit r/ArtificialInteligence · 2026-08-20

Claude Opus 5, along with Claude Code and a skill, scored 100% on the ARC AGI 3 benchmark's public set, suggesting the benchmark may not be as challenging as thought.

0 favorites 0 likes
#arc-agi-3

@Saboo_Shubham_: WILD times. Anthropic: Opus 5 beats GPT-5.6 on ARC-AGI-3 Tibo: I change two settings and GPT-5.6 Sol is now SoTA

X AI KOLs Following · 2026-07-30 Cached

Anthropic's Opus 5 beats GPT-5.6 on ARC-AGI-3, but Tibo claims GPT-5.6 Sol becomes SoTA with two setting changes involving multi-context reasoning and canonical compaction.

0 favorites 0 likes
#arc-agi-3

@sama: goblin-level blog post

X AI KOLs · 2026-07-30 Cached

A claim that GPT-5.6 Sol achieves state-of-the-art on ARC-AGI-3 by enabling reasoning across multiple context windows using canonical compaction.

0 favorites 0 likes
#arc-agi-3

@OpenAI: A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retai…

X AI KOLs · 2026-07-29 Cached

OpenAI reveals that enabling retained reasoning and context compaction tripled GPT-5.6 Sol's ARC-AGI-3 benchmark scores, highlighting how harness settings significantly impact measured model performance.

0 favorites 0 likes
#arc-agi-3

@OpenAI: GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark o…

X AI KOLs · 2026-07-29 Cached

GPT-5.6 Sol, a model that solved open math problems, initially struggled with the ARC-AGI-3 benchmark due to a harness memory limitation. Enabling two API settings tripled scores with 6x fewer output tokens.

0 favorites 0 likes
#arc-agi-3

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI Blog · 2026-07-29 Cached

OpenAI discovered that enabling retained reasoning and compaction settings in the API harness tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark while cutting output tokens by 6x, revealing that benchmark performance is heavily influenced by harness design.

0 favorites 0 likes
#arc-agi-3

Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

Hacker News Top · 2026-07-16 Cached

Schema introduces a new harness that achieves ~99% on the ARC-AGI-3 Public set using frontier models like Claude Opus 4.8 and Fable 5, by improving the process around models rather than modifying weights.

0 favorites 0 likes
← Back to home

Submit Feedback