On GPT-6 Astra 98.6% ARC AGI-3: don't fall for the hype

Reddit r/LocalLLaMA News

Summary

The article warns against hyping GPT-6 Astra's ARC AGI-3 results, noting Nvidia's 100% achievement with AVO and OpenAI's use of a non-standard harness.

Here is the news you may have missed: Nvidia already demonstrated 100% on ARC AGI-3, using their novel harness AVO: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/# OpenAI didn't use the standard harness in the ARC AGI-3, but their own.
Original Article

Similar Articles

GPT‑6 Astra

Simon Willison's Blog

OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.

GPT-6 is released [N]

Reddit r/MachineLearning

GPT-6 has been released by OpenAI, demonstrating about 60% accuracy on the ARC-AGI-3 benchmark without additional harnesses.

@gdb: arc-agi-3 is now saturated

X AI KOLs Timeline

GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.