@gdb: arc-agi-3 现已饱和

X AI KOLs Timeline 模型

摘要

OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 基准测试中取得最先进的性能,得分63%,在96%的层级上超越人类表现,展示了先进的符号建模能力。

arc-agi-3 现已饱和
查看原文
查看缓存全文

缓存时间: 2026/09/04 12:23

arc-agi-3 已达饱和

ARC Prize (@arcprize): GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:

  • Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
  • It surpasses human performance on 96% of ARC-AGI-3 levels
  • It builds the most precise symbolic model of novel environments we’ve seen

Our analysis:

相似文章

GPT‑6 Astra

Simon Willison's Blog

OpenAI 发布 GPT‑6 Astra,在安全任务和长上下文处理方面表现出色,在 ARC-AGI 3 上达到 99.9%,尽管在某些基准测试上仍落后于 Claude Fable。

GPT-6 已发布 [N]

Reddit r/MachineLearning

OpenAI 发布了 GPT-6,在没有额外辅助工具的情况下,在 ARC-AGI-3 基准测试中展示了约 60% 的准确率。

GPT-5.6 系列 (2分钟阅读)

TLDR AI

OpenAI发布了GPT-5.6系列模型,其中Sol是突出模型,在ARC-AGI-3公开测试中取得13.33%的成绩,并成为首个赢得游戏的模型,展示了在陌生环境中自我定位能力的提升。