Claude Opus 5 + Claude Code + 1 Skill Scores 100% on ARC AGI 3 (public set)
Summary
Claude Opus 5, along with Claude Code and a skill, scored 100% on the ARC AGI 3 benchmark's public set, suggesting the benchmark may not be as challenging as thought.
Blog post: https://arc-skill.vercel.app/ Looks like the benchmark isn’t that hard after all.
Similar Articles
Claude Opus 4.8 scores over 1% on ARC-AGI 3 !!
Claude Opus 4.8 achieves a score of over 1% on the ARC-AGI 3 benchmark, demonstrating slight progress on a difficult AI reasoning test.
@cline: Claude Opus 5 takes #1 on SWE-Bench at 97%, and claims Fable 5 level intelligence at half the price. Incredible that in…
Anthropic released Claude Opus 5, achieving 97% on SWE-Bench and claiming Fable 5-level intelligence at half the price, marking rapid SOTA improvement.
Claude Opus 5 BENCHMARKS!
An article presenting benchmark results for the upcoming Claude Opus 5 AI model.
Claude Fable 5 gets 65 on Artificial Analysis
Claude Fable 5 achieved a score of 65 on the Artificial Analysis intelligence index.
Opus 5 ARC AGI score was benchmaxxed
Opus 5 achieved a high score on the ARC AGI benchmark, indicating advanced reasoning capabilities.