humanity's last exam current benchmarks thoughts?
Summary
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
Similar Articles
Fable passes the "When A.I. Passes This Test, Look Out" test
Claude Fable achieves 53% on the 'Humanity's Last Exam' benchmark, surpassing the expected end-of-2025 milestone earlier than projected, indicating rapid AI progress.
What happens after all AI hit % 100 on benchmarks
The article speculates on what will happen when all AI models achieve 100% on benchmarks, questioning how they will demonstrate superiority.
Human Bench 3.0, Humanity-2 AI-0
Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.
AI can finally pass the Turing Test better than a human, study warns
A new study published in PNAS shows that advanced LLMs like GPT-4.5 can pass the Turing Test, with participants finding them more human than actual humans, prompting a reevaluation of what the test measures.
One AI just scored 1753 on a test where 'human expert' is 1000. Here's why I don't fully trust that number
The article critiques a viral AI benchmark that claims Grok 4.6 scored 1753 vs 1000 for human experts, highlighting that the test uses preference-based comparisons between AI outputs rather than objective correctness, so polished-looking work may win without being truly better.