humanity's last exam current benchmarks thoughts?
Summary
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
Similar Articles
Fable passes the "When A.I. Passes This Test, Look Out" test
Claude Fable achieves 53% on the 'Humanity's Last Exam' benchmark, surpassing the expected end-of-2025 milestone earlier than projected, indicating rapid AI progress.
What happens after all AI hit % 100 on benchmarks
The article speculates on what will happen when all AI models achieve 100% on benchmarks, questioning how they will demonstrate superiority.
AI can finally pass the Turing Test better than a human, study warns
A new study published in PNAS shows that advanced LLMs like GPT-4.5 can pass the Turing Test, with participants finding them more human than actual humans, prompting a reevaluation of what the test measures.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.
@cline: 5 months ago the highest score on Artificial Analysis Intelligence Index was 51 (GPT-5.4 xhigh). This week DeepSeek V4-…
A tweet notes that DeepSeek V4-Flash scored 50 on the Artificial Analysis Intelligence Index, close to GPT-5.4's 51 from five months ago, and predicts local models will become the majority choice within two years.