Updated version of Humanity’s Last Exam: HLE-Diamond
Summary
An updated version of the Humanity's Last Exam, a benchmark for AI evaluation, named HLE-Diamond, has been announced and released.
Similar Articles
Human Bench 3.0, Humanity-2 AI-0
Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.
humanity's last exam current benchmarks thoughts?
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
New benchmark dropped
A new benchmark has been released, likely for evaluating AI or software performance.
Fable passes the "When A.I. Passes This Test, Look Out" test
Claude Fable achieves 53% on the 'Humanity's Last Exam' benchmark, surpassing the expected end-of-2025 milestone earlier than projected, indicating rapid AI progress.
Agents' Last Exam
Introduces Agents' Last Exam (ALE), a benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters with over 1000 tasks, revealing a large gap between benchmark performance and practical deployment.