Updated version of Humanity’s Last Exam: HLE-Diamond

Reddit r/singularity Tools

Summary

An updated version of the Humanity's Last Exam, a benchmark for AI evaluation, named HLE-Diamond, has been announced and released.

Link to tweet: https://x.com/CAIS/status/2102787839964729431 Link to blog: https://lastexam.ai/blog/hle-diamond
Original Article

Similar Articles

Human Bench 3.0, Humanity-2 AI-0

Reddit r/singularity

Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.

humanity's last exam current benchmarks thoughts?

Reddit r/singularity

Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.

New benchmark dropped

Reddit r/singularity

A new benchmark has been released, likely for evaluating AI or software performance.

Agents' Last Exam

Hugging Face Daily Papers

Introduces Agents' Last Exam (ALE), a benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters with over 1000 tasks, revealing a large gap between benchmark performance and practical deployment.