Tag
A skeptical deep dive finds numerous errors in the Humanity's Last Exam benchmark, with the official o3-mini grader incorrectly marking correct answers as wrong.
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
Claude Fable achieves 53% on the 'Humanity's Last Exam' benchmark, surpassing the expected end-of-2025 milestone earlier than projected, indicating rapid AI progress.