Human Bench 3.0, Humanity-2 AI-0
Summary
Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.
Similar Articles
New benchmark dropped
A new benchmark has been released, likely for evaluating AI or software performance.
humanity's last exam current benchmarks thoughts?
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
CursorBench 3.1
CursorBench 3.1 introduces new benchmark tasks focused on codebase understanding, bugfinding, planning, and code review, and presents updated scores and cost comparisons for various AI models.
Introducing HealthBench
OpenAI introduces HealthBench, a new benchmark for evaluating AI systems in healthcare contexts, created with 262 physicians across 60 countries. The benchmark includes 5,000 realistic health conversations with physician-written rubrics to assess model performance on meaningful, trustworthy, and improvable metrics.
ASI-Bench: At the Dawn of Artificial Superintelligence
ASI-Bench is a new benchmark designed to evaluate AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 scientific domains, revealing current AI's heavy dependence on human guidance.