@dawnsongtweets: Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is …
Summary
This article introduces Agents' Last Exam (ALE), a rolling benchmark designed to test whether AI agents can perform economically valuable work. Evaluations on frontier models like Fable 5 show 0% success on the hardest tasks, indicating that truly job-ready agents are not yet here.
View Cached Full Text
Cached at: 06/12/26, 04:54 AM
Everyone says the latest AI agents will be “job-ready” soon, especially after the release of Fable 5 this week. But is that really the case?
Over the past many months, my group and collaborators have been building Agents’ Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work.
My group and collaborators previously have created many of the benchmarks the field runs on, including MMLU, MATH, CyberGym, and ExploitGym. Today, I’m excited to share Agents’ Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains.
With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations. The result is both impressive and sobering.
Today’s agents can solve a meaningful fraction of professional tasks. But when we look at the hardest tasks, the ones requiring sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.
On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate. The age of useful agents is here.
The age of truly job-ready agents is not.
We hope Agents’ Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains.
Similar Articles
Agents' Last Exam
Introduces Agents' Last Exam (ALE), a benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters with over 1000 tasks, revealing a large gap between benchmark performance and practical deployment.
@rohanpaul_ai: Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest. This paper pr…
This paper introduces Agents' Last Exam, a benchmark that tests AI agents on real expert work across 55 digital work areas. Current best agents fail most tasks, averaging only 2.6% pass rate on the hardest tier, revealing a large gap between benchmark scores and real-world automation readiness.
@dair_ai: // Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 2…
Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks designed to evaluate AI agents on real-world workflows, with a current full pass rate of only 2.6% on its hardest tier.
How ready are AI agents for real-world work?
A discussion of how ready AI agents are for real-world work, covering their current abilities and the key open questions around reliability, permissions, failures, and human oversight.
Can Agents Use a Computer Yet? We've Got the Data (17 minute read)
a16z examines progress in computer-using AI agents, citing benchmark improvements that now surpass human-level performance on OSWorld-Verified and noting a shift from raw capability to reliable production deployment in enterprise workflows.