@rohanpaul_ai: Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest. This paper pr…
Summary
This paper introduces Agents' Last Exam, a benchmark that tests AI agents on real expert work across 55 digital work areas. Current best agents fail most tasks, averaging only 2.6% pass rate on the hardest tier, revealing a large gap between benchmark scores and real-world automation readiness.
View Cached Full Text
Cached at: 06/11/26, 03:39 PM
Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest.
This paper proposes a Agents’ Last Exam, a benchmark that asks AI agents to finish real expert work, and today’s agents mostly fail.
Even strong agents of today are nowhere near reliable on the hardest real workflows, which means benchmark success has not yet become broad workplace capability.
So this paper shifts the question from “can AI answer hard questions?” to “can AI complete real work that people get paid to do?”
Most of today’s AI benchmarks show impressive scores, but they do not prove that agents can finish useful work in real jobs.
Agents’ Last Exam tries to fix this by testing agents on long tasks from 55 digital work areas, including engineering, finance, medicine, law, media, and science.
The tasks come from experts’ real completed projects, and the agent must use normal computer tools like files, browsers, command lines, and desktop software to produce a finished result.
The authors tested many current agent systems and models, then scored their finished work with automatic checks or strict rubrics instead of loose human opinions.
The main result is that today’s best systems still struggle badly, with an average full pass rate of only 2.6% on the hardest tier.
Link – arxiv. org/abs/2606.05405
Title: “Agents’ Last Exam”
Similar Articles
@dair_ai: // Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 2…
Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks designed to evaluate AI agents on real-world workflows, with a current full pass rate of only 2.6% on its hardest tier.
@dawnsongtweets: Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is …
This article introduces Agents' Last Exam (ALE), a rolling benchmark designed to test whether AI agents can perform economically valuable work. Evaluations on frontier models like Fable 5 show 0% success on the hardest tasks, indicating that truly job-ready agents are not yet here.
@rohanpaul_ai: Arena just released a real-world agent leaderboard that ranks AI models by how well they complete actual user jobs, not…
Agent Arena is a new leaderboard that evaluates AI models on real-world agentic tasks such as coding, research, and file analysis, using signals like task success, steerability, and recovery, with GPT-5.5 High leading.
@OkhayIea: Everyone's racing to build "AI scientists." So we asked a blunt question: Can today's best coding agents beat the publi…
Introduces NatureBench, a cross-disciplinary benchmark of 90 tasks from Nature papers to test AI coding agents, finding the best agent (Claude Opus 4.7) surpasses SOTA on only 17.8% of tasks and often succeeds by reducing science to supervised ML rather than genuine discovery.
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.