Tag
This article introduces Agents' Last Exam (ALE), a rolling benchmark designed to test whether AI agents can perform economically valuable work. Evaluations on frontier models like Fable 5 show 0% success on the hardest tasks, indicating that truly job-ready agents are not yet here.