@rohanpaul_ai: This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed fee…
Summary
A paper evaluates eight leading AI models on long-horizon tasks, finding that even the best-performing model achieves only 27.3% of human performance, highlighting significant limitations for dependable long-horizon AI execution.
View Cached Full Text
Cached at: 08/29/26, 03:55 PM
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.
A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
Rohan Paul (@rohanpaul_ai): “Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work.”
- Chamath at Stanford
Similar Articles
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
This paper systematically evaluates seven frontier AI agents on long-horizon tasks, revealing they function more as engineering optimizers than autonomous researchers, with recommendations for improving training and experience management.
@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…
This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.
@dair_ai: Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficul…
AutoLab is a new benchmark evaluating 17 frontier models on 36 expert-curated long-horizon tasks (system optimization, model development, CUDA kernels, puzzles), finding that persistence—not initial attempt quality—is the dominant predictor of success. Claude-opus-4.6 led all categories, while most other models terminated prematurely or exhausted budgets with minimal progress.
@rohanpaul_ai: Univ of Texas paper shows AI agents can slowly become less reliable after deployment, even when the model itself does n…
A University of Texas paper introduces AgingBench, a benchmark that reveals AI agents can become less reliable after deployment due to memory and maintenance decay, even when the underlying model remains unchanged.
@rohanpaul_ai: Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest. This paper pr…
This paper introduces Agents' Last Exam, a benchmark that tests AI agents on real expert work across 55 digital work areas. Current best agents fail most tasks, averaging only 2.6% pass rate on the hardest tier, revealing a large gap between benchmark scores and real-world automation readiness.