@rohanpaul_ai: Current agent benchmarks may be ending before the real failures start. FM-Bench shows that a model can look strong afte…
Summary
FM-Bench is a benchmark for evaluating long-horizon AI agents, revealing that short-term performance does not guarantee long-term success in simulated management tasks over 20 years.
View Cached Full Text
Cached at: 09/03/26, 04:05 AM
Current agent benchmarks may be ending before the real failures start.
FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations.
FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future.
The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th.
Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever.
What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board.
So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world.
– arxiv. org/abs/2608.18423
Title: “FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents”
Similar Articles
@rohanpaul_ai: Today’s frontier agents are far less ready for real-world automation than their benchmark scores suggest. This paper pr…
This paper introduces Agents' Last Exam, a benchmark that tests AI agents on real expert work across 55 digital work areas. Current best agents fail most tasks, averaging only 2.6% pass rate on the hardest tier, revealing a large gap between benchmark scores and real-world automation readiness.
@rohanpaul_ai: Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart. Mercha…
MerchantBench is a benchmark that assesses AI agents by having them manage a simulated online store for a year, revealing challenges with sustained performance and continuous action.
@rohanpaul_ai: This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed fee…
A paper evaluates eight leading AI models on long-horizon tasks, finding that even the best-performing model achieves only 27.3% of human performance, highlighting significant limitations for dependable long-horizon AI execution.
@rohanpaul_ai: Univ of Texas paper shows AI agents can slowly become less reliable after deployment, even when the model itself does n…
A University of Texas paper introduces AgingBench, a benchmark that reveals AI agents can become less reliable after deployment due to memory and maintenance decay, even when the underlying model remains unchanged.
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
FM-Bench is a new benchmark for evaluating long-horizon decision-making of LLM agents managing a football club over 20 years, revealing that managerial behavior drives performance more than model scale or token spend.