CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Summary
CoffeeBench is a benchmark for evaluating LLM agents in a long-horizon multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.
View Cached Full Text
Cached at: 06/26/26, 10:06 AM
Paper page - CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Source: https://huggingface.co/papers/2606.16613
Abstract
CoffeeBench evaluates LLM agents in a multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.
AsLLM agentsbecome capable of increasinglylong-horizon tasks, evaluating their performance ineconomic systemsis becoming increasingly important. Unlike existing benchmarks that primarily evaluate a single agent interacting with a passive environment,economic systemsare inherently multi-agent, requiringautonomous agentsto communicate, negotiate, andtransactwhile pursuing their own objectives over extended periods. We introduce CoffeeBench, a benchmark for evaluatingLLM agentsin a long-horizonmulti-agent economycomposed of heterogeneous firms. In CoffeeBench, two farmers, two roasters, and two retailers autonomously operate their businesses over a 90-day simulation, each seeking to maximizecumulative net incomethroughcommunicationandtransactions while managing cash, inventory, and pricing. The evaluated model controls one coffee roaster, while the remaining firms are controlled by fixed reference agents. Across several recent open-weight and proprietary LLMs, all models outperform a passive baseline that takes no actions, with most achieving positive net income. Analysis ofagent behaviorreveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans. We release our code and agent trajectories to support future research.
View arXiv pageView PDFProject pageGitHub6Add to collection
Get this paper in your agent:
hf papers read 2606\.16613
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.16613 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.16613 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.16613 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
MerchantBench is a new benchmark for evaluating LLM agents' long-term coherence in e-commerce operations, using a 365-day order-level simulation with 98,843 real product records and 26 tools. Results show the best LLM achieves only 27.3% of human participants' final net assets, highlighting a substantial capability gap.
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
This paper introduces EcoAgent-Bench, a 304-task benchmark for evaluating LLM agents' economic decision-making under explicit budgets and priced actions, testing four cost-related decisions across five task families.
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Introduces PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows across five domains, demonstrating significant performance degradation compared to monolingual settings.
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Business Arena is a new benchmark that evaluates LLM agents in a realistic cross-border shop environment, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions.
New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]
Introduces a new benchmark for evaluating multi-agent coordination in LLMs, finding that most models struggle with long-horizon open-ended tasks, but Gemini 3.1 Pro performs comparably to trained MARL agents on the hardest setting.