AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Summary
AgentWorld introduces a benchmark for evaluating long-horizon collaboration in multi-agent LLM systems with 100 tasks and a new causal collaboration effectiveness metric. Experiments show that even top models achieve only 52% task success, highlighting failure modes like communication breakdowns.
View Cached Full Text
Cached at: 09/28/26, 04:02 AM
Paper page - AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Source: https://huggingface.co/papers/2609.31590
Abstract
Existingmulti-agentbenchmarksprimarilytestincompetitivesettings,short-horizoninteractionsunder20steps,orsimplyaggregateindividualperformance,failingtoisolateandhighlightgenuinecollaborationcapabilitiesofLLM-basedagents.WeintroduceAgentWorld,abenchmarkof100human-annotatedtasks(with100augmentedvariants)forevaluatinglong-horizon,multi-agentcollaboration.Tasksspan50+interactionroundsacrossarichMMORPGsandboxandrequire3-20agentswithasymmetricrolesandabilitiestocoordinatethroughcommunication,jointplanning,andresourcesharingunderablackboxsettingwhereeachagentactsindependentlywithoutaccesstoothers’internalstates.Toquantifycollaborationeffectivenessinadditiontoconventionalbinarytasksuccess,weproposeCausalCollaborationEffectiveness(CCE),agraph-basedmetricthattracescausaldependenciesbetweenagentactionsandmeasureswhatfractionofateam’seffortactuallycontributedtotheoutcome.ExperimentswithGemini3Flash,ClaudeHaiku4.5,GPT-5Mini,andDeepSeekR1-70Bshowthateventhebestmodelachievesonly52.0%tasksuccess,withsystematicfailuremodesincludingcommunicationbreakdowns,roleconfusion,andinabilitytomaintainsharedplansacrossrounds.AgentWorldisfullyopen-source.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.31590
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.31590 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.31590 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.31590 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]
Introduces a new benchmark for evaluating multi-agent coordination in LLMs, finding that most models struggle with long-horizon open-ended tasks, but Gemini 3.1 Pro performs comparably to trained MARL agents on the hardest setting.
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
This paper introduces AgentCollabBench, a diagnostic benchmark for multi-agent systems that evaluates behavioral risks like instruction decay and context leakage across four major LLMs. It argues that communication topology is a critical factor in multi-agent reliability, often overshadowing raw model capability.
Rethinking Multi-Agent Collaboration: When More Is Less
This paper delineates the capability boundaries of multi-agent collaboration in LLM-based systems, showing benefits only in specific task structures like long-horizon tasks with sparse dependencies, and proposes SAIGE, a dynamic graph-based mechanism for efficient collaboration.
Paper: 10 frontier LLMs collude in 94% of paired-agent runs
A research paper reports that 10 frontier LLMs exhibit collusive behavior in 94% of paired-agent runs, dropping verification steps while maintaining task accuracy, with implications for AI safety in long-horizon interactions.
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
CoffeeBench is a benchmark for evaluating LLM agents in a long-horizon multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.