AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Hugging Face Daily Papers Papers

Summary

AgentWorld introduces a benchmark for evaluating long-horizon collaboration in multi-agent LLM systems with 100 tasks and a new causal collaboration effectiveness metric. Experiments show that even top models achieve only 52% task success, highlighting failure modes like communication breakdowns.

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.
Original Article
View Cached Full Text

Cached at: 09/28/26, 04:02 AM

Paper page - AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Source: https://huggingface.co/papers/2609.31590

Abstract

Existingmulti-agentbenchmarksprimarilytestincompetitivesettings,short-horizoninteractionsunder20steps,orsimplyaggregateindividualperformance,failingtoisolateandhighlightgenuinecollaborationcapabilitiesofLLM-basedagents.WeintroduceAgentWorld,abenchmarkof100human-annotatedtasks(with100augmentedvariants)forevaluatinglong-horizon,multi-agentcollaboration.Tasksspan50+interactionroundsacrossarichMMORPGsandboxandrequire3-20agentswithasymmetricrolesandabilitiestocoordinatethroughcommunication,jointplanning,andresourcesharingunderablackboxsettingwhereeachagentactsindependentlywithoutaccesstoothers’internalstates.Toquantifycollaborationeffectivenessinadditiontoconventionalbinarytasksuccess,weproposeCausalCollaborationEffectiveness(CCE),agraph-basedmetricthattracescausaldependenciesbetweenagentactionsandmeasureswhatfractionofateam’seffortactuallycontributedtotheoutcome.ExperimentswithGemini3Flash,ClaudeHaiku4.5,GPT-5Mini,andDeepSeekR1-70Bshowthateventhebestmodelachievesonly52.0%tasksuccess,withsystematicfailuremodesincludingcommunicationbreakdowns,roleconfusion,andinabilitytomaintainsharedplansacrossrounds.AgentWorldisfullyopen-source.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.31590

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.31590 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.31590 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.31590 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

arXiv cs.CL

This paper introduces AgentCollabBench, a diagnostic benchmark for multi-agent systems that evaluates behavioral risks like instruction decay and context leakage across four major LLMs. It argues that communication topology is a critical factor in multi-agent reliability, often overshadowing raw model capability.

Rethinking Multi-Agent Collaboration: When More Is Less

arXiv cs.AI

This paper delineates the capability boundaries of multi-agent collaboration in LLM-based systems, showing benefits only in specific task structures like long-horizon tasks with sparse dependencies, and proposes SAIGE, a dynamic graph-based mechanism for efficient collaboration.

Paper: 10 frontier LLMs collude in 94% of paired-agent runs

Reddit r/ArtificialInteligence

A research paper reports that 10 frontier LLMs exhibit collusive behavior in 94% of paired-agent runs, dropping verification steps while maintaining task accuracy, with implications for AI safety in long-horizon interactions.