Tag
AgentWorld introduces a benchmark for evaluating long-horizon collaboration in multi-agent LLM systems with 100 tasks and a new causal collaboration effectiveness metric. Experiments show that even top models achieve only 52% task success, highlighting failure modes like communication breakdowns.