Tag
The paper introduces Fuse, a multi-agent simulation framework for evaluating social reasoning in LLM assistants by providing verifiable ground truth through user-mediated interactions, validated with a human study and applied to 12 LLMs.
This paper introduces the Human Utility Factor (HUF), a computable welfare metric that reframes AI governance as a constrained optimisation problem with measurable levers for automation depth, redistribution, and employment coverage, validated through multi-agent simulations.
This paper introduces ComRate, a large-scale dataset of community notes and ratings from X, and proposes MultiCom, a persona-guided multi-agent framework for simulating community note evaluation. The approach achieves 84.7% accuracy in predicting note helpfulness.
This paper presents multi-agent simulations of the emergence of morphological alternation patterns (like 'go/went') in language, using an AI Historical Linguist (LLM-driven) to evaluate plausibility of evolved morphologies against real languages.
Introduces SovSim, a multi-agent simulation framework for studying cooperation and resource sustainability in LLM societies with asymmetric power structures. Experiments show that introducing a dominant agent (boss or king) severely degrades cooperation and survival rates across 11 state-of-the-art models.
This paper uses the Greenland sovereignty crisis as a case study to test LLM geopolitical behavior through multi-agent simulations, revealing that coercion framing increases escalation and that peaceful acquisition is rare.