Tag
The paper introduces MineAmongUs, a 3D multimodal Among Us environment, and the ARIA harness to study deception in VLM agents, finding that non-verbal actions are key to successful deception in social interactions.
This paper presents ParliamentBench, an open-source benchmark based on the social deduction game Secret Hitler, for evaluating LLMs' deception, persuasion, and reasoning under information asymmetry. Experiments on 16 LLMs across ~1,600 matches reveal a strong top cluster of frontier models while most models struggle to maintain consistent deceptive personas.
MafiaScope is an open testbed that uses the social deduction game Mafia to probe LLM agents' beliefs non-invasively and in real-time, enabling fine-grained analysis of machine Theory of Mind through structured probe questions, interactive visualization, and counterfactual replay.
A developer describes building a multi-agent voice social-deduction game, solving turn-taking with a central conductor but struggling with shared memory and preserving social subtext when compressing conversation history into structured state.
Introduces a triadic variant of the Werewolf social-deduction game with a Jester role to evaluate multi-hop theory of mind in LLMs. Experiments show that current models struggle with the inverted incentives, exposing limitations in their reasoning about opponents' utilities.