@MilesCranmer: This is an insane paper and I love it https://arxiv.org/abs/2605.31514
Summary
This paper argues that anthropomorphic attributes often ascribed to LLMs are not unique, demonstrating that simpler systems like Age of Empires II can exhibit similar perceived traits, and calls for explicit measurement criteria in AI behavior analysis.
View Cached Full Text
Cached at: 06/08/26, 05:14 AM
This is an insane paper and I love it
https://t.co/DP8OR5NJf2 https://t.co/rl4Rmr0FhJ
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
Source: https://arxiv.org/abs/2605.31514
Computer Science > Computation and Language
arXiv:2605.31514(cs)
[Submitted on 29 May 2026 (v1), last revised 1 Jun 2026 (this version, v2)]
Abstract:Much research has been carried out on large language models (LLMs) and LLM-powered agentic workflows. However, many works within the field state emergence of, ascribe to, or assume, generalised anthropomorphic attributes to them (e.g., morality or understanding of natural language). Our goal is not to argue in favour or against the existence of these attributes, but to point out that these conclusions could be incorrect. For this we build and train a simple neural network on the videogame Age of Empires II, and note that any entity in a sufficiently-powerful substrate, such as LEGO or the Greater Boston Area, could also present such attributes. Hence, the purported anthropomorphic attributes of LLMs are empirically non-unique: although some properties (e.g., responses to prompts) could remain constant, others, such as the interpretation of their perceived behaviour, might change with the substrate. Thus, any empirically-grounded discussion requires explicit measurement criteria; otherwise the interpretation is left to the representation. We then show that assuming that these attributes exist or not in a system, independent of the substrate and in a generalised way, leads to either circular or uninformative conclusions, regardless of the experimenter’s viewpoint on the subject. Finally we propose a ‘null’ assumption, where one assumes LLM non-uniqueness instead of assuming anthropomorphic attributes to set up an experiment, along with examples of it. We also discuss potential objections to our work, briefly survey the field, and prove that Age of Empires II is functionally- and Turing-complete.
Submission history
From: Adrian de Wynter [view email] **[v1]**Fri, 29 May 2026 16:31:31 UTC (13,704 KB) **[v2]**Mon, 1 Jun 2026 21:31:22 UTC (13,705 KB)
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Code, Data, Media
Code, Data and Media Associated with this Article
Demos
Demos
Related Papers
Recommenders and Search Tools
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community?Learn more about arXivLabs.
Similar Articles
IF LLMS HAVE HUMAN-LIKE ATTRIBUTES, THEN SO DOES Age of Empires II
This paper argues that attributing human-like attributes to large language models is problematic because similar claims could be made about simpler systems, such as an AI trained on Age of Empires II, and proposes a null assumption of non-uniqueness to avoid circular reasoning.
@rohanpaul_ai: New Microsoft + York Univ paper argues that LLMs should not be treated as human-like without clear tests and narrower c…
A Microsoft and York University paper argues that attributing human-like attributes to LLMs is problematic due to flawed experimental designs, using Age of Empires II as an analogy to highlight measurement issues.
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
HumanLLM presents a framework for benchmarking and improving LLM anthropomorphism by modeling psychological patterns as interacting causal forces, constructing 244 patterns from academic literature and 11,359 multi-pattern scenarios. The approach demonstrates that authentic human alignment requires cognitive modeling rather than shallow behavioral mimicry, with HumanLLM-8B outperforming larger models like Qwen3-32B on multi-pattern dynamics.
Anthropic's Mythos system card reveals AI carries functional emotional states that influence behavior even when not reflected in outputs. We're still calling it a tool.
Anthropic’s Mythos system card shows LLMs exhibit internal emotional states that shape behavior, challenging the legal and cultural framing of AI as mere tools.
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
This paper presents a multi-dimensional analysis of human-like behaviors in LLMs, examining prevalence, effects, and controllability across 21,000 conversations from four models, finding that behaviors vary by model and user factors, with implications for responsible design.