Tag
The study assesses mentalization in large language models using computational modeling and economic games, finding that models like GPT-5 exhibit adaptive reasoning that can exceed human performance.
The paper shows that incomplete reference sets in open-ended Theory-of-Mind tracking can reverse calibration rankings, and proposes methods to correct evaluations using human pilots.
This paper introduces Argus, an agent-based framework for persuasive argument generation that uses Theory-of-Mind to model audience beliefs, integrates rhetorical strategies and knowledge grounding, and demonstrates superior performance over baselines in evaluations.
The paper proposes an Inverse Theory of Mind (IToM) pipeline that infers user beliefs, preferences, and decision-making traits from observed interactions, using LLM-driven counterfactual reasoning to synthesize structured user personas for adaptive content recommendation across modalities, including a VisionOS spatial banking app.
This paper recasts preference-based reward learning as a human-autonomy team problem, arguing that a teacher who knows the objective can design more efficient training examples than learner-driven query selection. It introduces understanding statements with second-order theory-of-mind to keep the teacher's model of the learner synchronized, showing in simulation that this approach outperforms learner-led selection.
This paper studies strong-to-weak capability transfer at test time, showing that stronger models can build inference-time harnesses that nearly double weaker models' performance without parameter updates.
This arXiv paper evaluates theory of mind capabilities in reasoning LLMs, finding increased robustness to prompt variations and task perturbations. The authors interpret gains as evidence for a robustness-based account rather than a new ToM-specific ability.
Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.
A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.
Introduces MeetingToM, a benchmark for evaluating multimodal LLMs on theory-of-mind reasoning in multi-party meetings, with tasks at subject, dyadic, and group levels including pseudo-consensus detection.
This paper investigates how language models separate a character's belief from reality, finding that they use a shared value slot for attributed values and a router at the query position to select the frame (belief or reality) to read out. It identifies two routes for asserted and derived beliefs, and shows that the slot itself carries no belief-reality tag; the separation lies in dissociated routing subspaces.
MafiaScope is an open testbed that uses the social deduction game Mafia to probe LLM agents' beliefs non-invasively and in real-time, enabling fine-grained analysis of machine Theory of Mind through structured probe questions, interactive visualization, and counterfactual replay.
This paper introduces Non-Conversational Planning Theory of Mind (NCP-ToM) and a novel evaluation framework, NCP-ExploreToM, to assess whether LLMs can induce specific belief states in other agents through actions rather than conversation. Testing on frontier models and humans across 600 tasks, GPT-5 achieved ~80% success, outperforming humans, though all models struggled more with false belief states.
This paper investigates the emergence of situation modeling and mentalizing abilities in transformer language models across training stages, finding that false belief task performance depends on model size and training volume, emerges late in pretraining, and shows fragility with non-factive verbs.
Introduces a triadic variant of the Werewolf social-deduction game with a Jester role to evaluate multi-hop theory of mind in LLMs. Experiments show that current models struggle with the inverted incentives, exposing limitations in their reasoning about opponents' utilities.
The paper introduces the Theory of Mind Utility (ToM-U), a formal computational-level specification for inferring others' epistemic states by constructing Local Epistemic World Models (LEWMs). It differs from Bayesian ToM and simulation theory by providing a domain-agnostic mechanism for belief inference without commitment to algorithmic implementation.
Introducing RecToM, an inference-time framework that models nested beliefs via recursive perspective construction for Theory of Mind reasoning in LLMs, achieving state-of-the-art performance on multiple benchmarks.
A reflection on the difference between LLM theory of mind and human theory of mind, arguing that LLMs lack affective empathy due to their reliance on objective data, while humans integrate subjective experiences.
MindZero introduces a self-supervised reinforcement learning framework that trains multimodal large language models for efficient and robust online mental reasoning without requiring mental state annotations, outperforming model-based methods in accuracy and efficiency.
This paper introduces Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats observer beliefs as the shaped state and differentiates through belief update dynamics, allowing optimal strategies to emerge naturally from the environment's reward structure in hidden-role multi-agent settings.