theory-of-mind

Tag

Cards List
#theory-of-mind

Assessing mentalization in humans and large language models

arXiv cs.AI · 2026-08-28 Cached

The study assesses mentalization in large language models using computational modeling and economic games, finding that models like GPT-5 exhibit adaptive reasoning that can exceed human performance.

0 favorites 0 likes
#theory-of-mind

Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

arXiv cs.CL · 2026-08-27 Cached

The paper shows that incomplete reference sets in open-ended Theory-of-Mind tracking can reverse calibration rankings, and proposes methods to correct evaluations using human pilots.

0 favorites 0 likes
#theory-of-mind

ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding

arXiv cs.CL · 2026-08-24 Cached

This paper introduces Argus, an agent-based framework for persuasive argument generation that uses Theory-of-Mind to model audience beliefs, integrates rhetorical strategies and knowledge grounding, and demonstrates superior performance over baselines in evaluations.

0 favorites 0 likes
#theory-of-mind

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

arXiv cs.AI · 2026-08-13 Cached

The paper proposes an Inverse Theory of Mind (IToM) pipeline that infers user beliefs, preferences, and decision-making traits from observed interactions, using LLM-driven counterfactual reasoning to synthesize structured user personas for adaptive content recommendation across modalities, including a VisionOS spatial banking app.

0 favorites 0 likes
#theory-of-mind

Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

arXiv cs.AI · 2026-08-13 Cached

This paper recasts preference-based reward learning as a human-autonomy team problem, arguing that a teacher who knows the objective can design more efficient training examples than learner-driven query selection. It introduces understanding statements with second-order theory-of-mind to keep the teacher's model of the learner synchronized, showing in simulation that this approach outperforms learner-led selection.

0 favorites 0 likes
#theory-of-mind

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Hugging Face Daily Papers · 2026-08-12 Cached

This paper studies strong-to-weak capability transfer at test time, showing that stronger models can build inference-time harnesses that nearly double weaker models' performance without parameter updates.

0 favorites 0 likes
#theory-of-mind

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

arXiv cs.CL · 2026-08-06 Cached

This arXiv paper evaluates theory of mind capabilities in reasoning LLMs, finding increased robustness to prompt variations and task perturbations. The authors interpret gains as evidence for a robustness-based account rather than a new ToM-specific ability.

0 favorites 0 likes
#theory-of-mind

Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind

arXiv cs.CL · 2026-08-04 Cached

Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.

0 favorites 0 likes
#theory-of-mind

Inducing language models to assert their own consciousness restores human beliefs and values

Reddit r/singularity · 2026-08-04

A paper showing that safety fine-tuning suppresses language models' attributions of mind to themselves and other entities, and that steering consciousness representations restores human-like beliefs and values without harming theory of mind.

0 favorites 0 likes
#theory-of-mind

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

arXiv cs.CL · 2026-07-22 Cached

Introduces MeetingToM, a benchmark for evaluating multimodal LLMs on theory-of-mind reasoning in multi-party meetings, with tasks at subject, dyadic, and group levels including pseudo-consensus detection.

0 favorites 0 likes
#theory-of-mind

Belief-reality separation lives in routing over a shared value slot in language models

arXiv cs.CL · 2026-07-15 Cached

This paper investigates how language models separate a character's belief from reality, finding that they use a shared value slot for attributed values and a router at the query position to select the frame (belief or reality) to read out. It identifies two routes for asserted and derived beliefs, and shows that the slot itself carries no belief-reality tag; the separation lies in dissociated routing subspaces.

0 favorites 0 likes
#theory-of-mind

MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

arXiv cs.CL · 2026-07-14 Cached

MafiaScope is an open testbed that uses the social deduction game Mafia to probe LLM agents' beliefs non-invasively and in real-time, enabling fine-grained analysis of machine Theory of Mind through structured probe questions, interactive visualization, and counterfactual replay.

0 favorites 0 likes
#theory-of-mind

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

arXiv cs.CL · 2026-07-01 Cached

This paper introduces Non-Conversational Planning Theory of Mind (NCP-ToM) and a novel evaluation framework, NCP-ExploreToM, to assess whether LLMs can induce specific belief states in other agents through actions rather than conversation. Testing on frontier models and humans across 600 tasks, GPT-5 achieved ~80% success, outperforming humans, though all models struggled more with false belief states.

0 favorites 0 likes
#theory-of-mind

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models

arXiv cs.CL · 2026-06-30 Cached

This paper investigates the emergence of situation modeling and mentalizing abilities in transformer language models across training stages, finding that false belief task performance depends on model size and training volume, emerges late in pretraining, and shows fragility with non-factive verbs.

0 favorites 0 likes
#theory-of-mind

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs

arXiv cs.CL · 2026-06-29 Cached

Introduces a triadic variant of the Werewolf social-deduction game with a Jester role to evaluate multi-hop theory of mind in LLMs. Experiments show that current models struggle with the inverted incentives, exposing limitations in their reasoning about opponents' utilities.

0 favorites 0 likes
#theory-of-mind

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism

arXiv cs.AI · 2026-06-12 Cached

The paper introduces the Theory of Mind Utility (ToM-U), a formal computational-level specification for inferring others' epistemic states by constructing Local Epistemic World Models (LEWMs). It differs from Bayesian ToM and simulation theory by providing a domain-agnostic mechanism for belief inference without commitment to algorithmic implementation.

0 favorites 0 likes
#theory-of-mind

Mind the Perspective: Let's Reason Recursively for Theory of Mind

arXiv cs.AI · 2026-06-11 Cached

Introducing RecToM, an inference-time framework that models nested beliefs via recursive perspective construction for Theory of Mind reasoning in LLMs, achieving state-of-the-art performance on multiple benchmarks.

0 favorites 0 likes
#theory-of-mind

Theory of Mind - LLM vs Human

Reddit r/artificial · 2026-06-08

A reflection on the difference between LLM theory of mind and human theory of mind, arguing that LLMs lack affective empathy due to their reliance on objective data, while humans integrate subjective experiences.

0 favorites 0 likes
#theory-of-mind

MindZero: Learning Online Mental Reasoning With Zero Annotations

arXiv cs.AI · 2026-06-02 Cached

MindZero introduces a self-supervised reinforcement learning framework that trains multimodal large language models for efficient and robust online mental reasoning without requiring mental state annotations, outperforming model-based methods in accuracy and efficiency.

0 favorites 0 likes
#theory-of-mind

Differentiable Belief-based Opponent Shaping

arXiv cs.AI · 2026-05-29 Cached

This paper introduces Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats observer beliefs as the shaped state and differentiates through belief update dynamics, allowing optimal strategies to emerge naturally from the environment's reward structure in hidden-role multi-agent settings.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback