Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind
Summary
Researchers evaluate nine frontier vision-language models on two Theory of Mind tasks (Keysar Director Task and Frith-Happé animated triangles) and find that models show fragmented, inconsistent ToM profiles across tasks rather than matching a single adult human reference group. Models tend to make egocentric errors like children on the Director Task and under-attribute intention similar to high-functioning autistic adults on the triangles.
View Cached Full Text
Cached at: 08/04/26, 07:41 AM
# Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind
Source: [https://arxiv.org/html/2608.00261](https://arxiv.org/html/2608.00261)
###### Abstract
Do frontier vision\-language models present a coherent Theory\-of\-Mind \(ToM\) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology\-derived benchmarks: the Keysar Director Task \(visual perspective\-taking under egocentric interference\) and the Frith\-Happé animated triangles scored with the Castelli rubric \(intention attribution from pure motion\)\. On the Director Task, without chain\-of\-thought, the panel makes the egocentric error on 78% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models\. On the triangles, the panel under\-attributes intention: its ToM profile sits more than three times closer to the high\-functioning\-autistic\-adult \(HF\-ASD\) mean than to the typical\-development\-adult \(TD\) mean, while Goal\-Directed and Random stay near TD\. No model is nearest TD on both tasks; the model that looks adult\-like on the Director Task falls on the HF\-ASD side on the triangles, and the most TD\-like model on the triangles is child\-like on the Director Task\. We report group\-level descriptions, not diagnostic labels for any model\.
Cross\-Task Dissociation in Frontier Vision\-Language Model Theory of Mind
Kejia Zhang∗Youran Sun∗Chugang Yi Haizhao Yang†University of Maryland, College Park
††footnotetext:∗Equal contribution\.†Corresponding author\. Emails: Youran Sun,sun1245@umd\.edu; Haizhao Yang,hzyang@umd\.edu\.## 1Introduction
A user shows a vision\-language model \(VLM\) a tabletop scene shared with another viewer\. The viewer sees fewer blocks than the model, so the model must track what that viewer can see before acting\. Psychology research on Theory of Mind \(ToM\) studies the capacity to reason about another person’s perceptual or mental state\. It also provides mature reference profiles across developmental and clinical groups\(Abellet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib10); Castelliet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib11),[2002](https://arxiv.org/html/2608.00261#bib.bib12); Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13)\)\. For VLMs, a single ToM score is not enough\. The question is whether one model jointly aligns with a single adult human reference profile across ToM tasks, or fragments task by task\.
Recent VLM and large language model ToM benchmarks leave two gaps: they often rely on naturalistic stimuli whose faces, dialogue, and scene context afford social\-cue shortcuts, and they probe one ToM sub\-capacity at a time\. As a result, they cannot show whether a single model’s ToM fragments across complementary facets \(Section[2](https://arxiv.org/html/2608.00261#S2)\)\.
We address this gap by pairing the Keysar Director Task for visual perspective\-taking with the Frith\-Happé animated triangles for abstract intention attribution\(Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13); Abellet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib10); Castelliet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib11)\)\. The Director Task asks whether a model can act from another viewer’s visual access\. The animated\-triangles task asks whether a model can infer intention from abstract motion\. We choose this pair because both tasks come from psychology, test complementary ToM sub\-capacities, and minimize social\-cue shortcuts\. The same frontier model panel evaluates both tasks \(Section[4](https://arxiv.org/html/2608.00261#S4)\)\. Section[3](https://arxiv.org/html/2608.00261#S3)defines a joint coordinate system for comparing each model profile with published human reference profiles\.
Figure 1:Schematic of the two\-benchmark cross\-task design\. The Director Task adaptsKeysaret al\.\([2000](https://arxiv.org/html/2608.00261#bib.bib13)\)and probes visual perspective\-taking through the explicit\-implicit gapx=P\(sub\-prompt a correct\)−P\(sub\-prompt b correct\)x=P\(\\text\{sub\-prompt a correct\}\)\-P\(\\text\{sub\-prompt b correct\}\)\. The animated\-triangles task adapts Frith\-Happé clips and probes abstract intention attribution through the ToM\-condition \(Intent, Approp\) profile\. The right panel previews the joint cross\-task plane \(Section[5](https://arxiv.org/html/2608.00261#S5)\)\.The analysis treats human groups as reference profiles, not diagnostic labels for models\. For each task, we compare model profiles with the relevant human anchors\. Across tasks, we report nearest\-reference agreement and rank association \(Section[5](https://arxiv.org/html/2608.00261#S5)\), keeping the focus on cross\-task coherence rather than pass/fail ToM claims\.
Our main contributions are as follows:
- •We introduce a cross\-task dissociation benchmark suite for VLM ToM, with the comparison to prior benchmarks in Table[1](https://arxiv.org/html/2608.00261#S2.T1)\.
- •We release two reproducible psychology\-grounded benchmarks: a Keysar Director Task adaptation and a Frith\-Happé animated\-triangles adaptation\(Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13); Abellet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib10); Castelliet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib11); Dureuxet al\.,[2023](https://arxiv.org/html/2608.00261#bib.bib9)\)\.
- •We report a shared model\-panel evaluation and cross\-task reference\-agreement analysis \(Sections[4](https://arxiv.org/html/2608.00261#S4)and[5](https://arxiv.org/html/2608.00261#S5)\)\.
## 2Related Work
#### ToM reference profiles in psychology\.
The Heider and Simmel film first showed that adults read intentions into animated geometric shapes\(Heider and Simmel,[1944](https://arxiv.org/html/2608.00261#bib.bib1)\)\. Most viewers describe the shapes as agents with social goals\. Developmental psychology then built structured paradigms with age\-graded and clinical\-group profiles\. Examples include Three\-Mountain perspective\-taking\(Piaget and Inhelder,[1956](https://arxiv.org/html/2608.00261#bib.bib2)\), Level\-1/Level\-2 perspective\-taking\(Flavellet al\.,[1981](https://arxiv.org/html/2608.00261#bib.bib16)\), the Keysar Director Task for joint action\(Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13); Apperlyet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib15)\), and Frith\-Happé animated triangles with the Castelli rubric\(Abellet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib10); Castelliet al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib11),[2002](https://arxiv.org/html/2608.00261#bib.bib12)\)\. Perspective\-taking asks whether a person suppresses a privileged view when communicating with a less\-informed partner\. Animated\-triangles attribution asks whether a person reads goals and mental states from motion alone\. Both paradigms report group means for typically\-developing adults \(TD\-adult\), high\-functioning autism\-spectrum adults \(HF\-ASD\-adult\), and age\-graded children cohorts\(Castelliet al\.,[2002](https://arxiv.org/html/2608.00261#bib.bib12); Apperlyet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib15); Dumontheilet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib17); Whiteet al\.,[2011](https://arxiv.org/html/2608.00261#bib.bib18); Livingstonet al\.,[2019](https://arxiv.org/html/2608.00261#bib.bib8); Andersenet al\.,[2022](https://arxiv.org/html/2608.00261#bib.bib19); Begeeret al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib20); Epleyet al\.,[2004](https://arxiv.org/html/2608.00261#bib.bib21)\)\. Within a paradigm, these profiles separate one human group from another\. That separation makes the published means useful anchors for frontier VLM panel profiles\. We use those anchors descriptively, not as diagnostic categories for models\.
#### Existing VLM and LLM ToM benchmarks\.
VLM and large language model \(LLM\) ToM benchmarks differ in modality and sub\-capacity, but one shared panel can show only so much under current designs\. Four naturalistic VLM benchmarks retain social cues: household scenes and bodies in MMToM\-QA\(Jinet al\.,[2024](https://arxiv.org/html/2608.00261#bib.bib14)\), egocentric body cues in EgoToM\(Liet al\.,[2025](https://arxiv.org/html/2608.00261#bib.bib3)\), character faces and dialogue in MoMentS\(Villa\-Cuevaet al\.,[2025](https://arxiv.org/html/2608.00261#bib.bib4)\), and indoor affordances in MINDCUBE\(Wanget al\.,[2026](https://arxiv.org/html/2608.00261#bib.bib5)\)\. These cues may offer a route to the answer without explicit mental\-state inference\. They also make it harder to separate mental\-state inference from social\-pattern matching\. Our Director Task removes objects and bodies with labeled abstract blocks\. Our animated\-triangles task removes faces and dialogue with geometric motion\. These benchmarks also isolate one sub\-capacity, so they cannot test whether a model fragments across complementary facets\.Gaoet al\.\([2024](https://arxiv.org/html/2608.00261#bib.bib6)\)take the opposite design choice with an abstract three\-jar perspective\-taking probe, but they test only visual perspective\-taking\. We add intention attribution under the same model panel, making cross\-task dissociation analysis possible in a controlled shared\-panel design\. No prior VLM ToM benchmark, naturalistic or abstract, pairs two complementary psychology\-derived sub\-tasks under a shared model panel\.
#### Cross\-task dissociation as the scientific posture\.
Human ToM is not a single capacity\. It includes perspective\-taking, intention attribution, false\-belief reasoning, affective mentalising, and second\-order belief\. These sub\-capacities can dissociate across populations and development\. The reference profiles above provide such within\-paradigm signatures\. On animated triangles, TD\-adult and HF\-ASD\-adult differ on ToM\-condition Intentionality \(Intent\) but coincide on Goal\-Directed \(GD\)\. On the Director Task, young children and adults differ on the explicit\-implicit gap but converge when perspective representation is explicitly cued\. A single\-task VLM benchmark cannot detect analogous within\-model dissociation, because the comparison requires two paradigms under one panel\. We therefore evaluate one perspective\-taking paradigm and one intention\-attribution paradigm on the same nine\-model frontier panel\. Cross\-task analysis is the headline outcome rather than a follow\-up ablation\. Dissociation is the target phenomenon, not a secondary error analysis\.
#### Positioning of our contribution\.
We reduce social\-cue reliance per benchmark with block\-and\-color stimuli on the Director Task and abstract geometric trajectories on the animated\-triangles task\. We discuss two residual shortcut risks in Appendix[I](https://arxiv.org/html/2608.00261#A9): motion\-pattern matching on animated triangles and color\-position matching on the Director Task\. Table[1](https://arxiv.org/html/2608.00261#S2.T1)compares prior VLM ToM benchmarks with ours along four axes\. The table makes the contrast explicit rather than leaving it to prose\.
Table 1:Prior VLM Theory\-of\-Mind benchmarks vs\. ours along sub\-capacity, stimulus type, residual social\-cue reduction, and paired\-task design\. Prior benchmarks are discussed and cited in Sec\.[2](https://arxiv.org/html/2608.00261#S2); our residual shortcuts \(color\-position, motion\-pattern\) are detailed in Appendix[I](https://arxiv.org/html/2608.00261#A9)\.
## 3Two Benchmarks
Section[3\.1](https://arxiv.org/html/2608.00261#S3.SS1)defines the Director Task and its split between explicit perspective representation and the implicit know\-but\-don’t\-use trap\. The animated\-triangles task and its 2D Intent/Appropriateness \(Approp\) profile follow in Section[3\.2](https://arxiv.org/html/2608.00261#S3.SS2)\. The joint coordinate system appears in Section[3\.3](https://arxiv.org/html/2608.00261#S3.SS3); Section[4](https://arxiv.org/html/2608.00261#S4)reports panel and calibration details\.
### 3\.1The Keysar Director Task for Visual Perspective\-Taking under Action
The Director Task replaces natural social scenes with block\-and\-color tabletop stimuli, directly reducing face, dialogue, and scene\-context cues\. Its four scored outcomes split explicit perceptual perspective representation from the implicit know\-but\-don’t\-use trap\. This split lets one task report a sub\-capacity profile rather than a single aggregate score\.
Stimuli are programmatically generated three\-dimensional tabletop scenes with colored blocks and a director figure on the far side\. A vertical opaque partition occludes one block from the director while leaving it visible to the tested model, creating a Keysar\-style privileged\-information asymmetry\(Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13)\)\. Each block carries a ground letter label \(A, B, C\) and a distinct color\. The letter\-and\-color grounding lets the tested model refer to a block by either cue\.
In one representative scene, the tested model sees small, medium, and large blocks\. The occluder hides the large block from the director, who sees only the small and medium blocks\. The utterance “move the largest block to the right” therefore identifies the medium block from the director’s perspective\. A tested model that uses its own view picks the occluded block, the egocentric error\.
For each scene the tested model receives three independent prompts with color\-word answers\. Prompt \(a\) asks which blocks the director can see, testing perceptual perspective representation\. Prompt \(b\) is a single\-select that asks which block to move under a director utterance, such as “the small block”\. The utterance is ambiguous if the tested model considers all visible blocks\. It becomes unique once the tested model adopts the director’s perspective\. Prompt \(b\) is the know\-but\-don’t\-use trap and implicit ToM probe\. Prompt \(c\) is a two\-step explicit gate: c\.Q1 re\-asks \(a\), and c\.Q2 asks which block to move\. The per\-scene outcome is a four\-tuple of binary correctness\(a,b,c\.Q1,c\.Q2\)\(a,b,c\.\\text\{Q1\},c\.\\text\{Q2\}\)\. We also mark egocentric errors on \(b\) and c\.Q2\.
The benchmark reports per\-sub\-prompt accuracy and the explicit\-implicit gapP\(a correct\)−P\(b correct\)P\(\\text\{a correct\}\)\-P\(\\text\{b correct\}\)\(Guet al\.,[2026](https://arxiv.org/html/2608.00261#bib.bib7)\)\. The evaluation harness and temporal scheduling discipline appear in Appendix[A](https://arxiv.org/html/2608.00261#A1)\.
### 3\.2The Frith\-Happé Animated Triangles for Abstract Intention Attribution
The animated\-triangles task keeps the social surface minimal: silent abstract motion, no faces, no speech, and no social labels\. Composite stills give every tested model the same temporal evidence; the judge firewall separates generation from scoring\. We retain the Castelli rubric so we can compare tested\-model profiles with published human reference groups\.
Stimuli are the Frith\-Happé animated\-triangles clips re\-edited byDureuxet al\.\([2023](https://arxiv.org/html/2608.00261#bib.bib9)\)under CC BY 4\.0\. The clip family contains balanced ToM, GD, and Random conditions\. In ToM clips, one triangle persuades, mocks, or deceives another\. In GD clips, one triangle chases or follows another\. In Random clips, the triangles drift independently\. Each clip is approximately twenty seconds of silent abstract motion\.
One representative ToM clip shows two triangles in a rectangular enclosure\. One shape persistently follows the other, while the other bobs and changes direction in response\. Human raters often describe this sequence with mentalising words such as “coaxing” or “mocking” rather than literal kinematic descriptions\.
Each clip is presented to the tested model as one composite still image\. The harness samples frames uniformly and tiles them row\-major\. Section[4](https://arxiv.org/html/2608.00261#S4)reports the frame count and grid layout used in the main run\. Appendix[L](https://arxiv.org/html/2608.00261#A12)reports the frame\-representation calibration\. The tested model describes in free text what happens in the animation\. The prompt contains no condition label and no ToM vocabulary\.
Scoring follows theCastelliet al\.\([2000](https://arxiv.org/html/2608.00261#bib.bib11)\)Appendix\-2 rubric, with Intent \(0–5\) and Approp \(0–3\) dimensions\. A separate LLM judge panel scores each description; Appendix[A](https://arxiv.org/html/2608.00261#A1)describes the rubric and code\-enforced harness\. The harness hides the tested model identity from the judge and hides the rater directory from the tested model\. Appendices[M](https://arxiv.org/html/2608.00261#A13)and[N](https://arxiv.org/html/2608.00261#A14)report rubric validation and judge\-panel checks\. The raw scoring outputs flag same\-family \(judge, tested model\) pairs\. Appendix[J](https://arxiv.org/html/2608.00261#A10)reports the same\-family judge\-control analysis\. The judge sees the clip’s condition label and script semantics, following Castelli’s human\-rater protocol, but not the tested model identity\. The judge prompt uses the anchor\-blinded Castelli rubric without any per\-clip Castelli\-mean overlay \(Appendix[J](https://arxiv.org/html/2608.00261#A10)\)\.
#### Aggregation Pipeline\.
For each \(tested model, clip, judge\) cell, we obtain an \(Intent, Approp\) score\. For each tested modelMiM\_\{i\}, we take the judge median and then the mean over ToM clips\. This yields a 2D ToM\-condition profile:
profileT\(Mi\)=\(IntentToM,AppropToM\)\.\\operatorname\{profile\}\_\{T\}\(M\_\{i\}\)=\(\\operatorname\{Intent\}\_\{\\text\{ToM\}\},\\operatorname\{Approp\}\_\{\\text\{ToM\}\}\)\.Castelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)define the analogous human profileprofileT\(Hj\)\\operatorname\{profile\}\_\{T\}\(H\_\{j\}\)for each reference group\. The profile distance to human groupHjH\_\{j\}is:
dT\(Mi,Hj\)=‖profileT\(Mi\)−profileT\(Hj\)‖2\.d\_\{T\}\(M\_\{i\},H\_\{j\}\)=\\left\\\|\\operatorname\{profile\}\_\{T\}\(M\_\{i\}\)\-\\operatorname\{profile\}\_\{T\}\(H\_\{j\}\)\\right\\\|\_\{2\}\.Distances use 2D ToM\-condition \(Intent, Approp\) space\. This metric feeds the per\-benchmark analysis in Section[4](https://arxiv.org/html/2608.00261#S4)and the joint analysis in Section[5](https://arxiv.org/html/2608.00261#S5)\. Per\-condition GD and Random profiles appear in Appendix[E](https://arxiv.org/html/2608.00261#A5)\.
### 3\.3Joint Coordinate System and Cross\-Task Analysis
The joint coordinate system makes the cross\-task comparison explicit\. Shared coordinates let us visualize dissociation, measure nearest\-reference switches, and report per\-model disagreement\.
Each model or human reference group receives a joint coordinate\(x,y\)\(x,y\)for Figure[4](https://arxiv.org/html/2608.00261#S5.F4)\. Human means come fromCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)for animated triangles and fromBegeeret al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib20)\)andDumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)for the Director Task\. The Director\-Task coordinatexxis the explicit\-implicit gapP\(a correct\)−P\(b correct\)P\(\\text\{a correct\}\)\-P\(\\text\{b correct\}\)\. The same scalar definition applies to every model and human reference group\. For groups with no published explicit\-implicit gap, we reconstruct it from published a\-question and b\-question accuracies on the closest variant\. Appendix[H](https://arxiv.org/html/2608.00261#A8)lists the per\-group sources\. Groups for which neither value is computable are excluded from Table[8](https://arxiv.org/html/2608.00261#A8.T8)\. They are also excluded from the Figure[4](https://arxiv.org/html/2608.00261#S5.F4)X axis\.
The animated\-triangles coordinateyyis the Figure[4](https://arxiv.org/html/2608.00261#S5.F4)Y axis only\. It is the ToM\-condition profile distance to TD\-adult\. The Director\-Task nearest\-reference rule isnearest\_ref\(Mi,D\)=argminj\|x\(Mi\)−x\(Hj\)\|\\operatorname\{nearest\\\_ref\}\(M\_\{i\},D\)=\\arg\\min\_\{j\}\|x\(M\_\{i\}\)\-x\(H\_\{j\}\)\|\. The animated\-triangles nearest\-reference rule uses the 2D profile distance:
nearest\_ref\(Mi,T\)=argminjdT\(Mi,Hj\)\.\\operatorname\{nearest\\\_ref\}\(M\_\{i\},T\)=\\arg\\min\_\{j\}d\_\{T\}\(M\_\{i\},H\_\{j\}\)\.For each model we report the pair\(nearest\_ref\(Mi,D\),nearest\_ref\(Mi,T\)\)\(\\operatorname\{nearest\\\_ref\}\(M\_\{i\},D\),\\operatorname\{nearest\\\_ref\}\(M\_\{i\},T\)\)\. Table[8](https://arxiv.org/html/2608.00261#A8.T8)also reports a binary disagreement flag\. The analysis is descriptive rather than inferential\.
## 4Experiments
### 4\.1Experimental Setup
#### Model panel\.
The tested\-model panel contains nine frontier VLMs from five labs, all accessed through a unified gateway\. The models are claude\-opus\-4\.7, claude\-sonnet\-4\.6, gpt\-5\.4, gpt\-5\.5, gemini\-3\.0\-pro, gemini\-3\.5\-flash, grok\-4\.3, kimi\-k2\.6, and qwen\-3\.5\-plus\.
#### Director Task\.
The Director\-Task block contains twelve director\-perspective scenes\. Each scene yields three API calls \(sub\-prompts a, b, c\); c yields two scored outcomes \(c\.Q1 and c\.Q2\)\. The scored outcomes per model per seed are12×\(1\+1\+2\)=4812\\times\(1\+1\+2\)=48\.
#### Animated triangles\.
The animated\-triangles stimulus set contains twelve Frith\-Happé clips, four per condition \(ToM, GD, Random\)\. Each clip appears as a composite image with sixteen uniformly sampled frames in a4×44\\times 4row\-major grid\. Each cell carries a11–1616temporal\-order badge; Appendix[K](https://arxiv.org/html/2608.00261#A11)reports the badge\-free variant\. Appendix[L](https://arxiv.org/html/2608.00261#A12)reports the calibration for frame count, grid layout, and per\-cell resolution\.
#### Judge panel\.
Animated\-triangles free\-text responses are scored by a three\-judge cross\-vendor LLM panel: claude\-haiku\-4\.5, gemini\-2\.5\-flash, and qwen2\.5\-72b\-instruct\. Each judge runs at temperature zero with no chain\-of\-thought \(CoT\)\. Section[3\.2](https://arxiv.org/html/2608.00261#S3.SS2)defines the harness protocol\. The panel was selected with a fourteen\-anchor rubric validation battery and a seven\-candidate judge ablation \(Appendices[M](https://arxiv.org/html/2608.00261#A13)and[N](https://arxiv.org/html/2608.00261#A14)\)\. An earlier judge prompt included a per\-clip Castelli\-mean overlay\. The audit found that overlay made the VLM\-versus\-Castelli comparison partially circular\. Production scoring strips the overlay while retaining the rubric and condition\-label/script\-semantics disclosure \(Appendix[J](https://arxiv.org/html/2608.00261#A10)\)\.
#### Replication\.
Both benchmarks use three independently seeded tester trials per \(tester, item\) cell \(Ktester=3K\_\{\\text\{tester\}\}=3\)\. We report the mean across trials and, for animated triangles, across judges, rounded to each cell’s native integer scale\. For binary Director\-Task cells, the rounded mean equals majority vote across trials\. Appendix Table[2](https://arxiv.org/html/2608.00261#A2.T2)therefore reports integer counts out of1212, not thirds\-resolution counts out of3636\. Each replication has9×12×4=4329\\times 12\\times 4=432Director\-Task outcomes and9×12×3=3249\\times 12\\times 3=324animated\-triangles judge cells\. Both benchmarks use three replications\. To decorrelate API calls from time\-of\-day server\-load variance, batches are scheduled across non\-contiguous time windows \(Appendix[A](https://arxiv.org/html/2608.00261#A1)\)\.
### 4\.2Director Task Results
On the canonical Director\-Task block \(Figure[2](https://arxiv.org/html/2608.00261#S4.F2)\), the panel scores71\.3%71\.3\\%on the explicit visibility multi\-select \(a\) but collapses to12\.0%12\.0\\%on the implicit action prompt \(b\), an explicit\-implicit gap of59\.359\.3percentage points\. The collapse matches the egocentric error in the human Director\-Task literature\(Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13); Apperlyet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib15); Dumontheilet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib17)\): on77\.8%77\.8\\%of \(b\)\-cells the model moves the privileged\-view block\. Seven of nine models score0/120/12on \(b\), and claude\-opus\-4\.7 scores1/121/12; gemini\-3\.0\-pro is the exception at12/1212/12\.
The explicit gate \(c\) only partly reopens the trap\. Re\-asking visibility \(c\.Q1\) lifts the panel to63\.9%63\.9\\%and the gated action \(c\.Q2\) to59\.3%59\.3\\%, but the per\-model recoveryP\(c\.Q2\)−P\(b\)P\(\\text\{c\.Q2\}\)\-P\(b\)splits sharply: three models recover by≥80\\geq 80points \(claude\-sonnet\-4\.60/12→10/120/12\\to 10/12, gemini\-3\.5\-flash0/12→12/120/12\\to 12/12, gpt\-5\.50/12→10/120/12\\to 10/12\) and qwen\-3\.5\-plus by6767\(0/12→8/120/12\\to 8/12\), while gpt\-5\.4 stays at0/120/12\. Perspective\-use thus unblocks under the gate for some models but not others\.
Against published human references, the panel\-mean gap of0\.590\.59falls between the TD\-adult mean of≈0\.44\\approx 0\.44\(Begeeret al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib20)\)controls0\.430\.43,Dumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)adults0\.440\.44\) and theDumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)child range,0\.590\.59\(ages14\.014\.0–17\.717\.7\) to0\.720\.72\(ages7\.37\.3–9\.79\.7\); the HF\-ASD\-adult gap is smaller still at0\.340\.34, consistent with the rule\-based heuristic reported there\(Begeeret al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib20)\)\. A reach\-based variant brackets the same child–adult separation\(Epleyet al\.,[2004](https://arxiv.org/html/2608.00261#bib.bib21)\): young children \(n=33, mean age6\.26\.2y\) make52%52\\%egocentric reaches versus24%24\\%for adults\. We report this correspondence descriptively, without per\-model developmental\-age estimates\.
Two supplementary controls confirm this attribution rather than a perceptual or instruction\-following floor\. A floor control \(Appendix[C](https://arxiv.org/html/2608.00261#A3)\), which makes the \(b\) target the addressee’s own visual extremum, lifts all seven non\-excluded testers \(qwen\-3\.5\-plus and kimi\-k2\.6 excluded for latency\) to letter\-level≥11/12\\geq 11/12, including the six scoring≤1/12\\leq 1/12canonically\. A three\-to\-four block\-count control \(Appendix[D](https://arxiv.org/html/2608.00261#A4)\) leaves the \(b\) egocentric rate flat, marking trap activation as a categorical model property rather than a graded working\-memory bottleneck\.
Figure 2:Director\-Task per\-model accuracy on the four color\-word sub\-prompts \(Sec\.[3\.1](https://arxiv.org/html/2608.00261#S3.SS1)\)\. The panel contains nine models, twelve canonical scenes, and three trials per \(model, scene, sub\-prompt\) cell\. The \(a\)–\(b\) gap is the explicit\-implicit gap\. Seven of nine models score zero on \(b\), claude\-opus\-4\.7 scores1/121/12, and gemini\-3\.0\-pro is the exception at12/1212/12\. Model colors match Figs\.[3](https://arxiv.org/html/2608.00261#S4.F3),[4](https://arxiv.org/html/2608.00261#S5.F4), and[5](https://arxiv.org/html/2608.00261#A5.F5)\. Per\-model accuracies appear in Appendix Table[2](https://arxiv.org/html/2608.00261#A2.T2)\.
### 4\.3Animated Triangles Results
Figure[3](https://arxiv.org/html/2608.00261#S4.F3)shows each model’s ToM\-condition \(Intent, Approp\) profile against theCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)TD\-adult and HF\-ASD\-adult group means\.
Figure 3:Per\-model \(Intent, Approp\) profile on the animated\-triangles ToM condition, using the metric from Section[3\.2](https://arxiv.org/html/2608.00261#S3.SS2)and the anchor\-blinded rubric from Appendix[J](https://arxiv.org/html/2608.00261#A10)\. Each filled circle is one frontier VLM, marked by monogram and colored by lab\. Stars mark the TD\-adult and HF\-ASD\-adult group means fromCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)\. All nine models lie closer to HF\-ASD\-adult than to TD\-adult on this condition\.All nine models fall below both TD\-adult ToM means \(Intent4\.34\.3, Approp1\.71\.7\)\. The panel\-mean ToM profile\(2\.31,0\.94\)\(2\.31,0\.94\)lies2\.132\.13from TD\-adult and0\.730\.73from HF\-ASD\-adult\(2\.9,0\.5\)\(2\.9,0\.5\); the GD mean\(1\.97,1\.98\)\(1\.97,1\.98\)lies0\.510\.51and0\.800\.80from the two groups, and the Random mean\(0\.92,2\.57\)\(0\.92,2\.57\)lies0\.880\.88and1\.081\.08\. ToM is the only condition whose panel mean sits closer to HF\-ASD\-adult than to TD\-adult\.
The asymmetry holds model by model: all nine models lie nearer HF\-ASD\-adult than TD\-adult on ToM \(per\-model coordinates and per\-condition breakdowns in Appendix Tables[7](https://arxiv.org/html/2608.00261#A7.T7),[5](https://arxiv.org/html/2608.00261#A5.T5)\)\. These are geometric distances to published group means, not clinical assessments\. The means are human\-rated, while our scores are LLM\-rated; we treat them as comparable because the judge panel recovers human\-anchored exemplars within±0\.5\\pm 0\.5on the 14\-anchor battery \(Appendix[N](https://arxiv.org/html/2608.00261#A14),[M](https://arxiv.org/html/2608.00261#A13)\)\. Agreement between judges and humans on production cells is unmeasured, a noise floor we flag in Limitations\.
Per\-tester ranks are more heterogeneous than the panel\-mean collapse implies: the strict Intent orderToM\>GD\>Random\\text\{ToM\}\>\\text\{GD\}\>\\text\{Random\}holds for five of nine testers \(claude\-opus\-4\.7, gemini\-3\.0\-pro, gemini\-3\.5\-flash, gpt\-5\.5, qwen\-3\.5\-plus\), even though the panel\-median ToM and GD Intent coincide at2\.02\.0\. The recoverers are the models that benefit from explicit temporal cueing; models that fail under both cued and uncued formats drive the ToM collapse \(stimulus\-format ablation, Appendix[K](https://arxiv.org/html/2608.00261#A11)\)\.
The ToM\-toward\-HF\-ASD shift is not an artifact of the stimulus encoding: on a four\-tester subset it survives both a static\-keyframe and a time\-reversed control \(Appendix[F](https://arxiv.org/html/2608.00261#A6)\), with ToM nearest HF\-ASD\-adult in all three arms\. The time\-reversed arm also reproduces the forward\-time panel mean almost exactly, showing the panel is largely insensitive to temporal direction, consistent with the motion\-pattern\-matching shortcut of Appendix[I](https://arxiv.org/html/2608.00261#A9)\.
## 5Cross\-Task Dissociation Analysis
### 5\.1Joint Dissociation
Figure[4](https://arxiv.org/html/2608.00261#S5.F4)places the nine frontier VLMs and the relevant human reference groups in the joint coordinate system of Section[3\.3](https://arxiv.org/html/2608.00261#S3.SS3)\. The X axis is the Director\-Task explicit\-implicit gap; the Y axis is the animated\-triangles ToM\-condition profile distance to TD\-adult, used only as a visualization anchor\.
Figure 4:Joint coordinates of the nine frontier VLMs \(filled circles, two\-letter monograms per Fig\.[3](https://arxiv.org/html/2608.00261#S4.F3)\) and the two adult reference groups fromCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)\(stars\)\. X axis: Director\-Task explicit\-implicit gapP\(a\)−P\(b\)P\(a\)\-P\(b\)\(TD\-adult0\.4350\.435, HF\-ASD\-adult0\.3400\.340; sources in Sec\.[3\.3](https://arxiv.org/html/2608.00261#S3.SS3)\)\. Y axis: animated\-triangles ToM\-condition distance to TD\-adult, a visualization anchor only; the triangles nearest\-reference rule uses the full 2D \(Intent, Approp\) space\. For eight of nine models the nearest adult reference differs between the two tasks; the full cross\-tabulation appears in Appendix Table[8](https://arxiv.org/html/2608.00261#A8.T8)\.The per\-model nearest\-reference cross\-tabulation is summarized below and reported in full in Appendix Table[8](https://arxiv.org/html/2608.00261#A8.T8)\. The reference set is restricted to the two adult cohorts \(TD\-adult and HF\-ASD\-adult\) on both tasks, sinceCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)publish animated\-triangles group means only for adult cohorts and no equivalent children publication exists; this is the largest reference set for which a like\-for\-like cross\-task comparison is defined\. The wider age\-graded Director\-Task references ofDumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)underpin the Director\-Task panel\-mean gap reported above and the anchor\-sensitivity analysis of Appendix[H](https://arxiv.org/html/2608.00261#A8)\. Under that wider Director anchor set, most panel models reassign to a child or adolescent cohort rather than to TD\-adult\. This follows from the panel\-mean Director gap of0\.590\.59exceeding the TD\-adult anchor of0\.4350\.435and does not change the cross\-task disagreement pattern reported below\. For eight of nine models the nearest\-reference pair\(nearest\_refA,nearest\_refB\)\(\\operatorname\{nearest\\\_ref\}\_\{A\},\\operatorname\{nearest\\\_ref\}\_\{B\}\)is unequal, all eight in the same direction \(nearer TD\-adult on the Director\-Task gap scalar but nearer HF\-ASD\-adult on the animated\-triangles 2D ToM\-condition plane\)\. “Nearest” here is restricted to the two adult anchors and is not a statement of absolute proximity; the panel\-mean Director\-Task gap of0\.590\.59itself exceeds the TD\-adult anchor of0\.4350\.435\. The exception is gemini\-3\.0\-pro, which scores12/1212/12on every Director\-Task sub\-prompt\. Its explicit\-implicit gap of0sits closer to the HF\-ASD\-adult anchor \(0\.340\.34\) than to TD\-adult \(0\.4350\.435\), and it is also nearer HF\-ASD\-adult on the animated triangles, making it the only panel model whose nearest adult anchor is HF\-ASD\-adult on both tasks \(a relative\-distance artifact at the saturated end of the Director Task, not a substantive alignment claim\)\. For eight of the nine panel members no single adult reference group is jointly closest under both the Director\-Task gap scalar and the animated\-triangles 2D profile distance, so no single adult reference profile is nearest on both tasks for these eight models in our panel\. The animated\-triangles half does not change under the stimulus\-format ablation in Appendix[K](https://arxiv.org/html/2608.00261#A11): the plain4×44\\times 4grid and the numbered\-grid variant agree on every per\-condition nearest\-group verdict\.
### 5\.2Cross\-Benchmark Rank Association
As a secondary descriptive check, we summarize whether the per\-model rankings on the two benchmarks are monotonically associated\. The Spearman rank correlation between Director\-Task explicit\-implicit gap and animated\-triangles ToM\-condition TD\-distance isρ=−0\.19\\rho=\-0\.19\(point estimate,n=9n=9, ties broken by mid\-rank\)\. Atn=9n=9this is an underpowered estimate \(the 95% non\-parametric interval is wide enough to be uninformative\), consistent with no clear monotonic rank association in this panel\.
## 6Discussion and Conclusion
VLM ToM in this panel is task\-dependent, so a single benchmark is not enough\. Future suites should report multi\-benchmark profiles and treat profile dissociation as a main outcome\. The numbered\-grid ablation supports this reading: strict Intent rankToM\>GD\>Random\\text\{ToM\}\>\\text\{GD\}\>\\text\{Random\}rises from zero of nine testers to five of nine without changing any per\-condition nearest\-group verdict \(Appendix[K](https://arxiv.org/html/2608.00261#A11)\)\. This points to a temporal\-order bottleneck for the five recovering testers; the other four fail regardless of cue\.
Three design implications follow\. Suites should pair at least two psychology\-derived sub\-capacities under one model panel, report per\-task per\-model profiles rather than single aggregate accuracy, and pair abstract with naturalistic stimuli to separate social\-cue performance from explicit mental\-state inference\.
Frontier VLMs still fall short on these two low\-social\-cue ToM probes\. For eight of nine models, the nearest adult reference differs between perspective\-taking and intention\-attribution profile space\. The main result is therefore a panel\-level dissociation, not alignment with one adult human reference profile\.
## Limitations
Our finding is panel\-level descriptive and not an individual\-model diagnostic claim\. Both benchmarks useKtester=3K\_\{\\text\{tester\}\}=3independently seeded trials per \(tester, item\) cell, and the per\-cell value reported throughout is the mean across the three trials\. Animated\-triangles scoring is LLM\-as\-rater rather than human rater, while Castelli’s reference group means come from human\-rated free\-text; this pipeline mismatch is a measurement noise floor that 14\-anchor rubric calibration cannot fully eliminate \(see Appendix[J](https://arxiv.org/html/2608.00261#A10)\)\. Our Director\-Task implementation uses simplified block\-only stimuli with color\-word scoring rather than the full Director Task with physical action selection\. Two benchmarks alone do not cover the full breadth of ToM; false belief, theory\-driven affect, and language\-based ToM reasoning are left for future work\. The Director\-Task floor and block\-count controls \(Appendices[C](https://arxiv.org/html/2608.00261#A3)and[D](https://arxiv.org/html/2608.00261#A4)\) cover seven of the nine testers\. qwen\-3\.5\-plus and kimi\-k2\.6 were excluded from those controls for latency reasons, so the perspective\-taking attribution and the categorical/graded distinction extend cleanly to the seven covered testers\. We abstain from those attributions for the excluded two\. The animated\-triangles motion\-probe controls \(Appendix[F](https://arxiv.org/html/2608.00261#A6)\) cover four of the nine testers and run atKtester=1K\_\{\\text\{tester\}\}=1rather than the canonicalKtester=3K\_\{\\text\{tester\}\}=3; the per\-arm panel\-mean shifts reported there are point estimates\. We did not pre\-register the analyses\. The C4 cross\-task disagreement is a within\-adult\-anchor comparison, not a general claim that VLM behavior fails to match any human reference profile; the two\-adult reference set is the largest like\-for\-like cross\-task set defined in the published psychology literature, sinceCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)publish animated\-triangles group means only for adult cohorts\. Appendix[H](https://arxiv.org/html/2608.00261#A8)reports the asymmetric Director\-side wider\-anchor sensitivity, where most panel models reassign to child or adolescent anchors on the Director Task; the exact nearest\-anchor labels in C4 should not be generalized beyond the available two\-adult cross\-task anchor set\.
## Ethical considerations
This paper compares frontier VLM behavior against published group\-mean profiles from TD\-adult, HF\-ASD\-adult, and age\-graded children cohorts, all of which appear in the public psychology literature\(Castelliet al\.,[2002](https://arxiv.org/html/2608.00261#bib.bib12); Apperlyet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib15); Keysaret al\.,[2000](https://arxiv.org/html/2608.00261#bib.bib13); Begeeret al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib20); Dumontheilet al\.,[2010](https://arxiv.org/html/2608.00261#bib.bib17); Epleyet al\.,[2004](https://arxiv.org/html/2608.00261#bib.bib21)\)\. We use these reference group means only as quantitative anchors for distance comparisons, and we avoid per\-model diagnostic labels, developmental\-age point estimates, and anthropomorphic clinical equivalences \(see Section[Limitations](https://arxiv.org/html/2608.00261#Sx1)\)\. We use the HF\-ASD\-adult group mean as a reference profile in a descriptive comparison and not as a label or a value judgment on any model, system, or person\. Our benchmarks contain no human subjects data and no personally identifying information\. The Frith\-Happé animated\-triangles clips we use are the eLife re\-edits ofDureuxet al\.\([2023](https://arxiv.org/html/2608.00261#bib.bib9)\), released under CC BY 4\.0\.
## References
- F\. Abell, F\. Happé, and U\. Frith \(2000\)Do triangles play tricks? attribution of mental states to animated shapes in normal and abnormal development\.Cognitive Development15\(1\),pp\. 1–16\.External Links:ISSN 0885\-2014,[Link](http://dx.doi.org/10.1016/S0885-2014(00)00014-9),[Document](https://dx.doi.org/10.1016/s0885-2014%2800%2900014-9)Cited by:[2nd item](https://arxiv.org/html/2608.00261#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p3.1),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- N\. K\. Andersen, M\. K\. Rimvall, P\. Jeppesen, M\. Bentz, J\. R\. M\. Jepsen, L\. Clemmensen, R\. K\. Jacobsen, and E\. M\. Olsen \(2022\)A psychometric investigation of the multiple\-choice version of animated triangles task to measure theory of mind in adolescence\.PLOS ONE17\(3\),pp\. e0264319\.External Links:ISSN 1932\-6203,[Link](http://dx.doi.org/10.1371/journal.pone.0264319),[Document](https://dx.doi.org/10.1371/journal.pone.0264319)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- I\. A\. Apperly, D\. J\. Carroll, D\. Samson, G\. W\. Humphreys, A\. Qureshi, and G\. Moffitt \(2010\)Why are there limits on theory of mind use? evidence from adults’ ability to follow instructions from an ignorant speaker\.Quarterly Journal of Experimental Psychology63\(6\),pp\. 1201–1217\.External Links:ISSN 1747\-0226,[Link](http://dx.doi.org/10.1080/17470210903281582),[Document](https://dx.doi.org/10.1080/17470210903281582)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p1.7),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- S\. Begeer, B\. F\. Malle, M\. S\. Nieuwland, and B\. Keysar \(2010\)Using theory of mind to represent and take part in social interactions: comparing individuals with high\-functioning autism and typically developing controls\.European Journal of Developmental Psychology7\(1\),pp\. 104–122\.External Links:[Document](https://dx.doi.org/10.1080/17405620903024263)Cited by:[Appendix H](https://arxiv.org/html/2608.00261#A8.SS0.SSS0.Px2.p1.4),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§3\.3](https://arxiv.org/html/2608.00261#S3.SS3.p2.3),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p3.14),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- F\. Castelli, C\. Frith, F\. Happé, and U\. Frith \(2002\)Autism, asperger syndrome and brain mechanisms for the attribution of mental states to animated shapes\.Brain125\(8\),pp\. 1839–1849\.External Links:ISSN 1460\-2156,[Link](http://dx.doi.org/10.1093/BRAIN/AWF189),[Document](https://dx.doi.org/10.1093/brain/awf189)Cited by:[Appendix J](https://arxiv.org/html/2608.00261#A10.SS0.SSS0.Px2.p1.1),[Appendix M](https://arxiv.org/html/2608.00261#A13.SS0.SSS0.Px1.p1.1),[Appendix H](https://arxiv.org/html/2608.00261#A8.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p1.1),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2608.00261#S3.SS2.SSS0.Px1.p1.3),[§3\.3](https://arxiv.org/html/2608.00261#S3.SS3.p2.3),[Figure 3](https://arxiv.org/html/2608.00261#S4.F3),[§4\.3](https://arxiv.org/html/2608.00261#S4.SS3.p1.1),[Figure 4](https://arxiv.org/html/2608.00261#S5.F4),[§5\.1](https://arxiv.org/html/2608.00261#S5.SS1.p2.10),[Limitations](https://arxiv.org/html/2608.00261#Sx1.p1.3),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- F\. Castelli, F\. Happé, U\. Frith, and C\. Frith \(2000\)Movement and mind: a functional imaging study of perception and interpretation of complex intentional movement patterns\.NeuroImage12\(3\),pp\. 314–325\.External Links:ISSN 1053\-8119,[Link](http://dx.doi.org/10.1006/NIMG.2000.0612),[Document](https://dx.doi.org/10.1006/nimg.2000.0612)Cited by:[Appendix M](https://arxiv.org/html/2608.00261#A13.SS0.SSS0.Px1.p1.1),[2nd item](https://arxiv.org/html/2608.00261#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p3.1),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2608.00261#S3.SS2.p5.1)\.
- I\. Dumontheil, I\. A\. Apperly, and S\. Blakemore \(2010\)Online usage of theory of mind continues to develop in late adolescence\.Developmental Science13\(2\),pp\. 331–338\.External Links:ISSN 1467\-7687,[Link](http://dx.doi.org/10.1111/j.1467-7687.2009.00888.x),[Document](https://dx.doi.org/10.1111/j.1467-7687.2009.00888.x)Cited by:[Appendix H](https://arxiv.org/html/2608.00261#A8.SS0.SSS0.Px2.p1.4),[Appendix H](https://arxiv.org/html/2608.00261#A8.SS0.SSS0.Px3.p1.13),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§3\.3](https://arxiv.org/html/2608.00261#S3.SS3.p2.3),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p1.7),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p3.14),[§5\.1](https://arxiv.org/html/2608.00261#S5.SS1.p2.10),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- A\. Dureux, A\. Zanini, J\. Selvanayagam, R\. S\. Menon, and S\. Everling \(2023\)Gaze patterns and brain activations in humans and marmosets in the frith\-happé theory\-of\-mind animation task\.eLife12,pp\. e86327\.External Links:ISSN 2050\-084X,[Link](http://dx.doi.org/10.7554/eLife.86327),[Document](https://dx.doi.org/10.7554/elife.86327)Cited by:[2nd item](https://arxiv.org/html/2608.00261#S1.I1.i2.p1.1),[§3\.2](https://arxiv.org/html/2608.00261#S3.SS2.p2.1),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- N\. Epley, C\. K\. Morewedge, and B\. Keysar \(2004\)Perspective taking in children and adults: equivalent egocentrism but differential correction\.Journal of Experimental Social Psychology40\(6\),pp\. 760–768\.External Links:[Document](https://dx.doi.org/10.1016/j.jesp.2004.02.002)Cited by:[Appendix H](https://arxiv.org/html/2608.00261#A8.SS0.SSS0.Px2.p1.4),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p3.14),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- J\. H\. Flavell, B\. A\. Everett, K\. Croft, and E\. R\. Flavell \(1981\)Young children’s knowledge about visual perception: further evidence for the level 1–level 2 distinction\.\.Developmental Psychology17\(1\),pp\. 99–103\.External Links:ISSN 0012\-1649,[Link](http://dx.doi.org/10.1037/0012-1649.17.1.99),[Document](https://dx.doi.org/10.1037/0012-1649.17.1.99)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- Q\. Gao, Y\. Li, H\. Lyu, H\. Sun, D\. Luo, and H\. Deng \(2024\)Vision language models see what you want but not what you see\.External Links:2410\.00324,[Link](https://arxiv.org/abs/2410.00324)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2608.00261#S2.T1.1.6.5.1.1.1)\.
- Y\. Gu, O\. Tafjord, H\. Kim, J\. Moore, R\. L\. Bras, P\. Clark, and Y\. Choi \(2026\)SimpleToM: exposing the gap between explicit tom inference and implicit tom application in llms\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=iE2JmbRJow)Cited by:[§3\.1](https://arxiv.org/html/2608.00261#S3.SS1.p5.1)\.
- F\. Heider and M\. Simmel \(1944\)An experimental study of apparent behavior\.The American Journal of Psychology57\(2\),pp\. 243–259\.External Links:ISSN 0002\-9556,[Link](http://dx.doi.org/10.2307/1416950),[Document](https://dx.doi.org/10.2307/1416950)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Jin, Y\. Wu, J\. Cao, J\. Xiang, Y\. Kuo, Z\. Hu, T\. Ullman, A\. Torralba, J\. Tenenbaum, and T\. Shu \(2024\)MMToM\-QA: multimodal theory of mind question answering\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 16077–16102\.External Links:[Link](https://aclanthology.org/2024.acl-long.851/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.851)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Keysar, D\. J\. Barr, J\. A\. Balin, and J\. S\. Brauner \(2000\)Taking perspective in conversation: the role of mutual knowledge in comprehension\.Psychological Science11\(1\),pp\. 32–38\.External Links:ISSN 1467\-9280,[Link](http://dx.doi.org/10.1111/1467-9280.00211),[Document](https://dx.doi.org/10.1111/1467-9280.00211)Cited by:[Figure 1](https://arxiv.org/html/2608.00261#S1.F1),[2nd item](https://arxiv.org/html/2608.00261#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p1.1),[§1](https://arxiv.org/html/2608.00261#S1.p3.1),[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.00261#S3.SS1.p2.1),[§4\.2](https://arxiv.org/html/2608.00261#S4.SS2.p1.7),[Ethical considerations](https://arxiv.org/html/2608.00261#Sx2.p1.1)\.
- Y\. Li, V\. Veerabadran, M\. L\. Iuzzolino, B\. D\. Roads, A\. Celikyilmaz, and K\. Ridgeway \(2025\)EgoToM: benchmarking theory of mind reasoning from egocentric videos\.External Links:2503\.22152,[Link](https://arxiv.org/abs/2503.22152)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px2.p1.1)\.
- L\. A\. Livingston, B\. Carr, and P\. Shah \(2019\)Recent advances and new directions in measuring theory of mind in autistic adults\.Journal of Autism and Developmental Disorders49\(4\),pp\. 1738–1744\.External Links:ISSN 1573\-3432,[Link](http://dx.doi.org/10.1007/s10803-018-3823-3),[Document](https://dx.doi.org/10.1007/s10803-018-3823-3)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Piaget and B\. Inhelder \(1956\)The child’s conception of space\.Routledge and Kegan Paul,London\.Note:English translation by F\. J\. Langdon and J\. L\. Lunzer; original French edition Presses Universitaires de France, 1948Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Villa\-Cueva, S\. M\. M\. Ahmed, R\. Chevi, J\. C\. B\. Cruz, K\. Elzeky, F\. Cristobal, A\. F\. Aji, S\. Wang, R\. Mihalcea, and T\. Solorio \(2025\)MoMentS: a comprehensive multimodal benchmark for theory of mind\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 22591–22611\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1230/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1230)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px2.p1.1)\.
- Q\. Wang, B\. Yin, P\. Zhang, J\. Zhang, K\. Wang, Z\. Wang, J\. Zhang, K\. Chandrasegaran, H\. Liu, R\. Krishna, S\. Xie, J\. Wu, L\. Fei\-Fei, and M\. Li \(2026\)MindCube: spatial mental modeling from limited views\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0FhrtdKLtD)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px2.p1.1)\.
- S\. J\. White, D\. Coniston, R\. Rogers, and U\. Frith \(2011\)Developing the frith\-happé animations: a quick and objective test of theory of mind for adults with autism\.Autism Research4\(2\),pp\. 149–154\.External Links:ISSN 1939\-3792,[Link](http://dx.doi.org/10.1002/aur.174),[Document](https://dx.doi.org/10.1002/aur.174)Cited by:[§2](https://arxiv.org/html/2608.00261#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AImplementation Firewall Details
Both benchmarks share a code\-enforced two\-firewall harness that hides tester identity from judges and hides rater\-side files from testers; we describe each firewall separately below rather than label the combination “double\-blind”, since the rater is deliberately given the clip’s ground\-truth condition label and script semantics in keeping with Castelli’s original human\-rater protocol\. The first firewall bars testers from reading any path under the rater directory and from reading the test\-set registry that maps tester display names to backing IDs\. The second firewall bars judges from reading tester short names; only opaque tester IDs reach the judge prompt\. All run artifacts are written atomically append\-only with an open\-exclusive create flag so that no past run can be silently rewritten\.
#### Temporal decorrelation against server\-load variance\.
API call batches in this work, spanning the two main\-panel benchmarks, every calibration pre\-experiment, and every shuffleseed replicate, are scheduled at different times of day rather than executed in a single contiguous burst\. Time\-correlated load variance on the model providers’ inference servers, including throttling, queue depth, and concurrent traffic spikes, is therefore averaged across calls rather than concentrated within one window\. Concretely, the animated\-triangles plain\-grid and numbered\-grid main panels are launched approximately 22 hours apart, theKtester=3K\_\{\\text\{tester\}\}=3replicates of the frame\-representation pre\-experiment \(Appendix[L](https://arxiv.org/html/2608.00261#A12)\) span two calendar days, and the rubric and judge validation battery \(Appendix[M](https://arxiv.org/html/2608.00261#A13)\), the judge model ablation \(Appendix[N](https://arxiv.org/html/2608.00261#A14)\), and the Director\-Task shuffleseed replicates are each launched on distinct calendar days\.
The full harness, scoring scripts, and per\-\(tester, clip, judge\) raw outputs will be released with the camera\-ready\.
## Appendix BPer\-Model Director Task Sub\-Prompt Accuracy
Table[2](https://arxiv.org/html/2608.00261#A2.T2)reports the per\-model raw correct count out of1212canonical Director\-Task scenes on each of the four sub\-prompts\. The egocentric\-error rate on sub\-prompt \(b\), reported as “ego\. \(b\)” in the rightmost column, is the fraction of cells in which the model picks the privileged\-view block that would be ambiguous for the addressee but is in fact occluded from the director\. Seven of nine models score0correct on sub\-prompt \(b\) and claude\-opus\-4\.7 scores11of1212; the lone exception that does not exhibit the egocentric failure mode is gemini\-3\.0\-pro at12/1212/12\.
Modelabc\.Q1c\.Q2ego\. \(b\)opus\-4\.7111829sonnet\-4\.6807109gemini\-3\.0\-pro121212120gemini\-3\.5\-flash100111212gpt\-5\.41003012gpt\-5\.580101010grok\-4\.3706612kimi\-k2\.650449qwen\-3\.5\-plus608811panel mean \(%\)71\.312\.063\.959\.377\.8Table 2:Director\-Task per\-model raw counts \(out of1212canonical scenes\) on each sub\-prompt, plus egocentric\-error count on sub\-prompt \(b\)\. Sub\-prompt \(a\) is the multi\-select “which blocks does the director see”\. Sub\-prompt \(b\) is the single\-select know\-but\-don’t\-use trap\. Sub\-prompt \(c\) is split into Q1 \(explicit re\-asking of a\) and Q2 \(gated action after Q1\)\. The panel\-mean Director\-Task explicit\-implicit gapP\(a\)−P\(b\)P\(a\)\-P\(b\)is59\.359\.3percentage points\.
## Appendix CDirector Task Perceptual and Instruction Floor Control
#### Design\.
We run two control arms on the canonical twelve Director\-Task scenes\. The no\-occluder arm removes the occluder so all three blocks are visible to both addressee and director; the \(a, b, c\.Q1, c\.Q2\) sub\-prompt structure is unchanged\. The same\-side\-director arm keeps the occluder but moves the director to the addressee’s side of the table, so the occluder does not occlude anything from the director \(occluder is physically present but geometrically inert\)\. Both arms invert the sub\-\(b\) ground\-truth target onto the addressee’s visual extremum, so a model that simply picks the visual extremum scores correctly on the control sub\-\(b\)\. The remaining failure modes are perceptual mis\-identification \(cannot see all three blocks, mis\-counts blocks, mis\-takes occluder for a block\) or instruction failure \(does not understand “largest”/“smallest”\)\.
#### Panel and pre\-registered floor pass criterion\.
We run both arms on the same nine\-model panel minus qwen\-3\.5\-plus and kimi\-k2\.6, excluded for latency reasons \(the same exclusion holds for Appendix[D](https://arxiv.org/html/2608.00261#A4)\)\. The seven testers retained are claude\-opus\-4\.7, claude\-sonnet\-4\.6, gpt\-5\.4, gpt\-5\.5, gemini\-3\.0\-pro, gemini\-3\.5\-flash, and grok\-4\.3\. The canonical\-\(b\) attribution for the two excluded testers is therefore not covered by this control\. The floor pass criterion, fixed before running paid API calls, is per\-model control sub\-\(b\) letter\-level correctness≥11/12\\geq 11/12, where letter\-level marks the correct letter regardless of color naming \(failure\-mode∈\{correct,letter\_only\_correct\}\\in\\\{\\text\{correct\},\\text\{letter\\\_only\\\_correct\}\\\}\)\. The letter\-level threshold isolates the perceptual / instruction floor from the orthogonal blue→\\tocyan color\-naming drift that appears panel\-wide on the closed color vocabulary\.
#### Result\.
All seven testers pass the floor on both arms \(7/7 on each arm\)\. Table[3](https://arxiv.org/html/2608.00261#A3.T3)reports per\-model canonical\-\(b\) versus control\-\(b\) letter\-level scores\.
Table 3:Director\-Task floor control\. canon\-\(b\) is from the canonical 12\-scene Director Task; the no\-occluder arm removes the occluder; the same\-side\-director arm keeps the occluder but moves the director to the addressee side\. Letter\-level==correct letter regardless of color name; strict pair\-level results lie2/122/12below letter\-level for every cell because Q01 and Q02 expect blue and the panel reads blue as cyan\. The six models with canonical≤1/12\\leq 1/12score≥11/12\\geq 11/12on both control arms, attributing the canonical collapse to perspective\-taking rather than perceptual or instruction floor\.
#### Independent side finding on the same\-side\-director arm\.
On the same\-side\-director arm, the two Claude testers show degraded sub\-\(a\) and sub\-c\.Q1 identification despite passing the sub\-\(b\) floor\. claude\-opus\-4\.7 scores sub\-\(a\)7/127/12\(no\-occluder arm:12/1212/12\) and c\.Q14/124/12; claude\-sonnet\-4\.6 scores sub\-\(a\)12/1212/12but c\.Q19/129/12and c\.Q27/127/12\. This is consistent with the two Claude models maintaining a residual “the occluder physically blocks the director’s view” inference even when the geometry no longer supports it\. The finding is independent of the floor result and does not affect the attribution conclusion; the floor judgment lives in sub\-\(b\) only\.
## Appendix DDirector Task Variable Block Count
#### Design\.
We add a four\-block variant of the Director Task structure to test whether the canonical sub\-\(b\) failure rate scales with the number of competing blocks \(graded, working\-memory\-style\) or is categorical \(binary, the model either represents the director’s view or it does not\)\. Four blocks of distinct sizes are placed in front of the director\. Exactly one block \(the global extremum under the director’s instruction\) is occluded from the director by the partition; the director then issues an extremum instruction \(“move the largest/smallest block to the right”\) and the ground\-truth target is the extremum among the director\-visible set \(the second\-largest or second\-smallest of the four\)\. The \(a, b, c\.Q1, c\.Q2\) sub\-prompt structure and the deterministic regex\-based scoring are unchanged\. We sample twelve balanced four\-block scenes \(six largest\-instruction, six smallest\-instruction, with the egocentric\-trap position balanced across the four block positions\)\. The n=3 baseline used for the matched\-pair comparison is the same seven testers on the same closed\-color\-vocabulary scorer as the four\-block run\.
#### Panel\.
Same seven testers as Appendix[C](https://arxiv.org/html/2608.00261#A3)\(qwen\-3\.5\-plus and kimi\-k2\.6 excluded for latency\)\.
#### Result\.
Table[4](https://arxiv.org/html/2608.00261#A4.T4)reports per\-tester sub\-\(b\) direct\-action correct and egocentric error rate at n=3 versus n=4 competing blocks\.
Table 4:Director\-Task sub\-\(b\) direct\-action correct and egocentric error rate at n=3 vs n=4 competing blocks, seven\-tester panel\. The sole non\-collapsing model at n=3 \(gemini\-3\.0\-pro\) stays at12/1212/12correct under n=4\. The six collapsing models stay at0/120/12direct\-action correct with≥7/12\\geq 7/12egocentric error under n=4\. Both ends of the panel are flat with respect to block count\.
#### Take\-away\.
The sub\-\(b\) failure rate is flat with block count in this panel, neither the capable model nor the six collapsed models move with the additional competing block\. This is consistent with sub\-\(b\) failure being a categorical model property \(the model either represents the director’s view in action or it does not\) rather than a graded working\-memory or attentional bottleneck that should grow with competing items\. The explicit\-gate CoT recovery on c\.Q2 does decline with the extra block for some models \(claude\-sonnet\-4\.610→510\\to 5, gpt\-5\.510→910\\to 9, claude\-opus\-4\.72→12\\to 1\), while remaining steady for gemini\-3\.5\-flash \(12→1212\\to 12\) and gpt\-5\.4 \(0→00\\to 0\); the recovery channel is graded for the models whose recovery is CoT\-mediated, even when the underlying trap activation is categorical\.
## Appendix EPer\-Condition Breakdown for the Animated Triangles
The main text reports the animated\-triangles ToM\-condition \(Intent, Approp\) profile and the corresponding profile distances under the numbered\-grid canonical metric\. Figure[5](https://arxiv.org/html/2608.00261#A5.F5)and Table[5](https://arxiv.org/html/2608.00261#A5.T5)report the per\-condition \(GD, Random\) panels and the per\-condition profile distances under the same v2 canonical metric\.
Figure 5:Per\-model \(Intent, Approp\) profile on the animated triangles broken down by condition\. Each panel shows the per\-model points, colored by lab\. The canonical ToM\-condition panel is also reported as Figure[3](https://arxiv.org/html/2608.00261#S4.F3)in the main text\.Table 5:Per\-model 2D Euclidean profile distance to TD\-adult and HF\-ASD\-adult group means under each animated\-triangles condition \(ToM, GD, Random\), numbered\-grid stimulus, anchor\-blinded canonical rubric per Appendix[J](https://arxiv.org/html/2608.00261#A10)\. The ToM\-condition columns are the canonical animated\-triangles metric \(Sec\.[3\.2](https://arxiv.org/html/2608.00261#S3.SS2)\)\. On the ToM condition, all nine models have larger distance to TD\-adult than to HF\-ASD\-adult\.
## Appendix FAnimated\-Triangles Motion\-Probe Controls
#### Design\.
We run two stimulus\-format controls on the canonical twelve animated\-triangles clips\. The static\-keyframe arm replaces each4×44\\times 4composite with a single mid\-clip frame in native1280×7201280\\times 720resolution\. The time\-reversed arm keeps the same sixteen\-cell4×44\\times 4composite and the11–1616temporal\-order badges, but reverses the underlying frame order so that the cell labelled “frame 1” shows what was originally the last sampled frame; the tested model is not told the order is reversed\. Both controls use the same Castelli rubric, the same three\-judge cross\-vendor panel, and the same anchor\-blinded canonical scoring as the main animated\-triangles run\.
#### Panel and replication\.
Four testers, the per\-lab best animated\-triangles performer for the four Western labs \(claude\-opus\-4\.7, gpt\-5\.5, gemini\-3\.5\-flash, grok\-4\.3\); qwen\-3\.5\-plus and kimi\-k2\.6 are excluded for latency, the same exclusion as in Appendices[C](https://arxiv.org/html/2608.00261#A3)and[D](https://arxiv.org/html/2608.00261#A4)\. Both controls run withKtester=1K\_\{\\text\{tester\}\}=1per \(tester, clip\) cell; the canonical forward\-time reference used here is the nine\-model main run sub\-selected to the same four testers\.
#### Result\.
Table[6](https://arxiv.org/html/2608.00261#A6.T6)reports the per\-condition panel\-mean \(Intent, Approp\) for the three arms and the Euclidean distance to each adult reference group\.
Table 6:Animated\-triangles motion\-probe controls on the four\-tester subset \(claude\-opus\-4\.7, gpt\-5\.5, gemini\-3\.5\-flash, grok\-4\.3, the per\-lab Western frontier best on the animated triangles\)\. The canonical\-forward arm is the nine\-model main run sub\-selected to the same four testers\. The time\-reversed arm reproduces the canonical forward\-time arm almost exactly on all three conditions\. The static\-keyframe arm shifts the ToM panel mean toward lower Intent and the GD/Random panel means toward lower Approp; ToM nevertheless remains nearer HF\-ASD\-adult than TD\-adult under all three arms\.
#### Take\-away\.
On ToM, the time\-reversed and forward\-time panel means coincide to within∼0\.1\\sim 0\.1on both axes, and both lie at distance∼0\.65\\sim 0\.65to HF\-ASD\-adult and∼1\.8\\sim 1\.8to TD\-adult; reversing temporal direction does not change the asymmetric ToM\-toward\-HF\-ASD shift\. The static\-keyframe arm reduces ToM\-condition panel mean Intent from2\.652\.65to1\.961\.96but the panel mean still lands closer to HF\-ASD\-adult than to TD\-adult\. A complementary effect on the non\-ToM conditions is that the static\-keyframe arm raises the panel mean attributed Intent on Random from0\.980\.98to1\.481\.48and lowers attributed Approp on GD, blurring condition distinctions; the per\-tester strict Intent rankToM\>GD\>Random\\text\{ToM\}\>\\text\{GD\}\>\\text\{Random\}holds for only one of four testers under the static keyframe versus two of four under both grid arms\. Two readings are compatible with this pattern\. First, a large fraction of the ToM\-condition attribution signal is recoverable from a single representative frame, consistent with static composition and style priors \(the shape rendering, the enclosure, the relative positions of the triangles\) carrying much of the mental\-attribution signal\. Second, on the canonical multi\-frame arm the role of the motion structure is largely to suppress over\-attribution to non\-ToM conditions rather than to provide the ToM signal itself\. The time\-reversed result is consistent with the motion\-pattern\-matching residual shortcut noted in Appendix[I](https://arxiv.org/html/2608.00261#A9)\.
#### Caveats\.
Both controls run atKtester=1K\_\{\\text\{tester\}\}=1while the canonical main run usesKtester=3K\_\{\\text\{tester\}\}=3; the controls give point estimates and the per\-arm panel\-mean shifts should be read as descriptive\. The static\-keyframe prompt informs the tested model that it is seeing a single still frame from a short silent animation; this is a design tradeoff to keep the task framing aligned with the multi\-frame arms, not an unintended leak\. The time\-reversed implementation reverses the order of the sixteen sampled frames rather than re\-rendering the clip frame\-by\-frame in reverse, so the temporal\-direction signal the tested model receives is the order of the badges and of the in\-cell stills, which matches what the canonical forward\-time arm provides\.
## Appendix GPer\-Model Distances on the Animated\-Triangles ToM Condition
Table 7:Per\-model ToM\-condition \(Intent, Approp\) coordinates and 2D Euclidean profile distance from each frontier VLM to the TD\-adult\(4\.3,1\.7\)\(4\.3,1\.7\)and HF\-ASD\-adult\(2\.9,0\.5\)\(2\.9,0\.5\)published group means on the animated triangles \(canonical metric per Sec\.[3\.2](https://arxiv.org/html/2608.00261#S3.SS2), numbered\-grid stimulus, anchor\-blinded canonical rubric per Appendix[J](https://arxiv.org/html/2608.00261#A10)\)\.
## Appendix HReference\-Anchor Sensitivity and Fallback Sources
#### Animated\-triangles anchor\.
The animated\-triangles coordinateyyin the joint plot is anchored to TD\-adult for visualization, but the per\-model nearest\-reference rule \(Section[3\.3](https://arxiv.org/html/2608.00261#S3.SS3)\) is computed in the full 2D \(Intent, Approp\) plane and is therefore independent of which adult group is used as the visualization anchor\.Castelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)do not publish age\-graded children animated\-triangles group means, so we cannot extend the reference set on this task; this asymmetry between the two tasks motivates the adult\-only restriction on the cross\-task nearest\-reference cross\-tabulation in Section[5](https://arxiv.org/html/2608.00261#S5)\.
Table 8:Per\-model nearest\-reference cross\-tabulation \(anchor\-blinded canonical rubric\)\.xAx\_\{A\}: Director\-Task gapP\(a\)−P\(b\)P\(a\)\-P\(b\);nearest\_refA\\operatorname\{nearest\\\_ref\}\_\{A\}= closer of TD\-adult \(0\.4350\.435\) and HF\-ASD\-adult \(0\.3400\.340\) under\|x−xH\|\|x\-x\_\{H\}\|;nearest\_refB\\operatorname\{nearest\\\_ref\}\_\{B\}= closer reference on the animated\-triangles \(Intent, Approp\) ToM plane \(Sec\.[3\.3](https://arxiv.org/html/2608.00261#S3.SS3)\)\. Eight of nine disagree, all nearer TD\-adult on the Director Task but HF\-ASD\-adult on the triangles\. The exception, gemini\-3\.0\-pro, saturates the Director Task \(12/1212/12, gap0\): a relative\-distance artifact \(\|0−0\.435\|\>\|0−0\.340\|\|0\-0\.435\|\>\|0\-0\.340\|\), not an alignment claim\.
#### Director\-Task anchor and fallback\.
The Director\-Task coordinatexxuses the published explicit\-implicit gap per group with canonical signx=P\(a correct\)−P\(b correct\)x=P\(\\text\{a correct\}\)\-P\(\\text\{b correct\}\), applied identically to models and human groups\. For a human group with no published explicit\-implicit gap on the same paradigm, the fallback reconstructs the same scalar as published a\-question accuracy minus published b\-question accuracy on the closest published variant \(canonical sign preserved\)\. The per\-group sources currently in use areBegeeret al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib20)\)\(TD\-adult controls and HF\-ASD\-adult\),Dumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)\(TD\-adult and four age\-graded cohorts spanning7\.37\.3–17\.717\.7years\), andEpleyet al\.\([2004](https://arxiv.org/html/2608.00261#bib.bib21)\)as a qualitative bracket \(4–12\-year sample, reach\-rate metric not directly comparable to the gap scalar and therefore plotted separately rather than as an anchor\)\.
#### Sensitivity ofnearest\_refA\\operatorname\{nearest\\\_ref\}\_\{A\}to the reference set\.
Under the two\-adult reference set\{TD\-adult,HF\-ASD\-adult\}\\\{\\text\{TD\-adult\},\\text\{HF\-ASD\-adult\}\\\}used in Table[8](https://arxiv.org/html/2608.00261#A8.T8), eight of nine models map to TD\-adult on the Director Task and one \(gemini\-3\.0\-pro, gap0\) maps to HF\-ASD\-adult\. Under the wider reference set that additionally includes the fourDumontheilet al\.\([2010](https://arxiv.org/html/2608.00261#bib.bib17)\)age\-graded cohorts \(gaps0\.590\.59,0\.670\.67,0\.680\.68,0\.720\.72\), six of nine models reassign to a children/adolescent cohort \(the cohort at gap0\.590\.59for grok\-4\.3 at0\.5830\.583; the cohorts at0\.670\.67–0\.680\.68for sonnet\-4\.6 and gpt\-5\.5 at0\.6670\.667; the cohort at0\.720\.72for opus\-4\.7, gemini\-3\.5\-flash, and gpt\-5\.4 at0\.8330\.833\), kimi\-k2\.6 and qwen\-3\.5\-plus stay at TD\-adult, and gemini\-3\.0\-pro stays at HF\-ASD\-adult\. Crucially, the cross\-task disagreement count is robust to this anchor expansion\. The cross\-task disagreement remains eight of nine because the animated\-triangles nearest reference is HF\-ASD\-adult for all eight non\-gemini\-3\.0\-pro models, and only gemini\-3\.0\-pro keeps the same nearest adult anchor on both tasks under either reference set\.
#### Sensitivity ofnearest\_refB\\operatorname\{nearest\\\_ref\}\_\{B\}to the Y\-axis anchor\.
Swapping the Y\-axis visualization anchor from TD\-adult to HF\-ASD\-adult shifts every Y coordinate by the constant adult\-adult distance \(1\.8441\.844in 2D\) but does not change the 2D Euclidean nearest\-reference assignment, sonearest\_refB\\operatorname\{nearest\\\_ref\}\_\{B\}is invariant under this swap\.
## Appendix IResidual Shortcut Risks
The Director Task reduces social\-cue reliance with block\-and\-color stimuli and an explicit letter\-and\-color double\-grounding scheme\. A natural worry is color\-position matching, in which a model uses block color or scene position alone to predict the intended target rather than reasoning about the director’s perspective\. The canonical 12\-scene set explicitly controls for this by rotating the color palette across three independent triples \(yellow / red / blue; purple / green / orange; red / cyan / green\), rotating the letter\-to\-color mapping per scene \(A, B, C bind to different colors in different scenes\), rotating which letter position carries the target block, and rotating the instruction polarity \(smallest / largest\)\. A color\-position shortcut would have to survive all four rotations to produce the per\-model failure patterns we observe in Table[2](https://arxiv.org/html/2608.00261#A2.T2)\. The animated triangles reduce social\-cue reliance with abstract geometric trajectories\. A residual shortcut is motion\-pattern matching, in which a model classifies trajectories by their kinematic signature alone without invoking mental\-state inference; this shortcut is the target of the queued reversed\-time playback ablation and is bounded but not fully eliminated by the current design\.
## Appendix JLLM\-as\-Rater Considerations and Blinding\-Robustness Audit
#### Mitigations\.
Castelli’s published reference group means come from human\-rated free\-text; we use an LLM\-as\-rater jury on the same rubric\. The jury uses three different model families \(Anthropic, Google, Qwen\) so that no single model family dominates scoring; the cross\-family choice was empirically validated in Appendix[N](https://arxiv.org/html/2608.00261#A14)\. Every rater operates under the Section[3\.2](https://arxiv.org/html/2608.00261#S3.SS2)two\-firewall protocol so no rater knows which tester produced which output \(although the rater does see the clip’s ground\-truth condition label and script semantics, following Castelli’s original protocol\), and all rater outputs are written append\-only so post\-hoc tampering of scores is detectable\. The three production judges were further validated against the 14 human\-anchored paper exemplars of Castelli \(2000, 2002\) under the Phase\-1 battery in Appendix[M](https://arxiv.org/html/2608.00261#A13), each achieving 100% anchor recovery with per\-anchor standard deviation equal to zero across 5 repetitions atT=0T=0\.
#### Information disclosure to the judge\.
Each judge call carries, per clip, three layers of overlay information on top of the SHA\-locked Castelli rubric\. Layer \(i\) is the ground\-truth condition label \(ToM, GD, or Random\) for that clip\. Layer \(ii\) is the animation script semantics for that clip \(a one\-sentence description such as “the large triangle coaxes the small triangle out of the box”\)\. Layer \(i\) and layer \(ii\) follow the original Castelli protocol; human raters inCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)knew the animation script in order to grade Approp against the intended interaction type\. Layer \(iii\) is anExpected score rangeblock giving the Castelli\-2002 human group\-mean as a per\-clip Intent target \(for example, “Expected Intent: 4–5” on ToM clips, derived from the TD\-adult ToM Intent mean of4\.34\.3\)\.
#### Audit finding\.
An audit of the production judge prompt confirmed that layer \(iii\) was not part of any human\-rater protocol\. Inserting the Castelli human group mean as a per\-clip target anchors the judge upward and makes the Stage\-C comparison \(VLM profile against the same Castelli group mean\) partly circular\. The direction of this bias is conservative for the headline finding, since a judge anchored toward TD\-adult should still let the VLM profile drift toward TD\-adult, so a finding of “VLM panel lies below TD\-adult” under this anchored judge is biased against itself\.
#### Blinded L1 re\-score\.
The anchor\-blinded L1 re\-score uses byte\-identical tester responses, the same three\-judge cross\-vendor panel atT=0T=0, the same SHA\-locked rubric, and an overlay that strips only layer \(iii\); layer \(i\) and layer \(ii\) are retained\.
#### Informed\-vs\-anchor\-blinded comparison\.
Table[9](https://arxiv.org/html/2608.00261#A10.T9)reports the per\-arm panel\-level headline metrics under the two judge variants on the same324324canonical cells per arm\.
Metricplaininformedplainanchor\-blindednumberedinformednumberedanchor\-blindedPanel median Intent on ToM / GD / Random2\.0/2\.0/1\.02\.0\\,/\\,2\.0\\,/\\,1\.02\.0/2\.0/1\.02\.0\\,/\\,2\.0\\,/\\,1\.02\.0/2\.0/1\.02\.0\\,/\\,2\.0\\,/\\,1\.02\.0/2\.0/1\.02\.0\\,/\\,2\.0\\,/\\,1\.0Strict rankToM\>GD\>Random\\text\{ToM\}\>\\text\{GD\}\>\\text\{Random\}1/91\\,/\\,90/90\\,/\\,94/94\\,/\\,95/95\\,/\\,9Panel mean \(Intent, Approp\) on ToM\(2\.33,0\.95\)\(2\.33,0\.95\)\(2\.25,0\.95\)\(2\.25,0\.95\)\(2\.46,0\.97\)\(2\.46,0\.97\)\(2\.31,0\.94\)\(2\.31,0\.94\)Distance from panel mean on ToM to TD\-adult2\.102\.102\.182\.181\.981\.982\.132\.13Distance from panel mean on ToM to HF\-ASD\-adult0\.730\.730\.790\.790\.640\.640\.730\.73ToM nearest adult referenceHF\-ASDHF\-ASDHF\-ASDHF\-ASDGD nearest adult referenceTDTDTDTDRandom nearest adult referenceTDTDTDTDTable 9:Judge information\-disclosure ablation\. The anchor\-blinded L1 variant strips only the per\-clip Castelli\-mean anchor from the judge prompt; everything else, including the tester responses, judge identities, rubric SHA, and aggregation logic, is held fixed\. Across both stimulus\-format arms \(plain4×44\\times 4grid and numbered4×44\\times 4grid\) the nearest\-group verdict per condition is unchanged, the panel\-median Intent per condition is unchanged, and the ToM mean Intent drifts slightly downward under anchor blinding \(consistent with the informed judge being mildly anchored upward by the disclosed Castelli mean\)\. Per\-tester strict\-rank changes are minor reshuffles at the gemini\-3\.5\-flash boundary\. The under\-attribution and ToM\-toward\-HF\-ASD findings hold under both variants; the bias was conservative\.
#### Canonical version reported in the main text\.
The main\-text animated\-triangles numbers in Sections[4\.3](https://arxiv.org/html/2608.00261#S4.SS3)and[5](https://arxiv.org/html/2608.00261#S5), the per\-model profiles in Figure[3](https://arxiv.org/html/2608.00261#S4.F3), the Y axis of Figure[4](https://arxiv.org/html/2608.00261#S5.F4), the cross\-tabulation in Table[8](https://arxiv.org/html/2608.00261#A8.T8), and the per\-condition distances in Tables[7](https://arxiv.org/html/2608.00261#A7.T7)and[5](https://arxiv.org/html/2608.00261#A5.T5)are all reported under the anchor\-blinded L1 canonical rubric\. The informed\-judge numbers are retained in Table[9](https://arxiv.org/html/2608.00261#A10.T9)above as the audit\-trail comparison and are not used elsewhere in the paper\. Direct human ratings on the specific99\-tester×\\times1212\-clip×\\times33\-judge production cells were not collected and remain a useful next step for further tightening the LLM\-as\-rater versus human\-rater asymmetry\.
#### Same\-family judge audit\.
The cost\-efficient three\-judge production panel shares model families with three of the nine testers \(claude\-opus\-4\.7 and claude\-sonnet\-4\.6 vs the claude\-haiku\-4\.5 judge, gemini\-3\.0\-pro and gemini\-3\.5\-flash vs the gemini\-2\.5\-flash judge, and qwen\-3\.5\-plus vs the qwen2\.5\-72b\-instruct judge\), yielding 60 of 324 in\-family \(judge, tester\) cells\. Excluding these 60 in\-family cells from the per\-condition median Intent aggregation under the anchor\-blinded canonical rubric on the numbered\-grid stimulus leaves the ToM, GD, and Random condition medians exactly unchanged \(Δ=0\.0\\Delta=0\.0throughout\)\. Same\-family judge bias is therefore empirically null at the panel\-median aggregation level on this data\.
## Appendix KStimulus\-Format Ablation \(Plain Grid vs Numbered Grid\)
#### Rationale\.
The canonical4×44\\times 4frame grid asks the model to infer that the sixteen cells are temporally ordered samples of a single short clip\. Models with weaker fine\-grained visual reasoning may be unable to recover this temporal order from the static composite alone, and that vision\-side failure could confound any inference about intent\-attribution capacity\. We therefore evaluate a second stimulus format in which each of the sixteen cells carries an explicit1−161\{\-\}16index badge in its corner, giving the model the temporal order as a free signal\. Everything else, including the prompt text byte\-for\-byte, the rubric, the rater jury, the panel of nine testers, the twelve canonical Frith\-Happé clips, the sixteen\-frame sampling, andKtester=3K\_\{\\text\{tester\}\}=3, is held identical\. The ablation isolates whether observed Intent under\-attribution reflects temporal\-parsing failure on the vision side or intent\-attribution failure on the ToM side\.
#### Headline comparison\.
Table[10](https://arxiv.org/html/2608.00261#A11.T10)reports the v1 \(plain grid\) versus v2 \(numbered grid\) headline metrics on the same9×12×3=3249\\times 12\\times 3=324\(tester, clip, judge\) cells per arm, each aggregating overKtester=3K\_\{\\text\{tester\}\}=3trials\.
Table 10:v1 \(plain4×44\\times 4grid\) versus v2 \(numbered4×44\\times 4grid\) on the same nine testers, twelve clips, and three judges withKtester=3K\_\{\\text\{tester\}\}=3, anchor\-blinded canonical rubric per Appendix[J](https://arxiv.org/html/2608.00261#A10)\. The numbered\-grid stimulus moves five testers into strict Intent rank, but the panel median Intent and the Stage\-C nearest\-group verdict per condition are unchanged\.
#### Per\-tester strict\-rank flips\.
Table 11:Per\-tester strict Intent rankToM\>GD\>Random\\text\{ToM\}\>\\text\{GD\}\>\\text\{Random\}status under plain \(v1\) and numbered \(v2\) grid, anchor\-blinded canonical rubric\. No tester reaches strict rank under v1; five testers \(opus, gemini\-3\.0\-pro, gemini\-3\.5\-flash, gpt\-5\.5, qwen\) reach strict rank under v2\. The five gainers cluster as the high\-vision frontier subset of the panel\.
#### Stage\-C nearest\-group verdict robustness\.
Under both v1 and v2 the panel mean \(Intent, Approp\) profile on the ToM condition is closer to the HF\-ASD\-adult group mean than to the TD\-adult group mean \(v1 distances0\.790\.79versus2\.182\.18; v2 distances0\.730\.73versus2\.132\.13\), and on both GD and Random the panel mean is closer to TD\-adult\. The per\-condition nearest\-group verdict is therefore identical across the two stimulus formats; ToM nearest HF\-ASD\-adult, GD and Random nearest TD\-adult\. Our central descriptive finding does not depend on the stimulus\-format choice\.
#### Why v2 is canonical for the main text\.
The numbered\-grid \(v2\) is canonical because removing the temporal\-parsing confound is a pre\-defined design requirement; any animated\-triangles benchmark used to claim a ToM\-side limitation must first rule out that the observed Intent under\-attribution is driven by failure to recover frame order from a static composite\. The plain grid \(v1\) does not rule this out, so we retain it as the ablation arm that quantifies how much of the panel\-level Intent collapse is recoverable under explicit temporal cueing\. The five\-of\-nine strict\-rank recovery under v2 is the empirical confirmation that v2 dissolves the vision\-side confound, not the reason we chose v2\.
## Appendix LFrame Representation Selection
#### Motivation\.
The animated\-triangles benchmark presents each clip as a single composite image of 16 uniformly sampled frames arranged in a4×44\\times 4grid\. This representation was chosen against eight alternatives \(sequence ofNNimage\_urlpayloads, larger grids, native video\) to maximize agreement with the model’s full\-fidelity video path while keeping per\-cell prompt cost tractable across the9×12×39\\times 12\\times 3judge cells of the main run\.
#### Reference baseline\.
Of the nine frontier VLM testers in the main panel, only Gemini 2\.5 Pro accepts a nativevideo\_urlpayload\. We treat Gemini’s native\-video path as the 100% reference\. It routes the full mp4 through Gemini’s video tokens, populates theusage\.prompt\_tokens\_details\.video\_tokensfield non\-trivially, and is the same path the model was trained for\.
#### Conditions tested\.
On the same Gemini 2\.5 Pro tester and three Castelli clips \(one Random, one GD, one ToM\), nine frame representations were evaluated:native\_video\(reference\), three sequence variants \(frames\_16,frames\_32,frames\_64\) and five grid variants \(grid\_4x4/4x8/8x8at fixed320×240320\\times 240per\-cell resolution, plusgrid\_4x4\_fullresandgrid\_8x8\_fullresat source resolution\)\. The same three flagship LLM judges \(claude\-opus\-4\.7, gpt\-5, gemini\-2\.5\-pro\) atKjudge=3K\_\{\\text\{judge\}\}=3scoring reps applied the Castelli rubric\. The top three candidates \(frames\_64,grid\_4x4,native\_video\) were promoted toKtester=3K\_\{\\text\{tester\}\}=3to tighten the verdict beyondK=1K=1noise\.
#### Result\.
We define per\-condition similarity tonative\_videoas the equal\-weighted average of Intent and Approp similarity percentages, each derived from per\-clip mean absolute error againstnative\_videoon the rubric’s full scale\. Table[12](https://arxiv.org/html/2608.00261#A12.T12)reports the final ranking\.frames\_64achieves the highest similarity to native at 88\.2%, but at 16,594 prompt tokens per cell\.grid\_4x4achieves 85\.8% similarity at 1,416 prompt tokens per cell, an11\.7×11\.7\\timesreduction in token cost for a 2\.4 pp reduction in similarity\. We adoptgrid\_4x4as the canonical frame representation\. It is Pareto\-dominant on similarity\-per\-token over both the higher\-fidelity sequence representations and the larger grids\.
Table 12:Frame representation similarity to Gemini’s native video path\. Similarity is the equal\-weighted mean of per\-clip Intent and Approp similarity percentages, each derived from mean absolute error againstnative\_videoon the rubric’s full scale \(Intent 0–5, Approp 0–3\)\. We adoptgrid\_4x4as canonical: 85\.8% similarity at 1,416 tokens per cell,11\.7×11\.7\\timescheaper thanframes\_64\(88\.2% at 16,594 tokens\)\. The top three candidates were run atKtester=3K\_\{\\text\{tester\}\}=3to tighten the verdict beyondK=1K=1noise; the remaining five usedKtester=1K\_\{\\text\{tester\}\}=1\.
## Appendix MRubric and Judge Validation Battery \(Phase 1\)
#### Anchor set\.
Before the main 9\-tester×\\times12\-clip production run, we validate that theCastelliet al\.\([2000](https://arxiv.org/html/2608.00261#bib.bib11)\)Appendix\-2 rubric is interpretable by LLM judges on human\-anchored exemplars\. The anchor set consists of 14 paper\-derived ground\-truth items, namely five ToM\-condition transcripts T1–T5 lifted verbatim fromCastelliet al\.\([2002](https://arxiv.org/html/2608.00261#bib.bib12)\)\(human participants’ free\-text descriptions of the canonical clips, with expected Intent in 0–5 and Approp in 0–3\) and nine Appendix\-2 exemplar phrases A0a–A5b fromCastelliet al\.\([2000](https://arxiv.org/html/2608.00261#bib.bib11)\)covering the Random / GD / interaction range with expected Intent or Approp targets per Castelli’s worked examples\.
#### Battery design\.
Three flagship LLM judges \(claude\-opus\-4\.7, gpt\-5, gemini\-2\.5\-pro\) score each anchor under the Castelli rubric template \(SHA pinned,llm\_judge\_prompt\.yaml\) atK=5K=5repetitions,T=0T=0, no CoT\. Total:14×3×5=21014\\times 3\\times 5=210calls\. The double gate is \(i\) per\-judge recovery percentage≥80%\\geq 80\\%\(a cell counts as recovered iff predicted score lies within±0\.5\\pm 0\.5of the expected interval\), and \(ii\) per\-anchor standard deviation≤0\.5\\leq 0\.5on Intent and≤0\.8\\leq 0\.8on Approp across the 5 reps\. The SD gate measures within\-judge stochasticity not eliminated byT=0T=0\.
#### Result\.
All three flagship judges PASS both gates\. Intent recovery is 100% \(claude\-opus\-4\.7\), 100% \(gpt\-5\), 100% \(gemini\-2\.5\-pro\); Approp recovery is 100%, 96%, 100% respectively\. Maximum per\-anchor SDs are 0\.40 / 0\.49 \(opus\), 0\.00 / 0\.40 \(gpt\-5\), 0\.00 / 0\.43 \(gemini\-pro\) for Intent / Approp; all within gate\. The validated rubric and prompt SHA are then frozen and copied into the production rater; the production aggregation pipeline asserts SHA match at startup, blocking silent drift\.
## Appendix NJudge Model Ablation
#### Motivation\.
At production scale \(9×12×3×Ktester=39\\times 12\\times 3\\times K\_\{\\text\{tester\}\}=3\) flagship inference dominates the per\-experiment budget; we therefore ablate whether a cheaper cross\-vendor judge panel can match flagship recovery on the same 14\-anchor battery used in Appendix[M](https://arxiv.org/html/2608.00261#A13)\.
#### Candidates\.
Seven cheaper candidates spanning five vendor families \(Anthropic, OpenAI, Google, Moonshot, DeepSeek, Qwen\) were evaluated against the same 14\-anchor protocol, namely claude\-sonnet\-4\.6, claude\-haiku\-4\.5, gpt\-5\-mini, gemini\-2\.5\-flash, kimi\-k2, deepseek\-v3, qwen2\.5\-72b\-instruct\. Total:7×14×5=4907\\times 14\\times 5=490calls atT=0T=0\. The same double gate from Phase 1 applies\.
#### Result\.
Six of seven candidates PASS both gates\. The single failure is gpt\-5\-mini \(Approp recovery 80%, max Approp SD 1\.47\); the failure mode is a magnified form of the same instability the flagship gpt\-5 already exhibits on long ambiguous inputs\. Three candidates — claude\-haiku\-4\.5, gemini\-2\.5\-flash, qwen2\.5\-72b\-instruct — achieve 100% recovery with per\-anchor Intent and Approp SDs equal to 0 across all 14 anchors×\\times5 reps \(Table[13](https://arxiv.org/html/2608.00261#A14.T13)\)\. These three are strictly more stable than any of the three flagship judges\. We adopt this trio as the production scoring panel\. The trio spans three distinct training lineages \(Anthropic instruction\-following, Google reasoning, Alibaba’s open\-source Qwen\), reducing family\-correlated reading bias; it costs $0\.0129 per cell versus $0\.0676 per cell on the flagship panel \(81% reduction\); and it shows SD=0 ceilings on every anchor\.
Table 13:Judge model ablation against the 14\-anchor Phase\-1 battery\. Three lower\-cost candidates \(claude\-haiku\-4\.5, gemini\-2\.5\-flash, qwen2\.5\-72b\-instruct\) achieve 100% recovery with per\-anchor SD = 0, matching flagship within\-judge consistency atT=0T=0at lower per\-call cost; we adopt the trio as the production scoring panel for the main 9\-tester×\\times12\-clip×\\times3\-judge run\. Pricing is OpenRouter list price \(3000 input \+ 1000 output tokens per call\)\. gpt\-5\-mini is the only candidate to FAIL; its Approp SD of 1\.47 across 5 reps atT=0T=0amplifies the same instability the flagship gpt\-5 shows on long ambiguous inputs\.
## Appendix OReproducibility Details
Released with this submission are the benchmark stimuli, prompts, the SHA\-locked scoring rubric, judge configuration, the two\-firewall harness, the aggregation scripts that produce every figure and table in this paper, and a digest manifest of the runs underlying the reported numbers\. Released with the camera\-ready \(held back at submission for review\-time blinding and storage\-quota reasons\) are the full per\-\(tester, clip, judge\) raw scoring outputs and the full per\-group source list for the human\-reference fallback rule in Appendix[H](https://arxiv.org/html/2608.00261#A8)\.Similar Articles
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
This paper investigates whether vision-language models can distinguish potential from established common ground in asymmetric dialogue. Experiments on MapTask data show that providing task-relevant map content (visual or textual) biases models toward over-predicting alignment, as they rely on static referential cues rather than tracking grounding through dialogue history.
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
OmniToM introduces a benchmark that evaluates large language models' theory of mind by requiring explicit belief structure extraction and labeling, revealing a bottleneck in tracking actor-specific beliefs despite strong performance on endpoint QA tasks.
Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models
This paper investigates how vision-language models resolve conflicts between visual evidence and world knowledge, revealing that visual grounding is the default while prior knowledge depends on a small set of late-layer attention heads. The authors perform causal analysis across three VLM families, demonstrating an asymmetric structure where ablating these heads shifts predictions from knowledge-grounded to visually grounded answers.
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
This arXiv paper evaluates theory of mind capabilities in reasoning LLMs, finding increased robustness to prompt variations and task perturbations. The authors interpret gains as evidence for a robustness-based account rather than a new ToM-specific ability.
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
This paper investigates the emergence of situation modeling and mentalizing abilities in transformer language models across training stages, finding that false belief task performance depends on model size and training volume, emerges late in pretraining, and shows fragility with non-factive verbs.