MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

arXiv cs.AI Papers

Summary

Introduces MultivationBench, a benchmark for evaluating multimodal large language models' sequential motivation reasoning using story-driven visual narratives based on Maslow's hierarchy and Reiss's desires. Results show all tested models struggle with dynamic motivation inference.

arXiv:2607.26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
Original Article
View Cached Full Text

Cached at: 07/31/26, 04:00 AM

# MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Source: [https://arxiv.org/html/2607.26465](https://arxiv.org/html/2607.26465)
Kawai Chung♠Chunkit Chan♠Yauwai Yim♠Yuxuan Liu♠Haochen Shi♠ Weiqi Wang♠Qing Zong♠Tianshi Zheng♠FU Yixuan♠Wong Kai Chung♠ Hao Liang♠Yifan Gao♣Xi Yang♠Janet Hui\-wen Hsiao♠ Yangqiu Song♠ ♠The Hong Kong University of Science and Technology ♣Amazon\.com [![[Uncaptioned image]](https://arxiv.org/html/2607.26465v1/github-logo.png)Code](https://github.com/HKUST-KnowComp/MultivationBench)\{kwchungac, yqsong\}@cse\.ust\.hk

###### Abstract

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied\. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real\-world behavioral drivers\. To address this gap, we introduceMulTivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story\-driven visual narratives\. The benchmark builds upon established psychological frameworks—Maslow’s hierarchy and Reiss’s basic desires—and requires models to integrate accumulated multimodal context to infer evolving motivations\. Results indicate thatMulTivationBenchpresents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human\-like social understanding\.

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Kawai Chung♠Chunkit Chan♠Yauwai Yim♠Yuxuan Liu♠Haochen Shi♠Weiqi Wang♠Qing Zong♠Tianshi Zheng♠FU Yixuan♠Wong Kai Chung♠Hao Liang♠Yifan Gao♣Xi Yang♠Janet Hui\-wen Hsiao♠Yangqiu Song♠♠The Hong Kong University of Science and Technology♣Amazon\.com[![[Uncaptioned image]](https://arxiv.org/html/2607.26465v1/github-logo.png)Code](https://github.com/HKUST-KnowComp/MultivationBench)\{kwchungac, yqsong\}@cse\.ust\.hk

![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/overall_after_re.png)Figure 1:A Maslow motivation example fromMultivationBench\. Given the accumulated story context and images up toI11I\_\{11\}, Debbie reaching for her phone is plausibly interpreted as reflectingCognitiveneeds; with additional context up toI13I\_\{13\}, which reveals family photographs and her social isolation at the dinner, the inferred motivation shifts toLove and Belonging\. This example shows how the inferred practical motivation changes as sequential multimodal context accumulates\.## 1Introduction

Motivation is the internal force that initiates, guides, and sustains goal\-directed behavior, fundamental to explaining human action\(Hagger and Chatzisarantis,[2005](https://arxiv.org/html/2607.26465#bib.bib19); Ryan and Deci,[2000](https://arxiv.org/html/2607.26465#bib.bib20)\)\. It helps explain why people begin, persist in, or stop a behavior as situations unfold\(Kazdin,[2000](https://arxiv.org/html/2607.26465#bib.bib21)\)\. Yet motivation is latent, which cannot be directly observed\. Instead, it must be inferred from behavior and context\. For this reason, people rarely interpret behavior from a single static moment alone\(Scholl and Tremoulet,[2000](https://arxiv.org/html/2607.26465#bib.bib72)\)\. As events unfold, people accumulate evidence across time, using later actions and the surrounding context to revise earlier interpretations\. This sequential reasoning process is critical in narratives\. The motivation for an action often emerges only after additional visual and situational information is disclosed\(Zacks and Swallow,[2007](https://arxiv.org/html/2607.26465#bib.bib36); Baldassanoet al\.,[2017](https://arxiv.org/html/2607.26465#bib.bib63); Bakeret al\.,[2017](https://arxiv.org/html/2607.26465#bib.bib35); Knoblich and Sebanz,[2008](https://arxiv.org/html/2607.26465#bib.bib25)\)\.

However, existing motivation reasoning benchmarks focus on static text\-based scenarios or isolated multimodal snapshots\(Sapet al\.,[2019b](https://arxiv.org/html/2607.26465#bib.bib30); Yonget al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib29)\)\. Such settings may reward locally plausible guesses, but do not assess whether a model can update its interpretation as a story develops\. Therefore, current benchmarks provide limited evidence on whether Multimodal Large Language Models \(MLLMs\)\(Caffagniet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib82)\)can perform sequential motivation reasoning from an accumulated visual and textual context\.

To bridge this gap, we presentMultivationBench, a large\-scale benchmark specifically designed to stress\-test multimodal sequential motivation reasoning in story\-driven narratives\. Our benchmark is grounded in two complementary psychological frameworks: Maslow’s Expanded Hierarchy of Needs\(Maslow,[1970](https://arxiv.org/html/2607.26465#bib.bib2)\), which captures broad categories of human needs, and Reiss’s 16 Basic Desires\(Reiss,[2004](https://arxiv.org/html/2607.26465#bib.bib4)\), which captures finer\-grained differences in motivational content\. This design enables evaluation at both a coarse\-grained, universal level and a fine\-grained, individual level, requiring models to reason over accumulated multimodal context, not just isolated observations\.

Figure[1](https://arxiv.org/html/2607.26465#S0.F1)illustrates the core challenge targeted byMultivationBench\. A sequence of images shows Debbie at a formal dinner event\. At imageI11I\_\{11\}, she reaches toward a phone\. Viewed in isolation, this moment is ambiguous\. A model might interpret the behavior as information seeking, aligning withCognitiveneeds in Maslow’s hierarchy\. However, later frames reveal family photograph on the phone\. The broader narrative establishes Debbie’s social isolation\. The explanation then shifts toward seeking emotional connection with absent loved ones, which aligns withLove and Belonging\. The challenge is not simply recognizing the action, but revising the inferred motivation as new evidence accumulates\. Models relying on surface\-level cues may choose simpler explanations, missing motivations that require a longer multimodal context\.

To benchmark this,MultivationBenchhas 1,000 unique stories, 4,023 visually grounded behavioral instances, and 16,092 evaluation questions\. Experiments on state\-of\-the\-art MLLMs show that sequential motivation reasoning remains challenging\. Several models improve from short to medium\-length stories, suggesting that moderate additional context can help in some cases\. However, this benefit does not persist for longer narratives, and story\-level consistency remains low across models, indicating that maintaining correct motivation reasoning across the whole narrative is still difficult\. Our contributions are summarized as follows:

- •To the best of our knowledge,MultivationBenchis the first human\-annotated benchmark for evaluating multimodal sequential motivation reasoning in story\-driven visual narratives\.
- •MultivationBenchgrounds motivation labels in two complementary psychological frameworks—Maslow’s Expanded Hierarchy of Needs and Reiss’s 16 Basic Desires—thereby enabling evaluation of both coarse\-grained needs and fine\-grained desires under accumulated multimodal context\.
- •We conduct systematic experiments on state\-of\-the\-art MLLMs under full multimodal, text\-only, and image\-only settings, and provide in\-depth analyses of local task accuracy, story\-level consistency, taxonomy granularity, and the gap between model and human performance\.

## 2Related Work

### 2\.1Text\-Based Social and Motivation Reasoning\.

Research on social and motivational reasoning has predominantly relied on text\-based evaluations\. Prior work has studied event\-centric intent and reaction inference\(Sapet al\.,[2019a](https://arxiv.org/html/2607.26465#bib.bib70); Rashkinet al\.,[2018b](https://arxiv.org/html/2607.26465#bib.bib64)\), character psychology in simple stories\(Rashkinet al\.,[2018a](https://arxiv.org/html/2607.26465#bib.bib65)\), and contextualized story explanation\(Mostafazadehet al\.,[2020](https://arxiv.org/html/2607.26465#bib.bib66)\)\. SocialIQA\(Sapet al\.,[2019b](https://arxiv.org/html/2607.26465#bib.bib30)\)evaluates social commonsense by presenting daily scenarios and asking models to reason about actions and their implications\. While it scales well, its reliance on purely textual inputs limits its ability to assess how models interpret non\-verbal social signals\. More recently, MotiveBench\(Yonget al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib29)\)has explicitly targeted motivation reasoning by integrating detailed character profiles and behavioral descriptions to infer psychological drives\. However, it relies exclusively on textual context and typically presents isolated, single\-turn scenarios\. This unimodal, static approach overlooks the critical role of visual cues and fails to test whether a model can track a character’s evolving motivation across a continuous narrative sequence\.

#### Multimodal Social Reasoning

Recent multimodal benchmarks have begun to incorporate visual evidence into socially grounded reasoning and temporal visual understanding\(Linet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib69); Villa\-Cuevaet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib71); Jinet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib15); Liet al\.,[2024b](https://arxiv.org/html/2607.26465#bib.bib84); Wuet al\.,[2024a](https://arxiv.org/html/2607.26465#bib.bib87); Renet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib85); Maazet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib86)\)\. While prior works address visual social commonsense, multimodal theory\-of\-mind reasoning, video understanding, temporal grounding, and visual goal inference, they are often restricted to short\-horizon, event\-centered, or perception\-oriented settings that lack the depth to model complex, evolving psychological states\. These approaches rely on surface\-level behavior\-to\-intent mapping rather than inferring latent desires through a grounded psychological framework, thereby limiting their ability to evaluate motivation reasoning that unfolds over accumulated multimodal narrative context\(Chenet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib46)\)\. Given the limitations in prior work, there is a need for evaluating motivation within realistic multimodal settings, capturing authentic social interactions beyond goals and beliefs alone\.

## 3MultivationBench

During the construction ofMultivationBench, we follow four core design principles: \(1\)Psychological grounding, employing authoritative frameworks to define a label space beyond surface\-level intents; \(2\)Sequential human\-centric dynamics, prioritizing narratives with rich interpersonal interactions where sequential visual cues are essential for tracking mental states; \(3\)Long\-term consistency evaluation, stratifying data into distinct intervals based on sequence length to evaluate a model’s ability to maintain psychological continuity over different memory horizons; and \(4\)Data integrity, assessing and mitigating potential data contamination risk\.

### 3\.1Data Sources

Following these principles, MULTIVATIONBENCH prioritizes sequential visual narratives over static scenes, since character motivations often emerge only through later visual and situational evidence\. We aggregate data from three complementary sources—SSID\(Malakanet al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib16)\),StoryReasoning\(Oliveira and Matos,[2024](https://arxiv.org/html/2607.26465#bib.bib17)\), andMovieBench\(Wuet al\.,[2024b](https://arxiv.org/html/2607.26465#bib.bib18)\)—to cover diverse narrative styles and temporal complexity\. SSID provides shorter social\-media\-style stories, while StoryReasoning and MovieBench offer longer cinematic narratives with richer character arcs\. We also conduct a contamination analysis and find low risk across all sources; details are in Appendix[B\.1](https://arxiv.org/html/2607.26465#A2.SS1)\.

### 3\.2Psychological Frameworks and Task Formulation

#### Maslow’s Expanded Hierarchy of Needs

We adopt the 8\-level version of Maslow’s Expanded Hierarchy of Needs\(Maslow,[1970](https://arxiv.org/html/2607.26465#bib.bib2)\)to represent universal levels of human needs that can motivate behavior\. The eight labels arePhysiological,Safety,Love and Belonging,Esteem,Cognitive,Aesthetic,Self\-Actualization, andTranscendence\. InMultivationBench, Maslow serves as a theory of need priority to capture whether a behavior is primarily driven by immediate survival and security concerns, by social and esteem needs, or by higher\-order growth and meaning\-oriented motives\.

#### Reiss’s Basic Desires and their relationship to Maslow

We additionally incorporate Reiss’s theory of 16 basic desires through the Reiss Motivation Profile \(RMP\)\(Reiss,[2004](https://arxiv.org/html/2607.26465#bib.bib4)\)\. The sixteen labels arePhysical Exercise,Eating,Order,Saving,Tranquility,Romance,Family,Acceptance,Social Contact,Independence,Vengeance,Honor,Power,Status,Curiosity, andIdealism\. Unlike Maslow, which organizes motivation as universal hierarchical levels of need, Reiss emphasizes individual differences in the relative strength of specific desires that shape behavior\. The two theories, therefore, serve different but complementary purposes in our benchmark: Maslow identifies the broad level of need a behavior serves, whereas Reiss distinguishes the more specific desire content that may motivate the same behavior\. Further discussion of the framework choice and full definitions of both theories are provided in Appendix[A](https://arxiv.org/html/2607.26465#A1)\.

#### Task formulation

Based on these two psychological frameworks,MultivationBenchevaluates motivation reasoning on visually grounded character behaviors in sequential visual narratives\. For each visually grounded behavior exhibited by a character at a particular image index, we use the accumulated story context and images up to that point to formulate four tasks: Maslow Definition, Maslow Practical Motivation, Reiss Definition and Reiss Practical Motivation\. TheDefinitiontasks ask models to classify the motivation of the behavior based on the standard theory labels, whereas thePractical Motivationtasks ask models to choose among context\-specific options derived from the same theories but instantiated in the current narrative situation\. Some behaviors can be supported by more than one valid motivation under the story context\. We, therefore, formulate all four tasks as multi\-label prediction problems, where each instance may have one or more ground\-truth annotations\.

![Refer to caption](https://arxiv.org/html/2607.26465v1/x1.png)Figure 2:AI–human pipeline for constructingMulTivationBench\. We summarize the stages here and defer prompts and implementation details in Appendix[D](https://arxiv.org/html/2607.26465#A4)\.

### 3\.3Benchmark Construction Overview

#### Data Construction

We constructMulTivationBenchthrough a compact AI–human pipeline \(Figure[2](https://arxiv.org/html/2607.26465#S3.F2)\)\. For each narrative, multiple MLLMs first extract a character\-centeredbehavior chain: a list of visually grounded behaviors aligned to image indices with candidate motivations\. Based on the verified behavior chain and the story context up to each behavior, we then generate practical motivation options under Maslow and Reiss’s definitions\. These candidate options are further refined through AI\-based cross\-review and validation to ensure logical consistency, multimodal grounding, and alignment with the psychological theories\. Finally, human filtering is applied to remove hallucinated or weakly grounded samples\. Detailed prompts for behavior extraction, option generation, AI review, and the human filtering protocol are described in Appendix[D](https://arxiv.org/html/2607.26465#A4)\.

#### Human Annotation and Agreement

We conduct a human annotation study to establish the ground\-truth labels for the four reasoning tasks\. Three annotators, all graduate students from English\-speaking universities, were recruited for the annotation process\. To ensure annotation quality, each annotator first completed a training phase covering the task definitions and representative examples\. We then evaluated their performance on the first 100 instances and provided detailed feedback on typical errors to calibrate their understanding\. Fleiss’s ϰ scores for the Maslow Definition, Maslow Practical Motivation, Reiss Definition, and Reiss Practical Motivation tasks are 80\.99%, 78\.87%, 74\.87%, and 75\.35%Fleiss \([1971](https://arxiv.org/html/2607.26465#bib.bib42)\)\. More details of the annotation process are provided in Appendix[D\.3](https://arxiv.org/html/2607.26465#A4.SS3)\.

### 3\.4Dataset Statistics

MulTivationBenchcontains 1,000 visual narratives with 4,023 visually grounded character behaviors\. Each behavior is paired with four questions, yielding 16,092 evaluation instances\. To evaluate long\-horizon coherence, we stratify stories by maximum narrative length: Short \(2–5 images\), Medium \(6–9\), and Long \(10\+\)\. Extended statistics are provided in Appendix[D\.5](https://arxiv.org/html/2607.26465#A4.SS5)\.

## 4Experiment

### 4\.1Experimental Setting

We assess the motivation reasoning capabilities of eight Multimodal Large Language Models\. For closed\-source models, we evaluate o4\-mini\(OpenAI,[2025](https://arxiv.org/html/2607.26465#bib.bib81)\), Gemini\-3\-Flash\(Google,[2025](https://arxiv.org/html/2607.26465#bib.bib12)\), and Grok\-4\.1\-Fast\(xAI,[2025](https://arxiv.org/html/2607.26465#bib.bib51)\)\. For open\-source models, we select the Llama\-4 family \(Scout\-17B and Maverick\-17B\)\(Meta AI,[2025](https://arxiv.org/html/2607.26465#bib.bib52)\), the Phi family \(Phi\-3\.5\-Vision\-instruct\(Abdinet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib54)\)and Phi\-4\-multimodal\-instruct\(Microsoftet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib55)\)\), and Nemotron\-Nano\-12B\-v2\-vl\(NVIDIA,[2025](https://arxiv.org/html/2607.26465#bib.bib56)\)\. All evaluations are conducted in a zero\-shot setting with the temperature parameter set to0to ensure reproducibility\. To mitigate position bias, we randomly shuffle the order of the answer options for every instance\(Zhenget al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib61)\)\. We also report human performance, measured using annotations from three computer science graduate students who completed the same evaluation tasks\. More details of the human evaluation are provided in Appendix[E\.4](https://arxiv.org/html/2607.26465#A5.SS4)\. For Phi\-3\.5\-Vision\-instruct, the Long\-story multimodal and image\-only results are not reported due to limited image context length\. A common\-subset comparability check is provided in Appendix[E\.5](https://arxiv.org/html/2607.26465#A5.SS5)\.

#### Metrics

We report two complementary metrics\.Exact Match \(EM\)measures strict set\-level accuracy: a prediction is counted as correct only when the predicted option set exactly matches the ground\-truth set\.Example\-based F1 \(F1\)measures instance\-level partial agreement in the multi\-label setting\. Following prior work on multi\-label evaluation, we evaluate each test instance separately and then average across instances\(Zhang and Zhou,[2014](https://arxiv.org/html/2607.26465#bib.bib73); Loza Mencíaet al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib75)\)\. This metric rewards partial overlap between the predicted and ground\-truth sets while penalizing both false positives and false negatives\. We provide the formal definition and an illustrative example in Appendix[E\.3](https://arxiv.org/html/2607.26465#A5.SS3)\.

Table 1:Overall performance onMulTivationBench\.Results are aggregated across all four task types and reported across story lengths \(Short, Medium, Long\) and input modalities \(Multimodal, Text\-Only, Image\-Only\)\. Metrics are reported as EM \(Exact Match\) accuracy and F1 \(Example\-based F1\) \(%\) score\. Bold indicates best; underline indicates second\-best; “–” indicates that no score is reported because the model could not process the corresponding context\-length subset due to image context window limitations\.

### 4\.2Experimental Analysis

#### Overall Performance

Table[1](https://arxiv.org/html/2607.26465#S4.T1)reports the overall performance of several MLLMs aggregated across all four task types\. Overall,MultivationBenchis challenging for all evaluated models\. The results also suggest a trade\-off between F1 score and EM accuracy\. Some models are better at retrieving plausible motivation sets, while others are more accurate at recovering the full gold label set\. For example, Grok\-4\.1\-Fast achieves the highest overall F1 score \(55\.0%\), whereas Phi\-4\-multimodal\-instruct obtains the highest EM accuracy \(39\.2%\)\.

#### Modality Ablation

To assess the contribution of different input modalities, we compare model performance under three settings:Multimodal,Text\-Only, andImage\-Only\. The prompt variations for these settings are detailed in Table[22](https://arxiv.org/html/2607.26465#A5.T22)\. The modality ablation results in Table[1](https://arxiv.org/html/2607.26465#S4.T1)show a consistent gap betweenText\-OnlyandMultimodalperformance\. Across several strong models,Text\-Onlyinputs preserve a relatively large portion of theMultimodalF1 score, while EM accuracy drops much more sharply\. For example, Gemini\-3\-Flash retains a relatively high Text\-Only F1, but its EM accuracy declines sharply from 35\.7% to 13\.6%\. A similar F1–EM divergence can also be observed across other strong models\. This pattern suggests that textual narratives often provide enough information for models to recover partially correct motivation sets, but not enough to identify the exact gold label set\. Visual information therefore appears important for refining and revising the inferred motivation by providing disambiguating evidence beyond text alone\. This finding demonstrates the importance of visual information in motivation reasoning\.

#### Impact of Narrative Length

To study the effect of narrative length, we evaluate model performance across stories of different lengths\. As shown in Table[1](https://arxiv.org/html/2607.26465#S4.T1), most models improve from the short to medium interval, suggesting that current state\-of\-the\-art MLLMs can generally handle moderate multimodal contexts\. However, this trend does not continue for longer narratives\. We can observe that, for many models, performance in the long setting declines relative to the medium setting\. Figure[3](https://arxiv.org/html/2607.26465#S4.F3)provides a finer\-grained view\. For story lengths between 2 and 13, performance remains relatively stable\. However, most models begin to decline more consistently after length 13\. This pattern suggests that the challenge of longer narratives is not simply the presence of more context, but the need to continuously trace and revise character motivations as new visual and textual evidence unfolds\. As narratives grow longer, later evidence can change the most plausible explanation for earlier behavior, requiring models to update their interpretation over the course of the story\. We further investigate this issue through the story\-level consistency experiment\. Appendix[E\.6](https://arxiv.org/html/2607.26465#A5.SS6)further provides a same\-behavior context\-ablation analysis that tests whether predictions for a fixed behavior change when earlier accumulated context is removed\.

![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/sequential_trend_combined_4.png)Figure 3:Sequential Motivation Reasoning Performance\.Performance across increasing story lengths\. Left: F1 score\(%\)\. Right: Exact Match \(EM\) Accuracy\(%\)
#### Story\-level Consistency

While the previous analyses examine instance\-level performance under different conditions, we next examine whether models can maintain correct motivation reasoning across an entire story\. FollowingChanet al\.\([2024c](https://arxiv.org/html/2607.26465#bib.bib14)\), we report a story\-levelconsistency score, which requires a model to answer every question within a story sequence correctly\. As shown in Table[2](https://arxiv.org/html/2607.26465#S4.T2), all models perform poorly on this metric\. Even the best Fullconsistency scoreis only0\.80%, indicating that current MLLMs rarely track character motivations correctly throughout an entire story\. This weakness also appears in per\-theory consistency: while Phi\-4\-multimodal\-instruct remains relatively balanced on Maslow \(Definition: 15\.4%,Motivation: 15\.2%\), other models show much larger gaps, such as Llama\-4\-Maverick\-17B \(15\.8% vs\. 2\.7%\)\. This shows that strong performance on individual task instances does not necessarily translate into stable motivation reasoning over the full narrative\. Because this all\-correct metric is intentionally strict, Appendix[E\.7](https://arxiv.org/html/2607.26465#A5.SS7)reports macro story EM/F1 as softer complementary story\-level metrics\.

ModelMaslow \(8\)Reiss \(16\)FullDef\.Mot\.Def\.Mot\.Cons\.Closed\-sourceGrok\-4\.1\-Fast13\.95\.810\.31\.90\.20Gemini\-3\-Flash11\.29\.79\.83\.40\.50o4\-mini15\.03\.56\.70\.90\.10Open\-sourceLlama\-4\-Scout\-17B11\.67\.99\.12\.50\.30Llama\-4\-Maverick\-17B15\.82\.79\.51\.50\.10Phi\-3\.5\-Vision4\.64\.73\.73\.90\.10Phi\-4\-multimodal\-instruct15\.415\.26\.44\.10\.80Nemotron\-Nano\-12B\-v2\-vl2\.90\.20\.60\.10\.00Table 2:Story\-level consistency onMulTivationBench\.Percentage of stories where all questions in a sequence are answered correctly\.Fullrequires correctness across all four task types\. Bold indicates the best result per column\.

### 4\.3Task\-based Analysis

#### Reasoning Task Comparison

To analyze how MLLMs perform under different formulations of motivation reasoning, we compare results onDefinitiontasks \(theoretical matching\) andPractical Motivationtasks \(multimodal inference\)\. Table[3](https://arxiv.org/html/2607.26465#S4.T3)reveals a model\-dependent divergence between the two settings rather than a uniform difficulty gap\. Some models perform substantially better onDefinitionthan onPractical Motivation\. For example, o4\-mini drops from 44\.1% EM accuracy on MaslowDefinitiontasks to 24\.7% on MaslowPractical Motivationtask, while Llama\-4\-Maverick\-17B declines from 47\.3% to 21\.7%\. This pattern suggests that these models retain substantial familiarity with psychological category definitions, but are less effective at applying those concepts to situated multimodal narratives\. In contrast, other models show the reverse trend\. Phi\-4\-multimodal\-instruct improves from 47\.5% EM accuracy on MaslowDefinitiontasks to 52\.9% on MaslowPractical Motivationtasks, while Gemini\-3\-Flash increases from 33\.1% to 43\.2%\. For these models, narrative context appears to provide useful grounding cues rather than introducing additional difficulty\. Taken together, this contrast suggests thatDefinitionandPractical Motivationtasks probe partially distinct capabilities: semantic familiarity with psychological categories versus the ability to ground them in sequential multimodal narratives\.

Table 3:Task\-level performance onMULTIVATIONBENCH\.EM accuracy and F1 score \(%\) for Maslow \(8\) and Reiss \(16\) across Definition and Practical Motivation tasks\. Bold indicates the best*model*result in each column\.
#### Taxonomy Granularity

A clear overall pattern in Table[3](https://arxiv.org/html/2607.26465#S4.T3)is the performance drop from the coarserMaslowtaxonomy to the more fine\-grainedReisstaxonomy, especially in thePractical Motivationsetting\. For example,Grok\-4\.1\-Fastachieves 63\.2% F1 score on MaslowPractical Motivation, but falls to 45\.1% on the corresponding Reiss task\. Similar declines can be observed across most models, indicating that current MLLMs can often identify a broad motivational region but struggle when the task requires finer discrimination among closely related psychological alternatives\. This pattern suggests that fine\-grained motivation reasoning remains a major challenge even when models achieve comparatively strong results on coarser taxonomies\.

#### Human Comparison

Human performance remains substantially above all evaluated models across task settings\. As shown in Table[3](https://arxiv.org/html/2607.26465#S4.T3), the gap is especially large on the Practical Motivation tasks\. For example, in the ReissPractical Motivationsetting, human Exact Match reaches 63\.9%, compared with 30\.7% for the best\-performing model \(Gemini\-3\-Flash\)\. A similarly large gap is observed in F1 score, with human performance reaching 72\.9% compared with 46\.0% for the strongest model\. These results indicate thatMulTivationBenchis far from saturated and that current MLLMs remain substantially limited in grounded and fine\-grained sequential motivation reasoning\.

![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/figure5_maslow_left_reiss_right_bold_dotted_new_2_std.png)Figure 4:Label\-wise F1 across the eight Maslow need categories \(left\) and the 16 Reiss desire categories \(right\) options in theMulTivationBench\. Bars show the mean F1 score aggregated across all evaluated models for each option; overlaid lines show per\-option trajectories for the selected representative models\.![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/er_3.png)Figure 5:Grouped bar chart comparing the frequency of False Positive \(Red\) versus False Negative \(Yellow\) errors among all models\.

### 4\.4How MLLMs Make Errors

To further understand why performance degrades as story context grows longer, we analyze the error patterns of MLLMs\. FollowingLeeet al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib60)\), we defineOver\-Interpretation\(False Positive\) as instances where a model hallucinates a motivation that is not supported by the ground truth, andUnder\-Interpretation\(False Negative\) as instances where it fails to recognize a valid motivation\.

#### Over\- vs\. Under\-interpretation

As illustrated in Figure[5](https://arxiv.org/html/2607.26465#S4.F5), we observe a consistent bias toward Over\-Interpretation across the majority of high\-capacity models\. Notably, Nemotron\-Nano\-12B\-v2\-vl exhibits the largest disparity, committing over 31,000 False Positive errors compared to approximately 7,000 False Negatives\. This suggests that as story information accumulates, the model struggles to discard earlier but less relevant cues when inferring a character’s current motivation, leading to substantial over\-prediction\. Similarly, closed\-source models such as Grok\-4\.1\-Fast and Gemini\-3\-Flash also show a strong False Positive bias, with nearly three times more False Positives than False Negatives\. These results suggest that the main failure mode of current MLLMs is not simply overlooking valid motivations, but failing to revise earlier interpretations as new evidence unfolds\.

#### Label\-wise Performance on Desire and Need Options

To further examine performance variation across motivation labels, Figure[4](https://arxiv.org/html/2607.26465#S4.F4)reports the average F1 score for each label across models\. A clear pattern is that MLLMs perform better on labels that can be inferred from a single salient cue or local scene, such asEating,Safety, andRomancereaching from 50% to 60%\. These labels are often supported by direct evidence in one moment, such as eating food, reacting to danger, or a couple kissing\. In contrast, performance drops on labels such asSelf\-Actualization,Transcendence,Status, andIndependencereaching from 20% to 30%, whose interpretation typically requires accumulated context about the character’s longer\-term goals, abstracted values, or social position\. This trend suggests that MLLMs can handle explicit scene\-level motivation cues, but struggle to identify correct abstracted motivations that require information accumulate across the narrative\.

## 5Conclusion

We introduceMulTivationBench, a benchmark for multimodal sequential motivation reasoning grounded in Maslow’s hierarchy and Reiss’s basic desires\. Results show that current MLLMs identify locally plausible motivations but struggle with consistent reasoning across full narratives\.

## 6Limitations

#### Passive benchmark to evaluate the ability of LLMs

WhileMulTivationBenchadvances the evaluation of multimodal sequential motivation reasoning, it remains an offline benchmark that evaluates models as passive observers of completed narratives\(Chanet al\.,[2024c](https://arxiv.org/html/2607.26465#bib.bib14); Maet al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib43)\)\. In our current setting, models infer motivations from accumulated visual and textual context, but they are not required to act, intervene, or predict how a character’s motivation may change under alternative future observations or environmental changes\. The active reasoning benchmark should treat the language model as an active agent that perceives the physical and social context, reasons about others’ mental states, communicates with other agents, and interacts with the environment to complete predefined tasks\(Yimet al\.,[2024a](https://arxiv.org/html/2607.26465#bib.bib80)\)\. A natural direction for future work is therefore to extend this benchmark toward world\-model\-based motivation reasoning\. Rather than only asking why a behavior occurred, future evaluations could test whether a model can anticipate how motivations evolve, update earlier inferences after new evidence, and simulate how different interventions may alter a character’s future behavior\. We believe this is an important next step toward more active and embodied forms of social intelligence\.

## 7Ethics Statement

In this work, we followed applicable data usage policies and took steps to minimize privacy, licensing, and content\-related risks\.MulTivationBenchis built upon three source datasets: MovieBench \(CC BY 4\.0\), StoryReasoning \(CC BY\-ND 4\.0\), and SSID \(CC BY\-NC\-ND 4\.0\)\. For StoryReasoning and SSID, which carry “No\-Derivatives” restrictions, we obtained explicit written permission from the respective dataset authors to release the derived character\-motivation and behavior\-focused Q&A annotations generated in this work\. In accordance with these permissions, our release contains only newly generated derived annotations, together with a download script and guidelines for obtaining the corresponding images and story context from the upstream datasets\. We do not redistribute original images, story texts, or any substantial portion of the source datasets\. The subset derived from SSID is further restricted to non\-commercial academic research and evaluation only, consistent with the terms granted by the author\. All released materials clearly attribute the corresponding original datasets and publications\. During dataset curation, we filtered out candidate stories and samples containing potentially offensive or inappropriate narrative content\. Because the benchmark is derived from fictional stories and publicly available media, and because we do not redistribute the original multimodal content, we do not anticipate material privacy risks to real individuals from the released materials\. To facilitate reproducibility while respecting licensing constraints, we release only the derived annotations, a download script, and documentation describing how authorized researchers can obtain the corresponding images and story context from the original datasets\.

## 8Acknowledgments

The authors of this paper were supported by the ITSP Platform Research Project \(ITS/189/23FP\) from ITC of Hong Kong, SAR, China, and the AoE \(AoE/E\-601/24\-N\), the RIF \(R6021\-20\) and the GRF \(16205322\) from RGC of Hong Kong, SAR, China\. We also thank the support from Amazon\.

## References

- M\. Abdin, J\. Aneja, H\. Awadalla, A\. Awadallah, A\. A\. Awan, N\. Bach, A\. Bahree, A\. Bakhtiari, J\. Bao, H\. Behl, A\. Benhaim, M\. Bilenko, J\. Bjorck, S\. Bubeck, M\. Cai, Q\. Cai, V\. Chaudhary, D\. Chen, D\. Chen, W\. Chen, Y\. Chen, Y\. Chen, H\. Cheng, P\. Chopra, X\. Dai, M\. Dixon, R\. Eldan, V\. Fragoso, J\. Gao, M\. Gao, M\. Gao, A\. Garg, A\. D\. Giorno, A\. Goswami, S\. Gunasekar, E\. Haider, J\. Hao, R\. J\. Hewett, W\. Hu, J\. Huynh, D\. Iter, S\. A\. Jacobs, M\. Javaheripi, X\. Jin, N\. Karampatziakis, P\. Kauffmann, M\. Khademi, D\. Kim, Y\. J\. Kim, L\. Kurilenko, J\. R\. Lee, Y\. T\. Lee, Y\. Li, Y\. Li, C\. Liang, L\. Liden, X\. Lin, Z\. Lin, C\. Liu, L\. Liu, M\. Liu, W\. Liu, X\. Liu, C\. Luo, P\. Madan, A\. Mahmoudzadeh, D\. Majercak, M\. Mazzola, C\. C\. T\. Mendes, A\. Mitra, H\. Modi, A\. Nguyen, B\. Norick, B\. Patra, D\. Perez\-Becker, T\. Portet, R\. Pryzant, H\. Qin, M\. Radmilac, L\. Ren, G\. de Rosa, C\. Rosset, S\. Roy, O\. Ruwase, O\. Saarikivi, A\. Saied, A\. Salim, M\. Santacroce, S\. Shah, N\. Shang, H\. Sharma, Y\. Shen, S\. Shukla, X\. Song, M\. Tanaka, A\. Tupini, P\. Vaddamanu, C\. Wang, G\. Wang, L\. Wang, S\. Wang, X\. Wang, Y\. Wang, R\. Ward, W\. Wen, P\. Witte, H\. Wu, X\. Wu, M\. Wyatt, B\. Xiao, C\. Xu, J\. Xu, W\. Xu, J\. Xue, S\. Yadav, F\. Yang, J\. Yang, Y\. Yang, Z\. Yang, D\. Yu, L\. Yuan, C\. Zhang, C\. Zhang, J\. Zhang, L\. L\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, and X\. Zhou \(2024\)Phi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219,[Link](https://arxiv.org/abs/2404.14219)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- Rational quantitative attribution of beliefs, desires and percepts in human mentalizing\.Nature Human Behaviour1,pp\. 0064\.External Links:[Document](https://dx.doi.org/10.1038/s41562-017-0064)Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- C\. Baldassano, J\. Chen, A\. Zadbood, J\. W\. Pillow, U\. Hasson, and K\. A\. Norman \(2017\)Discovering event structure in continuous narrative perception and memory\.Neuron95\(3\),pp\. 709–721\.e5\.Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- T\. Berg\-Kirkpatrick, D\. Burkett, and D\. Klein \(2012\)An empirical investigation of statistical significance in NLP\.InProceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning,Jeju Island, Korea,pp\. 995–1005\.External Links:[Link](https://aclanthology.org/D12-1091/)Cited by:[§B\.2](https://arxiv.org/html/2607.26465#A2.SS2.p3.1)\.
- S\. Bubeck, V\. Chandrasekaran, R\. Eldan, J\. Gehrke, E\. Horvitz, E\. Kamar, P\. Lee, Y\. T\. Lee, Y\. Li, S\. M\. Lundberg, H\. Nori, H\. Palangi, M\. T\. Ribeiro, and Y\. Zhang \(2023\)Sparks of artificial general intelligence: early experiments with GPT\-4\.CoRRabs/2303\.12712\.External Links:2303\.12712Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- D\. Caffagni, F\. Cocchi, L\. Barsellotti, N\. Moratelli, S\. Sarto, L\. Baraldi, L\. Baraldi, M\. Cornia, and R\. Cucchiara \(2024\)The revolution of multimodal large language models: a survey\.External Links:2402\.12451,[Link](https://arxiv.org/abs/2402.12451)Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p2.1)\.
- C\. Chan, J\. Cheng, X\. Liu, Y\. Yim, Y\. Jiang, Z\. Deng, H\. Li, Y\. Song, G\. Y\. Wong, and S\. See \(2024a\)Audience persona knowledge\-aligned prompt tuning method for online debate\.InECAI 2024 \- 27th European Conference on Artificial Intelligence, 19\-24 October 2024, Santiago de Compostela, Spain \- Including 13th Conference on Prestigious Applications of Intelligent Systems \(PAIS 2024\),U\. Endriss, F\. S\. Melo, K\. Bach, A\. J\. B\. Diz, J\. M\. Alonso\-Moral, S\. Barro, and F\. Heintz \(Eds\.\),Frontiers in Artificial Intelligence and Applications, Vol\.392,pp\. 3851–3858\.External Links:[Link](https://doi.org/10.3233/FAIA240948),[Document](https://dx.doi.org/10.3233/FAIA240948)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Chan, J\. Cheng, W\. Wang, Y\. Jiang, T\. Fang, X\. Liu, and Y\. Song \(2024b\)Exploring the potential of chatgpt on sentence level relations: A focus on temporal, causal, and discourse relations\.InFindings of the Association for Computational Linguistics: EACL 2024, St\. Julian’s, Malta, March 17\-22, 2024,Y\. Graham and M\. Purver \(Eds\.\),pp\. 684–721\.External Links:[Link](https://aclanthology.org/2024.findings-eacl.47)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Chan, C\. Jiayang, Y\. Yim, Z\. Deng, W\. Fan, H\. Li, X\. Liu, H\. Zhang, W\. Wang, and Y\. Song \(2024c\)NegotiationToM: A benchmark for stress\-testing machine theory of mind on negotiation surrounding\.InFindings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12\-16, 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Findings of ACL,pp\. 4211–4241\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-emnlp.244),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-EMNLP.244)Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p1.1),[§4\.2](https://arxiv.org/html/2607.26465#S4.SS2.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2607.26465#S6.SS0.SSS0.Px1.p1.1)\.
- C\. Chan, X\. Liu, T\. H\. Chan, J\. Cheng, Y\. Song, G\. Y\. Wong, and S\. See \(2023a\)Self\-consistent narrative prompts on abductive natural language inference\.InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics, IJCNLP 2023 \-Volume 1: Long Papers, Nusa Dua, Bali, November 1 \- 4, 2023,J\. C\. Park, Y\. Arase, B\. Hu, W\. Lu, D\. Wijaya, A\. Purwarianti, and A\. A\. Krisnadhi \(Eds\.\),pp\. 1040–1057\.External Links:[Link](https://doi.org/10.18653/v1/2023.ijcnlp-main.67),[Document](https://dx.doi.org/10.18653/V1/2023.IJCNLP-MAIN.67)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Chan, X\. Liu, J\. Cheng, Z\. Li, Y\. Song, G\. Y\. Wong, and S\. See \(2023b\)DiscoPrompt: path prediction prompt tuning for implicit discourse relation recognition\.InFindings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9\-14, 2023,A\. Rogers, J\. L\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),pp\. 35–57\.External Links:[Link](https://doi.org/10.18653/v1/2023.findings-acl.4),[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.4)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Chan, Y\. Yim, H\. Zeng, Z\. Zou, X\. Cheng, Z\. Sun, Z\. Deng, K\. Chung, Y\. Ao, Y\. Fan, C\. Jiayang, E\. Nie, G\. Y\. Wong, H\. Schmid, H\. Schütze, S\. See, and Y\. Song \(2025\)XToM: exploring the multilingual theory of mind for large language models\.CoRRabs/2506\.02461\.External Links:[Link](https://doi.org/10.48550/arXiv.2506.02461),[Document](https://dx.doi.org/10.48550/ARXIV.2506.02461),2506\.02461Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p1.1)\.
- R\. Chen, W\. Jiang, C\. Qin, and C\. Tan \(2025\)Theory of mind in large language models: assessment and enhancement\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 31539–31558\.External Links:[Link](https://aclanthology.org/2025.acl-long.1522/)Cited by:[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- Z\. Chen, J\. Wu, J\. Zhou, B\. Wen, G\. Bi, G\. Jiang, Y\. Cao, M\. Hu, Y\. Lai, Z\. Xiong, and M\. Huang \(2024\)ToMBench: benchmarking theory of mind in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 15959–15983\.External Links:[Link](https://doi.org/10.18653/v1/2024.acl-long.847),[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.847)Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p1.1)\.
- J\. Cheng, L\. Qiu, T\. H\. Chan, T\. Fang, W\. Wang, C\. Chan, D\. Ru, Q\. Guo, H\. Zhang, Y\. Song, Y\. Zhang, and Z\. Zhang \(2023\)StoryAnalogy: deriving story\-level analogies from large language models to unlock analogical understanding\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6\-10, 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),pp\. 11518–11537\.External Links:[Link](https://doi.org/10.18653/v1/2023.emnlp-main.706),[Document](https://dx.doi.org/10.18653/V1/2023.EMNLP-MAIN.706)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- DeepSeek\-AI, D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Bao, H\. Xu, H\. Wang, H\. Ding, H\. Xin, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Wang, J\. Chen, J\. Yuan, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, S\. Ye, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Zhao, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Xu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. Zhang \(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.External Links:2501\.12948,[Link](https://arxiv.org/abs/2501.12948)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Z\. Deng, C\. Chan, W\. Wang, Y\. Sun, W\. Fan, T\. Zheng, Y\. Yim, and Y\. Song \(2024\)Text\-tuple\-table: towards information integration in text\-to\-table generation via global tuple extraction\.CoRRabs/2404\.14215\.External Links:[Link](https://doi.org/10.48550/arXiv.2404.14215),[Document](https://dx.doi.org/10.48550/ARXIV.2404.14215),2404\.14215Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Z\. Deng, C\. Chan, T\. Zheng, W\. Fan, W\. Wang, and Y\. Song \(2025\)Structuring the unstructured: A systematic review of text\-to\-structure generation for agentic AI with a universal evaluation framework\.CoRRabs/2508\.12257\.External Links:[Link](https://doi.org/10.48550/arXiv.2508.12257),[Document](https://dx.doi.org/10.48550/ARXIV.2508.12257),2508\.12257Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- S\. Dong, I\. Shaheen, M\. Shen, R\. Mallick, and S\. A\. Bargal \(2025\)ViSTA: visual storytelling using multi\-modal adapters for text\-to\-image diffusion models\.CoRRabs/2506\.12198\.Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1)\.
- R\. Dror, G\. Baumer, S\. Shlomov, and R\. Reichart \(2018\)The hitchhiker’s guide to testing statistical significance in natural language processing\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Melbourne, Australia,pp\. 1383–1392\.External Links:[Document](https://dx.doi.org/10.18653/v1/P18-1128),[Link](https://aclanthology.org/P18-1128/)Cited by:[§B\.2](https://arxiv.org/html/2607.26465#A2.SS2.p3.1)\.
- B\. Efron \(1979\)Bootstrap methods: another look at the jackknife\.The Annals of Statistics7\(1\),pp\. 1–26\.External Links:[Document](https://dx.doi.org/10.1214/aos/1176344552)Cited by:[§B\.2](https://arxiv.org/html/2607.26465#A2.SS2.p3.1)\.
- J\. L\. Fleiss \(1971\)Measuring nominal scale agreement among many raters\.\.Psychological bulletin76\(5\),pp\. 378\.Cited by:[§3\.3](https://arxiv.org/html/2607.26465#S3.SS3.SSS0.Px2.p1.1)\.
- S\. Frieder, L\. Pinchetti, R\. Griffiths, T\. Salvatori, T\. Lukasiewicz, P\. C\. Petersen, A\. Chevalier, and J\. Berner \(2023\)Mathematical capabilities of chatgpt\.CoRRabs/2301\.13867\.External Links:2301\.13867Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Fu, Y\. Dai, Y\. Luo, L\. Li, S\. Ren, R\. Zhang, Z\. Wang, C\. Zhou, Y\. Shen, M\. Zhang, P\. Chen, Y\. Li, S\. Lin, S\. Zhao, K\. Li, T\. Xu, X\. Zheng, E\. Chen, R\. Ji, and X\. Sun \(2024\)Video\-mme: the first\-ever comprehensive evaluation benchmark of multi\-modal llms in video analysis\.CoRRabs/2405\.21075\.Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1)\.
- S\. Golchin and M\. Surdeanu \(2024\)Time travel in llms: tracing data contamination in large language models\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=2Rwq6c3tvr)Cited by:[§B\.1](https://arxiv.org/html/2607.26465#A2.SS1.p1.1)\.
- Google \(2025\)Introducing Gemini 3 Flash: benchmarks, global availability\.Note:Official announcementExternal Links:[Link](https://blog.google/products/gemini/gemini-3-flash/)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, Y\. Wang, and J\. Guo \(2024\)A survey on llm\-as\-a\-judge\.CoRRabs/2411\.15594\.External Links:[Link](https://doi.org/10.48550/arXiv.2411.15594),[Document](https://dx.doi.org/10.48550/ARXIV.2411.15594),2411\.15594Cited by:[§D\.1](https://arxiv.org/html/2607.26465#A4.SS1.p1.1)\.
- M\. S\. Hagger and N\. L\. Chatzisarantis \(2005\)The social psychology of exercise and sport\.McGraw\-Hill Education \(UK\)\.Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. Liu \(2025\)A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions\.ACM Trans\. Inf\. Syst\.43\(2\),pp\. 42:1–42:55\.External Links:[Link](https://doi.org/10.1145/3703155),[Document](https://dx.doi.org/10.1145/3703155)Cited by:[§D\.1](https://arxiv.org/html/2607.26465#A4.SS1.p2.1)\.
- Y\. Jiang, C\. Chan, M\. Chen, and W\. Wang \(2023\)Lion: adversarial distillation of proprietary large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6\-10, 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),pp\. 3134–3154\.External Links:[Link](https://doi.org/10.18653/v1/2023.emnlp-main.189),[Document](https://dx.doi.org/10.18653/V1/2023.EMNLP-MAIN.189)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Jiayang, C\. Chan, Q\. Zhuang, L\. Qiu, T\. Zhang, T\. Liu, Y\. Song, Y\. Zhang, P\. Liu, and Z\. Zhang \(2024a\)ECON: on the detection and resolution of evidence conflicts\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12\-16, 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),pp\. 7816–7844\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.447)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Jiayang, L\. Qiu, C\. Chan, X\. Liu, Y\. Song, and Z\. Zhang \(2024b\)EventGround: narrative reasoning by grounding to eventuality\-centric knowledge graphs\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20\-25 May, 2024, Torino, Italy,N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),pp\. 6622–6642\.External Links:[Link](https://aclanthology.org/2024.lrec-main.587)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Jiayang, Q\. Zhuang, H\. Li, C\. Chan, X\. Liu, L\. Qiu, and Y\. Song \(2025\)InteGround: on the evaluation of verification and retrieval planning in integrative grounding\.InFindings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4\-9, 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),pp\. 13587–13602\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.732/)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Jin, Y\. Wu, J\. Cao, J\. Xiang, Y\. Kuo, Z\. Hu, T\. D\. Ullman, A\. Torralba, J\. B\. Tenenbaum, and T\. Shu \(2024\)MMToM\-qa: multimodal theory of mind question answering\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 16077–16102\.Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p2.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- A\. E\. Kazdin \(Ed\.\) \(2000\)Encyclopedia of psychology\.Oxford University Press\.Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- H\. Kim, M\. Sclar, X\. Zhou, R\. L\. Bras, G\. Kim, Y\. Choi, and M\. Sap \(2023\)FANToM: A benchmark for stress\-testing machine theory of mind in interactions\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6\-10, 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),pp\. 14397–14413\.External Links:[Link](https://doi.org/10.18653/v1/2023.emnlp-main.890),[Document](https://dx.doi.org/10.18653/V1/2023.EMNLP-MAIN.890)Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p1.1)\.
- G\. Knoblich and N\. Sebanz \(2008\)Evolving intentions for social interaction: from entrainment to joint action\.Philosophical transactions of the Royal Society of London\. Series B, Biological sciences363,pp\. 2021–31\.External Links:[Document](https://dx.doi.org/10.1098/rstb.2008.0006)Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- S\. Lee, J\. Jeong, D\. Kim, Y\. Son, and Y\. Yu \(2025\)Mind the motions: benchmarking theory\-of\-mind in everyday body language\.CoRRabs/2511\.15887\.Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p2.1),[§4\.4](https://arxiv.org/html/2607.26465#S4.SS4.p1.1)\.
- H\. Li, Y\. Chen, J\. Luo, Y\. Kang, X\. Zhang, Q\. Hu, C\. Chan, and Y\. Song \(2023\)Privacy in large language models: attacks, defenses and future directions\.CoRRabs/2310\.10383\.External Links:[Link](https://doi.org/10.48550/arXiv.2310.10383),[Document](https://dx.doi.org/10.48550/ARXIV.2310.10383),2310\.10383Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- H\. Li, Y\. Chen, Z\. Zheng, Q\. Hu, C\. Chan, H\. Liu, and Y\. Song \(2025a\)Simulate and eliminate: revoke backdoors for generative large language models\.InThirty\-Ninth AAAI Conference on Artificial Intelligence, Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 \- March 4, 2025,T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),pp\. 397–405\.External Links:[Link](https://doi.org/10.1609/aaai.v39i1.32018),[Document](https://dx.doi.org/10.1609/AAAI.V39I1.32018)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- H\. Li, D\. Guo, D\. Li, W\. Fan, Q\. Hu, X\. Liu, C\. Chan, D\. Yao, Y\. Yao, and Y\. Song \(2024a\)PrivLM\-bench: A multi\-level privacy evaluation benchmark for language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 54–73\.External Links:[Link](https://aclanthology.org/2024.acl-long.4)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- K\. Li, Y\. Wang, Y\. He, Y\. Li, Y\. Wang, Y\. Liu, W\. Zun, J\. Xu, G\. Chen, P\. Luo, L\. Wang, and Y\. Qiao \(2024b\)MVBench: a comprehensive multi\-modal video understanding benchmark\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 22195–22206\.Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- Y\. Li, W\. Zhang, Y\. Yang, W\. Huang, Y\. Wu, J\. Luo, Y\. Bei, H\. P\. Zou, X\. Luo, Y\. Zhao, C\. Chan, Y\. Chen, Z\. Deng, Y\. Li, H\. Zheng, D\. Li, R\. Jiang, M\. Zhang, Y\. Song, and P\. S\. Yu \(2025b\)Towards agentic RAG with deep reasoning: A survey of rag\-reasoning systems in llms\.CoRRabs/2507\.09477\.External Links:[Link](https://doi.org/10.48550/arXiv.2507.09477),[Document](https://dx.doi.org/10.48550/ARXIV.2507.09477),2507\.09477Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- F\. Liang, T\. Zheng, C\. Chan, Y\. Yim, and Y\. Song \(2025\)LLM\-hanabi: evaluating multi\-agent gameplays with theory\-of\-mind and rationale inference in imperfect information collaboration game\.CoRRabs/2510\.04980\.External Links:[Link](https://doi.org/10.48550/arXiv.2510.04980),[Document](https://dx.doi.org/10.48550/ARXIV.2510.04980),2510\.04980Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Z\. Lin, C\. Chan, Y\. Song, and X\. Liu \(2024\)Constrained reasoning chains for enhancing theory\-of\-mind in large language models\.InPRICAI 2024: Trends in Artificial Intelligence \- 21st Pacific Rim International Conference on Artificial Intelligence, PRICAI 2024, Kyoto, Japan, November 18\-24, 2024, Proceedings, Part II,R\. Hadfi, P\. Anthony, A\. Sharma, T\. Ito, and Q\. Bai \(Eds\.\),Lecture Notes in Computer Science, Vol\.15282,pp\. 354–360\.External Links:[Link](https://doi.org/10.1007/978-981-96-0119-6%5C_34),[Document](https://dx.doi.org/10.1007/978-981-96-0119-6%5F34)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Z\. Lin, Z\. Xu, X\. Song, Y\. Wan, X\. Yao, T\. Lin, S\. Song, P\. Subbaraman, B\. Zhou, K\. Chang, and Y\. Sun \(2025\)V\-alphasocial: benchmark and self\-reflective chain\-of\-thought generation for visual social commonsense reasoning\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 19025–19047\.External Links:[Link](https://dblp.org/rec/conf/acl/LinXSWYLSSZCS25)Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p2.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- E\. Loza Mencía, M\. Kulessa, S\. Bohlender, and J\. Fürnkranz \(2023\)Tree\-based dynamic classifier chains\.Machine Learning112,pp\. 4129–4165\.External Links:[Document](https://dx.doi.org/10.1007/s10994-022-06162-3)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.SSS0.Px1.p1.1)\.
- N\. Lukas, A\. Salem, R\. Sim, S\. Tople, L\. Wutschitz, and S\. Z\. Béguelin \(2023\)Analyzing leakage of personally identifiable information in language models\.CoRRabs/2302\.00539\.External Links:2302\.00539Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Y\. Ma, W\. Xu, C\. Zhao, K\. Sun, Q\. Jin, X\. Yang, Z\. Zhao, C\. Fan, and Z\. Hu \(2025\)Storynizor: consistent story generation via inter\-frame synchronized and shuffled id injection\.Proceedings of the AAAI Conference on Artificial Intelligence39\(6\),pp\. 6027–6035\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v39i6.32644)Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1)\.
- Z\. Ma, J\. Sansom, R\. Peng, and J\. Chai \(2023\)Towards a holistic landscape of situated theory of mind in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 1011–1031\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.72/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.72)Cited by:[§6](https://arxiv.org/html/2607.26465#S6.SS0.SSS0.Px1.p1.1)\.
- M\. Maaz, H\. Rasheed, S\. Khan, and F\. Khan \(2024\)Video\-ChatGPT: towards detailed video understanding via large vision and language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12585–12602\.External Links:[Link](https://aclanthology.org/2024.acl-long.679/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.679)Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- A\. Maharana, D\. Hannan, and M\. Bansal \(2022\)StoryDALL\-e: adapting pretrained text\-to\-image transformers for story continuation\.InEuropean Conference on Computer Vision,pp\. 70–87\.Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1)\.
- Z\. M\. Malakan, S\. Anwar, G\. M\. Hassan, and A\. Mian \(2023\)Sequential storytelling image dataset \(SSID\)\.IEEE DataPort\.External Links:[Document](https://dx.doi.org/10.21227/dbr9-dq51)Cited by:[§3\.1](https://arxiv.org/html/2607.26465#S3.SS1.p1.1)\.
- K\. Mangalam, R\. Akshulakov, and J\. Malik \(2023\)EgoSchema: a diagnostic benchmark for very long\-form video language understanding\.InAdvances in Neural Information Processing Systems,Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1)\.
- A\. H\. Maslow \(1970\)Motivation and personality\.2nd edition,Harper & Row,New York\.Cited by:[§A\.2](https://arxiv.org/html/2607.26465#A1.SS2.p1.1),[§1](https://arxiv.org/html/2607.26465#S1.p3.1),[§3\.2](https://arxiv.org/html/2607.26465#S3.SS2.SSS0.Px1.p1.1)\.
- Meta AI \(2025\)The llama 4 herd: the beginning of a new era of natively multimodal intelligence\.Note:Covers Llama\-4 Scout\-17B and Maverick\-17BExternal Links:[Link](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- Microsoft, :, A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen, D\. Chen, D\. Chen, J\. Chen, W\. Chen, Y\. Chen, Y\. Chen, Q\. Dai, X\. Dai, R\. Fan, M\. Gao, M\. Gao, A\. Garg, A\. Goswami, J\. Hao, A\. Hendy, Y\. Hu, X\. Jin, M\. Khademi, D\. Kim, Y\. J\. Kim, G\. Lee, J\. Li, Y\. Li, C\. Liang, X\. Lin, Z\. Lin, M\. Liu, Y\. Liu, G\. Lopez, C\. Luo, P\. Madan, V\. Mazalov, A\. Mitra, A\. Mousavi, A\. Nguyen, J\. Pan, D\. Perez\-Becker, J\. Platin, T\. Portet, K\. Qiu, B\. Ren, L\. Ren, S\. Roy, N\. Shang, Y\. Shen, S\. Singhal, S\. Som, X\. Song, T\. Sych, P\. Vaddamanu, S\. Wang, Y\. Wang, Z\. Wang, H\. Wu, H\. Xu, W\. Xu, Y\. Yang, Z\. Yang, D\. Yu, I\. Zabir, J\. Zhang, L\. L\. Zhang, Y\. Zhang, and X\. Zhou \(2025\)Phi\-4\-mini technical report: compact yet powerful multimodal language models via mixture\-of\-loras\.External Links:2503\.01743,[Link](https://arxiv.org/abs/2503.01743)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- Y\. Mo, T\. Zheng, Q\. Zong, J\. Liu, B\. Xu, Y\. Yim, C\. Chan, J\. Bai, and Y\. Song \(2025\)DixitWorld: evaluating multimodal abductive reasoning in vision\-language models with multi\-agent dixit gameplay\.CoRRabs/2510\.10117\.External Links:[Link](https://doi.org/10.48550/arXiv.2510.10117),[Document](https://dx.doi.org/10.48550/ARXIV.2510.10117),2510\.10117Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- N\. Mostafazadeh, A\. Kalyanpur, L\. Moon, D\. W\. Buchanan, L\. Berkowitz, O\. Biran, and J\. Chu\-Carroll \(2020\)GLUCOSE: generalized and contextualized story explanations\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 4569–4586\.External Links:[Link](https://dblp.org/rec/conf/emnlp/MostafazadehKMB20)Cited by:[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- M\. Ning, B\. Zhu, Y\. Xie, B\. Lin, J\. Cui, L\. Yuan, D\. Chen, and L\. Yuan \(2023\)Video\-bench: a comprehensive benchmark and toolkit for evaluating video\-based large language models\.CoRRabs/2311\.16103\.Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1)\.
- NVIDIA \(2025\)NVIDIA nemotron nano v2 vl\.Note:Official report for Nemotron\-Nano\-12B\-v2\-vlExternal Links:[Link](https://research.nvidia.com/labs/adlr/files/NVIDIA-Nemotron-Nano-V2-VL-report.pdf)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- D\. A\. P\. Oliveira and D\. M\. d\. Matos \(2024\)StoryReasoning dataset: using chain\-of\-thought for scene understanding and grounded story generation\.arXiv preprint\.Note:\(Forthcoming\)Cited by:[§3\.1](https://arxiv.org/html/2607.26465#S3.SS1.p1.1)\.
- OpenAI \(2023\)GPT\-4 technical report\.CoRRabs/2303\.08774\.External Links:[Link](https://doi.org/10.48550/arXiv.2303.08774),[Document](https://dx.doi.org/10.48550/ARXIV.2303.08774),2303\.08774Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- OpenAI \(2025\)Introducing openai o3 and o4\-mini\.External Links:[Link](https://openai.com/index/introducing-o3-and-o4-mini/)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- H\. Rashkin, A\. Bosselut, M\. Sap, K\. Knight, and Y\. Choi \(2018a\)Modeling naive psychology of characters in simple commonsense stories\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2289–2299\.External Links:[Document](https://dx.doi.org/10.18653/V1/P18-1213),[Link](https://dblp.org/rec/conf/acl/KnightCSRB18)Cited by:[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- H\. Rashkin, M\. Sap, E\. Allaway, N\. A\. Smith, and Y\. Choi \(2018b\)Event2Mind: commonsense inference on events, intents, and reactions\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 463–473\.External Links:[Document](https://dx.doi.org/10.18653/V1/P18-1043),[Link](https://dblp.org/rec/conf/acl/SmithCSRA18)Cited by:[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- S\. Reiss \(2004\)Multifaceted nature of intrinsic motivation: the theory of 16 basic desires\.Review of General Psychology8\(3\),pp\. 179–193\.External Links:[Document](https://dx.doi.org/10.1037/1089-2680.8.3.179)Cited by:[§A\.3](https://arxiv.org/html/2607.26465#A1.SS3.p1.1),[§1](https://arxiv.org/html/2607.26465#S1.p3.1),[§3\.2](https://arxiv.org/html/2607.26465#S3.SS2.SSS0.Px2.p1.1)\.
- S\. Ren, L\. Yao, S\. Li, X\. Sun, and L\. Hou \(2024\)TimeChat: a time\-sensitive multimodal large language model for long video understanding\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 14313–14323\.Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- R\. M\. Ryan and E\. L\. Deci \(2000\)Self\-determination theory and the facilitation of intrinsic motivation, social development, and well\-being\.American Psychologist55\(1\),pp\. 68\.Cited by:[§A\.1](https://arxiv.org/html/2607.26465#A1.SS1.p3.1),[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- M\. Sap, R\. L\. Bras, E\. Allaway, C\. Bhagavatula, N\. Lourie, H\. Rashkin, B\. Roof, N\. A\. Smith, and Y\. Choi \(2019a\)ATOMIC: an atlas of machine commonsense for if\-then reasoning\.InThe Thirty\-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty\-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 \- February 1, 2019,pp\. 3027–3035\.External Links:[Link](https://doi.org/10.1609/aaai.v33i01.33013027),[Document](https://dx.doi.org/10.1609/AAAI.V33I01.33013027)Cited by:[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- M\. Sap, H\. Rashkin, D\. Chen, R\. L\. Bras, and Y\. Choi \(2019b\)SocialIQA: commonsense reasoning about social interactions\.CoRRabs/1904\.09728\.External Links:[Link](http://arxiv.org/abs/1904.09728),1904\.09728Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- B\. J\. Scholl and P\. D\. Tremoulet \(2000\)Perceptual causality and animacy\.Trends in Cognitive Sciences4\(8\),pp\. 299–309\.Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- H\. Shi, T\. Zheng, W\. Wang, B\. Xu, C\. Li, C\. Chan, T\. Fan, Y\. Song, and Q\. Yang \(2025\)INFERENCEDYNAMICS: efficient routing across llms through structured capability and knowledge profiling\.External Links:2505\.16303,[Link](https://arxiv.org/abs/2505.16303)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- D\. Song, S\. Lai, M\. Wang, S\. Chen, L\. Sun, and B\. Wang \(2025\)Both text and images leaked\! a systematic analysis of data contamination in multimodal LLM\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 10527–10542\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.556/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.556),ISBN 979\-8\-89176\-335\-7Cited by:[§B\.1](https://arxiv.org/html/2607.26465#A2.SS1.p2.3)\.
- T\. Susnjak \(2022\)ChatGPT: the end of online exam integrity?\.CoRRabs/2212\.09292\.External Links:2212\.09292Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- E\. Villa\-Cueva, S\. M\. M\. Ahmed, R\. Chevi, J\. C\. B\. Cruz, K\. Elzeky, F\. Cristobal, A\. F\. Aji, S\. Wang, R\. Mihalcea, and T\. Solorio \(2025\)MoMentS: a comprehensive multimodal benchmark for theory of mind\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 22591–22611\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1230),[Link](https://aclanthology.org/2025.findings-emnlp.1230/)Cited by:[§C\.2](https://arxiv.org/html/2607.26465#A3.SS2.p2.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- M\. Wang, H\. Ding, J\. Peng, Y\. Zhao, Y\. Chen, and Y\. Wei \(2025\)CharaConsist: fine\-grained consistent character generation\.CoRRabs/2507\.11533\.Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1)\.
- W\. Wang, T\. Fang, C\. Li, H\. Shi, W\. Ding, B\. Xu, Z\. Wang, J\. Bai, X\. Liu, C\. Jiayang, C\. Chan, and Y\. Song \(2024a\)CANDLE: iterative conceptualization and instantiation distillation from large language models for commonsense reasoning\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 2351–2374\.External Links:[Link](https://doi.org/10.18653/v1/2024.acl-long.128),[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.128)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- Y\. Wang, Y\. Wang, P\. Wu, J\. Liang, D\. Zhao, Y\. Liu, and Z\. Zheng \(2024b\)Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 9972–9987\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.556)Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1)\.
- H\. Wu, D\. Li, B\. Chen, and J\. Li \(2024a\)LongVideoBench: a benchmark for long\-context interleaved video\-language understanding\.InAdvances in Neural Information Processing Systems,Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.SSS0.Px1.p1.1)\.
- W\. Wu, M\. Liu, Z\. Zhu, X\. Xia, H\. Feng, W\. Wang, K\. Q\. Lin, C\. Shen, and M\. Z\. Shou \(2024b\)MovieBench: a hierarchical movie level dataset for long video generation\.arXiv preprint arXiv:2411\.15262\.Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1),[§3\.1](https://arxiv.org/html/2607.26465#S3.SS1.p1.1)\.
- xAI \(2025\)Grok 4\.1 model card\.Note:Official model card for Grok\-4\.1 variants, including FastExternal Links:[Link](https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf)Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- Y\. Yang, W\. Wang, B\. Xu, W\. Fan, Q\. Zong, C\. Chan, Z\. Deng, X\. Liu, Y\. Gao, C\. Yu, C\. Luo, Y\. Li, Z\. Li, Q\. Yin, B\. Yin, and Y\. Song \(2025\)SessionIntentBench: A multi\-task inter\-session intention\-shift modeling benchmark for e\-commerce customer behavior understanding\.CoRRabs/2507\.20185\.External Links:[Link](https://doi.org/10.48550/arXiv.2507.20185),[Document](https://dx.doi.org/10.48550/ARXIV.2507.20185),2507\.20185Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- A\. Yeh \(2000\)More accurate tests for the statistical significance of result differences\.InCOLING 2000 Volume 2: The 18th International Conference on Computational Linguistics,External Links:[Link](https://aclanthology.org/C00-2137/)Cited by:[§B\.2](https://arxiv.org/html/2607.26465#A2.SS2.p3.1)\.
- Y\. Yim, C\. Chan, T\. Shi, Z\. Deng, W\. Fan, T\. Zheng, and Y\. Song \(2024a\)Evaluating and enhancing llms agent based on theory of mind in guandan: A multi\-player cooperative game under imperfect information\.InIEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology, WI\-IAT 2024, Bangkok, Thailand, December 9\-12, 2024,pp\. 461–465\.External Links:[Link](https://doi.org/10.1109/WI-IAT62293.2024.00074),[Document](https://dx.doi.org/10.1109/WI-IAT62293.2024.00074)Cited by:[§6](https://arxiv.org/html/2607.26465#S6.SS0.SSS0.Px1.p1.1)\.
- Y\. Yim, C\. Chan, T\. Shi, Z\. Deng, W\. Fan, T\. Zheng, and Y\. Song \(2024b\)Evaluating and enhancing llms agent based on theory of mind in guandan: A multi\-player cooperative game under imperfect information\.CoRRabs/2408\.02559\.External Links:[Link](https://doi.org/10.48550/arXiv.2408.02559),[Document](https://dx.doi.org/10.48550/ARXIV.2408.02559),2408\.02559Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- X\. Yong, J\. Lian, X\. Yi, X\. Zhou, and X\. Xie \(2025\)MotiveBench: how far are we from human\-like motivational reasoning in large language models?\.InFindings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 20059–20089\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1029/)Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.26465#S2.SS1.p1.1)\.
- J\. Zacks and K\. Swallow \(2007\)Event segmentation\.Current directions in psychological science16,pp\. 80–84\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-8721.2007.00480.x)Cited by:[§1](https://arxiv.org/html/2607.26465#S1.p1.1)\.
- M\. Zhang and Z\. Zhou \(2014\)A review on multi\-label learning algorithms\.IEEE Transactions on Knowledge and Data Engineering26\(8\),pp\. 1819–1837\.External Links:[Document](https://dx.doi.org/10.1109/TKDE.2013.39)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.SSS0.Px1.p1.1)\.
- W\. Zhang, Y\. Li, Y\. Bei, J\. Luo, G\. Wan, L\. Yang, C\. Xie, Y\. Yang, W\. Huang, C\. Miao, H\. P\. Zou, X\. Luo, Y\. Zhao, Y\. Chen, C\. Chan, P\. Zhou, X\. Zhang, C\. Zhang, J\. Shang, M\. Zhang, Y\. Song, I\. King, and P\. S\. Yu \(2025\)From web search towards agentic deep research: incentivizing search with reasoning agents\.CoRRabs/2506\.18959\.External Links:[Link](https://doi.org/10.48550/arXiv.2506.18959),[Document](https://dx.doi.org/10.48550/ARXIV.2506.18959),2506\.18959Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.
- C\. Zheng, H\. Zhou, F\. Meng, J\. Zhou, and M\. Huang \(2024\)Large language models are not robust multiple choice selectors\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=shr9PXz7T0)Cited by:[§4\.1](https://arxiv.org/html/2607.26465#S4.SS1.p1.1)\.
- J\. Zhou, Y\. Shu, B\. Zhao, B\. Wu, Z\. Liang, S\. Xiao, M\. Qin, Y\. Xi, Y\. Xiong, B\. Zhang, T\. Huang, and Z\. Liu \(2025\)MLVU: benchmarking multi\-task long video understanding\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13691–13701\.Cited by:[§C\.3](https://arxiv.org/html/2607.26465#A3.SS3.p1.1)\.
- Y\. Zhou, D\. Zhou, M\. Cheng, J\. Feng, and Q\. Hou \(2024\)StoryDiffusion: consistent self\-attention for long\-range image and video generation\.CoRRabs/2405\.01434\.Cited by:[§C\.4](https://arxiv.org/html/2607.26465#A3.SS4.p1.1)\.
- Q\. Zong, J\. Liu, T\. Zheng, C\. Li, B\. Xu, H\. Shi, W\. Wang, Z\. Wang, C\. Chan, and Y\. Song \(2025\)CritiCal: can critique help LLM uncertainty or confidence calibration?\.CoRRabs/2510\.24505\.External Links:[Link](https://doi.org/10.48550/arXiv.2510.24505),[Document](https://dx.doi.org/10.48550/ARXIV.2510.24505),2510\.24505Cited by:[§C\.1](https://arxiv.org/html/2607.26465#A3.SS1.p1.1)\.

## Appendix AAppendix for the Motivation Frameworks

### A\.1Choice of Psychological Frameworks

We use Maslow’s expanded hierarchy and Reiss’s 16 basic desires as complementary operational taxonomies rather than as an exhaustive theory of human motivation\. Maslow provides a coarse\-grained need\-level structure covering broad categories such as safety, belonging, esteem, cognitive needs, self\-actualization, and transcendence\. This makes it suitable for evaluating whether models can identify the broad motivational region expressed by a behavior\. Reiss’s theory provides a finer\-grained desire\-level taxonomy, distinguishing motivations such as family, acceptance, social contact, independence, status, curiosity, and idealism\. This enables a more detailed evaluation of whether models can discriminate among closely related motivational contents\.

This design supports the central coarse\-to\-fine comparison in our benchmark\. The Maslow tasks evaluate broad need recognition, while the Reiss tasks evaluate fine\-grained desire discrimination\. We do not assume that Maslow’s hierarchy is the only valid or definitive theory of motivation, nor do we rely on a strong claim that needs must always be satisfied in a strict hierarchy\. Instead, we use the expanded hierarchy as an interpretable broad\-level label space\.

Other theories of motivation are also relevant\. For example, Self\-Determination Theory\(Ryan and Deci,[2000](https://arxiv.org/html/2607.26465#bib.bib20)\)emphasizes autonomy, competence, and relatedness as central dimensions of intrinsic motivation\. These dimensions are useful for studying intrinsic motivation, but they provide a smaller and more abstract label space than required for our current multi\-label benchmark over diverse everyday narrative behaviors\. We therefore view Self\-Determination Theory as a valuable direction for future extensions ofMultivationBench, especially for studying intrinsic motivation in greater detail\.

### A\.2Details of the Hierarchy of Needs

Maslow’s Expanded Hierarchy of Needs\(Maslow,[1970](https://arxiv.org/html/2607.26465#bib.bib2)\)is a motivational theory that explains how human needs may be prioritized and fulfilled, from basic survival to higher\-order psychological growth and meaning\. While Maslow originally discussed five levels, later formulations commonly expand the hierarchy; in this work we adopt an 8\-level version \(Physiological, Safety, Love & Belonging, Esteem, Cognitive, Aesthetic, Self\-Actualization, Transcendence\) to cover both lower\-level and higher\-level motives in a unified taxonomy\.

•Physiological Needs:Physiological needs refer to essential survival requirements such as food, water, warmth, and sleep\. These needs form the foundation of the hierarchy and must be sufficiently met before individuals can consistently pursue higher\-level goals\.

•Safety Needs:Safety needs emphasize stability and protection, including physical security, health, order, and freedom from fear\. They capture motivations related to predictability and safeguarding one’s future\.

•Love and Belonging Needs:Love and belonging needs reflect the drive for social connection through friendship, intimacy, trust, and affiliation with groups \(e\.g\., family, peers, community\)\. They involve motivations to be accepted, supported, and meaningfully connected to others\.

•Esteem Needs:Esteem needs involve both internal self\-worth \(e\.g\., competence, independence, achievement\) and external recognition \(e\.g\., respect, status, prestige\)\. Meeting these needs supports confidence and social standing\.

•Cognitive Needs:Cognitive needs reflect motivations for knowledge and understanding, including curiosity, exploration, and the pursuit of meaning and predictability\.

•Aesthetic Needs:Aesthetic needs involve appreciation and pursuit of beauty, balance, and form, capturing motivations oriented toward harmony and aesthetic experience\.

•Self\-Actualization Needs:Self\-actualization concerns realizing personal potential through growth, mastery, and self\-fulfillment\. It is often expressed via creativity, goal pursuit, and striving toward an ideal self\.

•Transcendence Needs:Transcendence needs describe motivations beyond the personal self, such as altruism, spiritual connection, and contributing to a larger purpose or the well\-being of others\.

### A\.3Details of Reiss’s Basic Desires and the RMP

The Reiss theory of 16 basic desires\(Reiss,[2004](https://arxiv.org/html/2607.26465#bib.bib4)\)provides a complementary view of intrinsic motivation, positing that behavior can be explained by preferences over 16 fundamental motivational drives\. The Reiss Motivation Profile \(RMP\) operationalizes this idea as an inventory: individuals differ in the intensity of each desire, and these differences shape behavior, decision\-making, and lifestyle\.

To support interpretability when discussing fine\-grained motivation types, we summarize the 16 desires using Maslow\-style groupings\.Physiological\-oriented desires such asEatingandPhysical Exercisecapture bodily drives and activity\.Safety\-oriented desires such asOrder,Saving, andTranquilityreflect preferences for structure, safeguarding one’s future, and peace of mind\.Love & Belonging\-oriented desires includeAcceptance,Family,Romance, andSocial Contact, spanning approval, close bonds, companionship, and affiliation\.Esteem\-oriented desires such asStatus,Power,Vengeance, andHonoremphasize recognition, prestige, influence, and status\-defense\.Cognitive\-oriented desires such asCuriositymotivate knowledge seeking, understanding, and exploration\.Self\-Actualization\-oriented desires such asIndependenceemphasize autonomy, mastery, and personal growth\. Finally,Transcendence\-oriented desires such asIdealismcapture value\-driven motives beyond the personal self, including prosocial principles and contributing to a larger purpose\.

## Appendix BAppendix forMultivationBench

### B\.1Verification of Potential Contamination

A potential concern is that some source materials underlyingMultivationBenchmay have been exposed to large language models during training, since widely used NLP and multimodal benchmarks are often incorporated into both pre\-training and post\-training pipelines\(Golchin and Surdeanu,[2024](https://arxiv.org/html/2607.26465#bib.bib58)\)\. This issue is particularly relevant becauseMultivationBenchintegrates visual narratives from diverse sources, including MovieBench, SSID, and StoryReasoning, whose root materials may partially overlap with publicly available web data\.

To assess this risk, we follow theSlot Guessing for Perturbed Captionmethodology proposed bySonget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib57)\)\. For each constituent dataset, we randomly sample 100 instances and evaluate the model under two conditions: \(1\) anOriginalcondition using the original story text, and \(2\) aPerturbedcondition using meaning\-preserving paraphrases\. We then report thePerformance Gap\(Δ\\Delta\) between the two conditions and theMemorization Ratio\(Φ\\Phi\)\. Under this diagnostic, a substantially negativeΔ\\Deltawould indicate stronger reliance on exact surface forms, and thus provide evidence more consistent with memorization than robust reasoning\.

As shown in Table[4](https://arxiv.org/html/2607.26465#A2.T4), the observed performance gaps are small across all three source datasets\. MovieBench exhibits the largest negative gap \(Δ=−0\.07\\Delta=\-0\.07\), suggesting a limited contamination signal, which is plausible given the prevalence of movie\-related text in public training corpora\. SSID shows only a minor negative gap \(Δ=−0\.03\\Delta=\-0\.03\)\. By contrast, StoryReasoning yields a positive gap \(Δ=\+0\.06\\Delta=\+0\.06\), indicating that model performance is not tied to the original wording and showing no evidence of strong phrase\-level memorization under this test\.

While no systematic approach can fully rule out contamination without access to private training sets, we include this analysis to explicitly examine the risk and to verify that benchmark performance is not solely driven by exact\-match recall\. Overall, the results suggest thatMultivationBenchis not predominantly exploiting memorized textual patterns, and instead serves primarily as a test of reasoning over visual narratives\.

Table 4:Contamination probe results\.We compare model performance under original and paraphrased story text\. More negativeΔ\\Deltavalues indicate greater sensitivity to exact wording and therefore stronger evidence consistent with memorization\. Overall, all threeMultivationBenchsources exhibit only small performance gaps, suggesting limited contamination signals under this diagnostic\.
### B\.2Analysis of Potential Generator\-Family Advantage

Table 5:Generator\-family analysis in the mainMultimodalsetting\. OverallΔ\\Deltais the difference between the mean performance of the generator\-family and non\-generator\-family groups\.Δg=\(Practical−Definition\)gen−\(Practical−Definition\)non\-gen\\Delta\_\{g\}=\(\\text\{Practical\}\-\\text\{Definition\}\)\_\{\\text\{gen\}\}\-\(\\text\{Practical\}\-\\text\{Definition\}\)\_\{\\text\{non\-gen\}\}measures whether the generator\-family models obtain an additional advantage specifically on thePracticaltasks\. While generator\-family models perform better overall, the gap analysis does not indicate a corresponding Practical\-specific benefit\.A natural concern is that overlap between the benchmark construction pipeline and some evaluated model families may favor those models at test time\. This concern can be split into two separate questions\. First, are generator\-family models simply stronger overall? Second, do they enjoy an additional advantage specifically on thePracticaltasks, where overlap with the construction pipeline is most plausible?

To address this, we examine both the overall group difference and the difference in thePractical\-vs\.\-Definitionperformance gaps:

Δg\\displaystyle\\Delta\_\{g\}=\(Practical−Definition\)gen\\displaystyle=\(\\text\{Practical\}\-\\text\{Definition\}\)\_\{\\text\{gen\}\}−\(Practical−Definition\)non\-gen\.\\displaystyle\\quad\-\(\\text\{Practical\}\-\\text\{Definition\}\)\_\{\\text\{non\-gen\}\}\.Under a strong family\-specific construction bias, we would expect not only a positive overall difference, but also a clearly positiveΔg\\Delta\_\{g\}in the mainMultimodalmultimodal setting\.

We estimate uncertainty with nonparametric bootstrap resampling, following standard statistical practice in NLP evaluation\(Efron,[1979](https://arxiv.org/html/2607.26465#bib.bib76); Droret al\.,[2018](https://arxiv.org/html/2607.26465#bib.bib77); Berg\-Kirkpatricket al\.,[2012](https://arxiv.org/html/2607.26465#bib.bib78); Yeh,[2000](https://arxiv.org/html/2607.26465#bib.bib79)\)\. We also run an exact permutation test over model\-family assignments\. The results are shown in Table[5](https://arxiv.org/html/2607.26465#A2.T5)\.

The analysis reveals a clear overall advantage for generator\-family models in the main multimodal setting\. The mean difference between the generator\-family and non\-generator\-family groups is \+5\.88 points in EM accuracy \(95% CI: \[5\.29, 6\.46\]\) and \+9\.42 points in F1 score \(95% CI: \[8\.94, 9\.89\]\)\. However, this overall advantage does not extend to a reliable extra gain on thePracticaltasks\. For EM accuracy, the estimated gap difference is small \(\+0\.77\), and its 95% confidence interval crosses zero \(\[\-0\.42, 1\.95\]\)\. For F1 score, the estimated gap difference is negative \(\-1\.99\), with a 95% confidence interval below zero \(\[\-2\.94, \-1\.02\]\)\.

The permutation test leads to the same conclusion\. The observedPractical\-vs\.\-Definitiongap difference is not unusual under random family assignment \(EM accuracy:p=0\.93p=0\.93; F1 score:p=0\.79p=0\.79\)\. In other words, the actual generator\-family grouping does not show an unusually large Practical\-specific effect\.

Overall, these results suggest that generator\-family models are stronger baselines in general, but they do not provide evidence for an additional family\-specific advantage on the model\-generatedPracticaltasks in the main multimodal setting\. Construction/evaluation overlap remains a reasonable caveat, but our analysis does not indicate that it is a major driver of the benchmark results through a robust Practical\-specific effect\. We do not claim that this fully rules out all forms of stylistic or distributional alignment; rather, it shows that such an effect is not supported by the present gap\-based analysis\.

## Appendix CAppendix for Related Works

### C\.1Related Works for Large language Models

Recent studies have demonstrated the remarkable capabilities of instruction\-following large language models \(LLMs\)OpenAI \([2023](https://arxiv.org/html/2607.26465#bib.bib100),[2025](https://arxiv.org/html/2607.26465#bib.bib81)\); Google \([2025](https://arxiv.org/html/2607.26465#bib.bib12)\); xAI \([2025](https://arxiv.org/html/2607.26465#bib.bib51)\); Jianget al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib120)\), showing strong zero\-shot and few\-shot performance across a wide range of natural language processing tasksBubecket al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib155)\); Chanet al\.\([2024b](https://arxiv.org/html/2607.26465#bib.bib128)\); Chenget al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib125)\); Wanget al\.\([2024a](https://arxiv.org/html/2607.26465#bib.bib134)\); Jiayanget al\.\([2024a](https://arxiv.org/html/2607.26465#bib.bib138)\); Shiet al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib139)\); Jiayanget al\.\([2024b](https://arxiv.org/html/2607.26465#bib.bib126)\); Chanet al\.\([2023a](https://arxiv.org/html/2607.26465#bib.bib135)\)\. Despite these impressive advances, existing studies have identified several reasoning challenges that remain difficult for current LLMs, including complex mathematical reasoningFriederet al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib121)\), theory of mind reasoningLinet al\.\([2024](https://arxiv.org/html/2607.26465#bib.bib133)\), uncertainty and confidence calibrationZonget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib149)\), retrieval\-augmented generationJiayanget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib144)\); Zhanget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib146)\); Liet al\.\([2025b](https://arxiv.org/html/2607.26465#bib.bib145)\), intention reasoningYanget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib150)\), analogical reasoningChenget al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib125)\), discourse relation classificationChanet al\.\([2023b](https://arxiv.org/html/2607.26465#bib.bib132)\), text\-to\-table generationDenget al\.\([2024](https://arxiv.org/html/2607.26465#bib.bib131),[2025](https://arxiv.org/html/2607.26465#bib.bib143)\), complex game scenariosYimet al\.\([2024b](https://arxiv.org/html/2607.26465#bib.bib156)\); Moet al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib148)\); Lianget al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib147)\), argument impact classificationChanet al\.\([2024a](https://arxiv.org/html/2607.26465#bib.bib137)\), and the associated ethical and privacy challengesLiet al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib116)\); Susnjak \([2022](https://arxiv.org/html/2607.26465#bib.bib123)\); Liet al\.\([2024a](https://arxiv.org/html/2607.26465#bib.bib129)\); Lukaset al\.\([2023](https://arxiv.org/html/2607.26465#bib.bib124)\); Liet al\.\([2025a](https://arxiv.org/html/2607.26465#bib.bib130)\)\. State\-of\-the\-art LLMs, including o4\-mini\(OpenAI,[2025](https://arxiv.org/html/2607.26465#bib.bib81)\), Gemini\-3\-Flash\(Google,[2025](https://arxiv.org/html/2607.26465#bib.bib12)\), Grok\-4\.1\-Fast\(xAI,[2025](https://arxiv.org/html/2607.26465#bib.bib51)\), Llama\-4 family \(Scout\-17B and Maverick\-17B\)\(Meta AI,[2025](https://arxiv.org/html/2607.26465#bib.bib52)\), and DeepSeekDeepSeek\-AIet al\.\([2025](https://arxiv.org/html/2607.26465#bib.bib153)\), have significantly improved reasoning performance through large\-scale pre\-training and reinforcement learning\. Nevertheless, existing evaluations have primarily focused on static reasoning tasks, isolated text understanding, or single\-step multimodal perception\. Relatively little attention has been devoted to evaluating whether these models can infer and continuously update human motivations as additional multimodal evidence unfolds throughout a narrative\. Since human motivation is an inherently latent and dynamic mental state, effective reasoning requires not only recognizing observable behaviors but also integrating accumulated visual and textual context to revise previous interpretations\. Our work complements existing LLM evaluation benchmarks by introducing a new dimension of social reasoning: multimodal sequential motivation reasoning\. Unlike previous benchmarks that emphasize static social commonsense, intention prediction, or theory\-of\-mind reasoning, MULTIVATIONBENCH evaluates whether models can consistently infer psychologically grounded motivations under evolving multimodal contexts\.

### C\.2Related Works for Multimodal Social and Theory\-of\-Mind Reasoning

Theory\-of\-mind \(ToM\) reasoning is a central testbed for evaluating whether models can infer others’ latent mental states, including beliefs, desires, intentions, emotions, and goals\. Prior benchmarks have substantially advanced this direction through information\-asymmetric conversations, cross\-cultural ToM, negotiation\-based interactions, and systematic social\-cognitive evaluations\(Kimet al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib28); Chanet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib98),[2024c](https://arxiv.org/html/2607.26465#bib.bib14); Chenet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib99)\)\. However, these works mainly focus on text\-only settings or structured interaction scenarios\.

Recent multimodal social\-reasoning benchmarks extend theory\-of\-mind evaluation beyond text\-only stories by requiring models to integrate visual evidence with social context\. MMToM\-QA evaluates multimodal theory\-of\-mind question answering over video and text, emphasizing goals, beliefs, emotions, and social relations\(Jinet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib15)\)\. V\-ALPHASOCIAL studies visual social commonsense reasoning and uses self\-reflective chain\-of\-thought generation to improve reasoning over social situations\(Linet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib69)\)\. MoMentS further introduces realistic short\-film scenarios for multimodal theory\-of\-mind evaluation\(Villa\-Cuevaet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib71)\), while Mind the Motions focuses on everyday body language and nonverbal cues as evidence for mental\-state inference\(Leeet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib60)\)\. These works show that visual cues are important for social cognition, but they primarily target goals, beliefs, social commonsense, body\-language cues, or general mental states\. They do not evaluate theory\-grounded motivation reasoning under accumulated story context, where a model must revise why a character acts as later multimodal evidence accumulates\.

### C\.3Related Works for Temporal Visual Understanding and Long\-Video Reasoning

Another line of work evaluates whether multimodal models can understand temporally extended visual inputs\. Benchmarks such as MVBench, EgoSchema, Video\-Bench, Video\-MME, LongVideoBench, and MLVU test temporal perception, long\-form video\-language understanding, interleaved video\-language QA, and multi\-task long\-video reasoning\(Liet al\.,[2024b](https://arxiv.org/html/2607.26465#bib.bib84); Mangalamet al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib95); Ninget al\.,[2023](https://arxiv.org/html/2607.26465#bib.bib97); Fuet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib96); Wuet al\.,[2024a](https://arxiv.org/html/2607.26465#bib.bib87); Zhouet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib89)\)\. Complementary model\-oriented work, including Video\-ChatGPT and TimeChat, studies video conversation and time\-sensitive long\-video understanding\(Maazet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib86); Renet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib85)\)\. Temporal Grounding Bridge further examines temporal extrapolation and grounding for multimodal large language models\(Wanget al\.,[2024b](https://arxiv.org/html/2607.26465#bib.bib88)\)\. Together, these studies show that temporal visual understanding is an active benchmark direction\. However, their main focus is usually what events happen, when they happen, how actions unfold, or where relevant temporal segments are located\. They generally do not ask why a character behaves in a particular way under a psychological motivation taxonomy\.

### C\.4Related Works for Visual Storytelling and Narrative Consistency

Visual storytelling research studies how models generate or maintain coherent narratives across images and videos\. StoryDALL\-E adapts pretrained text\-to\-image transformers for story continuation\(Maharanaet al\.,[2022](https://arxiv.org/html/2607.26465#bib.bib90)\), and ViSTA uses multimodal adapters to condition text\-to\-image diffusion on visual storytelling history\(Donget al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib91)\)\. StoryDiffusion targets long\-range consistency for image and video generation through consistent self\-attention\(Zhouet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib92)\), while Storynizor and CharaConsist focus on character\-consistent story generation and fine\-grained character consistency\(Maet al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib93); Wanget al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib94)\)\. MovieBench provides a hierarchical movie\-level dataset for long\-video generation, further reflecting the need for narrative structure and long\-range visual coherence\(Wuet al\.,[2024b](https://arxiv.org/html/2607.26465#bib.bib18)\)\. These works are relevant because they address long\-range stories, character continuity, and visual narrative coherence\. Their central objective, however, is generation or consistency maintenance, whereas our benchmark evaluates recognition and reasoning over existing visual narratives\.

### C\.5Position ofMultivationBench

Prior work therefore covers three adjacent capabilities: multimodal social reasoning, temporal video understanding, and coherent visual storytelling\.MultivationBenchconnects these directions but evaluates a different target capability\. It requires models to ground behavior in multimodal evidence, accumulate sequential narrative context, and assign psychologically grounded motivation labels from Maslow’s hierarchy and Reiss’s basic desires\. This setting asks not only what a character did or when an event occurred, but why the behavior is best explained by a particular motivation after the story context has accumulated\. In this sense,MultivationBenchcomplements existing benchmarks by isolating sequential motivation reasoning as a distinct form of multimodal social intelligence\.

## Appendix DData Construction Pipeline

### D\.1Detailed Data Pre\-processing and Prompts

To curate meaningful evaluation samples, we focused on visually grounded and psychologically salient character behaviors, rather than background\-only presence or trivial scene details\. We automated this filtering process using a multi\-model extraction pipeline designed to minimize single\-model bias\(Guet al\.,[2024](https://arxiv.org/html/2607.26465#bib.bib37)\)\. Two state\-of\-the\-art MLLMs, Grok\-4\.1 and Gemini\-3, independently scanned the images and story context to identify the main character and extract the majorCharacterandBehavior Chain—a sequence ofBehaviorsmapped to specific image indices with hypothesizedMotivations\.

Subsequently, a reasoning\-specialized model \(Grok\-4\.1\-fast\-reasoning\) acted as an adjudicator, reviewing the candidate outputs to select the most coherent chain that maintained strict temporal logic and character consistency\. This consensus\-based approach significantly reduces hallucination rates compared to single\-pass generation\(Huanget al\.,[2025](https://arxiv.org/html/2607.26465#bib.bib38)\)\. Finally, to further eliminate model hallucinations and ensure psychological validity, the automated outputs underwent a rigorous manual review by the authors\. We filtered samples based on three strict criteria:

- •Visual Grounding:Ensuring the described cue is clearly visible in the image frame\.
- •Psychological Salience:Verifying that the inferred motivation is relevant to the scene rather than trivial movement\.
- •Human Character Focus:Confirming that the chain consistently tracks the main character without drifting to background characters\.

#### Prompt Templates

We present the specific extractor and reviewer prompts used in this stage in Tables[10](https://arxiv.org/html/2607.26465#A5.T10)and[11](https://arxiv.org/html/2607.26465#A5.T11)\. The extractor template is designed to identify the main character, behavior, and motivation given a sequence of images and the full story context, outputting the structured Behavior Chain\. The reviewer prompt evaluates these generated outputs to select the optimal chain based on the criteria described above\.

### D\.2Details of Options Generation Protocol

Existing benchmarks mostly rely on manual construction, which is labor\-intensive and limits scalability\. To reduce financial and labor costs while ensuring theoretical rigor, an Automated Multi\-Stage framework was employed to formulate questions and generate options, as illustrated in Figure[2](https://arxiv.org/html/2607.26465#S3.F2)\. The detailed algorithm is provided in Algorithm[1](https://arxiv.org/html/2607.26465#alg1)\.

We focused our generation efforts on thePractical Motivationtasks \(Tasks 3 and 4\), since options for theDefinitiontasks \(Tasks 1 and 2\) correspond directly to standard theoretical definitions and do not require generation\. Specifically, we utilize a Motivation Analyst designed to generate practically grounded motivation options based on character behavior, story context, and visual cues\. The analyst operates in two modes corresponding to distinct psychological frameworks: Maslow’s Hierarchy of Needs and Reiss’s 16 Basic Desires, referring to the definitions shown in Tables[12](https://arxiv.org/html/2607.26465#A5.T12)and[13](https://arxiv.org/html/2607.26465#A5.T13)\.

Two distinct LLM\-based analysts \(Grok\-4\.1\-Fast and Gemini\-3\) independently generate sets of practical motivation options based on the verified behavior chains\. These models then participate in a cross\-model review mechanism, providing feedback on:

1. 1\.Logical soundness of the options\.
2. 2\.Correctness of the answer based on visual and narrative context\.
3. 3\.Strict alignment with Maslow’s and Reiss’s definitions to minimize hidden biases from a single model\.

The feedback is compiled by a Validator \(Grok\-4\.1\-Fast\-reasoning\) to select the highest\-quality candidates and form the final option sets\. The prompt templates for the Motivation Analyst, Reviewer, and Validator for both Maslow and Reiss tasks are presented in Tables[14](https://arxiv.org/html/2607.26465#A5.T14),[15](https://arxiv.org/html/2607.26465#A5.T15), and[16](https://arxiv.org/html/2607.26465#A5.T16)\.

### D\.3Details of the annotation process

In this section, we present the annotation instructions and templates used in our annotation pipeline\. The story context and the theory definitions in Table[12](https://arxiv.org/html/2607.26465#A5.T12)and Table[13](https://arxiv.org/html/2607.26465#A5.T13)are provided to guide the annotators\. The annotate framework is illustrated in Figure[6](https://arxiv.org/html/2607.26465#A5.F6)\.

### D\.4Details of the Four Task Types

In this section, we describe the four question types used inMulTivationBench\. For the Definition tasks, the question stem is: “Based on the visual information and the story provided, which \[theory label\(s\)\] is/are most strongly expressed or fulfilled by the behavior of \[character\]?” The options correspond to the standard theory definitions\. For the Practical Motivation tasks, the question stem is: “Based on the visual information and the story provided, what is/are the most likely motivation\(s\) behind \[character\]’s behavior?” The options are context\-specific motivations generated from the corresponding theory and the accumulated narrative context\.

### D\.5Detailed Statistics ofMultivationbench

In this section, we present the statistics ofMultivationBenchin Table[17](https://arxiv.org/html/2607.26465#A5.T17)\. In addition to the story\-level and behavior\-level statistics, we report the distribution of gold label\-set sizes in Table[9](https://arxiv.org/html/2607.26465#A5.T9)\.

## Appendix EAppendix for Experiment

### E\.1Prompts for the Four Task Types

In this section, we detail the specific prompts used for the four distinct question configurations in our benchmark\. Tables[19](https://arxiv.org/html/2607.26465#A5.T19),[18](https://arxiv.org/html/2607.26465#A5.T18),[21](https://arxiv.org/html/2607.26465#A5.T21), and[20](https://arxiv.org/html/2607.26465#A5.T20)provide concrete examples of the data instances passed to the Multimodal Large Language Model \(MLLM\) to elicit the final choice selections\.

These templates cover both the definition\-based tasks \(Tables[19](https://arxiv.org/html/2607.26465#A5.T19)and[21](https://arxiv.org/html/2607.26465#A5.T21)\), which require the model to map behaviors directly to standard theoretical definitions, and the Practical Motivation tasks \(Tables[18](https://arxiv.org/html/2607.26465#A5.T18)and[20](https://arxiv.org/html/2607.26465#A5.T20)\), which utilize context\-specific motivation options\. Crucially, to mitigate position bias—where models may preferentially select options based on their index rather than content—the order of the multiple\-choice options is randomized for every instance prior to inference\.

### E\.2Prompt Variations for Different Modalities

To evaluate the contribution of different modalities to the reasoning process, we adjust the prompt header and context block accordingly\. Table[22](https://arxiv.org/html/2607.26465#A5.T22)details how the introduction string and story context are modified for theMultimodal,Image\-Only, andText\-Onlysettings\.

### E\.3Metric Details

#### Example\-based F1

In our setting, Exact Match \(EM\) is a strict set\-level metric: it only counts a prediction as correct when the predicted option set exactly matches the gold set\. While useful, this criterion does not distinguish between a completely wrong prediction and a partially correct one\. We therefore also reportExample\-based F1, which captures instance\-level partial agreement between the predicted and gold option sets\.

LetGiG\_\{i\}denote the gold option set andG^i\\hat\{G\}\_\{i\}denote the predicted option set for test instanceii\. For each instance, we first compute precision and recall:

Pi=\|Gi∩G^i\|\|G^i\|,Ri=\|Gi∩G^i\|\|Gi\|\.P\_\{i\}=\\frac\{\|G\_\{i\}\\cap\\hat\{G\}\_\{i\}\|\}\{\|\\hat\{G\}\_\{i\}\|\},\\qquad R\_\{i\}=\\frac\{\|G\_\{i\}\\cap\\hat\{G\}\_\{i\}\|\}\{\|G\_\{i\}\|\}\.The instance\-level F1 score is then

F​1i=2​Pi​RiPi\+Ri=2​\|Gi∩G^i\|\|Gi\|\+\|G^i\|\.F1\_\{i\}=\\frac\{2P\_\{i\}R\_\{i\}\}\{P\_\{i\}\+R\_\{i\}\}=\\frac\{2\|G\_\{i\}\\cap\\hat\{G\}\_\{i\}\|\}\{\|G\_\{i\}\|\+\|\\hat\{G\}\_\{i\}\|\}\.The final example\-based F1 is obtained by averaging over allNNtest instances:

Example\-based F1=1N​∑i=1NF​1i\.\\text\{Example\-based F1\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}F1\_\{i\}\.
This metric gives partial credit when the predicted and gold option sets overlap, while still penalizing both false positives and false negatives\. In this way, it complements EM by distinguishing near\-miss predictions from fully incorrect ones\.

### E\.4Human Performance Evaluation

To measure human performance onMulTivationBench, we employ three graduate\-student annotators to complete the same benchmark evaluation tasks as the models\. Each benchmark instance is evaluated under four task settings: Maslow Definition, Maslow Practical Motivation, Reiss Definition, and Reiss Practical Motivation\. The questions and instructions are identical to the templates described in Appendix[D\.3](https://arxiv.org/html/2607.26465#A4.SS3)We compute the final human predictions by majority vote over the three annotators’ responses\. This results in 70\.6% EM and 81\.5% F1 on Maslow Definition, 78\.6% EM and 87\.6% F1 on Maslow Practical Motivation, 60\.7% EM and 74\.2% F1 on Reiss Definition, and 63\.9% EM and 72\.9% F1 on Reiss Practical Motivation\.

### E\.5Common\-Subset Comparability Check

Phi\-3\.5\-Vision\-instruct cannot process the Long\-story multimodal and image\-only subsets due to image context\-window limitations\. As a result, its Overall score in Table[1](https://arxiv.org/html/2607.26465#S4.T1)is computed only over the successfully processed subsets and is less directly comparable to models evaluated on all story lengths\.

To check whether this missing subset changes the relative model comparison, we recompute normal\-mode model performance on the exact intersection of task instances shared by all retained models\. The common subset contains 14,180 task instances\. As shown in Table[6](https://arxiv.org/html/2607.26465#A5.T6), all retained models preserve the same EM and F1 ranks under the common\-subset evaluation\. Phi\-3\.5\-Vision\-instruct obtains 29\.10 EM / 36\.21 F1 both before and after common\-subset filtering, indicating that the missing Long\-story subset does not change its relative position under this check\. We therefore keep the original full evaluation in the main paper while explicitly noting the context\-window caveat\.

Table 6:Common\-subset comparability check\. Each cell reports EM / F1\. “Full” reports the original retained\-instance score, while “Common” recomputes scores on the exact intersection of task instances shared by all retained models\.
### E\.6Same\-Behavior Context Ablation

To more directly test whether model predictions for the same behavior change when accumulated context is altered, we conduct a same\-behavior context\-ablation analysis\. Unlike the standard evaluation setting, where each behavior is evaluated with the accumulated multimodal context available up to its image index, this analysis keeps the target behavior fixed while removing earlier context units from the input\.

Specifically, we sample 150 final behavior points fromMultivationBench\. For each sampled behavior, we preserve the same target character, target behavior, current image, and answer options, but remove one, two, or three earlier text\-image units from the accumulated context\. We then re\-evaluate model predictions under these reduced\-context settings\. We report EM and example\-based F1 averaged over the Maslow and Reiss Practical Motivation tasks\.

Table 7:Same\-behavior context\-ablation results\. Each cell in the Full, Drop1, Drop2, and Drop3 columns reports EM / F1, averaged over the Maslow and Reiss Practical Motivation tasks\. Full uses the complete accumulated context for the sampled behavior\. Drop1, Drop2, and Drop3 remove one, two, and three earlier text\-image context units, respectively, while keeping the target character, target behavior, current image, and answer options fixed\.Δ\\DeltaF1 denotes Drop3 F1 minus Full F1\.As shown in Table[7](https://arxiv.org/html/2607.26465#A5.T7), removing earlier context does not produce a uniform degradation across models\. Average F1 changes from 44\.82 under full context to 44\.38, 44\.84, and 43\.99 after dropping one, two, and three earlier text\-image units, respectively\. This suggests that current MLLMs do not consistently exploit distant earlier context when predicting the motivation of a fixed later behavior\. Some models show stronger context sensitivity, such as Grok\-4\.1\-Fast, whose F1 decreases by 3\.44 points after dropping three earlier units\. Overall, this analysis indicates that motivation revision under changing accumulated context remains limited and model\-dependent\.

### E\.7Softer Story\-Level Metrics

The full\-story consistency metric in Table[2](https://arxiv.org/html/2607.26465#S4.T2)is intentionally strict: a story is counted as correct only if all questions in that story are answered exactly correctly\. This provides a stress\-test of whether models can maintain fully correct motivation reasoning across an entire narrative\. However, because this all\-correct criterion can compress differences among models, we additionally report softer story\-level metrics\.

For each model and story, we first average EM and example\-based F1 over all questions within that story\. We then macro\-average these per\-story scores across all stories\. We refer to these metrics as macro story EM and macro story F1\. Unlike full\-story consistency, these metrics preserve partial story\-level performance while still evaluating models at the story level rather than only at the individual\-question level\.

Table 8:Softer story\-level metrics \(%\)\. Full is the all\-correct story consistency score, where a story is counted as correct only if all questions in that story are answered exactly correctly\. Story EM and Story F1 first average EM/F1 over questions within each story and then macro\-average across stories\.As shown in Table[8](https://arxiv.org/html/2607.26465#A5.T8), the all\-correct full\-story consistency score ranges only from 0\.00% to 0\.80%, confirming that exact story\-level motivation reasoning remains extremely challenging for current MLLMs\. In contrast, macro story EM ranges from 8\.65% to 38\.51%, and macro story F1 ranges from 36\.21% to 54\.80%, revealing clearer differences among models\. These softer metrics therefore complement the strict consistency score: full\-story consistency measures whether a model can answer an entire narrative perfectly, while macro story EM/F1 provide a more graded view of story\-level performance\.

Table 9:Gold label\-set cardinality distribution\. Ckkdenotes behavior points with exactlykkgold labels\. Multi\. reports the number and percentage of behavior points with more than one gold label\.Algorithm 1Automated Multi\-Stage Option Generation with Cross\-Review1:Behavior Chain

ℬ=\{C​h​a​r,B​e​h​a​v​i​o​r​s​\[\],I​m​g​I​n​d​i​c​e​s​\[\]\}\\mathcal\{B\}=\\\{Char,Behaviors\[\],ImgIndices\[\]\\\}, Story

𝒮\\mathcal\{S\}
2:Set of Validated Options

𝒪f​i​n​a​l\\mathcal\{O\}\_\{final\}
3:

𝒪f​i​n​a​l←∅\\mathcal\{O\}\_\{final\}\\leftarrow\\emptyset
4:for

k←1k\\leftarrow 1to length\(

B​e​h​a​v​i​o​r​sBehaviors\)do

5:

bk←B​e​h​a​v​i​o​r​s​\[k\]b\_\{k\}\\leftarrow Behaviors\[k\]
6:

i​d​xk←I​m​g​I​n​d​i​c​e​s​\[k\]idx\_\{k\}\\leftarrow ImgIndices\[k\]
7:Step 1: Context Accumulation

8:

𝒞k←\\mathcal\{C\}\_\{k\}\\leftarrowGetContext\(

𝒮\\mathcal\{S\},

0​…​i​d​xk0\\dots idx\_\{k\}\)⊳\\trianglerightAccumulate text & images up to current event

9:

V​a​l​i​d←FalseValid\\leftarrow\\text\{False\}
10:whilenot

V​a​l​i​dValiddo

11:Step 2: Independent Generation

12:

O​pX←Op\_\{X\}\\leftarrowModelX\.Generate\(

C​h​a​r,bk,𝒞kChar,b\_\{k\},\\mathcal\{C\}\_\{k\}\)

13:

O​pY←Op\_\{Y\}\\leftarrowModelY\.Generate\(

C​h​a​r,bk,𝒞kChar,b\_\{k\},\\mathcal\{C\}\_\{k\}\)

14:Step 3: Cross\-Model Review

15:

F​e​e​dX→Y←Feed\_\{X\\to Y\}\\leftarrowModelX\.Review\(

O​pY,𝒞kOp\_\{Y\},\\mathcal\{C\}\_\{k\}\)

16:

F​e​e​dY→X←Feed\_\{Y\\to X\}\\leftarrowModelY\.Review\(

O​pX,𝒞kOp\_\{X\},\\mathcal\{C\}\_\{k\}\)

17:Step 4: Validator Adjudication

18:

S​e​l​e​c​t​i​o​n←Selection\\leftarrowAgentZ\.Validate\(

O​pX,O​pY,F​e​e​dX→Y,F​e​e​dY→XOp\_\{X\},Op\_\{Y\},Feed\_\{X\\to Y\},Feed\_\{Y\\to X\}\)

19:if

S​e​l​e​c​t​i​o​n≠NULLSelection\\neq\\text\{NULL\}then

20:

𝒪f​i​n​a​l\\mathcal\{O\}\_\{final\}\.append\(

S​e​l​e​c​t​i​o​nSelection\)

21:

V​a​l​i​d←TrueValid\\leftarrow\\text\{True\}
22:else

23:continue⊳\\trianglerightValidation failed; Re\-run generation

24:endif

25:endwhile

26:endfor

27:return

𝒪f​i​n​a​l\\mathcal\{O\}\_\{final\}

Table 10:The Extractor Prompt used in data preprocessing to identify main characters and validate behaviors\.Table 11:The Reviewer Prompt used to validate the accuracy of extracted behavior chains against visual and textual evidence\.Table 12:Maslow’s 8\-Level Hierarchy of NeedsTable 13:Reiss’s 16 Basic DesiresTable 14:The generic Motivation Analyst Prompt\. The \{definitions\} placeholder is dynamically populated with either Maslow’s Hierarchy or Reiss’s Basic Desires during generation\.Table 15:The Validator Prompt used to rigorously audit generated questions for hallucinations, formatting errors, and logical consistency\.Table 16:The Validator Prompt used to compare two candidate option sets and select the one with higher linguistic variety and psychological accuracy\.![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/anno1.png)\(a\)The initial step of the annotation framework\.
![Refer to caption](https://arxiv.org/html/2607.26465v1/latex/anno2.png)\(b\)The validation step of the annotation framework\.

Figure 6:Overview of the framework provided to the annotators\. Panel \(a\) illustrates the interface for initial annotations, and panel \(b\) shows the review and validation interface\.Table 17:Data statistics of our proposedMultivationBench\.\(Behav\.: average behaviors per story, Img\.: average image count per story, Ctx Tok\.: average context tokens, Total Behav\.: total annotated behaviors, and Total Qs: total questions\.\)Table 18:The raw prompt input for the Maslow Practical task, containing 8 context\-specific options generated based on Maslow’s hierarchy\.Table 19:The raw prompt input for the Maslow Definition task, including the story context, exact definitions, and formatting constraints\.Table 20:The raw prompt input for the Reiss Practical task, containing 16 options generated based on Reiss’s basic desires theory\.Table 21:The raw prompt input for the Reiss Definition task, asking the model to map behaviors directly to the 16 standard theoretical definitions\.Table 22:Prompt modifications for ablation studies\. The introduction and context visibility are adjusted to restrict the model’s input source for Image\-Only and Text\-Only evaluations\.

Similar Articles

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

arXiv cs.CL

This paper introduces CrossMath, a controlled multimodal reasoning benchmark that reveals a critical limitation in current vision-language models: they perform reasoning primarily in textual space rather than genuine vision-grounded reasoning, with visual input often degrading performance compared to text-only baselines. The authors propose fine-tuning approaches to mitigate this modality gap and improve multimodal reasoning capabilities.