M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

arXiv cs.CL Papers

Summary

This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.

arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues.To address these limitations, we introduce M$^3$R-Bench, a unified and evidence-grounded benchmark containing 1,000 image--text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M$^3$R-Bench provides joint annotations for metaphor occurrence, Target--Source mapping, sentiment, and stage-wise explanations following ``evidence identification--mapping establishment--sentiment inference.''Evaluations on M$^3$R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target--Source mappings, exposing a cross-modal evidence--mapping mismatch. To address this mismatch, we propose M$^3$R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M$^3$R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at https://github.com/hongshi4/M3R-Bench.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:52 AM

# A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Source: [https://arxiv.org/html/2608.05817](https://arxiv.org/html/2608.05817)
## M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench: A Unified Benchmark for Evidence\-Grounded Multimodal Metaphor Understanding

Hong Jiang1\\equalcontrib, Junnan Zhu2\\equalcontrib, Jingwang Huang1\\equalcontrib, Xiao Sun1, Yuming Yang1, Jiang Zhong1\\corresponding, Ruirui Chen3, Jingman Shi4, Hao Wu4, Nayu Liu5, Xinyi Jiang6, Kaiwen Wei1\\corresponding

###### Abstract

Metaphor enables the understanding of abstract concepts through cross\-domain mappings while conveying affective attitudes\. In multimodal scenarios, visual and textual information jointly construct Target–Source mappings, requiring both conceptual understanding and cross\-modal reasoning\. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence\-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual cues\. To address these limitations, we introduceM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench, a unified and evidence\-grounded benchmark containing 1,000 image–text instances with human\-verified annotations\. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench provides joint annotations for metaphor occurrence, Target–Source mapping, sentiment, and stage\-wise explanations following “evidence identification–mapping establishment–sentiment inference\.” Evaluations onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target–Source mappings, exposing a cross\-modal evidence–mapping mismatch\. To address this mismatch, we proposeM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner, which combines curriculum\-based reasoning supervision with task\-aware reinforcement learning to align model reasoning with metaphor interpretation\. Experiments show that, with only an 8B\-parameter backbone,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner outperforms larger proprietary MLLMs across four unified\-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT\-5\.5 by 28\.45 and 30\.11 points, respectively, while surpassing Claude\-Sonnet\-4\.6 by 8\.00 points in mean rubric score\. The dataset and code are available athttps://github\.com/hongshi4/M3R\-Bench\. ‘

## Introduction

![Refer to caption](https://arxiv.org/html/2608.05817v1/x1.png)Figure 1:Motivation of this work\. Existing research predicts metaphor and sentiment but misses evidence\-grounded Target–Source mappings\. In contrast,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench jointly evaluates metaphor identification, mapping, and sentiment with evidence\-grounded explanations\.Metaphor enables the understanding of abstract concepts by mapping structures from a source domain onto a target domain\(Lakoff and Johnson[1980](https://arxiv.org/html/2608.05817#bib.bib1)\)\. In real\-world scenarios such as advertising and social media, such mappings are often constructed through visual and textual modalities, forming multimodal metaphors\(Forceville and Urios\-Aparisi[2009](https://arxiv.org/html/2608.05817#bib.bib2)\)\. Understanding these expressions is essential for recovering conceptual mappings, affective attitudes, and communicative intentions\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3); Xuet al\.[2022](https://arxiv.org/html/2608.05817#bib.bib4)\)\.

Table 1:Comparison ofM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench with existing metaphor benchmarks\.Recent advances in multimodal large language models \(MLLMs\) have significantly improved visual perception and multimodal reasoning capabilities, providing new opportunities to study metaphor understanding beyond surface\-level classification toward interpretable reasoning\(Kunduet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib10); Zhenget al\.[2026](https://arxiv.org/html/2608.05817#bib.bib11)\)\. Existing multimodal metaphor benchmarks have provided valuable resources for this research direction\. For example, MultiMET\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3)\), MultiCMET\(Zhanget al\.[2023](https://arxiv.org/html/2608.05817#bib.bib5)\), and MultiMM\(Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\)introduce annotations for metaphor occurrence, Target–Source relations, sentiment, and related attributes\.

However, existing benchmarks typically evaluate metaphor identification, Target–Source mapping, and sentiment inference as separate tasks, rather than modeling the complete understanding process\. This fragmented setting cannot reveal whether correct predictions arise from valid cross\-modal reasoning or superficial correlations\. Moreover, most benchmarks lack evidence\-grounded explanations for verifying how conceptual mappings are supported by visual and textual cues\. As shown in Figure[1](https://arxiv.org/html/2608.05817#Sx1.F1), an MLLM may correctly identify the metaphor and its positive sentiment, yet replace the grounded Target–Source relation with theme\-level concepts such as low\-carbon development and environmental protection\. Although recent studies explore chain\-of\-thought prompting and explanation generation\(Xuet al\.[2024](https://arxiv.org/html/2608.05817#bib.bib25); Tianet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib13)\), these explanations are mainly auxiliary signals or unconstrained outputs rather than standardized evidence\-grounded evaluations\.

To address these limitations, we introduceMultimodalMetaphorUnderstandingBenchmark \(M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\), a unified evidence\-grounded benchmark for multimodal metaphor understanding\.M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench contains 1,000 image–text instances collected from existing resources, re\-annotated under a theory\-guided framework, and verified by human annotators\(Lakoff and Johnson[1980](https://arxiv.org/html/2608.05817#bib.bib1); Wilks[1975](https://arxiv.org/html/2608.05817#bib.bib19)\)\. Each instance provides annotations for metaphor occurrence, Target–Source mapping, sentiment, and stage\-wise explanations following “evidence identification–mapping establishment–sentiment inference\.” As shown in Table[1](https://arxiv.org/html/2608.05817#Sx1.T1),M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench jointly evaluates these dimensions with evidence\-grounded explanations, enabling holistic evaluation and fine\-grained diagnosis of cross\-modal metaphor reasoning\.

We conduct extensive evaluations of 13 baselines onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench, including traditional vision\-language models, task\-specific metaphor approaches, and representative open\- and closed\-source MLLMs\. The results reveal that existing methods remain unreliable for unified evidence\-grounded multimodal metaphor understanding\. In particular, MLLMs often overlook critical visual evidence, rely on superficial textual cues, and generate Target–Source mappings with inappropriate conceptual granularity\. On the rubric\-based evaluation, GPT\-5\.5 scores only 29\.66 on Visual Evidence and 43\.95 on Mapping Correctness\. These failures reveal a cross\-modal evidence–mapping mismatch, where models may produce plausible predictions without establishing the intended metaphorical relations\.

To address this cross\-modal evidence–mapping mismatch, we proposeM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner, which combines curriculum\-based reasoning Supervised Fine\-Tuning \(SFT\) with task\-aware reinforcement learning to progressively align evidence identification, Target–Source mapping, and sentiment inference\. Using only an 8B\-parameter backbone,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner achieves state\-of\-the\-art performance across four unified\-task metrics and three rubric\-based explanation metrics, outperforming larger general\-purpose MLLMs\. These results validate the effectiveness of evidence\-grounded reasoning for multimodal metaphor understanding\. In summary, the contributions of this work are as follows:

1. 1\.We identify fragmented evaluation and limited interpretability in existing multimodal metaphor benchmarks, and address these limitations by introducingM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench, a unified evidence\-grounded benchmark with joint annotations for metaphor identification, Target–Source mapping, sentiment, and reasoning explanations\.
2. 2\.We conduct systematic evaluations with 13 baselines onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench and reveal that existing methods often overlook critical visual evidence, rely on superficial textual cues, and produce inaccurate Target–Source mappings\.
3. 3\.We proposeM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner, a curriculum\-based reasoning framework with task\-aware reinforcement learning\. Experiments reveal thatM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner effectively alleviates cross\-modal evidence–mapping mismatch through evidence\-grounded reasoning\.

## Related Work

#### Multimodal Metaphor Benchmarks\.

Multimodal metaphor benchmarks have expanded from metaphor detection to richer annotations of conceptual domains, Target–Source relations, sentiment, and intent\. MultiMET\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3)\), MET\-Meme\(Xuet al\.[2022](https://arxiv.org/html/2608.05817#bib.bib4)\), and MultiCMET\(Zhanget al\.[2023](https://arxiv.org/html/2608.05817#bib.bib5)\)establish representative image–text settings, while more recent datasets further investigate cultural variation with MultiMM\(Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\), fine\-grained emotion with EmoMeta\(Luet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib7)\), explicit conceptual mappings with CM3D\(Zhanget al\.[2025b](https://arxiv.org/html/2608.05817#bib.bib8)\), and unified evaluation of multimodal large language models with M3UCD\(Zhenget al\.[2026](https://arxiv.org/html/2608.05817#bib.bib11)\)\. Visual\-metaphor resources such as MetaCLUE\(Akula and others[2023](https://arxiv.org/html/2608.05817#bib.bib9)\)and ImageMet\(Kunduet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib10)\)additionally introduce localization, understanding, generation, captioning, and question\-answering tasks\. Despite broader task coverage, existing benchmarks still treat metaphor understanding as isolated subtasks and lack evidence\-grounded explanations\. As a result, they cannot determine whether models recover metaphorical mappings from cross\-modal evidence or rely on superficial correlations\.

#### Metaphor Understanding Methods\.

Early approaches primarily relied on image–text feature fusion and task\-specific semantic knowledge, including conceptual\-domain and sentiment representations\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3); Xuet al\.[2022](https://arxiv.org/html/2608.05817#bib.bib4); Zhanget al\.[2023](https://arxiv.org/html/2608.05817#bib.bib5); Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\)\. Recent studies increasingly exploit multimodal large language models for explicit reasoning: C4MMD\(Xuet al\.[2024](https://arxiv.org/html/2608.05817#bib.bib25)\)transfers multimodal knowledge through staged chain\-of\-thought reasoning, CoC\(Zhanget al\.[2025a](https://arxiv.org/html/2608.05817#bib.bib12)\)elicits candidate Target–Source entities and their associations, CPMMIM\(Zhanget al\.[2025b](https://arxiv.org/html/2608.05817#bib.bib8)\)combines chain\-of\-thought prompting with hierarchical optimization, and ImaRA\(Tianet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib13)\)reasons through imaginative frames, domain incongruity, and cross\-domain attribute similarity\. However, existing methods often focus on partial subtasks or optimize metaphor identification, mapping, and sentiment independently\. Their reasoning signals are usually introduced through prompts or free\-form outputs rather than a unified process\. To address this limitation,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner introduces curriculum\-based reasoning supervision and task\-aware rewards to align evidence\-grounded metaphor reasoning with final predictions\.

![Refer to caption](https://arxiv.org/html/2608.05817v1/x2.png)Figure 2:Data distribution and statistics ofM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\.![Refer to caption](https://arxiv.org/html/2608.05817v1/x3.png)Figure 3:Construction ofM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench through data curation, consensus\-based annotation, human adjudication, and evidence\-grounded explanation generation\.

## M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench

We introduceM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench, a unified benchmark for interpretable multimodal metaphor understanding\. To the best of our knowledge, it is the first benchmark to jointly evaluate multimodal metaphor identification, Target–Source conceptual mapping, and metaphor\-aware sentiment inference under a unified full\-task setting, while providing stage\-wise explanations that capture the reasoning process from cross\-modal evidence to the final outputs\.

#### Task Definition\.

Given an image–text samplex=\(I,T\)x=\(I,T\), whereIIandTTdenote the image and text, respectively, a model jointly predicts:

y=\(m,t,s,e\),y=\(m,t,s,e\),\(1\)wherem∈\{Yes,No\}m\\in\\\{\\text\{Yes\},\\text\{No\}\\\}indicates whether the sample contains a multimodal metaphor;ttandssdenote the most central Target and Source in the metaphorical mapping, respectively; ande∈\{Positive,Neutral,Negative\}e\\in\\\{\\text\{Positive\},\\text\{Neutral\},\\text\{Negative\}\\\}represents the overall sentiment conveyed by the image–text pair\. For non\-metaphorical samples,t=s=Nonet=s=\\texttt\{None\}\.

#### Data Sources\.

Following MultiMET\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3)\), MultiCMET\(Zhanget al\.[2023](https://arxiv.org/html/2608.05817#bib.bib5)\), MultiMM\(Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\), and EmoMeta\(Luet al\.[2025](https://arxiv.org/html/2608.05817#bib.bib7)\), our candidate image–text samples are drawn from public English Twitter and Facebook posts retrieved using the hashtags\#metaphorand\#metaphorical; Chinese commercial and public\-service advertisements collected through Baidu and Bing keyword searches; and public advertising corpora, including Chinese advertisements from the 2021 iFlytek Advertising Image Classification Competition and English product and public\-service advertisements from Ye et al\.\(Yeet al\.[2021](https://arxiv.org/html/2608.05817#bib.bib21)\)\. Because these channels may contain overlapping content, we perform cross\-source deduplication and remove instances with corrupted images or incomplete image–text information\. Since the original resources differ in task definitions, label spaces, and annotation granularity, we do not inherit their annotations; instead, all retained instances are re\-annotated under a unified framework\. The detailed data sources and corresponding URLs are listed in Appendix\.

#### Label Annotation Scheme\.

To balance annotation efficiency and reliability, we adopt an MLLM\-led, human\-assisted annotation protocol\. Five representative MLLMs, including Grok\-4\.1\(xAI[2025](https://arxiv.org/html/2608.05817#bib.bib32)\), Kimi\-K2\.5\(Kimi Team and others[2026](https://arxiv.org/html/2608.05817#bib.bib28)\), GPT\-5\.5\(OpenAI[2026](https://arxiv.org/html/2608.05817#bib.bib30)\), Claude\-Opus\-4\.7\(Anthropic[2026a](https://arxiv.org/html/2608.05817#bib.bib33)\), and Gemini\-3\.1\-Pro\(Google DeepMind[2026](https://arxiv.org/html/2608.05817#bib.bib34)\), independently annotate metaphor occurrence, core Target–Source mappings, and sentiment\. Instances without consensus among at least three models are re\-annotated and adjudicated by annotators\. The process is guided by Conceptual Metaphor Theory \(CMT\)\(Lakoff and Johnson[1980](https://arxiv.org/html/2608.05817#bib.bib1)\)and Selectional Preference Violation \(SPV\)\(Wilks[1975](https://arxiv.org/html/2608.05817#bib.bib19)\)\. Three trained postgraduate annotators with computer science or psychology backgrounds independently label all instances, achieving a Fleiss’κ\\kappaof 0\.76\. Remaining disagreements are resolved by a senior annotator through re\-examination of multimodal evidence and rationales under CMT and SPV principles\. Annotation prompts, annotator backgrounds, and annotation training procedures are provided in Appendix\.

#### Annotation of Evidence\-grounded Chains\.

Given the gold\-standard labels, we construct a three\-stage evidence\-grounded chain for each test instance\. The first stage identifies whether an image–text pair contains a multimodal metaphor\. The second extracts the central Target–Source mapping and attribute projection, while the third integrates visual evidence, textual cues, and metaphorical meaning to infer the overall sentiment\. Each stage is automatically verified against the corresponding annotations\. For open\-ended Target and Source predictions, semantic consistency is measured by BERTScore F1\(Zhanget al\.[2020](https://arxiv.org/html/2608.05817#bib.bib20)\), with a threshold of 0\.78 calibrated on 200 additional metaphorical instances to achieve the highest agreement with human judgments \(0\.91\)\. Explanations passing automatic verification are further reviewed by three human annotators and minimally revised to ensure evidence grounding and avoid unsupported inferences\. Annotation details are provided in Appendix\.

#### Statistics ofM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\.

As shown in Figure[2](https://arxiv.org/html/2608.05817#Sx2.F2),M3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench contains 1,000 image–text instances and is nearly balanced with respect to metaphor occurrence, comprising 510 metaphorical and 490 non\-metaphorical samples\. The sentiment distribution includes 495 positive, 206 neutral, and 299 negative instances, while the data are drawn from social media \(364\), commercial advertisements \(222\), and public\-service advertisements \(414\)\. Each instance is annotated with metaphor occurrence, the core Target–Source mapping, overall sentiment, Visual Evidence, and Sentiment Justification\. The 510 metaphorical instances additionally provide Mapping Correctness explanations\.

Table 2:Experiment results onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. MD/SC use Accuracy and Macro\-F1; TP/SP use BERTScore Precision and F1\. Rubric is the mean of VE, MC, and SJ\. “–” denotes unsupported outputs; the best result within each model type isbolded\.![Refer to caption](https://arxiv.org/html/2608.05817v1/x4.png)Figure 4:Overview ofM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner, comprising curriculum\-based evidence–mapping reasoning supervision and task\-aware reinforcement learning\.

## Performance onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench

### Evaluation Setting

#### Evaluation Metrics\.

The unified prediction components are Metaphor Detection \(MD\), Target Prediction \(TP\), Source Prediction \(SP\), and Sentiment Classification \(SC\)\. Following prior benchmarks\(Zhanget al\.[2021](https://arxiv.org/html/2608.05817#bib.bib3),[2023](https://arxiv.org/html/2608.05817#bib.bib5); Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\), we report Accuracy and Macro\-F1 for MD and SC, and BERTScore Precision and F1\(Zhanget al\.[2020](https://arxiv.org/html/2608.05817#bib.bib20)\)for the open\-ended TP and SP components\. Explanation quality is assessed by Visual Evidence \(VE\), Mapping Correctness \(MC\), and Sentiment Justification \(SJ\)\. Gemini\-3\.1\-Pro scores each explanation against its gold evidence\-grounded rationale in\[0,1\]\[0,1\]\. Each dimension is masked when its associated prediction is incorrect\. Scores are scaled to\[0,100\]\[0,100\], and*Mean*averages VE, MC, and SJ\.

#### Evaluated Models\.

We evaluate representative models from four categories: general\-purpose vision\-language models \(CLIP\(Radford and others[2021](https://arxiv.org/html/2608.05817#bib.bib22)\), ViT\(Dosovitskiy and others[2021](https://arxiv.org/html/2608.05817#bib.bib23)\), and MAE\(Heet al\.[2022](https://arxiv.org/html/2608.05817#bib.bib24)\)\), task\-specific metaphor models \(C4MMD\(Xuet al\.[2024](https://arxiv.org/html/2608.05817#bib.bib25)\)and SEMD\(Yanget al\.[2025](https://arxiv.org/html/2608.05817#bib.bib6)\)\), open\-weight MLLMs \(Qwen3\-VL\-8B\-Instruct, Qwen3\-VL\-32B\-Instruct\(Bai and others[2025](https://arxiv.org/html/2608.05817#bib.bib26)\), InternVL3\.5\-8B\-HF, InternVL3\.5\-14B\-HF\(Wang and others[2025](https://arxiv.org/html/2608.05817#bib.bib27)\), and Kimi\-K2\.5\(Kimi Team and others[2026](https://arxiv.org/html/2608.05817#bib.bib28)\)\), and proprietary MLLMs \(GPT\-4o\(OpenAI[2024](https://arxiv.org/html/2608.05817#bib.bib29)\), GPT\-5\.5\(OpenAI[2026](https://arxiv.org/html/2608.05817#bib.bib30)\), and Claude\-Sonnet\-4\.6\(Anthropic[2026b](https://arxiv.org/html/2608.05817#bib.bib31)\)\)\.

### Experiment Results

We investigate the following two research questions\.

#### RQ1: How do existing models perform on unified evidence\-grounded multimodal metaphor understanding?

Table[2](https://arxiv.org/html/2608.05817#Sx3.T2)reports the performance of evaluated models onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. We find: \(1\) Existing non\-MLLM methods remain limited in unified metaphor understanding\. CLIP achieves only 54\.81% Accuracy and 52\.99% Macro\-F1 for metaphor identification, while SEMD improves to 59\.64% Accuracy and 57\.57% Macro\-F1 but does not support open\-ended Target–Source prediction\. \(2\) MLLMs achieve stronger but uneven performance across tasks\. Qwen3\-VL\-32B\-Instruct obtains the best metaphor identification Accuracy \(63\.59%\), GPT\-4o achieves the best Target prediction \(72\.39% BERTScore F1\), and GPT\-5\.5 achieves the highest sentiment Accuracy \(75\.70%\)\. However, no model consistently performs well across metaphor identification, conceptual mapping, and sentiment inference, revealing the challenge of unified evidence\-grounded metaphor understanding\.

#### RQ2: What are the primary sources of error in existing MLLMs?

For error analysis, we randomly sample 100 GPT\-4o errors, given its strongest overall Target–Source mapping performance among baselines\. We find that*Textual Shortcut Bias*is the dominant failure mode \(43%\), where models rely on slogans, exaggerations, or promotional language while ignoring visual evidence\.*Target–Source Granularity Mismatch*accounts for 26%, with models predicting abstract themes rather than the entities involved in metaphorical mappings\.*Visual Grounding Failure*contributes another 21%, especially for metaphors involving entity substitution, morphological fusion, spatial embedding, or symbolic transformation\. The remaining errors include sentiment misclassification \(8%\) and hallucinated reasoning \(2%\)\. Overall, the first three categories account for 90% of failures, revealing that insufficient evidence grounding and conceptual alignment remain primary challenges for current MLLMs\.

## M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner

Motivated by the findings above, we proposeM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner for multimodal metaphor understanding through curriculum\-based reasoning supervision and task\-aware reinforcement learning\. As illustrated in Figure[4](https://arxiv.org/html/2608.05817#Sx3.F4),M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner comprises two stages: curriculum\-based evidence–mapping reasoning and task\-aware reinforcement learning\.

### Curriculum\-Based Evidence–Mapping Reasoning

#### Reasoning Supervision Corpus Construction\.

We construct the reasoning supervision corpus following a procedure similar to that used forM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. Unlike benchmark construction, we retain only high\-confidence training samples for which at least three MLLMs agree on metaphor occurrence, the core Target–Source mapping, and overall sentiment\. The generated evidence\-grounded explanations are verified against the agreed labels, and samples with malformed outputs, unsupported inferences, or label–explanation inconsistencies are discarded\. This process yields 9,000 image–text instances with stage\-wise evidence\-grounded supervision, which are organized into three curriculum corpora of increasing complexity\.𝒞\(1\)\\mathcal\{C\}^\{\(1\)\}supervises metaphor identification and its evidence\-grounded justification\.𝒞\(2\)\\mathcal\{C\}^\{\(2\)\}additionally supervises the core Target–Source mapping and the corresponding attribute projection, while𝒞\(3\)\\mathcal\{C\}^\{\(3\)\}further incorporates overall sentiment and its justification based on the image–text evidence and metaphorical mapping\. Each corpus contains its own set of image–text instances\. For non\-metaphorical samples, Target and Source are set toNone, and the mapping explanation is omitted\. Each target places the evidence\-grounded reasoning within<think\>tags and the structured prediction within<label\>tags\.

#### Curriculum\-Based Evidence–Mapping Reason SFT\.

We adopt Qwen3\-VL\-8B\-Instruct\(Bai and others[2025](https://arxiv.org/html/2608.05817#bib.bib26)\)as the base model and perform supervised fine\-tuning sequentially on𝒞\(1\)→𝒞\(2\)→𝒞\(3\)\\mathcal\{C\}^\{\(1\)\}\\rightarrow\\mathcal\{C\}^\{\(2\)\}\\rightarrow\\mathcal\{C\}^\{\(3\)\}\. This curriculum progressively expands the prediction space and evidence\-grounded reasoning requirements, allowing each stage to build on the capabilities acquired previously\.

Across all stages, the backbone parameters are frozen and only the LoRA adapters\(Huet al\.[2022](https://arxiv.org/html/2608.05817#bib.bib35)\)are optimized\. The adapter learned at each stage initializes the subsequent stage:

θk⋆=SFT⁡\(θk−1⋆;𝒞\(k\)\),k=1,2,3,θ0⋆=θ0,\\theta\_\{k\}^\{\\star\}=\\operatorname\{SFT\}\\left\(\\theta\_\{k\-1\}^\{\\star\};\\mathcal\{C\}^\{\(k\)\}\\right\),\\qquad k=1,2,3,\\quad\\theta\_\{0\}^\{\\star\}=\\theta\_\{0\},\(2\)where𝒞\(k\)\\mathcal\{C\}^\{\(k\)\}denotes the stage\-kkcorpus andθk⋆\\theta\_\{k\}^\{\\star\}denotes the resulting model state\. The intermediate statesθ1⋆\\theta\_\{1\}^\{\\star\}andθ2⋆\\theta\_\{2\}^\{\\star\}transfer the progressively acquired capabilities to subsequent stages, whileθ3⋆\\theta\_\{3\}^\{\\star\}initializes the reinforcement\-learning stage\.

### Task\-Aware Reinforcement Learning

After curriculum\-based supervised fine\-tuning, we initialize the trainable policyπθ\\pi\_\{\\theta\}fromθ3⋆\\theta\_\{3\}^\{\\star\}and jointly optimize metaphor identification, Target prediction, Source prediction, and sentiment classification within a unified output space\. Given an inputxix\_\{i\}, the policy samples a group ofGGcandidate responses\{oi,g\}g=1G\\\{o\_\{i,g\}\\\}\_\{g=1\}^\{G\}\. To accommodate the heterogeneous output fields, we define the task\-aware reward for responseoi,go\_\{i,g\}as

Ri,g=\\displaystyle R\_\{i,g\}=\{\}λf​rf\+λm​rm\+λe​re\\displaystyle\\lambda\_\{f\}r\_\{f\}\+\\lambda\_\{m\}r\_\{m\}\+\\lambda\_\{e\}r\_\{e\}\(3\)\+𝕀​\[mi=Yes\]​\(λt​rt\+λs​rs\),\\displaystyle\+\\mathbb\{I\}\[m\_\{i\}=\\texttt\{Yes\}\]\\left\(\\lambda\_\{t\}r\_\{t\}\+\\lambda\_\{s\}r\_\{s\}\\right\),where all reward components are evaluated onoi,go\_\{i,g\}\. Here,rfr\_\{f\}verifies compliance with the prescribed<think\>,<label\>, and JSON formats;rmr\_\{m\}andrer\_\{e\}use exact matching for the normalized metaphor and sentiment labels, respectively\. For metaphorical samples,rtr\_\{t\}andrsr\_\{s\}are the BERTScore F1 values between the predicted and gold Target and Source concepts\(Zhanget al\.[2020](https://arxiv.org/html/2608.05817#bib.bib20)\)\. The indicator disables these mapping rewards for non\-metaphorical samples, while missing, malformed, or unparsable fields receive zero reward for the corresponding component\.

We optimizeπθ\\pi\_\{\\theta\}using Group Relative Policy Optimization \(GRPO\)\(Shaoet al\.[2024b](https://arxiv.org/html/2608.05817#bib.bib36)\)\. For responses sampled from the same input, GRPO estimates group\-relative advantages from their rewards and updates the policy with a clipped objective\. Both the trainable policyπθ\\pi\_\{\\theta\}and the frozen reference policyπref\\pi\_\{\\mathrm\{ref\}\}are initialized fromθ3⋆\\theta\_\{3\}^\{\\star\}\. During training,πref\\pi\_\{\\mathrm\{ref\}\}provides KL regularization, preventing the policy from deviating excessively from the evidence\-grounded reasoning acquired during supervised fine\-tuning\.

## M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner Evaluation

We trainM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner using Qwen3\-VL\-8B\-Instruct as the base model and evaluate it onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. Training details and hyperparameter settings are provided in the Appendix\.

Table 3:Main results onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. MD and SC report Macro\-F1, while TP and SP report BERTScore F1\.#### Experiment Results\.

We compareM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner with representative proprietary and open\-weight MLLMs\. As shown in Table[3](https://arxiv.org/html/2608.05817#Sx6.T3), all models are evaluated under the same unified full\-task setting\.M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner achieves the best performance across all four metrics, improving MD, TP, SP, and SC over its Qwen3\-VL\-8B\-Instruct base model by 20\.27, 7\.85, 6\.91, and 22\.88 points, respectively\. It also consistently outperforms larger open\-weight models and proprietary MLLMs\. These results demonstrate the effectiveness of the proposed curriculum\-based reasoning supervision and task\-aware reinforcement learning\.

Table 4:Ablation results onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench\. MD and SC report Macro\-F1; TP and SP report BERTScore F1\.
#### Ablation Study\.

Table[4](https://arxiv.org/html/2608.05817#Sx6.T4)verifies the contribution of each component\. With curriculum learning fixed, replacing label\-only supervision with evidence\-grounded CoT supervision improves MD, TP, SP, and SC by 6\.53, 3\.70, 4\.14, and 0\.06 points, respectively, demonstrating its benefit for evidence–mapping alignment\. With CoT supervision fixed, curriculum learning further yields gains of 5\.08, 3\.88, 4\.01, and 0\.87 points across the four metrics\. Task\-aware reinforcement learning consistently improves the resulting model\. Removing eitherRsemR\_\{\\mathrm\{sem\}\}orRSCR\_\{\\mathrm\{SC\}\}causes substantial degradation, confirming their complementary roles in evidence\-grounded multimodal metaphor understanding\.

Table 5:Evidence\-grounded rubric results on theM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench test set\. Mean averages VE, MC, and SJ\.
#### Rubric\-Based Explanation Evaluation\.

As shown in Table[5](https://arxiv.org/html/2608.05817#Sx6.T5),M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner achieves 58\.11, 49\.41, and 69\.79 on VE, MC, and SJ, respectively, obtaining the highest mean score of 59\.10\. Compared with the strongest proprietary baseline, Claude\-Sonnet\-4\.6, it improves VE, MC, SJ, and Mean by 14\.55, 1\.51, 7\.95, and 8\.00 points, respectively\. It also consistently surpasses CoT\-SFT \+ Curr\., demonstrating that task\-aware reinforcement learning further strengthens evidence\-grounded visual reasoning, Target–Source mapping, and sentiment justification\.

![Refer to caption](https://arxiv.org/html/2608.05817v1/x5.png)Figure 5:Qualitative comparison between GPT\-4o andM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner: text\-driven versus visually grounded Target–Source mapping\.
#### Case Study\.

Figure[5](https://arxiv.org/html/2608.05817#Sx6.F5)presents a representative example of the evidence–mapping reasoning mismatch\. The advertisement constructs a personification metaphor by depicting Lay’s potato chips as human\-like food characters\. Although GPT\-4o correctly identifies the metaphor and positive sentiment, it relies on the slogan “A campaign full of flavor” and derives an incorrect mapping from textual cues rather than visual evidence\. In contrast,M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner recognizes the anthropomorphic features as the key evidence, recovers the correct Target–Source mapping, and justifies the sentiment using visual\-textual cues\. This example demonstrates that correct labels do not necessarily indicate evidence\-grounded metaphor understanding\.

![Refer to caption](https://arxiv.org/html/2608.05817v1/x6.png)Figure 6:Error distributions of GPT\-4o andM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner on the same 100 GPT\-4o failure cases\.
#### Error Analysis\.

Figure[6](https://arxiv.org/html/2608.05817#Sx6.F6)comparesM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner with GPT\-4o on the 100 erroneous cases from RQ2\.M3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner corrects 47% of these failures, with the largest improvements in Target–Source granularity mismatch \(reduced from 26 to 6 cases\) and textual shortcut bias \(from 43 to 24 cases\)\. Visual grounding failures decrease from 21 to 15 cases, while sentiment errors drop from 8 to 6\. Hallucinated reasoning remains unchanged, suggesting a broader limitation of current VLMs\. Overall, evidence\-grounded training improves conceptual mapping and reduces textual shortcuts, while visual grounding remains the main challenge\.

## Conclusion

We introducedM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench, a unified evidence\-grounded benchmark that jointly evaluates metaphor identification, Target–Source mapping, and sentiment inference with stage\-wise reference explanations\. Evaluations onM3​R\\text\{M\}^\{3\}\\text\{R\}\-Bench reveal that existing models frequently overlook visual evidence, rely on superficial textual cues, and produce inaccurate conceptual mappings\. To address these limitations, we proposedM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner, which combines curriculum\-based reasoning supervision with task\-aware reinforcement learning\. Experiments show thatM3​R\\text\{M\}^\{3\}\\text\{R\}\-Reasoner achieves the best overall performance across all four task metrics and produces explanations that are more faithfully grounded in visual evidence and metaphorical mappings\. We hope this benchmark and method provide a foundation for advancing interpretable multimodal metaphor understanding\.

## References

- MetaCLUE: towards comprehensive visual metaphors research\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 23201–23211\.Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.20.20.6),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1)\.
- Anthropic \(2026a\)Claude Opus 4\.7 system card\.Note:Anthropic System CardExternal Links:[Link](https://www.anthropic.com/claude-opus-4-7-system-card)Cited by:[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1)\.
- Anthropic \(2026b\)Claude Sonnet 4\.6 system card\.Note:Anthropic System CardPublished February 17, 2026External Links:[Link](https://www.anthropic.com/claude-sonnet-4-6-system-card)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- S\. Baiet al\.\(2025\)Qwen3\-VL technical report\.Note:arXiv preprint arXiv:2511\.21631External Links:[Link](https://arxiv.org/abs/2511.21631)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1),[Curriculum\-Based Evidence–Mapping Reason SFT\.](https://arxiv.org/html/2608.05817#Sx5.SSx1.SSS0.Px2.p1.1)\.
- J\. Birke and A\. Sarkar \(2006\)A clustering approach for nearly unsupervised recognition of nonliteral language\.In11th Conference of the European Chapter of the Association for Computational Linguistics,Trento, Italy,pp\. 329–336\.External Links:[Link](https://aclanthology.org/E06-1042/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.5.5.6)\.
- A\. Dosovitskiyet al\.\(2021\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=YicbFdNTTy)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- C\. J\. Forceville and E\. Urios\-Aparisi \(Eds\.\) \(2009\)Multimodal metaphor\.Applications of Cognitive Linguistics, Vol\.11,Mouton de Gruyter,Berlin and New York\.External Links:[Document](https://dx.doi.org/10.1515/9783110215366)Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p1.1)\.
- Google DeepMind \(2026\)Gemini 3\.1 Pro model card\.Note:Model CardPublished February 19, 2026External Links:[Link](https://deepmind.google/models/model-cards/gemini-3-1-pro/)Cited by:[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1)\.
- K\. He, X\. Chen, S\. Xie, Y\. Li, P\. Dollár, and R\. Girshick \(2022\)Masked autoencoders are scalable vision learners\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 16000–16009\.Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Curriculum\-Based Evidence–Mapping Reason SFT\.](https://arxiv.org/html/2608.05817#Sx5.SSx1.SSS0.Px2.p2.7)\.
- Kimi Teamet al\.\(2026\)Kimi K2\.5: visual agentic intelligence\.Note:arXiv preprint arXiv:2602\.02276External Links:[Link](https://arxiv.org/abs/2602.02276)Cited by:[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1),[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- M\. Kundu, S\. Shekhar, and P\. Bhattacharyya \(2025\)Looking beyond the pixels: evaluating visual metaphor understanding in VLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 23137–23158\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1257),[Link](https://aclanthology.org/2025.findings-emnlp.1257/)Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p2.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1)\.
- G\. Lakoff and M\. Johnson \(1980\)Metaphors we live by\.University of Chicago Press,Chicago\.Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p4.3),[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1)\.
- C\. Liu, G\. Geigle, R\. Krebs, and I\. Gurevych \(2022\)FigMemes: a dataset for figurative language identification in politically\-opinionated memes\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Abu Dhabi, United Arab Emirates,pp\. 7069–7086\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.476),[Link](https://aclanthology.org/2022.emnlp-main.476/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.40.40.6)\.
- X\. Lu, Y\. Liu, D\. Zhang, Z\. Wu, J\. Ren, and F\. Xia \(2025\)EmoMeta: a multimodal dataset for fine\-grained emotion classification in chinese metaphors\.InCompanion Proceedings of the ACM on Web Conference 2025,pp\. 3080–3083\.External Links:[Document](https://dx.doi.org/10.1145/3701716.3735080),[Link](https://doi.org/10.1145/3701716.3735080)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.60.60.6),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Data Sources\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px2.p1.1)\.
- OpenAI \(2024\)GPT\-4o system card\.Note:arXiv preprint arXiv:2410\.21276External Links:[Link](https://arxiv.org/abs/2410.21276)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- OpenAI \(2026\)GPT\-5\.5 system card\.Note:OpenAI System CardPublished April 23, 2026External Links:[Link](https://openai.com/index/gpt-5-5-system-card/)Cited by:[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1),[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- A\. Radfordet al\.\(2021\)Learning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.External Links:[Link](https://proceedings.mlr.press/v139/radford21a.html)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Y\. Shao, X\. Yao, X\. Qu, C\. Lin, S\. Wang, W\. Huang, G\. Zhang, and J\. Fu \(2024a\)CMDAG: a chinese metaphor dataset with annotated grounds as CoT for boosting metaphor generation\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),Torino, Italia,pp\. 3357–3366\.External Links:[Link](https://aclanthology.org/2024.lrec-main.298/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.15.15.6)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo \(2024b\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.Note:arXiv preprint arXiv:2402\.03300External Links:[Link](https://arxiv.org/abs/2402.03300)Cited by:[Task\-Aware Reinforcement Learning](https://arxiv.org/html/2608.05817#Sx5.SSx2.p2.5)\.
- G\. J\. Steen, A\. G\. Dorst, J\. B\. Herrmann, A\. A\. Kaal, T\. Krennmayr, and T\. Pasma \(2010\)A method for linguistic metaphor identification: from MIP to MIPVU\.Converging Evidence in Language and Communication Research, Vol\.14,John Benjamins,Amsterdam\.External Links:[Document](https://dx.doi.org/10.1075/celcr.14)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.10.10.6)\.
- Y\. Tian, M\. Wang, N\. Xu, and W\. Mao \(2025\)ImaRA: an imaginative frame augmented method for low\-resource multimodal metaphor detection and explanation\.InFindings of the Association for Computational Linguistics: NAACL 2025,Albuquerque, New Mexico,pp\. 3953–3967\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.220),[Link](https://aclanthology.org/2025.findings-naacl.220/)Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p3.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1)\.
- W\. Wanget al\.\(2025\)InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.Note:arXiv preprint arXiv:2508\.18265External Links:[Link](https://arxiv.org/abs/2508.18265)Cited by:[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- Y\. Wilks \(1975\)A preferential, pattern\-seeking, semantics for natural language inference\.Artificial Intelligence6\(1\),pp\. 53–74\.External Links:[Document](https://dx.doi.org/10.1016/0004-3702%2875%2990016-8)Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p4.3),[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1)\.
- xAI \(2025\)Grok 4\.1 model card\.Note:Model CardPublished November 17, 2025External Links:[Link](https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf)Cited by:[Label Annotation Scheme\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px3.p1.1)\.
- B\. Xu, T\. Li, J\. Zheng, M\. Naseriparsa, Z\. Zhao, H\. Lin, and F\. Xia \(2022\)MET\-Meme: a multimodal meme dataset rich in metaphors\.InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 2887–2899\.External Links:[Document](https://dx.doi.org/10.1145/3477495.3532019),[Link](https://doi.org/10.1145/3477495.3532019)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.30.30.6),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p1.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1)\.
- Y\. Xu, Y\. Hua, S\. Li, and Z\. Wang \(2024\)Exploring chain\-of\-thought for multi\-modal metaphor detection\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 91–101\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.6),[Link](https://aclanthology.org/2024.acl-long.6/)Cited by:[Introduction](https://arxiv.org/html/2608.05817#Sx1.p3.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1),[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- S\. Yang, D\. Zhang, J\. Ren, Z\. Xu, X\. Zhang, Y\. Song, H\. Lin, and F\. Xia \(2025\)Cultural bias matters: a cross\-cultural benchmark dataset and sentiment\-enriched model for understanding multimodal metaphors\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 26301–26317\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1275),[Link](https://aclanthology.org/2025.acl-long.1275/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.55.55.6),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p2.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1),[Data Sources\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px2.p1.1),[Evaluation Metrics\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px1.p1.2),[Evaluated Models\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px2.p1.1)\.
- K\. Ye, N\. H\. Nazari, J\. Hahn, Z\. Hussain, M\. Zhang, and A\. Kovashka \(2021\)Interpreting the rhetoric of visual advertisements\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(4\),pp\. 1308–1323\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2019.2947440)Cited by:[Data Sources\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px2.p1.1)\.
- C\. Zhanget al\.\(2025\)Can MLLMs understand the deep implication behind chinese images?\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 14369–14402\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.700),[Link](https://aclanthology.org/2025.acl-long.700/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.50.50.6)\.
- D\. Zhang, X\. Lu, M\. Zhuang, S\. Yang, and H\. Chen \(2025a\)Multimodal metaphor recognition based on chain\-of\-cognition prompting\.Cognitive Systems Research91,pp\. 101356\.External Links:[Document](https://dx.doi.org/10.1016/j.cogsys.2025.101356),[Link](https://doi.org/10.1016/j.cogsys.2025.101356)Cited by:[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1)\.
- D\. Zhang, S\. Yin, J\. Yu, Z\. Wu, Z\. Li, C\. Xu, X\. Wang, and F\. Xia \(2025b\)Towards multimodal metaphor understanding: a chinese dataset and model for metaphor mapping identification\.ACM Transactions on Asian and Low\-Resource Language Information Processing24\(12\),pp\. 1–25\.External Links:[Document](https://dx.doi.org/10.1145/3773989),[Link](https://doi.org/10.1145/3773989)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.45.45.6),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1)\.
- D\. Zhang, J\. Yu, S\. Jin, L\. Yang, and H\. Lin \(2023\)MultiCMET: a novel chinese benchmark for understanding multimodal metaphor\.InFindings of the Association for Computational Linguistics: EMNLP 2023,Singapore,pp\. 6141–6154\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.409),[Link](https://aclanthology.org/2023.findings-emnlp.409/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.35.35.6),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p2.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1),[Data Sources\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px2.p1.1),[Evaluation Metrics\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px1.p1.2)\.
- D\. Zhang, M\. Zhang, H\. Zhang, L\. Yang, and H\. Lin \(2021\)MultiMET: a multimodal dataset for metaphor understanding\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),Online,pp\. 3214–3225\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.249),[Link](https://aclanthology.org/2021.acl-long.249/)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.25.25.6),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p2.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1),[Metaphor Understanding Methods\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px2.p1.1),[Data Sources\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px2.p1.1),[Evaluation Metrics\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px1.p1.2)\.
- T\. Zhang, V\. Kishore, F\. Wu, K\. Q\. Weinberger, and Y\. Artzi \(2020\)BERTScore: evaluating text generation with BERT\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SkeHuCVFDr)Cited by:[Annotation of Evidence\-grounded Chains\.](https://arxiv.org/html/2608.05817#Sx3.SS0.SSS0.Px4.p1.1),[Evaluation Metrics\.](https://arxiv.org/html/2608.05817#Sx4.SSx1.SSS0.Px1.p1.2),[Task\-Aware Reinforcement Learning](https://arxiv.org/html/2608.05817#Sx5.SSx2.p1.12)\.
- T\. Zheng, Y\. Yang, R\. Dong, B\. Ma, L\. Wang, X\. Zhou, S\. Miao, and T\. Osman \(2026\)M3UCD: a multi\-task multimodal metaphor understanding challenge dataset for LLMs\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 35030–35040\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i41.40808),[Link](https://doi.org/10.1609/aaai.v40i41.40808)Cited by:[Table 1](https://arxiv.org/html/2608.05817#Sx1.T1.65.65.6),[Introduction](https://arxiv.org/html/2608.05817#Sx1.p2.1),[Multimodal Metaphor Benchmarks\.](https://arxiv.org/html/2608.05817#Sx2.SS0.SSS0.Px1.p1.1)\.

Similar Articles

MetaphorVU: Towards Metaphorical Video Understanding

Hugging Face Daily Papers

This paper introduces MetaphorVU-Bench, the first systematic benchmark for metaphorical video understanding, and proposes MetaphorBoost, an inference-time enhancement framework that improves cross-domain mapping in multimodal large language models.