MetaphorVU: Towards Metaphorical Video Understanding
Summary
This paper introduces MetaphorVU-Bench, the first systematic benchmark for metaphorical video understanding, and proposes MetaphorBoost, an inference-time enhancement framework that improves cross-domain mapping in multimodal large language models.
View Cached Full Text
Cached at: 05/26/26, 02:41 AM
Paper page - MetaphorVU: Towards Metaphorical Video Understanding
Source: https://huggingface.co/papers/2605.25461 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Current multimodal large language models struggle with metaphorical video understanding due to poor cross-domain mapping, prompting the development of a new benchmark and enhancement framework.
Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies onmetaphorical video understandingnot only constrains the real-world applicability of MLLMs but also impedes the thorough assessment of their high-order cognitive capabilities. To bridge this gap, we propose MetaphorVU-Bench, the first systematic and comprehensive benchmark dedicated tometaphorical video understanding. Through experiments, we find current MLLMs struggle with accuratemetaphorical video understanding, lagging far behind human level, primarily due to defectivecross-domain mapping. Motivated by this finding, we construct ametaphor knowledge graphas mapping augmentation and proposeMetaphorBoost, aninference-time enhancementframework achieving consistent performance improvement. Our benchmark, analysis, and method provide useful insights and a foundation for future research on advancing MLLMs.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2605\.25461
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.25461 in a model README.md to link it from this page.
Datasets citing this paper1
#### lzq2021/MetaphorVU-Bench Viewer• Updated12 minutes ago • 861 • 1.73k • 3
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.25461 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ViMU: Benchmarking Video Metaphorical Understanding
ViMU is the first benchmark designed to evaluate video understanding models' ability to interpret metaphorical, ironic, and social meanings beyond literal visual comprehension, using hint-free open-ended and multiple-choice questions.
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.
A Reproducible Multi-Architecture Baseline for Token-Level Chinese Metaphor Identification under the MIPVU Framework
This paper establishes a reproducible multi-architecture baseline for token-level Chinese metaphor identification using the MIPVU framework and the PSU Chinese Metaphor Corpus. It compares encoder models like RoBERTa and MelBERT against the Qwen3.5-9B generative model, releasing code and data to facilitate future research.
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Introduces Video-MME-Logical, a controlled benchmark for evaluating video temporal-logical reasoning in multimodal large language models, revealing a substantial human-model gap.
Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation
A PhD proposal outlining a unified end-to-end framework for multilingual metaphor processing, integrating metaphor detection, translation evaluation, and joint modeling using linguistic theory and large language models.