Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

arXiv cs.CL Papers

Summary

A PhD proposal outlining a unified end-to-end framework for multilingual metaphor processing, integrating metaphor detection, translation evaluation, and joint modeling using linguistic theory and large language models.

arXiv:2608.04260v1 Announce Type: new Abstract: Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require reasoning beyond literal lexical meaning. Existing research has largely investigated metaphor detection, machine translation, and translation evaluation as separate tasks, while little work has explored how these components can be integrated into a unified computational framework. This PhD proposal aims to develop an end-to-end framework for multilingual metaphor processing consisting of three complementary research directions: (1) robust metaphor detection across languages, (2) metaphor-oriented translation evaluation for both human assessment and automatic quality estimation, and (3) joint modelling that connects metaphor detection with translation evaluation. The proposed research will combine linguistic theory with recent advances in large language models to develop new datasets, annotation methodologies, evaluation benchmarks, and automatic evaluation approaches for metaphor-aware machine translation. The expected outcome is a unified framework that improves both the development and evaluation of multilingual NLP systems when processing figurative language.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:45 AM

# Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation
Source: [https://arxiv.org/html/2608.04260](https://arxiv.org/html/2608.04260)
Jiahui Liang1, Lifeng Han2,3 1Centre for Linguistics, Humanities, Leiden University, NL 2LIACS, Leiden University, NL 3BDS, Leiden University Medical Centre, NL j\.h\.l\.jiahui@hum\.leidenuniv\.nl \| l\.han@lumc\.nl

###### Abstract

Metaphorical language remains a major challenge for multilingual natural language processing because successful interpretation and translation require reasoning beyond literal lexical meaning\. Existing research has largely investigated metaphor detection, machine translation, and translation evaluation as separate tasks, while little work has explored how these components can be integrated into a unified computational framework\. This PhD proposal aims to develop an end\-to\-end framework for multilingual metaphor processing consisting of three complementary research directions: \(1\) robust metaphor detection across languages, \(2\) metaphor\-oriented translation evaluation for both human assessment and automatic quality estimation, and \(3\) joint modelling that connects metaphor detection with translation evaluation\. The proposed research will combine linguistic theory with recent advances in large language models to develop new datasets, annotation methodologies, evaluation benchmarks, and automatic evaluation approaches for metaphor\-aware machine translation\. The expected outcome is a unified framework that improves both the development and evaluation of multilingual NLP systems when processing figurative language\.

Towards End\-to\-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

Jiahui Liang1, Lifeng Han2,31Centre for Linguistics, Humanities, Leiden University, NL2LIACS, Leiden University, NL3BDS, Leiden University Medical Centre, NLj\.h\.l\.jiahui@hum\.leidenuniv\.nl \| l\.han@lumc\.nl

## 1Introduction

Metaphors are widespread in our daily usage and play a central role in human cognition and communication\. Rather than being merely a rhetorical device used for stylistic effect, metaphor is regarded as a fundamental mechanism through which people conceptualise and understand the worldLakoff and Johnson \([1980](https://arxiv.org/html/2608.04260#bib.bib125)\); Kövecses \([2010](https://arxiv.org/html/2608.04260#bib.bib99)\)\. Specifically, the process is realized by conceptualising the abstract target domain through mappings from the concrete source domainLakoff and Johnson \([1980](https://arxiv.org/html/2608.04260#bib.bib125)\)\. For example, in the typical conceptual metaphor TIME IS MONEY, time is understood as a valuable resource\.

These conceptual mappings are grounded in embodied physical and cultural experiencesLakoff and Johnson \([1980](https://arxiv.org/html/2608.04260#bib.bib125)\); Gibbs \([2008](https://arxiv.org/html/2608.04260#bib.bib21)\), while their interpretation often depends on contextual informationLi and Chen \([2025a](https://arxiv.org/html/2608.04260#bib.bib31)\)and shared cultural knowledgeLakoff and Johnson \([1980](https://arxiv.org/html/2608.04260#bib.bib125)\)\. Because embodied experiences and conceptualizations may differ across individuals and culturesKövecses \([2010](https://arxiv.org/html/2608.04260#bib.bib99)\), metaphorical meaning is often non\-compositional and cannot be derived from the literal meanings of individual words alone\. Instead, it requires the integration of linguistic, conceptual, contextual, and cultural information\. These challenges have motivated increasing research on computational metaphor processing, including metaphor detection, interpretation, generation and translationRai and Chakraverty \([2020](https://arxiv.org/html/2608.04260#bib.bib26)\); Geet al\.\([2023](https://arxiv.org/html/2608.04260#bib.bib25)\)\. Among these tasks, metaphordetectionand metaphortranslationhave developed largely asindependentresearch directionsLianget al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib20)\); Liang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\)\. Metaphor detection aims to distinguish metaphorical from literal language and serves as a prerequisite for many downstream applicationsShutova \([2010](https://arxiv.org/html/2608.04260#bib.bib24)\), whereas metaphor translation focuses on conveying metaphorical meaning across languages while accounting for linguistic and cultural differencesKarakantaet al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib30)\)\. Despite substantial advances in neural language models, both tasks remain challenging\. Successful metaphor detection requires models to recognize non\-literal meaning beyond surface lexical patterns, while metaphor translation must further address variations in linguistic form, communicative function, and cultural embeddedness across languagesKarakantaet al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib30)\)\. However, relativelylittle work has connected these two areas\. Existing translation evaluation methods generally assume that metaphor instances have already been identified manually, while advances in metaphor detection have rarely been incorporated into translation evaluation or quality estimation\.

Recent evidence also suggests that large language models may rely more heavily on superficial lexical cues than on genuine metaphor understanding, indicating that figurative language remains difficult even for state\-of\-the\-art systemsSanchez\-Bayona and Agerri \([2025](https://arxiv.org/html/2608.04260#bib.bib23)\)\. Furthermore, progress in metaphor translation is constrained by the limited availability of high\-quality parallel metaphor corpora and systematic evaluation resourcesWanget al\.\([2024](https://arxiv.org/html/2608.04260#bib.bib39)\); Hanet al\.\([2026b](https://arxiv.org/html/2608.04260#bib.bib38)\)\.

Recent work has also highlighted shortcomings in current evaluation practice\. General\-purpose MT metrics, such as BLEUBanerjee and Lavie \([2005](https://arxiv.org/html/2608.04260#bib.bib14)\)or COMETReiet al\.\([2020](https://arxiv.org/html/2608.04260#bib.bib13)\), are not designed to capture metaphor\-specific phenomena, and human evaluation protocols rarely distinguish different types of metaphor translation errors\. Emerging resources, including MMTE\(Wanget al\.,[2024](https://arxiv.org/html/2608.04260#bib.bib39)\), have begun addressing this gap, while our MetaHOPE framework extends the human\-oriented HOPE evaluation framework with metaphor\-specific error categories and annotation guidelinesLiang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\)\. Nevertheless, reliable metaphor\-oriented evaluation and its integration with automatic metaphor processing remain open research problems\.

This PhD proposes an end\-to\-end computational framework for multilingual metaphor processing that unifies three complementary research directions: \(1\) metaphor detection, \(2\) metaphor\-oriented translation evaluation integrating MetaHOPE, and \(3\) joint modelling of metaphor detection and translation evaluation\. Rather than treating these components as isolated tasks, the proposed research investigates how they can mutually reinforce one another to improve both the evaluation and the development of MT and LLM systems\. The long\-term objective is to establish reliable methodologies, resources and benchmarks for metaphor\-aware multilingual NLP\.

## 2Background and Related Work

### 2\.1Cognitive Metaphor

Metaphor has traditionally been regarded as a rhetorical device used to enrich literary expression\. However, this view was fundamentally challenged by Conceptual Metaphor Theory \(CMT\), which argues that metaphor is not merely a property of language but a fundamental mechanism of human thought and cognition\(Lakoff and Johnson,[1980](https://arxiv.org/html/2608.04260#bib.bib125)\)\. According to CMT, abstract concepts are systematically understood through more concrete source domains, giving rise to conceptual mappings such asARGUMENT IS WARandTIME IS MONEY\. Linguistic metaphorical expressions are therefore surface manifestations of deeper conceptual structures rather than isolated stylistic devices\.

Subsequent research has further demonstrated that metaphor is a multidimensional phenomenon involving cognition, language, communication, and culture\. While conceptual mappings provide the cognitive basis for metaphorical expressions, their linguistic realisations are influenced by discourse context and communicative intentionSemino \([2008](https://arxiv.org/html/2608.04260#bib.bib92)\)\. Deliberate Metaphor Theory \(DMT\), for example, distinguishes between metaphors that are intentionally employed to direct readers’ attention towards the source domain and those that have become conventionalised in everyday language\(Steen,[2017](https://arxiv.org/html/2608.04260#bib.bib103)\)\. This distinction is particularly relevant for computational processing because not all metaphorical expressions require the same level of semantic interpretation\.

Metaphor is especially prevalent in news discourse, political communication and other public\-facing genres, where it serves multiple communicative functions beyond stylistic ornamentation\. Previous studies have shown that metaphor can simplify complex events, frame public opinion, evoke emotional responses, and communicate ideological positions\(Steenet al\.,[2010b](https://arxiv.org/html/2608.04260#bib.bib76)\)\. Consequently, successful metaphor processing requires models to capture not only lexical meaning but also discourse\-level and pragmatic information\.

Cross\-linguistic variation further increases the complexity of metaphor processing\. Although some conceptual metaphors appear to be widely shared across cultures, their linguistic realisations often differ considerably because of language\-specific conventions, cultural experiences, and socio\-historical backgrounds\(Kövecses,[2010](https://arxiv.org/html/2608.04260#bib.bib99)\)\. As a result, multilingual NLP systems must account for both universal conceptual mappings and culture\-specific metaphorical expressionsHanet al\.\([2026a](https://arxiv.org/html/2608.04260#bib.bib41)\)\. These challenges motivate the need for computational approaches capable of recognising, interpreting, translating and evaluating metaphorical language across languages\.

### 2\.2Computational Metaphor Detection

Automatic metaphor detection has become one of the most active research topics in figurative language processingShutova \([2015](https://arxiv.org/html/2608.04260#bib.bib22)\); Rai and Chakraverty \([2020](https://arxiv.org/html/2608.04260#bib.bib26)\); Hanet al\.\([2026a](https://arxiv.org/html/2608.04260#bib.bib41)\)\. Early computational approaches mainly relied on hand\-coded knowledge, lexical resources, and selectional preferences to distinguish metaphorical from literal languageShutova \([2015](https://arxiv.org/html/2608.04260#bib.bib22)\)\. However, these rule\-based methods were often difficult to scale and generalise across domains and languages, motivating the development of data\-driven approaches\.

With the development of corpus\-based metaphor research, manual metaphor identification procedures also became increasingly standardised\. The Metaphor Identification Procedure \(MIP\)\(Group,[2007](https://arxiv.org/html/2608.04260#bib.bib80)\)and its extended version, MIPVU\(Steenet al\.,[2010b](https://arxiv.org/html/2608.04260#bib.bib76)\), provide systematic dictionary\-based guidelines for determining whether a lexical unit is used metaphorically by comparing its contextual meaning with a more basic meaning \. Based on this procedure, the VU Amsterdam Metaphor Corpus \(VUAMC\)Steenet al\.\([2010a](https://arxiv.org/html/2608.04260#bib.bib75)\)was constructed as one of the largest manually annotated metaphor corpora and has since become a widely adopted benchmark for metaphor detection research\. The availability of MIPVU\-annotated corpora has enabled the development and evaluation of supervised metaphor detection models under consistent annotation standards\.

Advances in deep learning have substantially improved metaphor detection performance\. Contextual language models, including BERT and subsequent Transformer architectures, enable richer semantic representations by modelling contextual information surrounding metaphorical expressions\. These models consistently outperform earlier statistical approaches on standard benchmark datasets and demonstrate improved robustness across multiple domains\. This transition was reflected in the VUA Metaphor Detection Shared Tasks\. While the top\-performing systems in 2018 were mainly based on recurrent neural networks and feature engineering, the 2020 shared task was dominated by BERT\-based models, with the best F1 score improving from 0\.651 to 0\.769Leonget al\.\([2018](https://arxiv.org/html/2608.04260#bib.bib95),[2020](https://arxiv.org/html/2608.04260#bib.bib96)\)\.

More recently, researchers have begun to explore large language models \(LLMs\) for automatic metaphor detection\. Instead of relying exclusively on task\-specific supervised training, LLMs can identify metaphorical expressions through zero\-shot and few\-shot prompting\. Recent studies have investigated the use of Llama, Qwen, Mistral, DeepSeek, Gemma, GPT\-4, and other LLMs for metaphor identification, demonstrating their potential for detecting conventional metaphorical expressions while also revealing limitations in achieving robust metaphor understanding\(Lianget al\.,[2025](https://arxiv.org/html/2608.04260#bib.bib20); Reimann and Scheffler,[2025](https://arxiv.org/html/2608.04260#bib.bib18); Hanet al\.,[2026a](https://arxiv.org/html/2608.04260#bib.bib41)\)\. Building on these advances, recent work has explored task\-specific prompt design, MIPVU\-inspired reasoning, chain\-of\-thought prompting, and retrieval\-augmented generation \(RAG\) to further improve metaphor detection\(Reimann and Scheffler,[2025](https://arxiv.org/html/2608.04260#bib.bib18); Fuoliet al\.,[2026](https://arxiv.org/html/2608.04260#bib.bib27)\)\. Related studies have also extended these approaches to multilingual, cross\-lingual, and multimodal metaphor detection\(Laiet al\.,[2023](https://arxiv.org/html/2608.04260#bib.bib17); Hülsing and Im Walde,[2024](https://arxiv.org/html/2608.04260#bib.bib16); Xuet al\.,[2024](https://arxiv.org/html/2608.04260#bib.bib15)\)\. Nevertheless, current results remain mixed, with the effectiveness of LLM\-based approaches depending strongly on the task, prompting strategy, and evaluation dataset\(Reimann and Scheffler,[2025](https://arxiv.org/html/2608.04260#bib.bib18); Fuoliet al\.,[2026](https://arxiv.org/html/2608.04260#bib.bib27); Hanet al\.,[2026a](https://arxiv.org/html/2608.04260#bib.bib41)\)\.

Despite these advances, several challenges remain\. Current systems often struggle with conventional metaphors, culturally specific metaphorical expressions, and implicit figurative meanings that require discourse\-level reasoning or common\-sense knowledge\. Moreover, most existing work evaluates metaphor detection as an isolated classification task, while relatively little attention has been paid to how detected metaphorical expressions can support downstream applications such as machine translation and translation quality evaluation\.

Building upon our previous work on GPT\-4\-based metaphor detection in English news textsLianget al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib20)\), this limitation motivates thefirst work packageof this PhD, which extends metaphor identification towardsmultilingualsettings, improved prompting strategies, and more robust detection methods that can support downstream translation and evaluation\.

### 2\.3Metaphor Translation

Metaphor translation has long been recognised as one of the most challenging problems in translation studies because metaphorical expressions often convey meaning beyond their literal lexical content\. Unlike literal translation, successful metaphor translation requires preserving not only semantic equivalence but also conceptual mappings, communicative functions, stylistic effects, and cultural connotations\. Classical translation scholars have therefore argued that metaphor translation should be viewed as a process of conceptual transfer rather than simple lexical substitution\.Newmark \([1988](https://arxiv.org/html/2608.04260#bib.bib119)\), for example, proposed a taxonomy of metaphor translation procedures ranging from reproducing the same metaphor in the target language to replacing or paraphrasing metaphorical expressions according to communicative needs\. Similarly,van den Broeck \([1981](https://arxiv.org/html/2608.04260#bib.bib120)\)andSchäffner \([2004](https://arxiv.org/html/2608.04260#bib.bib117)\)emphasised that translation strategies should consider both linguistic form and underlying conceptual structures, highlighting the importance of cultural adaptation in cross\-linguistic metaphor transfer\.

The rapid development of neural machine translation \(NMT\) has substantially improved translation quality for general\-purpose textsJohnsonet al\.\([2017](https://arxiv.org/html/2608.04260#bib.bib5)\); Kocmiet al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib32)\)\. Transformer\-based systems and multilingual large language models \(LLMs\) are increasingly capable of generating fluent and contextually appropriate translations without explicit linguistic rules\. Nevertheless, metaphorical language remains a persistent challenge because successful translation often requires pragmatic reasoning, cultural knowledge, and an understanding of implicit conceptual mappingsHanet al\.\([2026b](https://arxiv.org/html/2608.04260#bib.bib38)\)\. Existing MT systems may produce literal translations that distort metaphorical meaning, omit metaphorical imagery, or replace metaphors with more conventional expressions, thereby reducing their communicative impactLiang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\)\.

Recent research has begun to investigate metaphor translation from a computational perspective\.Wanget al\.\([2024](https://arxiv.org/html/2608.04260#bib.bib39)\)introduced MMTE, a multilingual benchmark specifically designed for evaluating machine translation of metaphorical language across several language pairs\. Their study demonstrated that although modern LLM\-based systems outperform earlier neural MT models, substantial performance gaps remain, particularly for culturally specific and conceptually complex metaphors\. Other recent studies have similarly compared human translators with NMT systems and LLMs, showing that although modern LLMs often produce fluent and semantically adequate translations, they remain less reliable in preserving metaphorical expressions, conceptual mappings, and rhetorical effects, particularly for culturally specific or novel metaphors\(Wanget al\.,[2024](https://arxiv.org/html/2608.04260#bib.bib39); Karakantaet al\.,[2025](https://arxiv.org/html/2608.04260#bib.bib30); Li and Chen,[2025b](https://arxiv.org/html/2608.04260#bib.bib6)\)\. These findings suggest that improvements in general machine translation quality do not necessarily translate into reliable metaphor translation\.

Although the quality of MT and LLM systems continues to improve rapidly, relatively little work has investigated how metaphor translation errors should be analysed systematically\. Existing studies typically report overall translation quality or metaphor preservation rates, providing limited diagnostic information about the specific linguistic and conceptual problems underlying translation failures\. Consequently, metaphor translation research increasingly requires evaluation frameworks capable of identifying fine\-grained error types beyond conventional MT quality metrics\. This observation motivatesthe second work packageof this PhD, which investigates metaphor\-oriented translation evaluation and analysis\.

### 2\.4Translation Evaluation

Automatic evaluation has become an indispensable component of modern machine translation research\. Traditional corpus\-level metrics such as BLEU\(Papineniet al\.,[2002](https://arxiv.org/html/2608.04260#bib.bib19)\), METEOR\(Banerjee and Lavie,[2005](https://arxiv.org/html/2608.04260#bib.bib14)\), TERSnoveret al\.\([2006](https://arxiv.org/html/2608.04260#bib.bib8)\), and LEPORHanet al\.\([2012](https://arxiv.org/html/2608.04260#bib.bib9),[2013](https://arxiv.org/html/2608.04260#bib.bib7)\)primarily measure lexical or n\-gram overlap between system outputs and reference translations\. Although these metrics remain useful for benchmarking general translation quality, they are often insensitive to semantic equivalence and perform poorly when evaluating figurative language, where multiple valid translations may differ substantially in lexical realisation\.

Recent evaluation methods have increasingly shifted towards semantic and learned metrics\. COMET\(Reiet al\.,[2020](https://arxiv.org/html/2608.04260#bib.bib13)\)and BLEURT\(Sellamet al\.,[2020](https://arxiv.org/html/2608.04260#bib.bib12)\), for example, employ pretrained language models to estimate translation quality based on semantic similarity rather than surface\-form matching\. These approaches generally correlate more strongly with human judgement than traditional automatic metrics and have become standard evaluation tools in contemporary MT research\. Nevertheless, they remain general\-purpose metrics and do not explicitly assess whether metaphorical meaning, conceptual mappings, or rhetorical effects have been preserved during translation\.

To address the limitations of automatic metrics, human\-centred evaluation frameworks have received growing attention\. The Multidimensional Quality Metrics \(MQM\) frameworkLommelet al\.\([2024](https://arxiv.org/html/2608.04260#bib.bib124),[2014](https://arxiv.org/html/2608.04260#bib.bib107)\); Gladkoffet al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib123)\)provides a hierarchical taxonomy for classifying translation errors according to dimensions such as accuracy, fluency, terminology, and style\. HOPEGladkoff and Han \([2022](https://arxiv.org/html/2608.04260#bib.bib40)\)further simplified human evaluation by introducing a penalty\-based annotation scheme that improves annotation efficiency and inter\-annotator consistency while retaining diagnostic capability\. However, neither MQM nor HOPE explicitly considers metaphor as an independent evaluation dimension\.

Recent metaphor\-specific evaluation research has therefore begun to bridge this gap\. MMTEWanget al\.\([2024](https://arxiv.org/html/2608.04260#bib.bib39)\)proposed benchmark datasets and dedicated evaluation metrics for metaphor translation, demonstrating that conventional automatic metrics frequently overestimate translation quality for figurative language\. However, MMTE corpus has isolated sentences without context and it only contains verbal metaphors\. Building upon these developments, our previous workLiang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\)introduced MetaHOPE, a metaphor\-oriented extension of the HOPE framework that incorporatesfive fine\-grained error categoriesto capture different aspects of metaphor translation quality, including communicative impact \(IMP\), adaptation to target language \(RAM\), semantic accuracy \(MIS\), stylistic preservation \(STL\), and naturalness \(PRF\)· MetaHOPE used context\-aware translation before extracting the metaphor\-containing segments for evaluation\. It also produces a open\-source parallel corpus during the post\-editing phase that can be useful for future metaphor studies and benchmarks\. Preliminary pilot studies further showed that refined annotation guidelines substantially improve inter\-annotator agreement, supporting the feasibility of reliable human evaluation for metaphor translation\.

Despite these advances, metaphor evaluation remains largely dependent on manual annotation, making large\-scale evaluation both expensive and time\-consuming\. Furthermore, current automatic metrics rarely incorporate information obtained during metaphor detection, and human evaluation frameworks generally assume that metaphorical expressions have already been identified\. These limitations motivate thefinal work packageof this PhD, which investigates automatic metaphor translation quality estimation by integrating metaphor detection, human evaluation, and LLM\-based evaluation into a unified computational framework\.

### 2\.5Research Questions from the Gap

Research Questions

- •RQ1: How can metaphorical expressions be reliably detected across languages?
- •RQ2: How should metaphor translation quality be evaluated?
- •RQ3: Can metaphor detection and translation evaluation be integrated into a unified computational framework?

![Refer to caption](https://arxiv.org/html/2608.04260v1/x1.png)Figure 1:Thesis Proposal Framework on Metaphor Detection, Translation Evaluation, and Joint Modelling\.Multilingual Metaphor ProcessingWP1Metaphor DetectionWP2Metaphor TranslationWP3Human EvaluationWP4Automatic EvaluationDetectionFrameworkMT & LLMTranslationBenchmarkMetaHOPECorpusGuidelinesAutomatic MQELLM\-as\-a\-JudgeJoint FrameworkFigure 2:Overview of the proposed PhD research programme\. The four work packages form an end\-to\-end framework for multilingual metaphor processing, spanning metaphor detection, translation, human evaluation, and automatic evaluation\.

## 3Proposed Research Framework

We design the thesis framework in Figure[1](https://arxiv.org/html/2608.04260#S2.F1), which includes three components 1\) metaphor detection, 2\) metaphor translation evaluation, and 3\) joint modelling on detection and translation evaluation\.

Step I detection task: we plan to investigate rule\-based, simple LLM\-prompting with examples and chain\-of\-thoughts \(CoTs\), and RAG\.

Step\-II translation and evaluation task: we divide it into reference\-free evaluation and reference\-based\. The reference\-free evaluation includes 1\) automatic quality estimation \(QE\) using source text and system output, 2\) human native\-speaker assessing the translation quality, and 3\) post\-editing based metrics\. The reference\-based evaluation needs corpus with multiple\-references for better mapping between hypothesis translation and the references\. For this, we can use the created references from post\-editing based evaluation from the reference\-free step\.

Step\-III joint task: we investigate the possibility of joint modelling on both metaphor detection task and metaphor translation evaluation, which can be reference\-based or reference\-free quality estimation\.

## 4WPs

WP1 includes Detection, datasets, prompting, fine\-tuning, and multilingual models\.

WP2 includes: metaphor translation, MT/LLM, and Translation Benchmark\.

WP3 is on Human Evaluation, Integrating MetaHOPELiang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\), annotation, corpus, guidelines, and reliability discussion\.

WP4 is of Automatic Eval, auto MQE, llm\-as\-a\-judge, and end\-to\-end Joint framework \( Detection, Translation, Evaluation /Automatic QE, LLM judge\)\.

## 5Methodology

In our earlier work on LLM\-based metaphor detection for English news textsLianget al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib20)\), we discovered that GPT\-4 achieved substantially lower performance than SOTA supervised metaphor detection models on conventional metaphor detection\. We also observed that model performance was highly sensitive to prompt formulation\. While one\-shot prompting consistently improved performance over zero\-shot prompting, providing additional examples did not lead to further improvements\. Our error analysis further showed that prepositional metaphors \(e\.g\., under in under strain\) were particularly difficult to detect, with many prompts failing to identify such cases in the zero\- to five\-shot settings\. Although detection improved in the ten\-shot setting, our findings indicate that the observed gains cannot be straightforwardly interpreted as improved metaphor understanding and require further investigation\. These findings motivate further research on prompt optimization, alternative output formats, fine\-tuning open\-source LLMs, and RAG\. Building upon our earlier work, the first work package \(WP1\) extends metaphor identification towardsmultilingualand morerobustdetection methods\. In particular, we will explore newer and more diverse open\-source LLMs such as GPT \-5, DeepSeek, Huanyuan, and Gemma 4 from 2025 onwards\. Furthermore, we will investigate chain of thoughts \(CoTs\) and RAG methods on this task, naming itMetaDetect\.

Regarding Metaphor Translation and Evaluation \(WP2 and WP3\), we are conducting the roll\-out of larger size testing set annotation based on the MetaHOPE frameworkLiang and Han \([2026](https://arxiv.org/html/2608.04260#bib.bib1)\)using 200\+ segments with their context using 5 annotators for both translation directions on English\-Chinese\. The resulting outputs from this will be a statistically robust analysis of the current state\-of\-the\-art LLMs regarding metaphor\-related word translations\. In addition, it will create a bilingual corpus to be shared publicly for research on English\-Chinese pairs with multiple references \(from multiple annotations\)\. We believe the multi\-reference parallel corpus will also serve as an valuable resource to the metaphor study field\.

For WP4 on automatic evaluation, we are attending the WMT2026 shared task on Test Suites111[https://www2\.statmt\.org/wmt26/testsuite\-subtask\.html](https://www2.statmt.org/wmt26/testsuite-subtask.html), where we used the AlphaMWEHanet al\.\([2026b](https://arxiv.org/html/2608.04260#bib.bib38),[2020](https://arxiv.org/html/2608.04260#bib.bib4)\)multilingual parallel corpus with MWE annotations including metaphors and idioms\. We have received 30\+ MT system submissions on the translation of this corpus that are being analysed using both automatic metrics and human evaluations\. The automatic metrics being deployed include BLEU, CharF\+\+, LEPOR, and BERTscore, covering both lexical\-based and word\-embedding metrics\.

Our next step on the thesis moving forward is about end\-to\-end metaphor detection and translation evaluation modelling, i\.e\., part of WP4\. For the end\-to\-end models, we will investigate “LLMs as Agents”Plaatet al\.\([2025](https://arxiv.org/html/2608.04260#bib.bib3)\), applying different LLMs for detection and translation evaluation from two joint phases in a collective manner\. LLM\-as\-a\-judge will also be explored as the evaluation phase with human\-in\-the\-loop, as suggested from the literatureWitet al\.\([2026](https://arxiv.org/html/2608.04260#bib.bib2)\)\.

Overall, the four WPs are coherently connected: outputs of WP1 feed into WP2; WP2 provides resources for WP3; and WP3 supervises WP4\.

## 6Research Plans

The Thesis Research Plan is listed in Figure[3](https://arxiv.org/html/2608.04260#S6.F3)\.

Milestone 1Metaphor DetectionMilestone 2TranslationMilestone 3Automatic EvaluationMilestone 4Joint Framework\+ ThesisPhase 1Phase 2Phase 3Phase 4Figure 3:Proposed research timeline and milestones for the PhD project\.
## 7Expected Contributions

Expected outputs include:

- •Resources - –multilingual corpus - –benchmark - –annotation guidelines
- •Methods - –MetaHOPE - –automatic evaluator - –joint framework
- •Scientific contributions - –end\-to\-end metaphor processing paradigm

## 8Conclusion

This PhD proposal presents a research agenda towards end\-to\-end multilingual metaphor processing by integrating metaphor detection, metaphor translation, human evaluation, and automatic translation evaluation into a unified computational framework\. While these research areas have each advanced considerably in recent years, they have largely been investigated independently\. As a result, current multilingual NLP systems still face substantial challenges when processing metaphorical language, particularly across languages and cultures\.

The proposed research addresses this gap through four complementary work packages\. The first investigates robust multilingual metaphor detection using both linguistic knowledge and recent large language models\. The second examines metaphor translation and develops benchmark resources for analysing metaphor translation strategies across machine translation and LLM systems\. The third develops a reliable human evaluation methodology for metaphor translation, including annotation guidelines, corpora, and empirical studies of annotation reliability\. Finally, the fourth explores automatic metaphor translation evaluation by integrating metaphor detection, human evaluation, and LLM\-based quality estimation into a unified framework\.

Beyond developing individual methods and datasets, this research aims to establish stronger connections between metaphor detection and translation evaluation\. Information produced during metaphor detection has the potential to support more accurate translation quality estimation, while human evaluation can provide supervision for developing metaphor\-aware automatic evaluation models\. By investigating these interactions, the proposed research seeks to move beyond isolated task\-specific solutions towards a coherent framework for multilingual metaphor processing\.

The expected outcomes include new multilingual resources, reproducible annotation methodologies, benchmark datasets, and automatic evaluation techniques that facilitate future research on figurative language processing\. More broadly, the proposed work contributes to the development of multilingual NLP systems that better capture the linguistic, conceptual, and cultural dimensions of metaphor, thereby supporting more reliable machine translation and evaluation of figurative language\.

## Current Progress

- •Completed: GPT\-4 metaphor detection study; MetaHOPE framework\.
- •In progress: Large\-scale MetaHOPE annotation; WMT 2026 Test Suites evaluation\.
- •Planned: Joint metaphor detection and evaluation models; LLM\-as\-a\-judge framework\.

## References

- METEOR: an automatic metric for MT evaluation with improved correlation with human judgments\.InProceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization,J\. Goldstein, A\. Lavie, C\. Lin, and C\. Voss \(Eds\.\),Ann Arbor, Michigan,pp\. 65–72\.External Links:[Link](https://aclanthology.org/W05-0909/)Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p4.1),[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p1.1)\.
- M\. Fuoli, W\. Huang, J\. Littlemore, S\. Turner, and E\. Wilding \(2026\)Metaphor identification using large language models: a comparison of rag, prompt engineering, and fine\-tuning\.Applied Corpus Linguistics,pp\. 100204\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.
- M\. Ge, R\. Mao, and E\. Cambria \(2023\)A survey on computational metaphor processing techniques: from identification, interpretation, generation to application\.Artificial Intelligence Review56\(Suppl 2\),pp\. 1829–1895\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1)\.
- R\. W\. Gibbs \(2008\)Metaphor and thought: the state of the art\.InThe Cambridge Handbook of Metaphor and Thought,Jr\. Gibbs \(Ed\.\),Cambridge Handbooks in Psychology,pp\. 3–14\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511816802.002)Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1)\.
- S\. Gladkoff, L\. Han, and K\. Gasova \(2025\)Non\-linear scoring model for translation quality evaluation\.arXiv preprint arXiv:2511\.13467\.Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p3.1)\.
- S\. Gladkoff and L\. Han \(2022\)HOPE: a task\-oriented and human\-centric evaluation framework using professional post\-editing towards more effective mt evaluation\.InProceedings of the Thirteenth Language Resources and Evaluation Conference,pp\. 13–21\.Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p3.1)\.
- P\. Group \(2007\)MIP: a method for identifying metaphorically used words in discourse\.Metaphor and symbol22\(1\),pp\. 1–39\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p2.1)\.
- A\. L\. F\. Han, D\. F\. Wong, and L\. S\. Chao \(2012\)LEPOR: a robust evaluation metric for machine translation with augmented factors\.InProceedings of COLING 2012: Posters,Mumbai, India,pp\. 441–450\.External Links:[Link](https://aclanthology.org/C12-2044/)Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p1.1)\.
- A\. L\.\-F\. Han, D\. F\. Wong, L\. S\. Chao, L\. He, Y\. Lu, J\. Xing, and X\. Zeng \(2013\)Language\-independent model for machine translation evaluation with reinforced factors\.InMachine Translation Summit XIV,pp\. 215–222\.Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p1.1)\.
- L\. Han, G\. Jones, and A\. Smeaton \(2020\)AlphaMWE: construction of multilingual parallel corpora with MWE annotations\.InProceedings of the Joint Workshop on Multiword Expressions and Electronic Lexicons,S\. Markantonatou, J\. McCrae, J\. Mitrović, C\. Tiberius, C\. Ramisch, A\. Vaidya, P\. Osenova, and A\. Savary \(Eds\.\),online,pp\. 44–57\.External Links:[Link](https://aclanthology.org/2020.mwe-1.6/)Cited by:[§5](https://arxiv.org/html/2608.04260#S5.p3.1)\.
- L\. Han, D\. Lindevelt, S\. Puts, E\. van Mulligen, and S\. Verberne \(2026a\)Dutch metaphor extraction from cancer patients’ interviews and forum data using llms and human in the loop\.CL4Health WS at LREC2026, Palma, Spain\.Cited by:[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p4.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.
- L\. Han, N\. H\. Mohamed, M\. Rassem, G\. J\. Jones, A\. F\. Smeaton, and G\. Nenadic \(2026b\)Towards a resource for multilingual lexicons: an mt assisted and human\-in\-the\-loop multilingual parallel corpus with multi\-word expression annotation\.Language Resources and Evaluation60\(2\),pp\. 33\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p2.1),[§5](https://arxiv.org/html/2608.04260#S5.p3.1)\.
- A\. Hülsing and S\. S\. Im Walde \(2024\)Cross\-lingual metaphor detection for low\-resource languages\.InProceedings of the 4th Workshop on Figurative Language Processing \(FigLang 2024\),pp\. 22–34\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.
- M\. Johnson, M\. Schuster, Q\. V\. Le, M\. Krikun, Y\. Wu, Z\. Chen, N\. Thorat, F\. Viégas, M\. Wattenberg, G\. Corrado, M\. Hughes, and J\. Dean \(2017\)Google’s multilingual neural machine translation system: enabling zero\-shot translation\.Transactions of the Association for Computational Linguistics5,pp\. 339–351\.External Links:[Link](https://aclanthology.org/Q17-1024/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00065)Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p2.1)\.
- A\. Karakanta, M\. Nas, and A\. G\. Dorst \(2025\)Metaphors in literary machine translation: close but no cigar?\.InProceedings of Machine Translation Summit XX: Volume 1,pp\. 276–286\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p3.1)\.
- T\. Kocmi, E\. Artemova, E\. Avramidis, R\. Bawden, O\. Bojar, K\. Dranch, A\. Dvorkovich, S\. Dukanov, M\. Fishel, M\. Freitag,et al\.\(2025\)Findings of the wmt25 general machine translation shared task: time to stop evaluating on easy test sets\.InProceedings of the Tenth Conference on Machine Translation,pp\. 355–413\.Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p2.1)\.
- Z\. Kövecses \(2010\)Metaphor and culture\.Acta Universitatis Sapientiae, Philologica2\(2\),pp\. 197–220\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p1.1),[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p4.1)\.
- H\. Lai, A\. Toral, and M\. Nissim \(2023\)Multilingual multi\-figurative language detection\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 9254–9267\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.
- G\. Lakoff and M\. Johnson \(1980\)Metaphors we live by\.Vol\.1,University of Chicago press Chicago\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p1.1),[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p1.1)\.
- C\. W\. Leong, B\. B\. Klebanov, C\. Hamill, E\. Stemle, R\. Ubale, and X\. Chen \(2020\)A report on the 2020 vua and toefl metaphor detection shared task\.InProceedings of the second workshop on figurative language processing,pp\. 18–29\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p3.1)\.
- C\. W\. Leong, B\. B\. Klebanov, and E\. Shutova \(2018\)A report on the 2018 vua metaphor detection shared task\.InProceedings of the workshop on figurative language processing,pp\. 56–66\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p3.1)\.
- Z\. Li and L\. Chen \(2025a\)Mind vs\. machine: comparative analysis of metaphor\-related word translation by human and ai systems\.Training, Language and Culture9\(1\),pp\. 10–27\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1)\.
- Z\. Li and L\. Chen \(2025b\)Mind vs\. machine: comparative analysis of metaphorrelated word translation by human and ai systems\.Training, Language and Culture9\(1\),pp\. 10–27\.Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p3.1)\.
- J\. Liang, A\. G\. Dorst, J\. Prokic, and S\. Raaijmakers \(2025\)Using gpt\-4 for conventional metaphor detection in english news texts\.Computational Linguistics in the Netherlands Journal14,pp\. 307–341\.External Links:[Link](https://hdl.handle.net/1887/4259126)Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p6.1),[§5](https://arxiv.org/html/2608.04260#S5.p1.1)\.
- J\. Liang and L\. Han \(2026\)MetaHOPE: a metaphor\-oriented evaluation framework for analysing mt and llm translation errors\.arXiv preprint arXiv:2607\.00848\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§1](https://arxiv.org/html/2608.04260#S1.p4.1),[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p2.1),[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p4.1),[§4](https://arxiv.org/html/2608.04260#S4.p3.1),[§5](https://arxiv.org/html/2608.04260#S5.p2.1)\.
- A\. Lommel, S\. Gladkoff, A\. Melby, S\. E\. Wright, I\. Strandvik, K\. Gasova, A\. Vaasa, A\. Benzo, R\. Marazzato Sparano, M\. Foresi, J\. Innis, L\. Han, and G\. Nenadic \(2024\)The multi\-range theory of translation quality measurement: MQM scoring models and statistical quality control\.InProceedings of the 16th Conference of the Association for Machine Translation in the Americas \(Volume 2: Presentations\),M\. Martindale, J\. Campbell, K\. Savenkov, and S\. Goel \(Eds\.\),Chicago, USA,pp\. 75–94\.External Links:[Link](https://aclanthology.org/2024.amta-presentations.6/)Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p3.1)\.
- A\. Lommel, H\. Uszkoreit, and A\. Burchardt \(2014\)Multidimensional quality metrics \(mqm\): a framework for declaring and describing translation quality metrics\.Tradumàtica\(12\),pp\. 0455–463\.Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p3.1)\.
- P\. Newmark \(1988\)A textbook of translation\.Prentice Hall,London\.Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p1.1)\.
- K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu \(2002\)Bleu: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,P\. Isabelle, E\. Charniak, and D\. Lin \(Eds\.\),Philadelphia, Pennsylvania, USA,pp\. 311–318\.External Links:[Link](https://aclanthology.org/P02-1040/),[Document](https://dx.doi.org/10.3115/1073083.1073135)Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p1.1)\.
- A\. Plaat, M\. Van Duijn, N\. Van Stein, M\. Preuss, P\. Van der Putten, and K\. J\. Batenburg \(2025\)Agentic large language models, a survey\.Journal of Artificial Intelligence Research84\.Cited by:[§5](https://arxiv.org/html/2608.04260#S5.p4.1)\.
- S\. Rai and S\. Chakraverty \(2020\)A survey on computational metaphor processing\.ACM Computing Surveys \(CSUR\)53\(2\),pp\. 1–37\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p1.1)\.
- R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. Lavie \(2020\)COMET: a neural framework for MT evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2685–2702\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.213/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.213)Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p4.1),[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p2.1)\.
- S\. Reimann and T\. Scheffler \(2025\)Using large language models to perform mipvu\-inspired automatic metaphor detection\.InThe 2nd Workshop on Analogical Abstraction in Cognition, Perception, and Language \(Analogy\-Angle II\),pp\. 10\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.
- E\. Sanchez\-Bayona and R\. Agerri \(2025\)Metaphor and large language models: when surface features matter more than deep understanding\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 17462–17477\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p3.1)\.
- C\. Schäffner \(2004\)Metaphor and translation: some implications of a cognitive approach\.Journal of pragmatics36\(7\),pp\. 1253–1269\.Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p1.1)\.
- T\. Sellam, D\. Das, and A\. Parikh \(2020\)BLEURT: learning robust metrics for text generation\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 7881–7892\.External Links:[Link](https://aclanthology.org/2020.acl-main.704/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.704)Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p2.1)\.
- E\. Semino \(2008\)Metaphor in discourse\.Cambridge University Press Cambridge\.Cited by:[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p2.1)\.
- E\. Shutova \(2010\)Models of metaphor in nlp\.InProceedings of the 48th annual meeting of the association for computational linguistics,pp\. 688–697\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p2.1)\.
- E\. Shutova \(2015\)Design and evaluation of metaphor processing systems\.Computational Linguistics41\(4\),pp\. 579–623\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p1.1)\.
- M\. Snover, B\. Dorr, R\. Schwartz, L\. Micciulla, and J\. Makhoul \(2006\)A study of translation edit rate with targeted human annotation\.InProceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers,Cambridge, Massachusetts, USA,pp\. 223–231\.External Links:[Link](https://aclanthology.org/2006.amta-papers.25/)Cited by:[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p1.1)\.
- G\. J\. Steen, A\. G\. Dorst, J\. B\. Herrmann, A\. A\. Kaal, and T\. Krennmayr \(2010a\)Vu amsterdam metaphor corpus\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p2.1)\.
- G\. J\. Steen, A\. G\. Dorst, T\. Krennmayr, A\. A\. Kaal, and J\. B\. Herrmann \(2010b\)A method for linguistic metaphor identification\.Cited by:[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p2.1)\.
- G\. Steen \(2017\)Deliberate metaphor theory: basic assumptions, main tenets, urgent issues\.Intercultural Pragmatics14\(1\),pp\. 1–24\.Cited by:[§2\.1](https://arxiv.org/html/2608.04260#S2.SS1.p2.1)\.
- R\. van den Broeck \(1981\)The limits of translatability exemplified by metaphor translation\.Poetics Today2\(4\),pp\. 73–87\.External Links:[Document](https://dx.doi.org/10.2307/1772487)Cited by:[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p1.1)\.
- S\. Wang, G\. Zhang, H\. Wu, T\. Loakman, W\. Huang, and C\. Lin \(2024\)MMTE: corpus and metrics for evaluating machine translation quality of metaphorical language\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 11343–11358\.Cited by:[§1](https://arxiv.org/html/2608.04260#S1.p3.1),[§1](https://arxiv.org/html/2608.04260#S1.p4.1),[§2\.3](https://arxiv.org/html/2608.04260#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2608.04260#S2.SS4.p4.1)\.
- T\. Wit, L\. Han, C\. Heipon, D\. Lindevelt, A\. Stiggelbout, and S\. Verberne \(2026\)Measuring the practice of shared\-decision making \(option12\): an investigation into open\-sourced smaller llms \(os\-sllms\) for better privacy and sustainability\.External Links:2607\.06127,[Link](https://arxiv.org/abs/2607.06127)Cited by:[§5](https://arxiv.org/html/2608.04260#S5.p4.1)\.
- Y\. Xu, Y\. Hua, S\. Li, and Z\. Wang \(2024\)Exploring chain\-of\-thought for multi\-modal metaphor detection\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 91–101\.Cited by:[§2\.2](https://arxiv.org/html/2608.04260#S2.SS2.p4.1)\.

## Appendix AInitial Annotation Guidelines

Two files: “Annotation Guideline for The Categorization of Potentially Universal Metaphors and Culture” and “Annotation Guidelines for Metaphor Translation Strategies and Quality Evaluation” both available at[https://github\.com/Jiahui84/MetaHOPE](https://github.com/Jiahui84/MetaHOPE)

Similar Articles

MetaphorVU: Towards Metaphorical Video Understanding

Hugging Face Daily Papers

This paper introduces MetaphorVU-Bench, the first systematic benchmark for metaphorical video understanding, and proposes MetaphorBoost, an inference-time enhancement framework that improves cross-domain mapping in multimodal large language models.