Tag
This research investigates whether fine-tuning large language models on cultural data improves figurative language understanding and vice versa, finding that while poetry fine-tuning enhances idiom comprehension, cultural fine-tuning can reduce proverb accuracy, highlighting a non-straightforward relationship.
This paper studies how multilingual training helps identify figurative language in proverbs across seven languages, introducing a multidimensional annotation framework and finding that about 50% of translated multilingual data is sufficient for near-optimal performance.
This paper introduces ProverbIT, a novel Italian benchmark of 100 multiple-choice questions to test LLMs' ability to complete proverbs. Evaluating 13 models, it finds that performance drops significantly in multiple-choice formats without correct answers, suggesting reliance on memorized patterns rather than deep semantic understanding.
This paper introduces VIVID, the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese, comprising 1,636 idioms and proverbs. Evaluation of eight state-of-the-art models reveals significant gaps, with Vietnamese-specialized models drastically underperforming multilingual systems and even top models achieving less than 50% correctness on average.
This paper investigates how large language models handle the combination of negation and figurative language, finding that this combination poses a particular challenge and that performance depends heavily on prompt style. The authors develop new annotations for the Fig-QA dataset and analyze embedding spaces to uncover additional linguistic factors like tense and concreteness.
IdiomX is a large-scale multilingual benchmark for idiom understanding, retrieval, and interpretation, containing over 190K examples across English, Arabic, and French, with four tasks for evaluating language models on idiomatic expressions.
This paper explores cross-lingual transfer of internal representations for figurative language generation in multilingual LLMs, showing that activation directions learned in one language can effectively steer generation in other languages.