Tag
This research investigates whether fine-tuning large language models on cultural data improves figurative language understanding and vice versa, finding that while poetry fine-tuning enhances idiom comprehension, cultural fine-tuning can reduce proverb accuracy, highlighting a non-straightforward relationship.
This paper introduces a benchmark dataset of 1,516 expert-verified Bangla sentences for disambiguating culturally entangled homographs (words that are both names and common nouns). It shows that LLMs suffer from dominant-meaning bias and proposes contrastive chain-of-thought prompting and distillation to reduce this bias.
This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.
This paper introduces the Proverb Aligned Narrative Dataset (PAND) for studying proverb-conditioned story generation in LLMs, finding a 'decompression gap' where models produce fluent stories but fail to faithfully capture the proverbs' moral and causal structure.
This paper proposes a self-supervised framework using multilingual self-consistency and a self-critique mechanism to transfer cultural knowledge across languages, achieving a 5.03% average improvement on English queries in the BLEnD benchmark by surfacing latent cultural knowledge from local-language representations.