ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

arXiv cs.CL Papers

Summary

This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.

arXiv:2608.03358v1 Announce Type: new Abstract: Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.
Original Article
View Cached Full Text

Cached at: 08/05/26, 07:45 AM

# ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models
Source: [https://arxiv.org/html/2608.03358](https://arxiv.org/html/2608.03358)
###### Abstract

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception\. We introduce culture\-conditioned visual emotion understanding, a task that predicts the culture\-specific emotional perception of a given image and explains the underlying rationale\. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority\-supported culture\-level emotion labels, and imbalanced cultural coverage\. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture\-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non\-Western content\. Evaluations of 16 open\- and closed\-source Multimodal Large Language Models \(MLLMs\) under a zero\-shot setting reveal that the task remains challenging, with the best model achieving below 50% accuracy\. To address this limitation, we introduce a retrieval\-augmented culture\-conditioned emotion understanding framework, which leverages a concept\-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training\. The framework improves both culturally aligned emotion prediction and grounded explanation generation\. Our benchmark and code will be publicly released\.

## Introduction

As Multimodal Large Language Models \(MLLMs\) become integrated into conversational assistants, human\-AI interactions are evolving from text\-centric toward multimodal exchanges, where assistants can retrieve and generate images in their responses\. Images, as a rich medium for human communication, can evoke diverse emotional perceptions across users, shaping their experiences and interactions with AI systems\. Thus, enabling MLLMs to understand such perceptions is essential for building more empathetic and user\-aligned assistants\. However, existing visual emotion understanding methods\(Aslanet al\.[2022](https://arxiv.org/html/2608.03358#bib.bib16); Yanget al\.[2023](https://arxiv.org/html/2608.03358#bib.bib14); Bhattacharyya and Wang[2025](https://arxiv.org/html/2608.03358#bib.bib15)\)assume that images evoke universal emotions across viewers, overlooking the cultural factors that shape emotional perception\(Kitayama and Cohen[2019](https://arxiv.org/html/2608.03358#bib.bib19)\)\. As shown in Figure[1](https://arxiv.org/html/2608.03358#Sx1.F1), the same image may evokesadnessamong Chinese viewers, who interpret the solitary moon\-watching person as loneliness, butcontentmentamong English and Arabic viewers\.

![Refer to caption](https://arxiv.org/html/2608.03358v1/x1.png)Figure 1:Example from ArtECulture\. Viewers from different cultures perceive different emotions from the same image and provide corresponding culture\-specific explanations\.![Refer to caption](https://arxiv.org/html/2608.03358v1/x2.png)Figure 2:Proposed paradigms:retrieval\-augmented culture\-conditioned emotion understandingandsupervised fine\-tuning\.Recent studies have started to incorporate cultural factors into emotion understanding\. Early works introduce benchmarks, e\.g\., ArtELingo\(Mohamedet al\.[2022a](https://arxiv.org/html/2608.03358#bib.bib10)\)and ArtELingo\-28\(Mohamedet al\.[2024](https://arxiv.org/html/2608.03358#bib.bib11)\), which provide culturally diverse emotional interpretations of images across languages and cultures, with each image paired with an emotion label and an explanatory caption\. However, these tasks assume that the target emotion is given and focus on emotion\-conditioned caption generation rather than emotion prediction\. More recent works, including CuLEmo\(Belayet al\.[2025](https://arxiv.org/html/2608.03358#bib.bib13)\)and CEDAR\(Daiet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib12)\), investigate culture\-conditioned emotion prediction\. Nevertheless, they mainly study textual contexts rather than visual perception: CuLEmo is purely text\-based, while CEDAR’s limited multimodal subset \(∼400\\sim 400images\) uses images as illustrations of textual narratives, capturing emotions inferred from descriptions rather than those perceived from visual content\. Thus, we define the task ofculture\-conditioned visual emotion understanding, which requires predicting the emotion perceived by viewers from a given culture when observing an image and generating the corresponding cultural rationale\.

Evaluating this task requires a benchmark with two key properties\. First, each image should have culture\-level emotion labels that capture the dominant emotional perception within each culture, enabling evaluation of culturally shared patterns and cross\-cultural differences\. Second, the benchmark should maintain balanced cultural coverage of visual content to avoid evaluation bias toward dominant cultures\. ArtELingo is the closest existing resource, providing large\-scale culturally diverse image\-emotion\-caption pairs that can potentially support our task\. However, it does not satisfy these requirements\. Its individual\-level annotations do not yield culture\-level labels, as only 40\.22% of Arabic images and 43\.13% of English images have a majority\-perceived emotion\. Moreover, 79\.75% of its images depict Western art, resulting in imbalanced cultural coverage\.

We therefore construct ArtECulture, the first benchmark satisfying both requirements, by reorganizing ArtELingo and adding culturally diverse artworks with new annotations across English, Chinese, and Arabic cultures\. Specifically, we retain ArtELingo samples with dominant emotional perceptions within each culture and collect additional non\-Western artworks annotated by native speakers to achieve balanced cultural coverage\. Finally, ArtECulture contains 6,792 artworks and 92,062 culture\-specific emotion\-explanation annotations across English, Chinese, and Arabic cultures, with a balanced distribution of Western \(55\.89%\) and non\-Western \(44\.11%\) content\.

Using ArtECulture, we evaluate1616open\- and closed\-source MLLMs and obtain two findings\. First, the task remains challenging, as the best\-performing model achieves an overall accuracy below50%50\\%\. Second, current MLLMs exhibit English\-centric affective priors, with every model performing best on English and nearly all performing worst on Arabic\. Motivated by these findings and the observation that cultural emotional perception is highly context\-dependent and requires culture\-specific knowledge, we propose a training\-free retrieval\-augmented framework that enhances MLLMs with concept\-level cultural emotion knowledge extracted from the training set of ArtECulture\. Specifically, we construct a cultural emotion knowledge base, where each entry associates a visual concept with culture\-specific emotions and rationales\. Given an image and target culture, the framework retrieves relevant concepts and their associated cultural knowledge to guide MLLMs in emotion prediction and explanation generation\. As a comparison, we also fine\-tune open\-source MLLMs on ArtECulture\. The two paradigms show a complementary trade\-off: fine\-tuning achieves the highest prediction accuracy but degrades explanation quality possibly because of the multiple\-reference training setup, whereas retrieval augmentation improves both prediction and explanations while enabling deployment on closed\-source models\.

Our contributions are summarized as follows:

- •We study culture\-conditioned visual emotion understanding, a new task that requires predicting the emotion an image evokes in a given culture and explaining why the emotion arises\. We present ArtECulture, the first benchmark for this task, in which every image carries an emotion label for each culture, with a balanced distribution of Western and non\-Western content\.
- •We evaluate1616MLLMs on ArtECulture and show that the task is far from solved and that every model is strongest on English, revealing pronounced English\-centric affective priors in current MLLMs\.
- •We build a cultural emotion knowledge base and propose a training\-free retrieval\-augmented pipeline that injects it into MLLMs\. Comparing it with supervised fine\-tuning reveals a complementary trade\-off between prediction accuracy and explanation quality\.

## Related Work

Visual emotion understandingpredicts the emotions evoked by visual content\. Existing datasets pair web or artistic images with discrete emotion labels\(Machajdik and Hanbury[2010](https://arxiv.org/html/2608.03358#bib.bib21); Borthet al\.[2013](https://arxiv.org/html/2608.03358#bib.bib22); Penget al\.[2015](https://arxiv.org/html/2608.03358#bib.bib23); Youet al\.[2016](https://arxiv.org/html/2608.03358#bib.bib24); Pandaet al\.[2018](https://arxiv.org/html/2608.03358#bib.bib25); Mertenset al\.[2024](https://arxiv.org/html/2608.03358#bib.bib26); Zhanget al\.[2025](https://arxiv.org/html/2608.03358#bib.bib40)\), explanations\(Mohammad and Kiritchenko[2018](https://arxiv.org/html/2608.03358#bib.bib27); Mathewset al\.[2016](https://arxiv.org/html/2608.03358#bib.bib43); Achlioptaset al\.[2021](https://arxiv.org/html/2608.03358#bib.bib28); Mohamedet al\.[2022b](https://arxiv.org/html/2608.03358#bib.bib29)\), or emotion\-related attributes\(Youet al\.[2017](https://arxiv.org/html/2608.03358#bib.bib44); Yanget al\.[2023](https://arxiv.org/html/2608.03358#bib.bib14)\)\. Beyond early hand\-crafted features\(Aslanet al\.[2022](https://arxiv.org/html/2608.03358#bib.bib16)\), recent work adapts MLLMs via instruction tuning\(Xieet al\.[2024](https://arxiv.org/html/2608.03358#bib.bib30); Yanget al\.[2024](https://arxiv.org/html/2608.03358#bib.bib31); Chenet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib38); Wuet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib39); Rhaet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib42)\)or benchmarks them on evoked emotions\(Bhattacharyya and Wang[2025](https://arxiv.org/html/2608.03358#bib.bib15)\)and emotional intelligence\(Huet al\.[2025](https://arxiv.org/html/2608.03358#bib.bib32)\)\. However, these efforts collapse annotations from different cultural backgrounds into a single label per image, overlooking the cultural dependence of emotional perception\.

Culture\-aware emotion understandingincorporates cultural factors into emotion understanding\. The first line studies affective captioning: ArtELingo\(Mohamedet al\.[2022a](https://arxiv.org/html/2608.03358#bib.bib10)\)and ArtELingo\-28\(Mohamedet al\.[2024](https://arxiv.org/html/2608.03358#bib.bib11)\)extend ArtEmis with multicultural annotations for affective captioning, where emotions are provided as inputs rather than prediction targets\. The second line predicts emotions under cultural or linguistic conditions from text\. Multilingual resources\(Muhammadet al\.[2025a](https://arxiv.org/html/2608.03358#bib.bib33),[b](https://arxiv.org/html/2608.03358#bib.bib34)\)cover emotion detection in2828languages, but each text is tied to a single language and lacks cross\-cultural judgments of the same stimulus\. CuLEmo\(Belayet al\.[2025](https://arxiv.org/html/2608.03358#bib.bib13)\)examines cultural perception with text\-only inputs, and CEDAR\(Daiet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib12)\)uses a limited image subset illustrating textual narratives\. These benchmarks focus on single\-label evaluation without explanations and leave model adaptation unexplored\. In contrast, we target culture\-conditioned visual emotion perception, requiring both emotion prediction and cultural rationales, and investigate how MLLMs can be enhanced for this task\.

Cultural understanding of MLLMsexamines whether MLLMs hold knowledge about diverse cultures\. Benchmarks\(Liuet al\.[2021](https://arxiv.org/html/2608.03358#bib.bib35); Romeroet al\.[2024](https://arxiv.org/html/2608.03358#bib.bib3); Nayaket al\.[2024](https://arxiv.org/html/2608.03358#bib.bib5); Vayaniet al\.[2024](https://arxiv.org/html/2608.03358#bib.bib36); Schneideret al\.[2025](https://arxiv.org/html/2608.03358#bib.bib9); Tanet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib4); Wanget al\.[2025](https://arxiv.org/html/2608.03358#bib.bib41)\)test factual cultural knowledge about food, clothing, rituals, and landmarks, revealing consistent gaps on non\-Western cultures, which follow\-up studies mitigate via cultural training data\(Nyandwiet al\.[2025](https://arxiv.org/html/2608.03358#bib.bib37)\)or retrieval\(Liet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib2); Lewiset al\.[2020](https://arxiv.org/html/2608.03358#bib.bib20)\)\. Unlike these works on objective cultural facts, we study subjective emotional perceptions elicited by visual content across cultures, with a knowledge base encoding concept\-level emotional conventions rather than factual knowledge\.

## ArtECulture Benchmark

We construct ArtECulture in two steps: re\-curating ArtELingo into a base pool with majority\-supported culture\-level labels, and collecting non\-Western artworks with new annotations to rebalance the regional distribution\.

Table 1:Detailed statistics of ArtECulture\.### Re\-curation of ArtELingo

Each ArtELingo image is annotated by multiple annotators from each culture\. We retain an image only when majority agreement is achieved separately within all three cultures, with the majority emotion in each culture assigned as its culture\-level label\. This yields4,9284\{,\}928images with emotion\-explanation annotations from all three cultures, comprising3,7963\{,\}796Western \(77\.03%77\.03\\%\) and1,1321\{,\}132non\-Western \(22\.97%22\.97\\%\) artworks\. However, the pool remains Western\-centric, requiring non\-Western artwork augmentation\.

### Non\-Western Artwork Augmentation

##### Collection\.

We collect additional non\-Western artworks from the public benchmark VULCA\-BENCH\(Yuet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib1)\), a multicultural art\-critique benchmark drawn from the open collections of authoritative museums\. In total, we obtain2,7832\{,\}783artworks spanning five major traditions, namely Chinese, Japanese, Korean, Indian, and Islamic art\.

##### Annotation\.

We recruit annotators from Prolific111https://www\.prolific\.com\.with the requirement that the target language is their first language\. Following ArtELingo\(Mohamedet al\.[2022a](https://arxiv.org/html/2608.03358#bib.bib10)\)and CEDAR\(Daiet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib12)\), native language serves as the criterion of cultural membership, and each culture\-level label reflects the dominant perception within the corresponding language community\. In each community, around70%70\\%of annotators reside in one dominant country \(69\.1%69\.1\\%in the UK for English,74\.6%74\.6\\%in China for Chinese, and73\.9%73\.9\\%in Egypt for Arabic\); full demographics are provided in the supplementary material\. As in ArtELingo, each annotator selects one emotion from nine categories \(*i\.e\.,*contentment,awe,amusement,sadness,fear,excitement,disgust,anger, andother\), and writes a short explanation in their native language\. For each image\-culture pair, we collect three independent annotations and adopt the majority emotion as the culture\-level label; if no majority is reached, two additional annotators are recruited and the vote is taken over the five annotations, discarding images that still lack a majority\. Importantly, annotations are collected independently within each culture to preserve cross\-cultural divergence\. This process retains1,8641\{,\}864of the2,7832\{,\}783collected images, yielding17,52617\{,\}526emotion\-explanation annotations across the three cultures\.

### Statistics and Insights

##### Dataset Statistics\.

Table[1](https://arxiv.org/html/2608.03358#Sx3.T1)summarizes the statistics of ArtECulture\. ArtECulture contains6,7926\{,\}792artworks, where4,9284\{,\}928are retained from ArtELingo and1,8641\{,\}864are newly collected, raising the non\-Western proportion from20\.25%20\.25\\%in ArtELingo to44\.11%44\.11\\%\. It provides92,06292\{,\}062emotion\-explanation annotations, averaging about4\.54\.5annotations per image per culture, with a roughly balanced distribution across English \(30,93230\{,\}932\), Chinese \(31,55931\{,\}559\), and Arabic \(29,57129\{,\}571\)\.

##### Agreement Analysis\.

Within each language community, nominal Krippendorff’sα\\alphais0\.3090\.309,0\.3240\.324, and0\.3060\.306for English, Chinese, and Arabic, comparable to ArtELingo\(Mohamedet al\.[2022a](https://arxiv.org/html/2608.03358#bib.bib10)\)\. On average, the majority emotion receives70\.81%70\.81\\%,76\.32%76\.32\\%, and70\.13%70\.13\\%of the votes per image, respectively\. In contrast, only924924images \(13\.60%13\.60\\%\) receive the same emotion across all three cultures, while1,3811\{,\}381\(20\.33%20\.33\\%\) exhibit three distinct emotions \(Fleiss’κ=0\.088\\kappa=0\.088treating three culture\-level labels as raters\)\. The contrast between within\- and cross\-culture agreement confirms the labels capture consistent perceptions within each community that genuinely diverge across cultures, motivating culture\-conditioned visual emotion understanding\.

##### Data Split\.

The samples are split into training, validation, and test sets in an 8:1:1 ratio, preserving the overall Western/non\-Western distribution\. Since each image carries labels from all three cultures, the same split is applied across cultures\. Detailed statistics and analysis of ArtECulture are provided in the supplementary material\.

## Method

### Problem Formulation

Suppose we have a set of images𝒱=\{v1,v2,⋯,vM\}\\mathcal\{V\}=\\\{v\_\{1\},v\_\{2\},\\cdots,v\_\{M\}\\\}and a set of cultural groups𝒞=\{c1,c2,⋯,cN\}\\mathcal\{C\}=\\\{c\_\{1\},c\_\{2\},\\cdots,c\_\{N\}\\\}, whereMMandNNare the number of images and cultures, respectively\. Eachcjc\_\{j\}is operationalized as a native\-speaker language community following prior work\(Mohamedet al\.[2022a](https://arxiv.org/html/2608.03358#bib.bib10); Daiet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib12)\)\. All cultures share a predefined set of discrete emotion categoriesℰ=\{el\}l=1L\\mathcal\{E\}=\\\{e\_\{l\}\\\}\_\{l=1\}^\{L\}, whereLLis the number of categories\. For each culturecjc\_\{j\}, every imageviv\_\{i\}is associated with a culture\-level emotion labeleij∈ℰe\_\{i\}^\{j\}\\in\\mathcal\{E\}and a set of reference explanations𝒮ij=\{si,kj\}k=1Kij\\mathcal\{S\}\_\{i\}^\{j\}=\\\{s\_\{i,k\}^\{j\}\\\}\_\{k=1\}^\{K\_\{i\}^\{j\}\}that support the assigned emotion label\. Given an imageviv\_\{i\}and target culturecjc\_\{j\}, the goal is to build a modelℱ\\mathcal\{F\}that predicts the perceived emotione^ij\\hat\{e\}\_\{i\}^\{j\}and generates its supporting explanations^ij\\hat\{s\}\_\{i\}^\{j\}\.

### Retrieval\-augmented Culture\-Conditioned Emotion Understanding

Existing MLLMs still struggle with culture\-conditioned emotion understanding without access to explicit cultural knowledge, as our experiments will demonstrate, motivating the use of external cultural knowledge\. Inspired by retrieval\-augmented generation\(Lewiset al\.[2020](https://arxiv.org/html/2608.03358#bib.bib20)\), we propose retrieval\-augmented culture\-conditioned emotion understanding, a training\-free framework that augments a frozen MLLM with retrieved cultural emotion knowledge\.

#### Cultural Emotion Knowledge Base\.

Since cultural differences in emotional responses are often reflected through key visual concepts \(*e\.g\.,*moon\) that carry culture\-specific meanings, effective cultural emotion knowledge should be concept\-grounded, capturing both associated emotions and their underlying cultural rationales\. Accordingly, we define each knowledge entry as a tuple\(concept,culture,dominant emotion,rationales\)\(\\textit\{concept\},\\,\\textit\{culture\},\\,\\textit\{dominant emotion\},\\,\\textit\{rationales\}\)\. We construct the knowledge base from the training set of ArtECulture through the following four steps\.

1\)MLLM\-Based Concept Extraction\.This step collects visual concepts that may carry cultural meanings\. We prompt an MLLM \(Gemini 3\.5 Flash\) to list salient concepts of each training image, covering objects, colors, scenes, and cultural motifs, e\.g\., “full moon”, and “crane”\. Merging concepts of all images gives a concept vocabulary𝒬=\{q1,…,qR\}\\mathcal\{Q\}=\\\{q\_\{1\},\\dots,q\_\{R\}\\\}, whereqrq\_\{r\}is therr\-th concept andRRis the vocabulary size\.

2\)Frequency\-Driven Concept\-Emotion Association\.This step mines concept\-emotion pairs based on co\-occurrence frequency statistics\. Let𝒩r\\mathcal\{N\}\_\{r\}denote the set of training images that contain conceptqrq\_\{r\}\. For a culturecjc\_\{j\}, the probability thatqrq\_\{r\}evokes an emotione∈ℰe\\in\\mathcal\{E\}is computed as the fraction of images in𝒩r\\mathcal\{N\}\_\{r\}whose culture\-level label undercjc\_\{j\}isee,

Prj​\(e\)=\|\{vi∈𝒩r∣eij=e\}\|\|𝒩r\|,P\_\{r\}^\{\\,j\}\(e\)=\\frac\{\\bigl\|\\\{v\_\{i\}\\in\\mathcal\{N\}\_\{r\}\\mid e\_\{i\}^\{j\}=e\\\}\\bigr\|\}\{\\bigl\|\\mathcal\{N\}\_\{r\}\\bigr\|\},\(1\)whereeij∈ℰe\_\{i\}^\{j\}\\in\\mathcal\{E\}denotes the cultural emotion label of imageviv\_\{i\}for culturecjc\_\{j\}\. The dominant emotion ofqrq\_\{r\}incjc\_\{j\}is then the most probable one,e¯rj=arg⁡maxe∈ℰ⁡Prj​\(e\)\\bar\{e\}\_\{r\}^\{\\,j\}=\\arg\\max\_\{e\\in\\mathcal\{E\}\}P\_\{r\}^\{\\,j\}\(e\)\.

Table 2:Accuracy of original models \(Orig\.\), retrieval\-augmented framework \(\+Know\.\), and SFT\. The best result for each model within each culture is inbold\. Subscripts denote absolute gains over Orig\.3\)MLLM\-Based Rationale Generation\.This step produces cultural rationales for each concept’s dominant emotion\. For each conceptqrq\_\{r\}and culturecjc\_\{j\}, we prompt Gemini 3\.5 Flash with three inputs: the concept, target culture, and emotion distribution𝐏rj=\{Prj​\(e\)\}e∈ℰ\\mathbf\{P\}\_\{r\}^\{\\,j\}=\\\{P\_\{r\}^\{\\,j\}\(e\)\\\}\_\{e\\in\\mathcal\{E\}\}\. Given these inputs, the model identifies the dominant emotion and generates rationales explaining why this emotion is associated with the concept in the given culture,

ℛrj=M​L​L​M​\(qr,cj,𝐏rj\),\\mathcal\{R\}\_\{r\}^\{\\,j\}=MLLM\\bigl\(q\_\{r\},\\,c\_\{j\},\\,\\mathbf\{P\}\_\{r\}^\{\\,j\}\\bigr\),\(2\)whereℛrj\\mathcal\{R\}\_\{r\}^\{j\}denotes the generated rationale set\.

4\)Reverse Emotion Validation for Knowledge Filtering\.Since MLLM\-produced rationales are prone to hallucinations, we filter the tuples to preserve trustworthy knowledge\. Our core validation strategy relies on reverse emotion reasoning: a faithful cultural rationale should independently recover the matched emotion it explains\. To eliminate bias inherited from the rationale\-generating model, we adopt two MLLM judges \(GPT\-5\.5 and Claude Opus 4\.8\)\. We supply each judge solely with the rationale setℛrj\\mathcal\{R\}\_\{r\}^\{j\}, and prompt them to infer the underlying emotion\. A tuple is flagged if either judge predicts an emotion different from the dominant emotione¯rj\\bar\{e\}\_\{r\}^\{\\,j\}, and all flagged tuples are further reviewed by native\-speaker annotators for final verification\. The remaining tuples constitute the cultural emotion knowledge base𝒦\\mathcal\{K\}, which contains8,4398\{,\}439tuples and28,14628\{,\}146rationales, distributed evenly across the three cultures \(2,7842\{,\}784,2,9022\{,\}902, and2,7532\{,\}753tuples with9,5389\{,\}538,9,6269\{,\}626, and8,9828\{,\}982rationales for English, Chinese, and Arabic, respectively\)\. The resulting candidate tuples are represented as\(qr,cj,e¯rj,ℛrj\)\(q\_\{r\},c\_\{j\},\\bar\{e\}\_\{r\}^\{j\},\\mathcal\{R\}\_\{r\}^\{j\}\)\. Detailed statistics are provided in the supplementary material\.

#### Retrieval\-augmented Inference\.

Built on the knowledge base, we design a training\-free inference pipeline that proceeds in three steps: extracting the concepts of the test image, retrieving the matched tuples from the knowledge base, and predicting with the retrieved tuples as context\.

1\)Visual concept extraction\.Given a test imagevˇ\\check\{v\}and target culturecjc\_\{j\}, we prompt the frozen MLLM to extract salient concepts, obtaining a concept set𝒫ˇ=\{pˇ1,…,pˇn\}\\check\{\\mathcal\{P\}\}=\\\{\\check\{p\}\_\{1\},\\dots,\\check\{p\}\_\{n\}\\\}, wherenndenotes the number of extracted concepts\.

Table 3:Generalization to the unseen CEDAR benchmark: accuracy of Qwen3\.6\-27B with the two paradigms\. The best result is inboldand the second best isunderlined\.2\)Exact\-to\-semantic knowledge retrieval\.We then use the extracted concepts to query the knowledge base\. Since the MLLM may express a concept using different wording or granularity from the knowledge base, such as “bamboo grove” versus “bamboo”, exact matching may miss useful tuples\. Therefore, we perform two\-stage retrieval\. Given the target culturecjc\_\{j\}, for each conceptpˇh\\check\{p\}\_\{h\}, we first search𝒦j\\mathcal\{K\}\_\{j\}, the subset of knowledge tuples associated withcjc\_\{j\}, for an exact concept match\. If no exact match exists, we compute the semantic similarity betweenpˇh\\check\{p\}\_\{h\}and each tuple concept using a sentence embedding model, and select the most similar tuple if its similarity exceeds a thresholdτ\\tau; otherwise, we discardpˇh\\check\{p\}\_\{h\}\. The retrieved tuples form𝒯vˇj\\mathcal\{T\}\_\{\\check\{v\}\}^\{\\,j\}, which serves as the cultural prior of the imagevˇ\\check\{v\}under the culturecjc\_\{j\}\.

3\)Knowledge\-conditioned emotion understanding\.Finally, we convert retrieved tuples𝒯vˇj\\mathcal\{T\}\_\{\\check\{v\}\}^\{\\,j\}into a text sequence for prompt injection\. Conditioned on the image, target culture, and the available cultural knowledge, the frozen MLLM predicts the emotione^\\hat\{e\}and generates the explanations^\\hat\{s\},

\(e^,s^\)=M​L​L​M​\(vˇ,cj,𝒯vˇj\)\.\(\\hat\{e\},\\,\\hat\{s\}\)=MLLM\\bigl\(\\check\{v\},\\,c\_\{j\},\\,\\mathcal\{T\}\_\{\\check\{v\}\}^\{\\,j\}\\bigr\)\.\(3\)If no tuple is retrieved, the MLLM performs prediction without the injected cultural knowledge\.

### Supervised Fine\-tuning

In this paradigm, instead of retrieving knowledge at inference time, we directly fine\-tune open\-source MLLMs on the ArtECulture training split, allowing the cultural emotion knowledge to be implicitly absorbed into model parameters\. To leverage all available explanations, we construct a separate training instance for each explanation\. The input consists of the imageviv\_\{i\}and a prompt specifying the target culturecjc\_\{j\}and requesting an emotion label and explanation, while the target is the corresponding sequencey=\[eij;si,kj\]y=\[e\_\{i\}^\{j\};s\_\{i,k\}^\{j\}\]\. The model is fine\-tuned with the standard autoregressive objective and directly outputs the emotion and corresponding explanation for the given image and target culture at inference\.

## Experiments

### Experimental Setup

For emotion prediction, we report accuracy against culture\-level labels\. For explanation generation, we employ two MLLM judges, Claude Sonnet 5 \(C5\) and GLM\-4\.5V \(G4V\), to score each explanation onemotion alignment, which measures how well the explanation supports the predicted emotion of the target culture, on a five\-point scale\. Implementation details are in the supplementary material\.

Table 4:Ablation results on Qwen3\.6\-27B \(Accuracy %\)\.Table 5:Accuracy of Qwen3\.6\-27B under all combinations of prompt language \(rows\) and target culture \(columns\)\.Table 6:Explanation quality judged by Claude Sonnet 5 \(C5\) and GLM\-4\.5V \(G4V\) for Orig\., \+Know\., and SFT\.Table 7:Pairwise human evaluation of emotion alignment on100100randomly sampled test images for Qwen3\.6\-27B\.
### Culture\-Conditioned Emotion Prediction

##### Effect of the two paradigms\.

Table[2](https://arxiv.org/html/2608.03358#Sx4.T2)compares the two paradigms for the1616MLLMs in per\-culture and overall accuracy\. For reference, we also incorporate the zero\-shot performance, denoted as Orig\.\. 1\) Based on the zero\-shot performance, we noted that all MLLMs struggle with the task: even the best one, Gemini 3\.5 Flash, attains an overall accuracy of only49\.88%49\.88\\%, and most fall below48%48\\%\. Moreover, the performance is highly imbalanced across cultures\. Every model achieves its highest accuracy on English, while accuracy on Arabic is consistently the lowest for nearly all models, at best merely41\.53%41\.53\\%, suggesting that the affective conventions absorbed during pre\-training mainly come from English\-dominated corpora\. The finer breakdown performance of each culture on Western and non\-Western images reveals an asymmetric pattern, as detailed in the supplementary material\. 2\) \+Know\. yields consistent and substantial accuracy gains across all models without extra training, demonstrating that the knowledge base compensates for missing parametric cultural knowledge\. Notably, the gains of \+Know\. are highly uneven across cultures, concentrating on Chinese, moderate on Arabic, and marginal on English\. This pattern is consistent with the English\-centric priors: external knowledge is redundant for English but fills a substantial gap for the other cultures\. A possible reason why Chinese benefits more than Arabic is that the rationales are written by an MLLM whose cultural coverage is likely richer for Chinese than for Arabic\. 3\) SFT brings much larger gains, showing in\-domain training is currently the most effective way to acquire cultural knowledge for emotion prediction\.

##### Generalization of the two paradigms\.

For generalization evaluation, we evaluate both paradigms using Qwen3\.6\-27B on the multimodal subset of CEDAR\(Daiet al\.[2026](https://arxiv.org/html/2608.03358#bib.bib12)\), an unseen benchmark whose image styles and question formats differ from ArtECulture\. As shown in Table[3](https://arxiv.org/html/2608.03358#Sx4.T3), both paradigms improve the overall accuracy\. Retrieval augmentation raises accuracy on Chinese and Arabic but reduces it on English, likely because the model is already better aligned with English emotional conventions, leaving limited room for additional gains from retrieved knowledge\. In such cases, imperfectly matched retrieved knowledge may occasionally introduce noise\. SFT improves all three cultures and achieves the best overall accuracy, indicating that the fine\-tuned model learns transferable cultural emotion knowledge rather than dataset\-specific patterns\.

![Refer to caption](https://arxiv.org/html/2608.03358v1/x3.png)Figure 3:Case study of Orig\., \+Know\., and SFT with Qwen3\.6\-27B on three test images from different cultures\. Each case includes the ground\-truth emotion \(GT\), predictions, and explanations\. Arabic and Chinese outputs are translated into English\.
##### Ablation study\.

To validate our proposed retrieval\-augmented paradigm, we compare \+Know\. with three variants on Qwen3\.6\-27B in Table[4](https://arxiv.org/html/2608.03358#Sx5.T4): 1\)\+Internal\-Know\., which relies solely on its internal parametric knowledge with chain\-of\-thought \(CoT\) prompting for emotion reasoning without the external knowledge base; 2\)\+Random\-Know\., which replaces the retrieved tuples with randomly sampled ones; and 3\)\+Cross\-Culture\-Know\., which injects tuples retrieved for a mismatched culture\. We draw three observations\. 1\) \+Internal\-Know\. improves the overall accuracy by only0\.740\.74points and even reduces the Arabic accuracy, suggesting that chain\-of\-thought prompting can elicit useful internal reasoning but remains insufficient to provide culture\-specific emotion knowledge, particularly for underrepresented cultures\. 2\) \+Random\-Know\. reaches49\.39%49\.39\\%, as injecting any external knowledge biases the predictions toward the benchmark’s emotion prior, and \+Cross\-Culture\-Know\. performs slightly better at50\.07%50\.07\\%, since concepts often share the same dominant emotion across cultures\. 3\) \+Know\. outperforms \+Cross\-Culture\-Know\. by5\.555\.55points overall and10\.3110\.31points on Chinese, showing that a substantial part of the gain comes from the culture\-specific knowledge itself and cannot be attributed to the injection mechanism alone\.

##### Effect of prompt language\.

Since the main experiments use prompts in the target culture’s language, cross\-cultural performance gaps may partly come from the prompt language rather than from differences in cultural emotion knowledge\. To examine this, we evaluate Qwen3\.6\-27B under all nine combinations of prompt language and target culture in Table[5](https://arxiv.org/html/2608.03358#Sx5.T5)\. Accuracy is consistently highest under the English culture condition regardless of the prompt language, while changing the prompt language produces smaller and less uniform effects\. These patterns are consistent with an English\-centric language bias and a Western\-centric cultural knowledge bias in current MLLMs, indicating that the performance gap across cultures cannot be explained solely by prompt language\.

### Culture\-Conditioned Explanation Generation

We evaluate explanation generation with the two MLLM judges on all1616MLLMs, and a human study on a subset\.

##### Automatic evaluation\.

Table[6](https://arxiv.org/html/2608.03358#Sx5.T6)reports the evaluation results\. We highlight two findings\. 1\) \+Know\. improves emotion alignment for all 16 models under G4V and most models under C5, showing that retrieval\-augmented knowledge enhances both prediction accuracy and explanation quality\. This is because retrieved cultural knowledge provides emotion\-related associations that guide explanations toward the predicted emotion\. 2\) In contrast to its superiority in emotion prediction, SFT yields the lowest alignment scores among the three settings\. We investigate this degradation and find that the average explanation length of open\-source models drops from30\.5030\.50to15\.1215\.12words after SFT\. This is likely caused by the multiple\-reference training setup: each image\-culture pair is associated with multiple explanations, while each explanation is used as an independent target during SFT\. Under token\-level likelihood optimization, the model tends to converge to short, generic explanations that are compatible with multiple references, sacrificing image\- and culture\-specific details\. We further measure the agreement between two judges\. Ranking the1616models within each setting, the Spearman correlation between the two judges ranges from0\.8600\.860to0\.9050\.905\(allp<0\.001p<0\.001\), confirming the reliability of the automatic evaluation\.

##### Human evaluation\.

We further conduct a pairwise human study using Qwen3\.6\-27B\. For100100randomly sampled test images, native speakers compare two anonymized explanations in random order, and each pair is rated by three annotators per culture\. As shown in Table[7](https://arxiv.org/html/2608.03358#Sx5.T7), the human judgments follow the same overall trend as the automatic evaluation in Table[6](https://arxiv.org/html/2608.03358#Sx5.T6): \+Know\. outperforms Orig\. while SFT falls behind both\. Specifically, \+Know\. receives more wins than losses on Chinese \(41\.67%41\.67\\%against26\.00%26\.00\\%\) and Arabic \(38\.33%38\.33\\%against30\.67%30\.67\\%\)\. On English, \+Know\. and Orig\. are comparable, and most comparisons end in a tie \(65\.33%65\.33\\%\)\. This is consistent with the accuracy results in Table[2](https://arxiv.org/html/2608.03358#Sx4.T2): the model already captures English emotional conventions relatively well, so external knowledge brings limited gains\.

##### Case study\.

Figure[3](https://arxiv.org/html/2608.03358#Sx5.F3)compares Orig\., \+Know\., and SFT with Qwen3\.6\-27B on three test images, one per culture\. Orig\. mainly relies on visible content and overlooks the underlying cultural meanings, interpreting the guillotine scene as dark humor, the misty cliffs as mysterious grandeur, and the praying figure as fear of death\. In contrast, \+Know\. incorporates retrieved cultural knowledge and yields more culture\-aware interpretations, recognizing the guillotine as a symbol of violence and mortality, the ink\-wash scene as reflecting the traditional Chinese appreciation of peaceful living, and the monk’s meditation with earthy tones as spiritual calmness in Arab culture\. SFT predicts the correct emotion in all three cases, but its explanations become short and generic \(e\.g\., “the depiction of rocks and trees is very detailed”\), lacking cultural reasoning for the emotion\. This reveals the trade\-off: SFT improves prediction accuracy at the cost of generating less culturally grounded explanations\.

## Conclusion

We study culture\-conditioned visual emotion understanding, which requires predicting the emotion an image evokes in a given culture and explaining the underlying rationale\. To support this task, we construct ArtECulture, the first benchmark that provides culture\-level emotion labels for three cultures with a regionally balanced visual distribution\. We propose a training\-free retrieval\-augmented framework built on a concept\-level cultural emotion knowledge base and compare it with supervised fine\-tuning, revealing a complementary trade\-off between prediction accuracy and explanation quality\. Future work includes extending the benchmark to more cultures, adopting distribution\-level labels that preserve within\-group disagreement, and combining the strengths of both paradigms\.

## References

- P\. Achlioptas, M\. Ovsjanikov, K\. Haydarov, M\. Elhoseiny, and L\. Guibas \(2021\)ArtEmis: affective language for visual art\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 11564–11574\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- S\. Aslan, G\. Castellano, V\. Digeno, G\. Migailo, R\. Scaringi, and G\. Vessio \(2022\)Recognizing the emotions evoked by artworks through visual features and knowledge graph\-embeddings\.InImage Analysis and Processing,pp\. 129–140\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- T\. D\. Belay, A\. H\. Ahmed, A\. G\. II, I\. Ameer, G\. Sidorov, O\. Kolesnikova, and S\. M\. Yimam \(2025\)CULEMO: cultural lenses on emotion \- benchmarking llms for cross\-cultural emotion understanding\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 18894–18909\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p2.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1)\.
- S\. Bhattacharyya and J\. Z\. Wang \(2025\)Evaluating vision\-language models for emotion recognition\.InFindings of the Association for Computational Linguistics,pp\. 1798–1820\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- D\. Borth, R\. Ji, T\. Chen, T\. Breuel, and S\. Chang \(2013\)Large\-scale visual sentiment ontology and detectors using adjective noun pairs\.InProceedings of ACM International Conference on Multimedia,pp\. 223–232\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- T\. Chen, T\. Furusawa, Y\. Hirakawa, R\. Shimizu, F\. Mo, and T\. Wada \(2026\)MultiEmo\-bench: multi\-label visual emotion analysis for multi\-modal large language models\.ArXiv\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- C\. Dai, Y\. Shen, Z\. Gao, J\. Li, Y\. Jiang, Y\. Wang, L\. Liu, Z\. Ge, and J\. Hu \(2026\)Tears or cheers? benchmarking llms via culturally elicited distinct affective responses\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 38171–38196\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p2.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1),[Annotation\.](https://arxiv.org/html/2608.03358#Sx3.SSx2.SSS0.Px2.p1.7),[Problem Formulation](https://arxiv.org/html/2608.03358#Sx4.SSx1.p1.16),[Generalization of the two paradigms\.](https://arxiv.org/html/2608.03358#Sx5.SSx2.SSSx2.Px2.p1.1)\.
- H\. Hu, Y\. Zhou, L\. You, H\. Xu, Q\. Wang, Z\. Lian, F\. R\. Yu, F\. Ma, and L\. Cui \(2025\)EmoBench\-m: benchmarking emotional intelligence for multimodal large language models\.CoRRabs/2502\.04424\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- S\. Kitayama and D\. Cohen \(2019\)Handbook of cultural psychology\.Guilford Press\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p1.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. Kiela \(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InAdvances in Neural Information Processing Systems,pp\. 9459–9474\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1),[Retrieval\-augmented Culture\-Conditioned Emotion Understanding](https://arxiv.org/html/2608.03358#Sx4.SSx2.p1.1)\.
- J\. Li, Y\. Yuan, W\. Li, M\. Aliannejadi, D\. Hershcovich, A\. Søgaard, I\. Vulić, W\. Zhang, P\. P\. Liang, Y\. Deng, and S\. Belongie \(2026\)RAVENEA: a benchmark for multimodal retrieval\-augmented visual culture understanding\.InThe International Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- F\. Liu, E\. Bugliarello, E\. M\. Ponti, S\. Reddy, N\. Collier, and D\. Elliott \(2021\)Visually grounded reasoning across languages and cultures\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 10467–10485\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- J\. Machajdik and A\. Hanbury \(2010\)Affective image classification using features inspired by psychology and art theory\.InProceedings of the International Conference on Multimedia,pp\. 83–92\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- A\. P\. Mathews, L\. Xie, and X\. He \(2016\)SentiCap: Generating Image Descriptions with Sentiments\.InAAAI Conference on Artificial Intelligence,pp\. 3574–3580\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- L\. P\. Mertens, E\. Yargholi, H\. P\. O\. de Beeck, J\. V\. den Stock, and J\. Vennekens \(2024\)FindingEmo: an image dataset for emotion recognition in the wild\.InAdvances in Neural Information Processing Systems,Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- Y\. Mohamed, M\. Abdelfattah, S\. Alhuwaider, F\. Li, X\. Zhang, K\. Church, and M\. Elhoseiny \(2022a\)ArtELingo: A million emotion annotations of wikiart with emphasis on diversity over language and culture\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 8770–8785\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p2.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1),[Annotation\.](https://arxiv.org/html/2608.03358#Sx3.SSx2.SSS0.Px2.p1.7),[Agreement Analysis\.](https://arxiv.org/html/2608.03358#Sx3.SSx3.SSS0.Px2.p1.12),[Problem Formulation](https://arxiv.org/html/2608.03358#Sx4.SSx1.p1.16)\.
- Y\. Mohamed, F\. F\. Khan, K\. Haydarov, and M\. Elhoseiny \(2022b\)It is okay to not be okay: overcoming emotional bias in affective image captioning by contrastive data collection\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 21231–21240\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- Y\. Mohamed, R\. Li, I\. S\. Ahmad, K\. Haydarov, P\. Torr, K\. Church, and M\. Elhoseiny \(2024\)No culture left behind: artelingo\-28, a benchmark of wikiart with captions in 28 languages\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 20939–20962\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p2.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1)\.
- S\. Mohammad and S\. Kiritchenko \(2018\)WikiArt emotions: an annotated dataset of emotions evoked by art\.InProceedings of the International Conference on Language Resources and Evaluation,Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- S\. H\. Muhammad, N\. Ousidhoum, I\. Abdulmumin, J\. P\. Wahle, T\. Ruas, M\. Beloucif, C\. de Kock, N\. Surange, D\. Teodorescu, I\. S\. Ahmad, D\. I\. Adelani, A\. F\. Aji, F\. D\. M\. A\. Ali, I\. Alimova, V\. Araujo, N\. Babakov, N\. Baes, A\. Bucur, A\. Bukula, G\. Cao, R\. T\. Cardenas, R\. Chevi, C\. I\. Chukwuneke, A\. Ciobotaru, D\. Dementieva, M\. S\. Gadanya, R\. Geislinger, B\. Gipp, O\. Hourrane, O\. Ignat, F\. I\. Lawan, R\. Mabuya, R\. Mahendra, V\. Marivate, A\. Panchenko, A\. Piper, C\. H\. P\. Ferreira, V\. Protasov, S\. Rutunda, M\. Shrivastava, A\. C\. Udrea, L\. D\. A\. Wanzare, S\. Wu, F\. V\. Wunderlich, H\. M\. Zhafran, T\. Zhang, Y\. Zhou, and S\. M\. Mohammad \(2025a\)BRIGHTER: bridging the gap in human\-annotated textual emotion recognition datasets for 28 languages\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 8895–8916\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1)\.
- S\. H\. Muhammad, N\. Ousidhoum, I\. Abdulmumin, S\. M\. Yimam, J\. P\. Wahle, T\. Lima Ruas, M\. Beloucif, C\. De Kock, T\. D\. Belay, I\. S\. Ahmad, N\. Surange, D\. Teodorescu, D\. I\. Adelani, A\. F\. Aji, F\. D\. M\. Ali, V\. Araujo, A\. A\. Ayele, O\. Ignat, A\. Panchenko, Y\. Zhou, and S\. Mohammad \(2025b\)SemEval\-2025 task 11: bridging the gap in text\-based emotion detection\.InProceedings of the International Workshop on Semantic Evaluation,pp\. 2558–2569\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p2.1)\.
- S\. Nayak, K\. Jain, R\. Awal, S\. Reddy, S\. V\. Steenkiste, L\. A\. Hendricks, K\. Stanczak, and A\. Agrawal \(2024\)Benchmarking vision language models for cultural understanding\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 5769–5790\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- J\. D\. D\. Nyandwi, Y\. Song, S\. Khanuja, and G\. Neubig \(2025\)Grounding multilingual multimodal LLMs with cultural knowledge\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 24187–24231\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- R\. Panda, J\. Zhang, H\. Li, J\. Lee, X\. Lu, and A\. K\. Roy\-Chowdhury \(2018\)Contemplating visual emotions: understanding and overcoming dataset bias\.InComputer Vision European Conference,pp\. 594–612\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- K\. Peng, T\. Chen, A\. Sadovnik, and A\. C\. Gallagher \(2015\)A mixed bag of emotions: model, predict, and transfer emotion distributions\.InIEEE Conference on Computer Vision and Pattern Recognition,pp\. 860–868\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- H\. Rha, J\. H\. Yeo, Y\. Kim, and Y\. M\. Ro \(2026\)Emotion\-coherent reasoning for multimodal llms via emotional rationale verifier\.InProceedings of the AAAI Conference on Artificial Intelligence and Conference on Innovative Applications of Artificial Intelligence and Sixteenth Symposium on Educational Advances in Artificial Intelligence,Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- D\. Romero, C\. Lyu, H\. A\. Wibowo, S\. Góngora, A\. Mandal, S\. Purkayastha, J\. Ortiz\-Barajas, E\. Villa\-Cueva, J\. Baek, S\. Jeong, I\. Hamed, Z\. X\. Yong, Z\. W\. Lim, P\. M\. Silva, J\. Dunstan, M\. Jouitteau, D\. L\. Meur, J\. Nwatu, G\. Batnasan, M\. Otgonbold, M\. Gochoo, G\. Ivetta, L\. Benotti, L\. A\. Alemany, H\. Maina, J\. Geng, T\. T\. Torrent, F\. Belcavello, M\. Viridiano, J\. C\. B\. Cruz, D\. J\. Velasco, O\. Ignat, Z\. Burzo, C\. Whitehouse, A\. Abzaliev, T\. Clifford, G\. Caulfield, T\. Lynn, C\. S\. Palacios, V\. Araujo, Y\. Kementchedjhieva, M\. Mihaylov, I\. A\. Azime, H\. B\. Ademtew, B\. F\. Balcha, N\. A\. Etori, D\. I\. Adelani, R\. Mihalcea, A\. L\. Tonja, M\. C\. B\. Cabrera, G\. Vallejo, H\. Lovenia, R\. Zhang, M\. Estecha\-Garitagoitia, M\. Rodríguez\-Cantelar, T\. Ehsan, R\. Chevi, M\. F\. Adilazuarda, R\. Diandaru, S\. Cahyawijaya, F\. Koto, T\. Kuribayashi, H\. Song, A\. Khandavally, T\. Jayakumar, R\. Dabre, M\. F\. M\. Imam, K\. R\. Y\. Nagasinghe, A\. Dragonetti, L\. F\. D’Haro, O\. Niyomugisha, J\. Gala, P\. A\. Chitale, F\. Farooqui, T\. Solorio, and A\. F\. Aji \(2024\)CVQA: culturally\-diverse multilingual visual question answering benchmark\.InAdvances in Neural Information Processing Systems,Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- F\. Schneider, C\. Holtermann, C\. Biemann, and A\. Lauscher \(2025\)GIMMICK: globally inclusive multimodal multitask cultural knowledge benchmarking\.InFindings of the Association for Computational Linguistics,pp\. 9605–9668\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- B\. C\. Z\. Tan, W\. Zheng, Z\. Liu, N\. F\. Chen, H\. Lee, K\. T\. W\. Choo, and R\. K\. Lee \(2026\)BLEnD\-vis: benchmarking multimodal cultural understanding in vision language models\.InProceedings of the Conference of the European Chapter of the Association for Computational Linguistics,pp\. 4647–4669\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- A\. Vayani, D\. Dissanayake, H\. Watawana, N\. Ahsan, N\. Sasikumar, O\. Thawakar, H\. B\. Ademtew, Y\. Hmaiti, A\. Kumar, K\. Kuckreja, M\. Maslych, W\. A\. Ghallabi, M\. Mihaylov, C\. Qin, A\. M\. Shaker, M\. Zhang, M\. K\. Ihsani, A\. Esplana, M\. Gokani, S\. Mirkin, H\. Singh, A\. Srivastava, E\. Hamerlik, F\. Izzati, F\. A\. Maani, S\. Cavada, J\. Chim, R\. Gupta, S\. Manjunath, K\. Zhumakhanova, F\. H\. Rabevohitra, A\. H\. Amirudin, M\. Ridzuan, D\. N\. A\. Kareem, K\. More, K\. Li, P\. Shakya, M\. Saad, A\. Ghasemaghaei, A\. Djanibekov, D\. Azizov, B\. Jankovic, N\. Bhatia, Á\. Cabrera, J\. Obando\-Ceron, O\. Otieno, F\. Farestam, M\. Rabbani, S\. Baliah, S\. Sanjeev, A\. Shtanchaev, M\. Fatima, T\. Nguyen, A\. Kareem, T\. Aremu, N\. Xavier, A\. Bhatkal, H\. O\. Toyin, A\. Chadha, H\. Cholakkal, R\. M\. Anwer, M\. Felsberg, J\. Laaksonen, T\. Solorio, M\. Choudhury, I\. Laptev, M\. Shah, S\. H\. Khan, and F\. S\. Khan \(2024\)All languages matter: evaluating lmms on culturally diverse 100 languages\.IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 19565–19575\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- Y\. Wang, Y\. Liu, F\. Yu, C\. Huang, K\. Li, Z\. Wan, W\. Che, and H\. Chen \(2025\)CVLUE: A new benchmark dataset for chinese vision\-language understanding evaluation\.InAAAI Conference on Artificial Intelligence, Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence,pp\. 8196–8204\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p3.1)\.
- D\. Wu, D\. Yang, J\. Yao, H\. Zhang, C\. Ma, Y\. Zhou, and S\. Zhao \(2026\)MVEI & emobserver: empowering mllm\-oriented visual emotional intelligence via emotion statement judgement\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- H\. Xie, C\. Peng, Y\. Tseng, H\. Chen, C\. Hsu, H\. Shuai, and W\. Cheng \(2024\)EmoVIT: revolutionizing emotion insights with visual instruction tuning\.IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 26586–26595\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- J\. Yang, Q\. Huang, T\. Ding, D\. Lischinski, D\. Cohen\-Or, and H\. Huang \(2023\)EmoSet: A large\-scale visual emotion dataset with rich attributes\.InIEEE/CVF International Conference on Computer Vision,pp\. 20326–20337\.Cited by:[Introduction](https://arxiv.org/html/2608.03358#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- Q\. Yang, M\. Ye, and B\. Du \(2024\)EmoLLM: multimodal emotional understanding meets large language models\.CoRR\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- Q\. You, H\. Jin, and J\. Luo \(2017\)Visual sentiment analysis by attending on local image regions\.InProceedings of AAAI Conference on Artificial Intelligence,pp\. 231–237\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- Q\. You, J\. Luo, H\. Jin, and J\. Yang \(2016\)Building a large scale dataset for image emotion recognition: the fine print and the benchmark\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 308–314\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.
- H\. Yu, D\. Yang, H\. He, F\. Zhang, and Q\. Yi \(2026\)VULCA\-bench: a multi\-cultural art critique benchmark for vision\-language models\.ArXiv\.Cited by:[Collection\.](https://arxiv.org/html/2608.03358#Sx3.SSx2.SSS0.Px1.p1.1)\.
- F\. Zhang, Z\. Cheng, C\. Deng, H\. Li, Z\. Lian, Q\. Chen, H\. Liu, W\. Wang, Y\. Zhang, R\. Zhang,et al\.\(2025\)MME\-emotion: a holistic evaluation benchmark for emotional intelligence in multimodal large language models\.ArXiv\.Cited by:[Related Work](https://arxiv.org/html/2608.03358#Sx2.p1.1)\.

Similar Articles