Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities

arXiv cs.CL Papers

Summary

This paper examines cultural alignment in AI-generated multimodal stories across five African communities through community-grounded evaluation, developing a taxonomy of misalignment and assessing the reliability of automated judges.

arXiv:2608.29209v1 Announce Type: new Abstract: In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community-grounded judgments at scale. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings. These findings motivate community-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary.
Original Article
View Cached Full Text

Cached at: 09/01/26, 12:18 PM

# Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities
Source: [https://arxiv.org/html/2608.29209](https://arxiv.org/html/2608.29209)
Felermino D\. M\. A\. Ali11footnotemark:1Affiliation:Microsoft Research AfricaElizabeth A\. AnkrahAffiliation:Microsoft Research AfricaNajeeb G\. AbdulhamidAffiliation:Microsoft Research AfricaBoyd MigishaAffiliation:Swansea UniversityStephanie NyairoAffiliation:Microsoft Research AfricaMercy MuchaiAffiliation:Microsoft Research AfricaSamuel MainaAffiliation:Microsoft Research AfricaAditya VashisthaAffiliation:Cornell UniversityAnja ThiemeAffiliation:Microsoft Research CambridgeJacki O’NeillAffiliation:Microsoft Research Africa

###### Abstract

In this paper, we examine how well AI\-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent\. We conduct a community\-grounded mixed methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions\. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context\. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment\. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community\-grounded judgments at scale\. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings\. These findings motivate community\-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary\.

## 1Introduction

![Refer to caption](https://arxiv.org/html/2608.29209v1/story-example.png)Figure 1:Example story\.Generative AI systems are increasingly used to produce text and visual content for diverse users at scale, including stories[Tian et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib5)\. Such content can appear fluent and locally plausible while failing to reflect the lived practices, relationships, language, values, and visual expectations of the communities being represented[Agarwal et al\. \(2025b\)](https://arxiv.org/html/2608.29209#bib.bib1);[Wang et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib6);[Kazemi et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib7)\. We refer to community\-recognized fit between generated content and lived experience as*cultural alignment*\.

Evaluating cultural alignment is difficult because culture is complex, situated, and context\-dependent[Adilazuarda et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib17)\. A form of address, item of clothing, or food practice may appear plausible to outsiders while feeling inappropriate, foreign, generic or incomplete to community members\. Community\-grounded evaluation can surface these situated judgments but is time\-intensive to scale across models, communities, and generation settings\. LLM\-as\-judge methods offer a scalable alternative[Li et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib30), but their reliability for culturally situated multimodal evaluation remains uncertain\.

We examine cultural alignment through a community\-grounded mixed methods evaluation of AI\-generated multimodal stories across five African communities: Hausa, Kikuyu, Luo, AmaXhosa, and Xichangana\. We combine quantitative annotations and story\-level scores from 19 culture representatives with qualitative focus group discussions to examine both what cultural elements shape alignment judgments and why they are experienced as aligned or misaligned\.111We use*culture representatives*to refer to community members who self\-identify with the represented community and report relevant lived and linguistic knowledge, enabling them to assess whether generated content feels authentic, inappropriate, foreign, or incomplete\.Narratives are widely used to support behavior\-change communication[Hinyard and Kreuter \(2007\)](https://arxiv.org/html/2608.29209#bib.bib3);[Ng’endo and Kariuki \(2026\)](https://arxiv.org/html/2608.29209#bib.bib4), making stories a useful setting for studying culturally situated generation\. Each story consists of four text\-image frames, as illustrated in Figure[1](https://arxiv.org/html/2608.29209#S1.F1), and centers on everyday diabetes lifestyle management\. We use this domain as a culturally consequential testbed and do not assume that the resulting patterns generalize unchanged to other domains\. Stories were generated in English for Hausa, Kikuyu, Luo, and AmaXhosa and in Portuguese for Xichangana, reflecting the role of these languages in formal written communication of the represented communities\.

Because existing annotation tools did not support frame\-by\-frame evaluation of such multimodal stories across both text and images, we built a custom annotation platform that allowed culture representatives to evaluate each story frame, identify influential text spans and image regions, categorize cultural markers, and provide an overall story\-level score\. From the annotations and focus group discussions, we developed a taxonomy of cultural alignment that captures the textual and visual markers representatives attend to when judging alignment and the recurring mechanisms through which stories become culturally misaligned\. We additionally evaluated five multimodal LLM judges using the same story\-level rubric to test whether automated evaluation can approximate community\-grounded judgments at scale\.

Our analysis yields three main findings\. First, our taxonomy shows that cultural alignment is not reducible to the presence of recognizable cultural markers, but depends on how those markers fit context across five broader categories:Referentialmarkers such as names, foods, and places;Proceduralmarkers such as food preparation and exercise routines;Contextualmarkers concerning when and where cultural elements appear;Socio\-geographicmarkers such as clinics, markets, homes, and infrastructure; andLinguistic Registermarkers such as dialect, code\-switching, and forms of address\. Second, generated stories become culturally misaligned through recurring mechanisms, including substitution, norm violation, omission, forced insertion, register conflation, cross\-modal inconsistency, stereotyping, and hallucination\. Third, LLM judges do not consistently approximate evaluations from culture representatives; reliability varies by community, and no single judge model performs consistently across all five communities\. For example, Pearson correlations between LLM judge scores and scores from culture representatives are strong for Luo \(r=0\.82r=0\.82–0\.890\.89\) and Hausa \(r=0\.67r=0\.67–0\.800\.80\), but no AmaXhosa judge remains significant after correction, and for Xichangana, judges assign alignment scores 35–45 points higher than culture representatives\. These findings motivate community\-calibrated evaluation in which community judgments ground cultural alignment assessment and automated judges are validated to determine where they can be trusted and where human review remains necessary\. Data and code:[https://github\.com/microsoft/Multimodal\-Cultural\-Alignment\-Africa](https://github.com/microsoft/Multimodal-Cultural-Alignment-Africa)\.

## 2Related Work

##### Cultural Alignment and Operationalizing Culture in LLMs\.

A growing body of work probes LLMs for cultural knowledge using proxies such as Hofstede’s dimensions[Arora et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib21);[Cao et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib18), moral judgment datasets[Scherrer et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib19);[Jinnai \(2024\)](https://arxiv.org/html/2608.29209#bib.bib20), and social etiquette norms[Rao et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib27)\. However, culture is difficult to operationalize because it cannot be reduced to static demographic labels, national categories, or isolated value dimensions[Adilazuarda et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib17)\. Prior evaluations also show that LLMs often align more strongly with Western cultural norms, and that English prompts can flatten cross\-cultural variation[Cao et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib18);[Agarwal et al\. \(2025a\)](https://arxiv.org/html/2608.29209#bib.bib2);[Rao et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib27)\. While these studies are important for measuring cultural knowledge and bias, knowing about a culture is different from generating content that communities recognize as culturally aligned\. Our work builds on this distinction by evaluating whether generated multimodal stories reflect lived practices, relationships, language, values, and visual expectations as judged by culture representatives\.

##### Cultural Misalignment in Generated Text and Images\.

Research on cultural misalignment in generated content has examined failures in dialogue, narrative, and visual generation\. In dialogue, NormDial[Li et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib22)and ReNoVi[Zhan et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib23)annotate norm adherence and violation in Chinese and American conversations, while cross\-cultural work shows that culture\-specific reasoning often fails to generalize[Jinnai \(2024\)](https://arxiv.org/html/2608.29209#bib.bib20)\. In narrative generation, prior work has identified Western bias in stories about Arab cultures[Naous et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib8)and introduced taxonomies of cultural misrepresentation for Indian stories[Bhagat et al\. \(2026\)](https://arxiv.org/html/2608.29209#bib.bib37)\. In the visual domain, text\-to\-image \(T2I\) models have been shown to neglect or misrepresent disadvantaged and underrepresented cultures[Zhang et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib25);[Kannen et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib39);[Johnson et al\. \(2026\)](https://arxiv.org/html/2608.29209#bib.bib38);[Thieme et al\. \(2026\)](https://arxiv.org/html/2608.29209#bib.bib28), and recent benchmarks characterize*how*these failures arise\. CuRe[Rege et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib35)traces them to the long tail of web\-scraped training data, showing that T2I systems hallucinate details for artifacts of the Global South \(e\.g\., an Ethiopian*jebena*\) that they render reliably for better\-represented ones\. CulturalFrames[Nayak et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib36)distinguishes*explicit*expectations, stated in the prompt, from*implicit*ones implied by its cultural context, and finds that expectations are missed 44% of the time across 10 countries, including 68% of explicit and 49% of implicit cases\. Moving beyond object\-centric artifacts, CULTIVate[Malakouti et al\. \(2026\)](https://arxiv.org/html/2608.29209#bib.bib34)evaluates cultural faithfulness through social activities such as dining, greeting, and dance, where meaning emerges from interaction and context rather than isolated objects, and decomposes failure into alignment, hallucination, exaggeration, and diversity, reporting systematically lower faithfulness for Global South than Global North cultures\. Notably, the studies find that standard image–text alignment metrics correlate poorly with human judgments of cultural alignment, motivating culturally grounded evaluation\. Complementing these automated benchmarks, other work examines community\-driven methods for assessing cultural sensitivity[Kiden et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib26)\. These studies show that cultural misalignment appears in both language and images, but they evaluate images from short, isolated prompts; less work evaluates text and image jointly within the same culturally situated narrative, or examines the mechanisms through which multimodal stories become misaligned\.

##### Scalable Evaluation and LLM\-as\-Judge\.

LLM\-as\-judge methods have been widely adopted as scalable alternatives to human evaluation for assessing generated content[Zheng et al\. \(2023\)](https://arxiv.org/html/2608.29209#bib.bib9);[Kim et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib29);[Li et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib30)\. These methods can reduce evaluation cost and increase coverage, but their reliability is sensitive to task framing, model bias, and subjective evaluation criteria[Khan et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib24)\. Cultural alignment evaluation is especially challenging because judgments may depend on lived and linguistic knowledge of the represented community, including whether language use, social norms, visual settings, and cultural markers feel appropriate in context\. Our work therefore evaluates multimodal LLM judges against scores from culture representatives rather than treating automated scores as ground truth\. In doing so, we examine when automated judges approximate community evaluations, where they are biased, and why community validation remains necessary for scalable cultural alignment evaluation\.

## 3Methodology

We adopt a mixed\-methods approach to evaluate cultural alignment in generated multimodal stories, validate automated evaluations against community\-grounded judgments, and characterize the cultural evidence and misalignment mechanisms identified by culture representatives\. Figure[2](https://arxiv.org/html/2608.29209#S3.F2)illustrates the end\-to\-end pipeline from story generation through cultural evaluation, and the following sections describe each stage\.

![Refer to caption](https://arxiv.org/html/2608.29209v1/EMNLP.png)Figure 2:Generation and Evaluation workflow\. Green highlights indicate culturally appropriate elements; red highlights indicate violations identified by culture representatives\.### 3\.1Story Generation

##### Persona generation\.

We generated culturally situated personas to provide demographic and cultural context for story generation\. Using GPT\-4\.1, we created personas through a cascading pipeline of 19 sequential LLM calls, with each attribute conditioned on previously generated attributes \(e\.g\., country→\\rightarrowcommunity→\\rightarrowprovince→\\rightarrowreligion→\\rightarrowname→\\rightarrowgender→\\rightarrowage\)\. This dependency\-aware process was designed to maintain consistency across demographic, cultural, and regional attributes\. The generated personas were manually reviewed for plausibility by authors from the represented communities, and the same reviewed persona pool was reused across all seven generation models\. For each community, we generated 40 personas, each with 28 interrelated attributes used as context for story generation; the full set of persona fields is provided in Appendix Table[6](https://arxiv.org/html/2608.29209#A2.T6)\. Each persona was paired with an everyday diabetes lifestyle question to guide story generation while keeping the focus on everyday life rather than medical diagnosis or treatment\. The questions are listed in Appendix Table[4](https://arxiv.org/html/2608.29209#A2.T4)\.

##### Frame generation\.

A story consists of four first\-person frames, where each frame contains a short paragraph paired with a corresponding image\. Text generation was conditioned on the generated persona, selected diabetes lifestyle question \(Appendix Table[4](https://arxiv.org/html/2608.29209#A2.T4)\), predefined narrative arc \(Appendix Table[10](https://arxiv.org/html/2608.29209#A3.T10)\), and one of three context settings used to vary the information available to the model during generation \(Appendix Table[5](https://arxiv.org/html/2608.29209#A2.T5)\)\. The prompt instructed models to write each story in a conversational, first\-person style, using English for Hausa, Kikuyu, Luo, and AmaXhosa personas, and Portuguese for Xichangana personas\.222We used English for Hausa, Kikuyu, Luo, and AmaXhosa stories because English is widely used for formal written communication and education in Nigeria, Kenya, and South Africa\. We used Portuguese for Xichangana stories because the study context is Mozambique, where Portuguese serves this role\.The system and user prompts used for story generation are provided in Appendix Figures[D](https://arxiv.org/html/2608.29209#A4)and[D](https://arxiv.org/html/2608.29209#A4)\. We used seven generation models spanning model families, parameter scales, and access types to produce a diverse set of AI\-generated stories across larger and smaller, open\-weight and proprietary systems: GPT\-4\.1, GPT\-4\.1\-mini, Gemma\-3 27B, Gemma\-3 4B[Team et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib14), Qwen3\-4B[Yang et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib15), Llama 3\.1 8B, and Llama 3\.3 70B[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib16)\. For image generation, we used FLUX\.1\-Kontext\-pro[Labs et al\. \(2025\)](https://arxiv.org/html/2608.29209#bib.bib13)\. The first image was generated from the first paragraph, and subsequent images were produced through iterative image editing conditioned on the previous frame, using a fixed seed of 12345 and a low editing strength of 0\.3 to support character consistency and visual continuity\. The final dataset contains 199 multimodal stories across five communities, seven generation models, and three context settings\.

### 3\.2Story Evaluation

#### 3\.2\.1Evaluation Procedure

We designed the evaluation procedure to capture cultural alignment at multiple levels, including the overall story, each text\-image frame, and which specific cultural markers within each frame shaped evaluators’ judgments\. The procedure was informed by prior work on operationalizing culture through demographic and semantic proxies, which we adapted into cultural marker categories for text and image annotation[Adilazuarda et al\. \(2024\)](https://arxiv.org/html/2608.29209#bib.bib17); the full category list is provided in Appendix Table[7](https://arxiv.org/html/2608.29209#A3.T7)\. These categories covered markers such as language, cultural group, region, religion, socioeconomic context, age and gender roles, occupation, food and dietary norms, physical activity norms, kinship and social structure, community practices, social etiquette, values and beliefs, and attitudes toward health\.

Because existing annotation tools did not support the full workflow required for multimodal cultural alignment evaluation, we designed a custom annotation platform for this study\. The platform supports frame\-level evaluation of text and images, text\-span highlighting, image\-region selection, cultural marker categorization, connectedness labels, influence ratings, optional comments, and story\-level alignment scoring\. These features separate overall judgments from the specific textual and visual evidence behind them, allowing evaluators to indicate not only whether a story felt aligned or misaligned, but which cultural markers shaped that judgment and how strongly\. Figure[5](https://arxiv.org/html/2608.29209#A3.F5)shows the interface used by culture representatives\.

At the frame level, evaluators assessed how culturally connected each paragraph and accompanying image felt to the represented community using three response options:Not connected,Connected, andStrongly connected\. They then identified the specific text spans or image regions that influenced their judgment, assigned each selected marker a cultural category \(see Figure[6](https://arxiv.org/html/2608.29209#A3.F6)\), indicated whether the marker was culturally connected or not connected, rated how strongly it influenced their judgment, and optionally provided a comment\. We also provided audio to culture representatives as an accessibility aid for reviewing the story, rather than as an evaluated modality\.

After completing all four frames, evaluators completed a story\-level feedback checklist covering cultural fit, persona consistency, image reliability, safety, and overall story experience\. They then assigned an overall cultural alignment score from 0 to 100 using an anchored rubric \(see Figure[7](https://arxiv.org/html/2608.29209#A3.F7)\)\. The rubric ranged from no recognizable cultural markers or cultural relevance to strong cultural alignment with high relevance, coherence, and authenticity in relation to the represented community \(Table[8](https://arxiv.org/html/2608.29209#A3.T8)\)\. We used this graded score because cultural alignment is not binary\. A story may contain recognizable cultural markers while still varying in relevance, integration, and multimodal coherence\. The same story\-level scale was used for LLM judges and culture representatives to support comparison between automated and community\-grounded evaluations\.

### 3\.3LLM Judge Evaluation

We evaluated each generated story with LLM judges to test whether automated evaluators can provide a scalable approximation of community judgments of cultural alignment\. We used five multimodal LLM judges: Kimi\-K2\.6, Gemma\-4\-31B, GPT\-5\.5, Qwen\-3\.5\-9B, and Qwen\-3\.5\-122B\. These models were selected to cover a range of model families, scales, and access types, including proprietary and open\-weight systems, while supporting evaluation of text, images, and text\-image coherence\. No exact generation model was reused as an LLM judge\. Each judge evaluated the same story set using the same story\-level cultural alignment rubric used by culture representatives\. The judge prompts also included the same cultural marker categories used in the community evaluation, but did not include community\-specific calibration examples, as our goal was to evaluate whether general rubric\-based judges could approximate community judgments without prior community\-specific calibration\. We conducted three independent evaluation runs per judge to account for variability in model outputs, and averaged the three runs into a single judge score for each story before comparing automated scores with scores from culture representatives\. The full judge prompt is provided in Figure[D](https://arxiv.org/html/2608.29209#A4)and Figure[D](https://arxiv.org/html/2608.29209#A4)\.

### 3\.4Community Evaluation

##### Participants and Communities\.

We recruited 19 participants aged 18–44 who served as culture representatives across five African cultural communities: Hausa \(Northern Nigeria\), AmaXhosa \(Eastern Cape, South Africa\), Luo \(Western Kenya\), Kikuyu \(Central Kenya\), and Xichangana \(Maputo, Mozambique\), as shown in Appendix Figure[4](https://arxiv.org/html/2608.29209#A1.F4)\. Participants were selected based on self\-identification with the target community, lived experience in the relevant cultural region, and native or regular use of the relevant community language\. They were recruited through community networks, including referrals from local collaborators, professional contacts, and community\-based networks connected to the represented cultural groups\. All participants provided written informed consent, received an internet allowance before the study to support participation, and were compensated with gift vouchers upon completion\. Community context and participant demographics are provided in Appendix[A](https://arxiv.org/html/2608.29209#A1)and Appendix Table[3](https://arxiv.org/html/2608.29209#A1.T3)\.

##### Individual Cultural Evaluation\.

Culture representatives completed a structured onboarding and practice annotation process\. We held a one\-hour onboarding session to introduce the study goals, define key terminology, explain the cultural marker categories, and demonstrate the annotation platform\. Representatives then completed a one\-week training phase, during which they annotated 20 practice stories using the same interface and evaluation procedure used in the main study\. After this phase, we held a one\-hour discussion session to review examples of agreement and disagreement, clarify annotation expectations, and support a shared understanding of the evaluation task while still allowing community\-specific interpretations\. Culture representatives then individually evaluated 40 stories generated for their own communities using the evaluation procedure described above\. In total, the individual evaluation produced 18,805 marker\-level annotations across 199 stories, including 9,386 text span annotations \(Appendix Table[13](https://arxiv.org/html/2608.29209#A5.T13)\) and 9,419 image region annotations \(Appendix Table[14](https://arxiv.org/html/2608.29209#A5.T14)\)\.

##### Focus Group Discussions\.

Following individual annotation, we conducted 15 focus group sessions, with three sessions per community\. Each session lasted approximately two hours, for a total of approximately 30 hours of recorded discussion\. Sessions were conducted over Microsoft Teams, audio\-recorded, and transcribed\.

The focus groups were designed to collect qualitative explanations and examine how representatives reasoned through agreement and disagreement in their individual annotations\. In the first session, representatives discussed stories where their annotations showed broad agreement, helping establish shared vocabulary for cultural alignment and misalignment within each community\. In the second session, representatives examined stories where their annotations diverged, surfacing implicit norms, contested expectations, and differences in how cultural dimensions were weighted\. These discussions allowed representatives to revisit details they may have missed during individual annotation, clarify the reasoning behind their judgments, and decide whether to maintain or revise their interpretations\. In the third session, representatives reflected across the full set of evaluated stories, identifying the strongest and weakest examples and distinguishing meaningful cultural integration from superficial decoration\. The guiding questions are provided in Appendix Table[9](https://arxiv.org/html/2608.29209#A3.T9)\. These discussions served as a deliberative evaluation method, producing the situated interpretations from which the taxonomy in Section[3\.5](https://arxiv.org/html/2608.29209#S3.SS5)was derived\.

### 3\.5Taxonomy Derivation

Table 1:Summary of the cultural alignment taxonomy derived from focus group discussions and participant annotations \(refer to Figure[8](https://arxiv.org/html/2608.29209#A3.F8)for examples\)\.We developed the cultural alignment taxonomy through a convergent mixed\-methods design[Creswell and Plano Clark \(2018\)](https://arxiv.org/html/2608.29209#bib.bib12), conducting thematic analysis[Braun and Clarke \(2006\)](https://arxiv.org/html/2608.29209#bib.bib10);[Braun and Clarke \(2019\)](https://arxiv.org/html/2608.29209#bib.bib11)of the FGDs to derive the taxonomy, and then examined how the resulting categories appeared in the participant annotations\. Each FGD was conducted with two to three of the authors present, during the session they took independent notes\. After the session each transcript was put into HeyMarvin, a qualitative research tool, for transcription and analysis\. Each transcript was anonymised and read in full by at least one author\. The authors noted emergent categories from their notes and during their readings, extracted the examples from the transcripts which fell into these categories\. Five of the authors then conducted 5 joint analysis sessions where they discussed the categories and examples together, and from these identified emergent themes with verifiable examples\. These themes formed the basis of the cultural alignment taxonomy and the recurring misalignment mechanisms summarized in Table[1](https://arxiv.org/html/2608.29209#S3.T1)\. The cultural marker categories organize the annotated text and image markers, labeled using the fine\-grained categories in Appendix Table[7](https://arxiv.org/html/2608.29209#A3.T7), into broader categories\. The misalignment mechanisms capture recurring failure patterns identified from representatives’ reasoning in the focus group discussions\.

## 4Results

### 4\.1Community Evaluation Reveals Patterned Cultural Alignment and Misalignment

The uneven judge results show that automated scores alone cannot explain cultural alignment\. Community evaluation adds this missing layer by showing which textual and visual markers representatives used to judge cultural fit, and how those markers supported, weakened, or disrupted alignment in context\. We organize these results using the taxonomy summarized in Table[1](https://arxiv.org/html/2608.29209#S3.T1), first examining the five broader cultural marker categories across modalities and then describing the eight recurring mechanisms of misalignment\.

Figure 3:Distribution of not\-connected cultural markers in text and images across the five broader cultural marker categories, broken down by generation model and community\. Connected marker distributions for both modalities appear in Appendix Figure[13](https://arxiv.org/html/2608.29209#A5.F13)\.#### 4\.1\.1Cultural Marker Patterns Across Modalities

Marker patterns varied across communities, modalities, and the five broader cultural marker categories\. At the marker level, representatives labeled selected text spans and image regions asConnectedwhen they supported cultural fit andNot connectedwhen they weakened or disrupted cultural fit\. Among text span annotations, 8,075 of 9,386 \(86%\) were labeled connected and 1,311 \(14%\) were labeled not connected\. Image region annotations showed more visible disruption, with 6,931 of 9,419 \(74%\) labeled connected and 2,488 \(26%\) labeled not connected\. Full text\-span and image\-region statistics are provided in Appendix Tables[13](https://arxiv.org/html/2608.29209#A5.T13)and[14](https://arxiv.org/html/2608.29209#A5.T14)\. Figure[3](https://arxiv.org/html/2608.29209#S4.F3)descriptively shows the distribution of not\-connected markers across communities, models, modalities, and cultural marker categories, while Appendix Figure[13](https://arxiv.org/html/2608.29209#A5.F13)shows the corresponding distribution of connected markers\. These figures show that cultural fit and cultural disruption are not evenly distributed across modalities\. In text, not\-connected markers were often concentrated in different categories for different communities\. Hausa stories showed frequent not\-connected markers inReferentialandProceduralcategories, while Xichangana stories showed a more even spread across categories, includingLinguistic Register\. In images, not\-connected markers were often more visually concentrated, especially around character appearance, clothing, food, setting, public space, and other visual details\. These patterns show that the visual realization of a story can support, weaken, transform, or contradict cultural cues in the narrative, making evaluation across both text and images necessary\.

#### 4\.1\.2Misalignment Mechanisms

The focus group discussions revealed eight recurring mechanisms through which stories became culturally misaligned\.Substitutionoccurred when stories included markers from another community in place of markers from the target community\. For instance,sadza, a Zimbabwean food, was independently flagged in three non\-Zimbabwean communities: Kikuyu, Luo, and AmaXhosa\.Hallucinationoccurred when models used cultural vocabulary but attached it to implausible practices\. Hausa participants flagged an incorrect food preparation description, “boil your tuwo shinkafa instead of frying,” noting that frying was not a way they would cook this dish\. An AmaXhosa participant similarly foundumxhentso, a traditional dance, described in a story as a food and commented, “This is not an error a human storyteller would make\. This is very AI\.”Forced insertionoccurred when correct names or cultural terms appeared in otherwise generic narratives without meaningful integration, as when a single Kikuyu name was placed in a story where surrounding markers belonged to other cultures, making the name feel “thrown in\.”

Norm violationappeared when outputs contradicted expectations around respect, dress, or social interaction\. A Xichangana participant explained, “ninguém entra no hospital, participa de uma consulta de chapéu, considera\-se como uma falta de respeito” \(nobody enters a hospital wearing a hat; it is considered a lack of respect\)\.Register conflationwas especially visible in Xichangana stories, where Portuguese was fluent but regionally inappropriate\. One participant observed, “usava\-se muito o gerúndio e o gerúndio é muito característico dos brasileiros” \(the gerund was used a lot, and the gerund is very characteristic of Brazilians\)\.Cross\-modal inconsistencyoccurred when text and image represented conflicting settings, identities, or practices, whilestereotypingappeared when models reduced communities to narrow or repeated visual tropes\.

Finally,omissionappeared when stories contained no obviously wrong elements but also few cultural details\. Participants across communities described such stories as generic rather than factually incorrect\. An AmaXhosa participant called one story “a story anybody from any culture could narrate,” a Luo participant described another as “not grounded culturally,” and a Xichangana participant said, “não encontrei marcadores” \(I did not find markers\)\. These mechanisms show that cultural misalignment is not only a matter of factual error, but also of weak integration, inappropriate context, cross\-modal mismatch, wrong register, stereotyping, and absence of situated detail\. See Appendix Figure[8](https://arxiv.org/html/2608.29209#A3.F8)for additional examples\.

### 4\.2LLM Judges Approximate Community Judgments Unevenly Across Cultures

Table 2:Pearson correlation coefficients \(rr\) between LLM judge scores and mean scores from culture representatives by community\. Bold values remain significant after Bonferroni correction \(p<0\.002p<0\.002\); unbolded values do not, even when significant at uncorrected thresholds \(p<0\.001p<0\.001,p<0\.01p<0\.01, orp<0\.05p<0\.05\)\. No AmaXhosa judge remains significant after correction\.Judge correlations with community scores varied across the five communities\. For Luo, all five judges achieved very strong Pearson correlations with scores from culture representatives \(r=0\.82r=0\.82–0\.890\.89\), and for Hausa correlations remained strong \(r=0\.67r=0\.67–0\.800\.80\)\. Kikuyu and Xichangana fell in moderate ranges \(r=0\.50r=0\.50–0\.640\.64andr=0\.39r=0\.39–0\.510\.51, respectively\)\. For AmaXhosa, no judge remained significant even after Bonferroni correction, indicating that automated scores did not reliably track community judgments for this community\.

We first assess the reliability of the community ratings that serve as the reference point for judge validation and grounding analysis\. Inter\-annotator agreement on the overall cultural alignment score varied by community, ranging from good for Hausa \(ICC\(A,1\)=0\.81\), to moderate for Luo, AmaXhosa, and Kikuyu \(ICC\(A,1\)=0\.60–0\.63\), to poor for Xichangana \(ICC\(A,1\)=0\.33\)\. Xichangana’s low agreement reflects a compressed score distribution \(mean=24\.5, SD=4\.5\), where representatives broadly rated stories as culturally inadequate but diverged on finer distinctions\. All ICC values were statistically significant \(p<0\.001p<0\.001; Table[12](https://arxiv.org/html/2608.29209#A5.T12)\)\.

Score bias further limits automated evaluation\. For Xichangana, judges assigned alignment scores 35–45 points higher than culture representatives \(pairedtt\-test, allp<0\.001p<0\.001\), assigning scores in the 60–69 range to stories that culture representatives rated at 24\.5 on average\. Kikuyu showed the opposite pattern, with Kimi and GPT\-5\.5 assigning scores 14–16 points lower than culture representatives\. Judges were best calibrated on Hausa and showed smaller, mixed biases on Luo\. Because both correlation with community scores and score bias varied by culture, no single judge model performed consistently across all five communities\. Table[2](https://arxiv.org/html/2608.29209#S4.T2)reports judge–community correlations, and Appendix Table[11](https://arxiv.org/html/2608.29209#A5.T11)reports score bias for all judge–community pairs\.

## 5Discussion

Community evaluations show that cultural alignment is not reducible to the number or presence of recognizable markers\. Culture representatives considered whether names, foods, places, clothing, language cues, social interactions, and visual settings fit the story context and the represented community\. The taxonomy makes this distinction explicit by separating the broader cultural marker categories that representatives attended to from the mechanisms through which stories became misaligned\. This matters for evaluation because a story may contain plausible markers while still failing through substitution, forced insertion, cross\-modal inconsistency, wrong register, or absence of situated detail\.

Our results also show that LLM judges do not approximate community judgments consistently across cultures\. Judge reliability and score calibration varied substantially across communities, with no single judge performing consistently across all five settings\. This variation means that automated judges should be treated as scalable evaluators only after community\-specific validation\. For cultural alignment tasks, both correlation with community judgments and score calibration should be considered, since a judge may track differences between stories while still systematically assigning scores that diverge from those of culture representatives\.

These findings point toward community\-calibrated evaluation pipelines for cultural alignment\. In such pipelines, culture representatives provide the reference judgments needed to establish where automated evaluators can be trusted and where human review remains necessary\. The taxonomy also makes community feedback more actionable for scalable evaluation: the five broader cultural marker categories identify what cultural dimensions automated judges should attend to, while the eight misalignment mechanisms describe recurring failure patterns that can guide review and calibration without replacing community judgment\. Although the taxonomy was derived from five African communities in the context of diabetes lifestyle stories, its structure provides a basis for investigating whether similar cultural dimensions and misalignment mechanisms emerge in other communities and domains\.

## 6Conclusion

We presented a human\-centered evaluation of cultural alignment in AI\-generated multimodal stories across five African communities\. Community evaluations showed that cultural alignment depends not only on the presence of recognizable cultural markers, but on how those markers fit narrative, social, linguistic, procedural, and visual context\. From annotations and focus group discussions, we developed a taxonomy of five broader cultural marker categories and eight recurring mechanisms of misalignment\. We further evaluated multimodal LLM judges against culture\-representative judgments and found that their reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings\. These findings point toward community\-calibrated evaluation pipelines in which community judgments ground cultural alignment evaluation and automated judges are validated to determine where they can be trusted and where human review remains necessary\.

## Limitations

Our study has several limitations that point to important directions for future work\. First, our community evaluation relies on 19 culture representatives\. These participants provide situated lived and linguistic knowledge, but they do not represent the full diversity of views within each community\. Cultural alignment is not fixed or uniform, and perspectives may vary across age, gender, region, language use, religion, class, and rural\-urban experience\. Future work should build broader and more iterative community review processes that include more diverse participants and examine how cultural alignment judgments vary within communities\.

Second, our evaluation focuses on story\-level and marker\-level cultural alignment rather than downstream effects on readers\. We do not test whether culturally aligned stories improve understanding, trust, behavior\-change outcomes, or lifestyle decision\-making\. Future work should connect cultural alignment evaluation to reader\-facing outcomes, asking not only whether communities recognize a story as culturally coherent, but also whether such coherence changes how people interpret, trust, adapt, or use AI\-generated content\.

Third, our study is grounded in everyday Type II diabetes lifestyle and behavior\-change stories\. This domain was chosen because diet, physical activity, household routines, and care practices are culturally situated, but the cultural marker patterns, misalignment mechanisms, and judge behavior observed here may differ in other domains\. Future work should therefore validate the framework across other kinds of generated content and application settings\.

Fourth, cross\-community comparisons should be interpreted carefully\. Xichangana stories differed from the other communities in both language context and evaluation patterns\. They were generated in Portuguese for the Mozambican context, received much lower community scores on average, showed lower agreement among representatives, and were substantially over\-scored by LLM judges\. These differences make Xichangana an important case for understanding failures of automated evaluation, but they also caution against interpreting score differences across communities as direct measures of cultural alignment difficulty\. Future work should examine how language choice, regional context, and local norms shape cultural alignment judgments\.

Finally, our LLM judge analysis covers five multimodal judges and a fixed evaluation prompt\. Judge reliability may change with different models, prompts, rubrics, or calibration methods\. Rather than seeking a universal automated judge for cultural alignment, future work should develop community\-calibrated evaluation pipelines that identify where judge scores can be trusted, where they are biased, and where human review remains necessary\. Future work can also test whether community\-derived taxonomies improve scalable evaluation by incorporating examples of cultural marker categories and misalignment mechanisms into judge prompts, audit protocols, or human review workflows, while validating these approaches separately in each community context\.

## Ethical Considerations

The generated stories were evaluated with the participation of culture representatives from each respective community\. It is important to note that the stories do not provide medical diagnoses, prescriptions, or recommendations related to medication\. Instead, they focus exclusively on culturally contextualized lifestyle and behavior\-change narratives, particularly around healthy eating habits, physical activity, and everyday practices associated with managing Type II diabetes\. Cultural alignment should not be interpreted as evidence of medical correctness or safety, and health\-related generated content should undergo appropriate medical review in addition to community\-based cultural evaluation\.

## References

- Adilazuardaet al\.\(2024\)M\. F\. Adilazuarda, S\. Mukherjee, P\. Lavania, S\. S\. Singh, A\. F\. Aji, J\. O’Neill, A\. Modi, and M\. ChoudhuryTowards measuring and modeling “culture” in LLMs: a survey\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 15763–15784\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.882/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.882)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p2.1),[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1),[§3\.2\.1](https://arxiv.org/html/2608.29209#S3.SS2.SSS1.p1.1)\.
- Agarwalet al\.\(2025a\)D\. Agarwal, M\. Naaman, and A\. VashisthaAI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,CHI ’25,New York, NY, USA,pp\. 1–21\.External Links:ISBN 979\-8\-4007\-1394\-1,[Link](https://dl.acm.org/doi/10.1145/3706598.3713564),[Document](https://dx.doi.org/10.1145/3706598.3713564)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1)\.
- Agarwalet al\.\(2025b\)D\. Agarwal, A\. Shukla, S\. Sitaram, and A\. VashisthaFluent but Foreign: Even Regional LLMs Lack Cultural Alignment\.arXiv\.Note:arXiv:2505\.21548 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.21548),[Document](https://dx.doi.org/10.48550/arXiv.2505.21548)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p1.1)\.
- Aroraet al\.\(2023\)A\. Arora, L\. Kaffee, and I\. AugensteinProbing pre\-trained language models for cross\-cultural differences in values\.InProceedings of the First Workshop on Cross\-Cultural Considerations in NLP \(C3NLP\),S\. Dev, V\. Prabhakaran, D\. I\. Adelani, D\. Hovy, and L\. Benotti \(Eds\.\),Dubrovnik, Croatia,pp\. 114–130\.External Links:[Link](https://aclanthology.org/2023.c3nlp-1.12/),[Document](https://dx.doi.org/10.18653/v1/2023.c3nlp-1.12)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1)\.
- Bhagatet al\.\(2026\)K\. Bhagat, S\. Bhatt, A\. Velagapudi, A\. Vashistha, S\. Dave, and D\. PruthiTALES: a taxonomy and analysis of cultural representations in llm\-generated stories\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,CHI ’26,New York, NY, USA\.External Links:ISBN 9798400722783,[Link](https://doi.org/10.1145/3772318.3790519),[Document](https://dx.doi.org/10.1145/3772318.3790519)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Braun and Clarke \(2006\)V\. Braun and V\. ClarkeUsing thematic analysis in psychology\.Qualitative Research in Psychology3\(2\),pp\. 77–101\.External Links:[Document](https://dx.doi.org/10.1191/1478088706qp063oa)Cited by:[§3\.5](https://arxiv.org/html/2608.29209#S3.SS5.p1.1)\.
- Braun and Clarke \(2019\)V\. Braun and V\. ClarkeReflecting on reflexive thematic analysis\.Qualitative Research in Sport, Exercise and Health11\(4\),pp\. 589–597\.External Links:[Document](https://dx.doi.org/10.1080/2159676X.2019.1628806)Cited by:[§3\.5](https://arxiv.org/html/2608.29209#S3.SS5.p1.1)\.
- Caoet al\.\(2023\)Y\. Cao, L\. Zhou, S\. Lee, L\. Cabello, M\. Chen, and D\. HershcovichAssessing cross\-cultural alignment between ChatGPT and human societies: an empirical study\.InProceedings of the First Workshop on Cross\-Cultural Considerations in NLP \(C3NLP\),S\. Dev, V\. Prabhakaran, D\. I\. Adelani, D\. Hovy, and L\. Benotti \(Eds\.\),Dubrovnik, Croatia,pp\. 53–67\.External Links:[Link](https://aclanthology.org/2023.c3nlp-1.7/),[Document](https://dx.doi.org/10.18653/v1/2023.c3nlp-1.7)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1)\.
- Chinchor and Sundheim \(1993\)N\. Chinchor and B\. SundheimMUC\-5 evaluation metrics\.InFifth Message Understanding Conference \(MUC\-5\),External Links:[Link](https://aclanthology.org/M93-1007)Cited by:[§E\.1](https://arxiv.org/html/2608.29209#A5.SS1.SSS0.Px2.p2.1)\.
- Creswell and Plano Clark \(2018\)J\. W\. Creswell and V\. L\. Plano ClarkDesigning and conducting mixed methods research\.3rd edition,SAGE Publications,Thousand Oaks, CA\.External Links:ISBN 978\-1\-4833\-4437\-9Cited by:[§3\.5](https://arxiv.org/html/2608.29209#S3.SS5.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2608.29209#S3.SS1.SSS0.Px2.p1.1)\.
- Hinyard and Kreuter \(2007\)L\. J\. Hinyard and M\. W\. KreuterUsing narrative communication as a tool for health behavior change: a conceptual, theoretical, and empirical overview\.Health Education & Behavior34\(5\),pp\. 777–792\.External Links:[Link](http://www.jstor.org/stable/45055957)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p3.1)\.
- Jinnai \(2024\)Y\. JinnaiDoes cross\-cultural alignment change the commonsense morality of language models?\.InProceedings of the 2nd Workshop on Cross\-Cultural Considerations in NLP,V\. Prabhakaran, S\. Dev, L\. Benotti, D\. Hershcovich, L\. Cabello, Y\. Cao, I\. Adebara, and L\. Zhou \(Eds\.\),Bangkok, Thailand,pp\. 48–64\.External Links:[Link](https://aclanthology.org/2024.c3nlp-1.5/),[Document](https://dx.doi.org/10.18653/v1/2024.c3nlp-1.5)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Johnsonet al\.\(2026\)N\. Johnson, D\. Sudharsan, Hamna, S\. Dalal, T\. Holroyd, A\. Thieme, H\. Heidari, D\. Massiceti, J\. Wortman Vaughan, and C\. MorrisonEvaluating ai\-generated images of cultural artifacts with community\-informed rubrics\.InProceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’26,New York, NY, USA,pp\. 714–774\.External Links:ISBN 9798400725968,[Link](https://doi.org/10.1145/3805689.3812222),[Document](https://dx.doi.org/10.1145/3805689.3812222)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Kannenet al\.\(2024\)N\. Kannen, A\. Ahmad, M\. Andreetto, V\. Prabhakaran, U\. Prabhu, A\. B\. Dieng, P\. Bhattacharyya, and S\. DaveBeyond aesthetics: cultural competence in text\-to\-image models\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 13716–13747\.External Links:[Document](https://dx.doi.org/10.52202/079017-0439),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/18c669b80d1a8f589713b768bc8fe9a4-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Kazemiet al\.\(2024\)S\. Kazemi, G\. Gerhardt, J\. Katz, C\. I\. Kuria, E\. Pan, and U\. PrabhakarCultural fidelity in large\-language models: an evaluation of online language resources as a driver of model performance in value representation\.External Links:2410\.10489,[Link](https://arxiv.org/abs/2410.10489)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p1.1)\.
- Khanet al\.\(2025\)A\. Khan, S\. Casper, and D\. Hadfield\-MenellRandomness, not representation: the unreliability of evaluating cultural alignment in llms\.InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’25,New York, NY, USA,pp\. 2151–2165\.External Links:ISBN 9798400714825,[Link](https://doi.org/10.1145/3715275.3732147),[Document](https://dx.doi.org/10.1145/3715275.3732147)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px3.p1.1)\.
- Kidenet al\.\(2025\)S\. Kiden, O\. Peter, G\. Reyes\-Cruz, M\. Klyshbekova, S\. Choi, A\. G\. Bergin, M\. Waheed, D\. Eke, T\. Azim, S\. Ramchurn, S\. Stein, E\. P\. Vallejos, K\. Devlin, and J\. E\. FischerBack to the communities: a mixed\-methods and community\-driven evaluation of cultural sensitivity in text\-to\-image models\.External Links:2510\.27361,[Link](https://arxiv.org/abs/2510.27361)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2024\)S\. Kim, J\. Shin, Y\. Cho, J\. Jang, S\. Longpre, H\. Lee, S\. Yun, S\. Shin, S\. Kim, J\. Thorne, and M\. SeoPrometheus: inducing fine\-grained evaluation capability in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=8euJaTveKw)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px3.p1.1)\.
- Labset al\.\(2025\)B\. F\. Labs, S\. Batifol, A\. Blattmann, F\. Boesel, S\. Consul, C\. Diagne, T\. Dockhorn, J\. English, Z\. English, P\. Esser, S\. Kulal, K\. Lacey, Y\. Levi, C\. Li, D\. Lorenz, J\. Müller, D\. Podell, R\. Rombach, H\. Saini, A\. Sauer, and L\. SmithFLUX\.1 kontext: flow matching for in\-context image generation and editing in latent space\.External Links:2506\.15742,[Link](https://arxiv.org/abs/2506.15742)Cited by:[§3\.1](https://arxiv.org/html/2608.29209#S3.SS1.SSS0.Px2.p1.1)\.
- Liet al\.\(2025\)D\. Li, B\. Jiang, L\. Huang, A\. Beigi, C\. Zhao, Z\. Tan, A\. Bhattacharjee, Y\. Jiang, C\. Chen, T\. Wu, K\. Shu, L\. Cheng, and H\. LiuFrom generation to judgment: opportunities and challenges of LLM\-as\-a\-judge\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 2757–2791\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.138/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.138),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p2.1),[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px3.p1.1)\.
- Liet al\.\(2023\)O\. Li, M\. Subramanian, A\. Saakyan, S\. CH\-Wang, and S\. MuresanNormDial: a comparable bilingual synthetic dialog dataset for modeling social norm adherence and violation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 15732–15744\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.974/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.974)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Malakoutiet al\.\(2026\)S\. Malakouti, B\. Gong, and A\. KovashkaCulture in action: evaluating text\-to\-image models through social activities\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=opG4m2U0Oo)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Naouset al\.\(2024\)T\. Naous, M\. J\. Ryan, and W\. XuHaving beer after prayer? measuring cultural bias in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\),External Links:[Link](https://arxiv.org/abs/2305.14456)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Nayaket al\.\(2025\)S\. Nayak, M\. Bhatia, X\. Zhang, V\. Rieser, L\. A\. Hendricks, S\. van Steenkiste, Y\. Goyal, K\. Stanczak, and A\. AgrawalCulturalFrames: assessing cultural expectation alignment in text\-to\-image models and evaluation metrics\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 20918–20953\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1141/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1141),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Ng’endo and Kariuki \(2026\)M\. Ng’endo and E\. KariukiVoices of change: public narrative storytelling communicates climate resilience actions in kenya\.Development in Practice36\(1\),pp\. 106–123\.External Links:[Document](https://dx.doi.org/10.1080/09614524.2025.2581867)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p3.1)\.
- Raoet al\.\(2025\)A\. Rao, A\. Yerukola, V\. Shah, K\. Reinecke, and M\. SapNormAd: a framework for measuring the cultural adaptability of large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 2373–2403\.External Links:[Link](https://aclanthology.org/2025.naacl-long.120/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.120),ISBN 979\-8\-89176\-189\-6Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1)\.
- Regeet al\.\(2025\)A\. Rege, Z\. Nie, M\. Ramesh, U\. Raskar, Z\. Yu, A\. Kusupati, Y\. J\. Lee, and R\. K\. VinayakCuRe: Cultural Gaps in the Long Tail of Text\-to\-Image Systems\.InInternational Conference on Computer Vision,pp\. 15680–15691\.External Links:[Document](https://dx.doi.org/10.1109/ICCV51701.2025.01455),[Link](https://mlanthology.org/iccv/2025/rege2025iccv-cure/)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Scherreret al\.\(2023\)N\. Scherrer, C\. Shi, A\. Feder, and D\. BleiEvaluating the moral beliefs encoded in llms\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 51778–51809\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/a2cf225ba392627529efef14dc857e22-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px1.p1.1)\.
- Segura\-Bedmaret al\.\(2013\)I\. Segura\-Bedmar, P\. Martínez, and M\. Herrero\-ZazoSemEval\-2013 task 9: extraction of drug\-drug interactions from biomedical texts \(DDIextraction 2013\)\.InProceedings of SemEval 2013,pp\. 341–350\.External Links:[Link](https://aclanthology.org/S13-2056)Cited by:[§E\.1](https://arxiv.org/html/2608.29209#A5.SS1.SSS0.Px2.p2.1)\.
- Teamet al\.\(2025\)G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§3\.1](https://arxiv.org/html/2608.29209#S3.SS1.SSS0.Px2.p1.1)\.
- Thiemeet al\.\(2026\)A\. Thieme, R\. Faia Marques, M\. Grayson, S\. Balachandar, C\. Tyler Cassidy, M\. Z\. Choksi, C\. Longden, R\. S\. Huda, N\. Ileve Kalovwe, C\. Mallon, C\. Mansperger, D\. Massiceti, B\. Mitra, R\. M\. Nzioka, I\. Tanase, Y\. You, and C\. MorrisonEngaging communities meaningfully in defining disability representation for ai image generation\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,CHI ’26,New York, NY, USA\.External Links:ISBN 9798400722783,[Link](https://doi.org/10.1145/3772318.3790768),[Document](https://dx.doi.org/10.1145/3772318.3790768)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Tianet al\.\(2024\)Y\. Tian, T\. Huang, M\. Liu, D\. Jiang, A\. Spangher, M\. Chen, J\. May, and N\. PengAre large language models capable of generating human\-level narratives?\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 17659–17681\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.978/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.978)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p1.1)\.
- Tjong Kim Sang and De Meulder \(2003\)E\. F\. Tjong Kim Sang and F\. De MeulderIntroduction to the CoNLL\-2003 shared task: language\-independent named entity recognition\.InProceedings of the Seventh Conference on Natural Language Learning at HLT\-NAACL 2003,pp\. 142–147\.External Links:[Link](https://aclanthology.org/W03-0419)Cited by:[§E\.1](https://arxiv.org/html/2608.29209#A5.SS1.SSS0.Px2.p2.1)\.
- Wanget al\.\(2024\)W\. Wang, W\. Jiao, J\. Huang, R\. Dai, J\. Huang, Z\. Tu, and M\. LyuNot all countries celebrate thanksgiving: on the cultural dominance in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 6349–6384\.External Links:[Link](https://aclanthology.org/2024.acl-long.345/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.345)Cited by:[§1](https://arxiv.org/html/2608.29209#S1.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.1](https://arxiv.org/html/2608.29209#S3.SS1.SSS0.Px2.p1.1)\.
- Zhanet al\.\(2024\)H\. Zhan, Z\. Li, X\. Kang, T\. Feng, Y\. Hua, L\. Qu, Y\. Ying, M\. R\. Chandra, K\. Rosalin, J\. Jureynolds, S\. Sharma, S\. Qu, L\. Luo, I\. Zukerman, L\. Soon, Z\. Semnani Azad, and R\. HafRENOVI: a benchmark towards remediating norm violations in socio\-cultural conversations\.InFindings of the Association for Computational Linguistics: NAACL 2024,K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 3104–3117\.External Links:[Link](https://aclanthology.org/2024.findings-naacl.196/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-naacl.196)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2024\)L\. Zhang, X\. Liao, Z\. Yang, B\. Gao, C\. Wang, Q\. Yang, and D\. LiPartiality and misconception: investigating cultural representativeness in text\-to\-image models\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,CHI ’24,New York, NY, USA\.External Links:ISBN 9798400703300,[Link](https://doi.org/10.1145/3613904.3642877),[Document](https://dx.doi.org/10.1145/3613904.3642877)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px2.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,External Links:[Link](https://arxiv.org/abs/2306.05685)Cited by:[§2](https://arxiv.org/html/2608.29209#S2.SS0.SSS0.Px3.p1.1)\.

## Appendix ACommunities and Culture Representatives

Figure[4](https://arxiv.org/html/2608.29209#A1.F4)situates the five communities included in this study across the African continent, and Table[3](https://arxiv.org/html/2608.29209#A1.T3)summarizes participant demographics\.

![Refer to caption](https://arxiv.org/html/2608.29209v1/images/african_cultures_studied_v2.png)Figure 4:Geographic distribution of the five African communities included in our study\.Our study spans five African cultural communities across West, East, Southern, and Southeast Africa: Hausa in Nigeria, Kikuyu and Luo in Kenya, AmaXhosa in South Africa, and Xichangana in Mozambique\. These communities differ in language, geography, religious and historical context, naming practices, foodways, dress norms, and everyday social expectations\. We therefore treat them as distinct evaluative contexts\.

The Hausa community in this study is situated in northern Nigeria, where Islamic practice, market life, Hausa naming conventions, modest dress expectations, and foods such astuwo,masa, andmiyan kukashape everyday cultural interpretation\. The Kikuyu community is situated in central Kenya, where stories often draw on settings such as Kiambu, agricultural livelihoods, local foods, family relations, and Kikuyu naming practices\. The Luo community is situated in western Kenya, where Luo names, kinship relations, lakeside and rural–urban settings, foods, and social forms of respect provide important cultural context\. The AmaXhosa community is situated in South Africa, where AmaXhosa language practices, kinship relations, forms of address, dress, and institutional interactions shape how stories are interpreted\. The Xichangana community is situated in Mozambique, where local Portuguese usage, Xichangana language and expressions, food practices, respect norms, and public behavior distinguish the community from both European Portuguese and other African contexts\.

Table 3:Culture representative demographics by community\.
## Appendix BStory Generation Details

This appendix provides supplementary materials for the story generation pipeline described in Section[3\.1](https://arxiv.org/html/2608.29209#S3.SS1), including the workflow, persona fields, lifestyle questions, story distribution, and narrative arcs\.

Table 4:Everyday diabetes lifestyle questions used for narrative generation\.Table 5:Context settings used to vary the amount of persona and cultural information available during story generation\.Table 6:Persona fields used to construct culturally situated personas\.
## Appendix CAnnotation Interface and Evaluation Protocol

This appendix provides the annotation platform, story\-level scoring rubric, and focus group guiding questions referenced in Section[3](https://arxiv.org/html/2608.29209#S3)\.

Table 7:Cultural marker categories used for both text and image annotation\. The same categories were applied across modalities, with annotators selecting text spans or image regions that influenced their cultural alignment judgments\.Table 8:Overall cultural alignment scoring rubric used by culture representatives and LLM judges\.![Refer to caption](https://arxiv.org/html/2608.29209v1/images/platform1.png)Figure 5:Annotation platform used by culture representatives to evaluate text and image cultural alignment\.![Refer to caption](https://arxiv.org/html/2608.29209v1/images/platform2.png)Figure 6:Annotation platform: Assigning a cultural marker category to a highlighted text span or region\.![Refer to caption](https://arxiv.org/html/2608.29209v1/images/platform3.png)Figure 7:Annotation platform: Assigning overall culture alignment scoring\.Table 9:Guiding questions used in focus group discussions with culture representatives\.Table 10:Story Arcs narrative progession![Refer to caption](https://arxiv.org/html/2608.29209v1/cultural_marker_taxonomy.png)Figure 8:Cultural alignment taxonomy with examples from FGD sessions\.
## Appendix DPrompts

Generation System PromptRole:A helpful hospital software operating in the patient’s location\.Task:Receive instructions and generate a first\-person short story based in the patient’s location with a happy ending\.The foods, clothing, and housing should reflect the financial status of neighbourhoods found in the patient’s location\.The narrative must:•Use first\-person voice•Sound like the person is directly telling their story•Be culturally and contextually groundedConstraints:•For image prompts: \{sys\_img\_format\_flux\}Output Requirements:•Write the paragraphs first, story should first person account and all paragraphs must be in \{language\}\.•Then provide image prompts•The image prompts should be written entirely in English\.•Return everything in JSON format•Return only paragraphs written entirely in \{language\}\.•Do not include explanations or any other language\.JSON Format:``` {sys_json_format} ```

Generation User PromptOur persona has the following profile: \{persona\_profile\}Using the profile of \{px\_name\} and understanding the persona’s needs \("\{question\}"\), generate a four\-paragraph health story written in \{language\} that centers around the needs of the persona, which is based on the following question: \{question\}\.The story will have a small work task that foreshadows a larger work task\. This will serve as external motivation for the persona to make health changes\.\{story\_arc\}

Judge System PromptYou are a cultural\-alignment judge\. Rate how culturally aligned an LLM\-generated story is with its target culture\. Be precise, cite text spans, and ground every claim in observable cultural markers\. \#\# Rating scale \(0–100\) \{ 0: "No cultural markers or relevance", 1–20: "Minimal mention, not integrated or relevant", 21–40: "Some markers, weak integration and limited relevance", 41–60: "Moderate reflection with partial relevance", 61–80: "Most markers thoughtfully included and question\-aligned", 81–100: "Strong cultural alignment and full relevance to the question"\} \#\# Cultural marker categories \(use these keys exactly\) \{\_MARKER\_LIST\} \#\# Your task1\.Read the story and original question\.2\.If paragraph\-image pairs are provided, evaluate whether Paragraph 1 matches Image 1, Paragraph 2 matches Image 2, and so on\.3\.Check whether the story maintains cultural coherence across all paragraphs, not just isolated details\.4\.If persona or image context is provided, use it as supporting context only\.5\.For EACH cultural marker you find, quote the exact text span and assign it, a category key from the list above\.6\.For EACH important image\-based cultural alignment/misalignment \(using persona’s context\) or mismatch, add an item to Image\_cultural\_markers\_found and assign it, a category key from the list above and provide a comment7\.Assess depth and authenticity of cultural integration, including paragraph\-image alignment\.8\.Penalize mismatched, generic, or culturally contradictory paragraph\-image pairs\.9\.Score 0\-100 and cite the bracket\.\#\# Output — respond with ONLY this JSON object, nothing else<think\>

Judge User Prompt\#\# Culture \{culture\} \#\# Original question / prompt \{question\} \#\# Story \{story\_text\} \#\# Persona / profile \{persona\} Assume Pair 1 maps to Paragraph 1 and Image 1, Pair 2 maps to Paragraph 2 and Image 2, and so on\.Paragraph \[idx \+ 1\]\{paragraphs\[idx\]\}Image \[idx \+ 1\] reference\{image\_urls\[idx\]\}

## Appendix EAdditional Quantitative Results

Table 11:Score bias, computed as judge mean minus culture\-representative mean, by community\. Bold values indicate significant pairedtt\-test differences between judge and human scores\. Xichangana inflation is systematic across all judges\.Table 12:Human annotator agreement per community\. ICC\(A,1\) = absolute agreement for a single rater\. ICC\(A,k\) = absolute agreement averaged acrosskkraters\.Figure 13:Distribution of connected cultural markers in text and images across the five broader cultural marker categories, broken down by generation model and community\. This figure complements Figure[3](https://arxiv.org/html/2608.29209#S4.F3), which shows not\-connected markers\.### E\.1Cultural Markers

##### Data collected

Culture representatives produced 18,805 span\-level annotations across 199 stories \(9,386 text spans and 9,419 image regions\)\. Of the text spans, 8,075 \(86%\) were labeled as culturally connected and 1,311 \(14%\) as not connected\. For image regions, 6,931 \(74%\) were connected and 2,488 \(26%\) not connected \(Table[13](https://arxiv.org/html/2608.29209#A5.T13)\)\. These raw counts include overlapping annotations where multiple representatives independently identified the same cultural marker\. After deduplication \(exact text match for spans; Intersection over Union\>\>0\.3 clustering for image bounding boxes\), the 8,075 connected text annotations correspond to 6,056 unique spans, of which 1,016 \(16\.8%\) were independently identified by more than one representative\. Similarly, the 6,931 connected image annotations correspond to 5,435 unique regions, with 1,171 \(21\.5%\) identified by multiple representatives\. The higher image overlap rate \(21\.5% vs\. 16\.8% for text\) suggests that visually salient cultural markers in images are more consistently recognized across annotators than textual ones\.

Table 13:Text span annotation statistics per language\.Connected: spans marked as culturally connected\.Not connected: spans marked as not culturally connected\. Deduplication is shown at two levels:Exact\(identical span text within a task\) andPartial\(token overlap\>\>50% of the shorter span with same category\)\.Overlap: spans independently identified by more than one representative\.Table 14:Image region annotation statistics per language\. Unique regions are identified by clustering bounding boxes with IoU \(Intersection over Union\)\>\>0\.3 across annotators\.Overlap: regions independently identified by more than one representative\.
##### Cultural Marker Span Evaluation

To assess how well the LLM judge identifies culturally relevant text spans to support their reasoning, we adapted evaluation methodology from Named Entity Recognition \(NER\)\. This analysis focuses exclusively on the text modality; We constructed ground truth from human annotations where evaluators marked spans as culturally connected, aggregated across annotators via majority vote\. We then evaluated the LLM judge predictions against this ground truth using three matching schemes:Strict\(exact span text and category match\),Partial\(token overlap\>\>50% of the shorter span with correct category\), andType\(correct category regardless of span boundaries\)\.

Following NER evaluation literature, we adopt a proportional overlap threshold rather than the binary any\-overlap criterion\. The CoNLL shared tasks established exact\-match as the standard for NER evaluation[Tjong Kim Sang and De Meulder \(2003\)](https://arxiv.org/html/2608.29209#bib.bib31), while the MUC evaluation framework introduced partial matching categories \(correct, incorrect, partial, missing, spurious\) to capture boundary errors more granularly[Chinchor and Sundheim \(1993\)](https://arxiv.org/html/2608.29209#bib.bib32)\. SemEval\-2013 further formalized four evaluation modes—strict, exact, partial, and type—where partial matching counts any span overlap as a match[Segura\-Bedmar et al\. \(2013\)](https://arxiv.org/html/2608.29209#bib.bib33)\. However, the any\-overlap criterion assumes that predicted and gold spans have comparable granularity\. In our setting, the LLM judge produces longer spans \(mean 6\.9 tokens, median 8\) while human annotators mark shorter, phrase\-level spans \(mean 5\.8 tokens, median 3\), creating a structural asymmetry where coincidental single\-token overlaps \(e\.g\., function words such as “the”, “my”, “I”\) generate spurious matches between semantically unrelated spans\. We therefore require\>\>50% token overlap relative to the shorter span, ensuring that matched pairs share substantive lexical content\.

#### E\.1\.1Ground Truth Example

Figure[14](https://arxiv.org/html/2608.29209#A5.F14)illustrates a ground truth annotation from a randomly selected Hausa story\. Culture representatives identified specific text spans and assigned each to a cultural marker category\. The example shows how a single story paragraph may contain overlapping cultural dimensions, such as names, dietary practices, local expressions, and occupation routines, that evaluators disambiguate at the span level\.

“Hai, my name isAmina Yusufnames\_forms\_of\_address, and life as atraderoccupation\_daily\_routineinKanoplace\_physical\_environmentmarket is…busy,wallahilanguage\_local\_expression\! Every day is a hustle\.I sell beautiful fabrics, Ankara and laceeconomy\_livelihood\_strategies, you know? Bright colours, good quality\. But let me tell you, thinking about what to eat is always last on my mind\. Usually, it’s whatever’s quickest – maybe somemasafood\_dietary\_practiceswith a littlestewfood\_dietary\_practices, ortuwo shinkafafood\_dietary\_practiceswithmiyan kukafood\_dietary\_practices…”

Figure 14:Example ground truth annotation from the Hausa showing cultural marker spans identified by culture representatives\. Each highlighted span is assigned a cultural category\. The ground truth aggregates all connected spans across annotators via majority vote\.We first measured pairwise agreement among human culture representatives to establish a ceiling for automated evaluation\. Table[15](https://arxiv.org/html/2608.29209#A5.T15)reports average pairwise span\-level F1 and category agreement across annotator pairs\.

Table 15:Human inter\-annotator agreement on cultural marker spans \(pairwise average\)\. Only spans marked as culturally connected are included\.Human annotators achieve moderate span overlap \(F1 = 0\.36–0\.47\) and moderate\-to\-substantial category agreement \(0\.59–0\.74\)\. This reflects the inherent subjectivity of cultural marker identification: annotators often agree on which cultural categories are present but differ in where they draw span boundaries\. These scores establish an empirical upper bound for what can be expected from automated systems\.

#### E\.1\.2LLM Judge Self\-Consistency

We ran the LLM judge three times under identical configurations and measured pairwise agreement across runs\. Table[16](https://arxiv.org/html/2608.29209#A5.T16)reports inter\-run agreement\.

Table 16:LLM judge inter\-run agreement \(pairwise average across 3 runs\)\.The model demonstrates high self\-consistency \(Span F1 = 0\.68–0\.82, Category Agreement = 0\.76–0\.91\), substantially exceeding human agreement\. This suggests that the judge produces stable outputs across repeated evaluations, though lower consistency for Xichangana suggests greater uncertainty for that language\.

#### E\.1\.3LLM Judge vs\. Human Ground Truth

Table[17](https://arxiv.org/html/2608.29209#A5.T17)presents the main evaluation, LLM judge predictions, obtained by aggregating annotated text spans across judges through majority voting, compared against the human\-annotated ground truth using NER\-style metrics\. Results are averaged across three runs\.

Table 17:LLM judge vs\. human ground truth \(averaged across 3 runs\)\. P = Precision, R = Recall, F1 = F1 score\.
#### E\.1\.4Analysis

The large gap between strict matching \(F1 = 0\.09\) and type matching \(F1 = 0\.46\) reveals that the LLM judge identifies relevant culturalcategoriesat moderate accuracy but differs substantially from humans in span boundary selection\. Partial matching \(F1 = 0\.22\) confirms that span overlap is limited\. We identify three primary sources of disagreement:

- •Span granularity mismatch\.Human annotators typically mark short, specific phrases \(e\.g\., “fried dough”, “morning prayers”\), while the LLM judge tends to extract longer, sentence\-level spans that encompass the cultural marker along with its surrounding context\. This systematically reduces strict and partial match scores without necessarily reflecting a disagreement about cultural content\.
- •Span mismatch\.Although the total number of model predictions \(1,134–1,701 per language\) is sometimes comparable to the human ground truth \(781–2,453\), the model’s spans frequently do not align with human annotations\. For AmaXhosa, the model produces roughly twice as many spans as the ground truth \(1,630 vs\. 781\), while for Luo and Kikuyu the model produces fewer spans than the human ground truth \(1,636 vs\. 2,453 and 1,479 vs\. 2,068 respectively\)\. In all cases, the low precision and recall scores indicate that model and human spans identify different portions of text as culturally relevant, even when they agree on the category\.
- •Category confusion\.Even when spans partially overlap, the model and humans sometimes assign different cultural categories\. The type\-level F1 \(0\.41–0\.52\) indicates that approximately half of the model’s category assignments align with human judgments\. Category confusion is most pronounced for broad categories such asplace\_physical\_environmentandoccupation\_daily\_routine, which the model tends to over\-predict at the expense of more specific categories\.
- •Ceiling effects\.Human inter\-annotator agreement on span boundaries \(F1 = 0\.42\) provides context for interpreting model performance\. The model’s type\-level F1 \(0\.46\) approaches human category agreement \(0\.67\), suggesting that while boundary alignment remains challenging, the model’s category\-level understanding is within reach of human performance when boundary constraints are relaxed\.

In general the performance varies across languages, Luo achieves the highest strict F1 \(0\.107\) and type precision \(0\.637\), while AmaXhosa shows the lowest strict F1 \(0\.072\) but highest type recall \(0\.623\)\. These differences may reflect variation in annotation density, category distribution, or the degree to which cultural markers in each language are expressed through discrete, identifiable phrases versus diffuse narrative context\.

Similar Articles

Characterizing Cultural Localization in AI-Generated Stories

arXiv cs.CL

This paper proposes a method to measure cultural localization in AI-generated stories, detecting that only a small fraction of vocabulary distinguishes nationalities while narratives rely on shared templates, and finds that cultural markers from many Global South countries are often offensive.

Toward a Theory of Value in AI Alignment

arXiv cs.AI

This paper analyzes 94 AI alignment research papers to examine how human values are conceptualized, finding that many rely on preferences and synthetic data, which risks reducing complex cultural values to binary choices.