Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory
Summary
This paper applies Cultural Consensus Theory to analyze LLM alignment with cultural norms across single and multi-cultural settings using World Values Survey data, showing that models either fail to form cohesive consensus or over-regularize consensus, and offering actionable diagnostics for evaluating true human diversity versus algorithmic homogenization.
View Cached Full Text
Cached at: 08/12/26, 08:31 AM
# Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory Source: [https://arxiv.org/abs/2608.09937](https://arxiv.org/abs/2608.09937) [View PDF](https://arxiv.org/pdf/2608.09937) > Abstract:Recent work in NLP has probed large language models for their understanding of cultural norms across countries\. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country\. In this work, we leverage cultural consensus theory \(CCT\) from cultural anthropology to model such multidimensional nuance\. Applying CCT to the World Values Survey \(WVS\) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over\-regularizing consensus\. Through explicit representation of intra\-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization\. ## Submission history From: John Lalor \[[view email](https://arxiv.org/show-email/ba44a2d5/2608.09937)\] **\[v1\]**Fri, 29 May 2026 17:46:47 UTC \(1,385 KB\)
Similar Articles
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
CulturALL introduces a 2,610-sample benchmark across 14 languages and 51 regions to evaluate LLMs on real-world, culturally grounded tasks; top model scores only 44.48%, highlighting large room for improvement.
CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries
Introduces CCBench, a framework for evaluating LLMs' cultural competence via health queries with personas across six cultures, finding that even top models achieve only 20-30% culturally appropriate responses.
Position: It's Time to Optimize LLMs for Self-Consistency
This position paper argues that many LLM failures stem from evaluating outputs independently and proposes a self-consistency framework that treats diverse techniques as special cases of consistency optimization.
The Culture Funnel: You Can't Align What isn't in the Data
This paper introduces the 'culture funnel' concept, demonstrating that cultural signals in LLM training data sharply decline during post-training stages. The authors release a 5.6M-sample tagged dataset to help preserve cultural grounding in model alignment.