Tag
Introduces CCBench, a framework for evaluating LLMs' cultural competence via health queries with personas across six cultures, finding that even top models achieve only 20-30% culturally appropriate responses.
This paper introduces CulturalNB, a dataset of Bengali cultural question-answer pairs, and evaluates nine LLMs for cross-lingual cultural bias. Findings show that English prompting increases global narrative substitution and reduces local perspectives, revealing that cultural failures in LLMs are grounding and prioritization issues, not just missing knowledge.
Empirical study showing frontier LLMs encode a culturally skewed baseline that privileges Western viewpoints when describing and judging global streetscapes, with non-Western prompts systematically deviating more from the default.