Tag
This paper introduces CulShield, the first public benchmark dedicated to evaluating and improving the cultural taboo safety of large language models across 77 countries and over 2,020 taboos, revealing a knowledge-behavior gap where models fail to apply known taboos in interaction.