Tag
This arXiv paper audits five frontier LLMs on native Bangla derogatory speech, finding that safety alignment fails to generalize to low-resource languages — models comprehend and generate unsafe content at high rates despite high-resource alignment. The authors propose a 'comprehension–containment decoupling' and show that reasoning and persona framing further break down safety filters.