Tag
OmniAlign is a unified multilingual aligner that supports both word-level and sentence-level alignment using a single lightweight model, achieving competitive performance on benchmarks and generalizing to unseen language pairs.
This arXiv paper audits five frontier LLMs on native Bangla derogatory speech, finding that safety alignment fails to generalize to low-resource languages — models comprehend and generate unsafe content at high rates despite high-resource alignment. The authors propose a 'comprehension–containment decoupling' and show that reasoning and persona framing further break down safety filters.