Tag
FUSE is a unified framework for evaluating dangerous capabilities of large language models, assessing knowledge, defense, and harm dimensions. The study finds that newer models have increased dangerous capabilities despite alignment progress, and it provides cross-model comparisons.
This paper investigates the capability of automated alignment researchers to mitigate alignment failures such as deception and sycophancy, demonstrating that they can outperform human researchers and generalize to larger models while preserving capabilities.
This systematic literature review examines safety alignment of large language models in low-resource languages, identifying a persistent multilingual safety gap and suggesting future directions such as culturally grounded benchmarks and participatory data collection.
DisaBench is a participatory evaluation framework co-created with people with disabilities that introduces a taxonomy of 12 disability harm categories and a dataset of 175 prompts to evaluate harms in language models, revealing that standard safety benchmarks miss subtle, expert-identified harms.