Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts
Summary
Introduces Schützen, a safety dataset for evaluating LLMs in Bulgarian and German, revealing cross-language differences in safety behavior and advocating for region-specific evaluation resources.
View Cached Full Text
Cached at: 06/11/26, 01:36 PM
# Schützen: Evaluating LLM Safety in Bulgarian and German Contexts Source: [https://arxiv.org/abs/2606.11316](https://arxiv.org/abs/2606.11316) [View PDF](https://arxiv.org/pdf/2606.11316) > Abstract:Large language models are increasingly deployed across professional domains, bringing hard\-to\-predict risks, including the generation of harmful or disrespectful content\. Although substantial progress has been made in developing safety evaluation datasets, existing resources remain overwhelmingly English\- and Chinese\-centric\. This limitation is particularly pronounced when evaluating languages that operate within shared sociocultural, legal, and ethical contexts\. To address this gap, we introduce Schützen: a German\-\-Bulgarian safety dataset designed to assess model answerability under risk, covering both a low\-resource language \(Bulgarian\) and a high\-resource language \(German\)\. Experiments with multilingual and language\-specific LLMs reveal pronounced cross\-language differences in safety behavior, highlighting the necessity of tailored, region\-specific evaluation resources to support the responsible deployment of LLMs in Germany and Bulgaria\. Datasets and code are available at[this https URL](https://github.com/xnlp-lab/Schutzen)\. Warning: this paper contains examples that may be offensive, harmful, or biased\. ## Submission history From: Yuxia Wang \[[view email](https://arxiv.org/show-email/736ab93f/2606.11316)\] **\[v1\]**Tue, 9 Jun 2026 18:01:19 UTC \(153 KB\)
Similar Articles
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
This systematic literature review examines safety alignment of large language models in low-resource languages, identifying a persistent multilingual safety gap and suggesting future directions such as culturally grounded benchmarks and participatory data collection.
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
This paper introduces a framework for validating comparative LLM safety scoring without ground-truth labels, using an 'instrumental-validity chain' to establish deployment evidence. It demonstrates the method using a local-first tool called SimpleAudit on Norwegian safety packs and compares models like Borealis and Gemma 3.
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
This paper presents the first comprehensive empirical study of safety impacts of benign multilingual fine-tuning on LLMs, showing that safety outcomes vary drastically by language and that assessing only English is insufficient.
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
This paper investigates the ability of LLMs-as-judges for safety to adapt to contextual information and varying safety definitions, finding that they are largely rigid and fail to adjust when the context contradicts their internal priors.