Sch\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts

arXiv cs.CL Papers

Summary

Introduces Schützen, a safety dataset for evaluating LLMs in Bulgarian and German, revealing cross-language differences in safety behavior and advocating for region-specific evaluation resources.

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disrespectful content. Although substantial progress has been made in developing safety evaluation datasets, existing resources remain overwhelmingly English- and Chinese-centric. This limitation is particularly pronounced when evaluating languages that operate within shared sociocultural, legal, and ethical contexts. To address this gap, we introduce Sch\"{u}tzen: a German--Bulgarian safety dataset designed to assess model answerability under risk, covering both a low-resource language (Bulgarian) and a high-resource language (German). Experiments with multilingual and language-specific LLMs reveal pronounced cross-language differences in safety behavior, highlighting the necessity of tailored, region-specific evaluation resources to support the responsible deployment of LLMs in Germany and Bulgaria. Datasets and code are available at https://github.com/xnlp-lab/Schutzen. Warning: this paper contains examples that may be offensive, harmful, or biased.
Original Article
View Cached Full Text

Cached at: 06/11/26, 01:36 PM

# Schützen: Evaluating LLM Safety in Bulgarian and German Contexts
Source: [https://arxiv.org/abs/2606.11316](https://arxiv.org/abs/2606.11316)
[View PDF](https://arxiv.org/pdf/2606.11316)

> Abstract:Large language models are increasingly deployed across professional domains, bringing hard\-to\-predict risks, including the generation of harmful or disrespectful content\. Although substantial progress has been made in developing safety evaluation datasets, existing resources remain overwhelmingly English\- and Chinese\-centric\. This limitation is particularly pronounced when evaluating languages that operate within shared sociocultural, legal, and ethical contexts\. To address this gap, we introduce Schützen: a German\-\-Bulgarian safety dataset designed to assess model answerability under risk, covering both a low\-resource language \(Bulgarian\) and a high\-resource language \(German\)\. Experiments with multilingual and language\-specific LLMs reveal pronounced cross\-language differences in safety behavior, highlighting the necessity of tailored, region\-specific evaluation resources to support the responsible deployment of LLMs in Germany and Bulgaria\. Datasets and code are available at[this https URL](https://github.com/xnlp-lab/Schutzen)\. Warning: this paper contains examples that may be offensive, harmful, or biased\.

## Submission history

From: Yuxia Wang \[[view email](https://arxiv.org/show-email/736ab93f/2606.11316)\] **\[v1\]**Tue, 9 Jun 2026 18:01:19 UTC \(153 KB\)

Similar Articles

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

arXiv cs.CL

Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.