german-language

Tag

Cards List
#german-language

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

arXiv cs.CL · 2026-08-19 Cached

This paper introduces MÖVE, a holistic evaluation framework for LLMs in the German public sector, examining governance dimensions like energy consumption, provider transparency, and knowledge of German-party positions, revealing trade-offs that necessitate context-specific model selection.

0 favorites 0 likes
#german-language

Authorship Verification of Transcribed German-Language Videos

arXiv cs.CL · 2026-08-03 Cached

This paper studies authorship verification on transcribed German-language videos, comparing traditional n-gram methods with transformer-based approaches across three self-compiled corpora. Traditional character and token n-gram methods outperformed modern transformers, achieving up to 88% accuracy and 90% AUC.

0 favorites 0 likes
#german-language

Soofi – Sovereign Open Source Foundation Models

Hacker News Top · 2026-07-20 Cached

Soofi introduces Soofi S, a 30B parameter Mixture-of-Experts open source foundation model trained on 27 trillion tokens, targeting industrial AI applications in German and English. The model is part of a European sovereign AI initiative.

0 favorites 0 likes
#german-language

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

arXiv cs.CL · 2026-06-03 Cached

This paper introduces ChristBERT, a family of domain-specific RoBERTa-based language models for German clinical NLP, and evaluates three domain adaptation strategies (continued pre-training, pre-training from scratch, and vocabulary adaptation) on medical named entity recognition and text classification tasks, achieving state-of-the-art results.

0 favorites 0 likes
#german-language

KletterMix: Climbing Toward High-Quality German Pretraining Data

Hugging Face Daily Papers · 2026-06-02

KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback