TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar
Summary
This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.
View Cached Full Text
Cached at: 09/21/26, 08:58 AM
# TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar Source: [https://arxiv.org/abs/2609.20832](https://arxiv.org/abs/2609.20832) [View PDF](https://arxiv.org/pdf/2609.20832) > Abstract:We introduce TatBLiMP, the first benchmark of linguistic minimal pairs for Tatar \(tt, ISO 639\-3 tat\), a Qypchaq Turkic language written in Cyrillic\. To our knowledge it is the first grammaticality evaluation for Tatar language models of any kind, since even the 101\-language MultiBLiMP does not include Tatar\. TatBLiMP covers 16 morphosyntactic phenomena in 1248 sentence pairs\. Each pair differs by a single morpheme, one grammatical and one ungrammatical\. A model passes a pair when it assigns higher probability to the grammatical member\. Scoring compares probabilities the model already assigns, so the benchmark needs no text generation and no parser, and it runs on base models and on mid\-training checkpoints\. TatBLiMP adapts the phenomenon inventory and single\-morpheme breaking operations of TurBLiMP to Tatar and adds one phenomenon specific to Tatar, bare\-noun number after numerals and quantifiers\. The grammatical member of every pair is an attested sentence from Tatar literary prose\. The ungrammatical member is produced by a deterministic single\-morpheme perturbation with the apertium\-tat transducer\. Every pair is ratified by a native speaker\. A plausibility principle governs construction, so the ungrammatical member is a plausible real\-world error rather than an arbitrary corruption\. Across from\-scratch Tatar models, cross\-lingual adaptations, and frontier multilingual LLMs, the benchmark tracks focused Tatar training rather than parameter scale\. A 478M from\-scratch model and a 125M monolingual model lead near 0\.97, a 7B adaptation trails, frontier LLMs of 30\-120B parameters fall to 0\.80\-0\.92, and a lightly tuned multilingual model is weakest\. We close with the benchmark's main limitation\. Its inherited taxonomy omits the morphophonology, vowel harmony and consonant assimilation, that is most salient to native speakers, and we sketch a native second layer that would add it\. ## Submission history From: Ilshat Saetov \[[view email](https://arxiv.org/show-email/6c4fa435/2609.20832)\] **\[v1\]**Thu, 23 Jul 2026 17:16:31 UTC \(69 KB\)
Similar Articles
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
Presents Tatoxa, a state-of-the-art system for text detoxification in the Tatar language, outperforming existing LLMs. Introduces a new dataset and shows that cross-lingual transfer performs worse than native data.
KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
This paper introduces KyrgyzLLM-Bench, a benchmark suite for evaluating large language models in the Kyrgyz language, comprising both natively authored and translated datasets, and provides a systematic evaluation of 26 models.
TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP
This paper introduces TajPersLexon, a lexical resource for Tajik-Persian cross-script NLP, and benchmarks hybrid models against neural baselines to demonstrate effective low-resource processing.
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
This paper introduces LingT2I, a 10-language, 33K-prompt benchmark for evaluating cross-lingual consistency in text-to-image generation, revealing linguistic inequality and language-dependent trade-offs across content generation and text rendering.
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
The paper presents 5-Dialects-BN, a manually annotated benchmark dataset for five Bangla dialects with aligned transliterations and translations, designed to improve evaluation of dialect-aware LLMs.