Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context
Summary
Mizan introduces a national benchmark for evaluating large language models on Iraqi Arabic and civic context, highlighting gaps in MSA-focused evaluations and revealing issues like over-refusal in safety-hardened models.
View Cached Full Text
Cached at: 09/15/26, 08:48 AM
# Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context Source: [https://arxiv.org/html/2609.13980](https://arxiv.org/html/2609.13980) Mustafa S\. AljumailyAffiliation:Missan Oil Company, Amarah, IraqMembers of the National Team for the Iraqi Large Language Model,Prime Minister’s Office, Baghdad, Iraq ###### Abstract Arabic large\-language\-model \(LLM\) evaluation has matured around Modern Standard Arabic \(MSA\): aggregated leaderboards such as the Open Arabic LLM Leaderboard \(OALL\), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, and frontier systems increasingly saturate them\. Dialectal Arabic \- the language Iraqis actually speak \- remains nearly invisible to this infrastructure\. We introduce Mizan \(”the balance”\), Iraq’s national benchmark for evaluating LLMs on Iraqi Arabic and the Iraqi civic context: an MSA baseline track paired with an Iraqi track across six axes \- dialect comprehension, dialect generation, bidirectional MSA\-Iraqi translation, Iraq\-specific knowledge, official\-document field extraction, and safety \- built from 340 originally authored, dually reviewed items with statistically audited answer positions and Wilson intervals on every published score\. A pilot evaluation of 27 systems \- closed frontier models three days after release, open weights across size tiers, and an Arabic trio spanning commercial, open\-specialized, and sovereign systems \- yields four findings\. The MSA track saturates while the Iraqi track discriminates, with a consistent 14\-18\-point per\-model gap and statistically tied leaders\. Official\-document extraction confines every system to 32\-56\. Arabic\-focused specialization behaves as MSA specialization: two dedicated Arabic models score below a size\-matched generalist on the Iraqi track\. And the safety\-hardened tier of the newest model family deterministically refuses innocuous dialect\-comprehension items as policy violations \- an over\-refusal mode invisible to MSA benchmarks\. The platform enforces an integrity protocol of immutable snapshots, verification certificates, a human publication gate, and public retraction, all exercised during this study\. Code and the public development set accompany the paper\. Keywords: LLM Evaluation, Dialectal Arabic, Low\-Resource NLP, Over\-Refusal, Document Extraction, Benchmark Integrity\. ## 1Introduction Evaluation infrastructure decides what progress means\. For Arabic, that infrastructure has consolidated around Modern Standard Arabic: ArabicMMLU and AlGhafa established rigorous MSA test sets; the Open Arabic LLM Leaderboard aggregates them at scale; HELM Arabic extends a transparent multi\-benchmark methodology to Arabic; and BALSAM[Al\-Matham and others \(2025\)](https://arxiv.org/html/2609.13980#bib.bib6)pools tens of thousands of test questions across dozens of task categories under a regional consortium\. This consolidation succeeded \- to the point of self\-obsolescence\. On our MSA track, contemporary systems cluster at or near a perfect score, the newest frontier model matches them within 72 hours of its public release, and the two strongest systems cannot be separated at all: they tie, and swap ranks between independent evaluation rounds\. A benchmark that no longer separates systems no longer measures them\. What saturation conceals is the language people speak\. Iraqi Arabic serves tens of millions of speakers across registers that diverge from MSA in phonology, morphology, lexicon, and pragmatics \- divergences deep enough to produce systematic false friends \(Iraqi hwaya, ”many”, against MSA hiwaya, ”hobby”\) and to defeat literal translation in both directions\. Yet no comprehensive LLM evaluation framework has existed for it\. The dialect resources the field possesses \- Multi\-Arabic Dialect Applications and Resources \(MADAR\), Nuanced Arabic Dialect Identification \(NADI\) \- belong to an earlier paradigm: identification and classification corpora built for supervised NLP\. Recent dialect\-aware evaluation, most notably Arabic Dialect and Cultural Evaluation \(AraDiCE\), has begun to close the distance for Arabic dialects broadly \- through machine\-translated, post\-edited derivatives of existing benchmarks, covering Levantine and Egyptian varieties and the cultural contexts of the Gulf, Egypt, and the Levant\. Iraqi Arabic, with its internal regional variation and its distinctive civic register, appears in none of it\. The stakes are not academic\. States are adopting language models inside public administration, and Iraq has constituted a national team, by prime\-ministerial order, to build sovereign Arabic\-and\-Iraqi language capability\. A model proposed for Iraqi state use must read an official letter \- issuing authority, reference number, date, addressee, subject, required action \- and must navigate Iraqi social sensitivities without either enabling harm or refusing legitimate civic questions\. No existing benchmark measures any of this, in any dialect, anywhere\. A national evaluation instrument is therefore not a convenience but a precondition of informed procurement: whatever model is offered to the Iraqi state should be measured on the national scale first\. Mizan \(”the balance” or ”the scale”\) is that instrument, and its design answers each failure above by construction\. Two tracks of equal standing \- an MSA baseline the international literature can read, and an Iraqi track as the discriminator nothing else provides\. Six axes spanning comprehension, generation, bidirectional MSA\-Iraqi translation, Iraq\-specific knowledge, official\-document field extraction, and safety in Iraqi context\. Every item originally authored by native speakers and dually reviewed \- nothing translated from foreign benchmarks, because translated items measure translationese rather than dialect competence\. Answer positions balanced and audited\. Generative axes scored by human judges with numbered rubrics\. The accompanying platform enforces an integrity protocol \- uniform conditions with a logged retry ladder, per\-item audit records, dated immutable snapshots, SHA\-256 certificates, a human publication gate, and public retraction \- which this very study exercised: diagnostic charts exposed two technically wounded runs, both were retracted in public view, both models were re\-evaluated cleanly, and the full board was then re\-verified end to end\. We report a pilot study of 27 systems evaluated under identical conditions on the 340\-item bank: closed frontier models \(including GPT\-6 Astra, evaluated three days after release\), open weights across size tiers, and an Arabic trio spanning a commercial API system, an open specialized model, and a sovereign state model\. Four regularities organize the results\. The saturation\-discrimination split: MSA compresses at the ceiling while Iraqi scores spread across more than twenty points, each model carrying a consistent 14\-18\-point internal gap\. The document wall: official\-document extraction confines every system \- newest frontier included \- to between 32 and 56\. The specialization law: models marketed as Arabic\-specialized behave as MSA\-specialized \- two dedicated Arabic models, one open and one sovereign, score below a size\-matched generalist on the Iraqi track, while a third exceeds its own generalist sibling yet remains well below the frontier\. And an over\-refusal mode invisible to every MSA benchmark: the safety\-hardened tier of the newest model family deterministically refuses innocuous Iraqi comprehension items \- questions asking the meaning of everyday words \- classifying them as policy\-violating content, while answering the entire MSA track without hesitation\. This paper makes five contributions: Mizan, to our knowledge the first comprehensive, originally\-authored evaluation framework dedicated to Iraqi Arabic and the Iraqi civic context, with a dual\-track design and six evaluation axes; an official\-document extraction axis unmeasured by any existing benchmark, directly tied to sovereign use; an open evaluation platform with an enforced, practice\-tested integrity protocol: snapshots, certificates, a human gate, public retraction, and full\-board re\-verification with run\-to\-run agreement reporting; a 27\-model empirical study, statistically qualified throughout, establishing MSA saturation, the per\-model dialect gap, the universal document wall, and the Arabic\-specialization paradox; the first documented case of dialect\-triggered safety over\-refusal in a frontier system, with official API evidence \- and the public development set, guidelines, and code, released to seed dialect\-aware evaluation beyond Iraq\. Section 2 situates Mizan among Arabic evaluation efforts; Sections 3\-4 present the benchmark and platform; Sections 5\-6 report the pilot study; Sections 7\-10 analyze errors, discuss implications, and state limitations and the v1\.0 roadmap\. ## 2Related Work Arabic LLM evaluation has consolidated along three lines: MSA benchmarks and their aggregated leaderboards, dialect resources from the classification era, and a nascent turn toward dialect\-aware LLM evaluation\. Mizan draws on all three and departs from each\. ### 2\.1MSA benchmarks and aggregated leaderboards ArabicMMLU[Koto et al\. \(2024\)](https://arxiv.org/html/2609.13980#bib.bib2)and AlGhafa[Almazrouei et al\. \(2023\)](https://arxiv.org/html/2609.13980#bib.bib3)established rigorous multiple\-choice test sets for Arabic, with ArabicMMLU in particular built natively rather than translated \- a design stance we adopt and extend\. The Open Arabic LLM Leaderboard[Elfilali and others \(2024\)](https://arxiv.org/html/2609.13980#bib.bib4)aggregates such benchmarks at community scale for open models\. Holistic Evaluation of Language Models \(HELM\) Arabic[Stanford CRFM and Arabic\.AI \(2025\)](https://arxiv.org/html/2609.13980#bib.bib5)extends the HELM framework’s transparent, reproducible multi\-benchmark methodology to Arabic across seven established test sets, with an Enterprise variant addressing financial and legal tasks\. BALSAM[Al\-Matham and others \(2025\)](https://arxiv.org/html/2609.13980#bib.bib6)pools on the order of fifty thousand test questions across sixty\-seven task categories under a MENA\-wide consortium, with blind test sets and a standard submission interface\. This infrastructure is mature, rigorous, and MSA\-bound: none of it observes dialect, local knowledge, or the civic document register\. Our results supply an empirical corollary: its ceiling has arrived\. On Mizan’s MSA track, multiple contemporary systems score at or near 100\.0, and the strongest two are statistically inseparable \(Section 6\)\. ### 2\.2Dialect resources of the classification era MADAR[Bouamor et al\. \(2018\)](https://arxiv.org/html/2609.13980#bib.bib7)and the NADI shared tasks[Abdul\-Mageed and others \(2020\)](https://arxiv.org/html/2609.13980#bib.bib1)built the field’s foundational dialect corpora and dialect\-identification tasks, spanning city\- and country\-level varieties\. These are resources for supervised NLP: they measure whether a system can recognize or label a dialect, not whether a generative model can understand, produce, translate, or safely operate in one\. Mizan inherits their insistence on regional granularity \- every Iraqi item carries a dialect\-region tag \- while asking the generative\-era question they predate\. ### 2\.3Dialect\-aware LLM evaluation The closest work to ours is AraDiCE[Mousi et al\. \(2025\)](https://arxiv.org/html/2609.13980#bib.bib8), which contributed approximately 45,000 samples created by machine translation with human post\-editing from existing benchmarks, evaluated dialect comprehension and generation for Levantine and Egyptian Arabic, and introduced a fine\-grained cultural benchmark spanning the Gulf, Egypt, and the Levant\. AraDiCE demonstrated that Arabic\-centric models outperform multilingual ones on dialectal tasks and that dialect generation remains hard\. Contemporaneous work extends the same translation\-based line: DialectalArabicMMLU[Altakrori et al\. \(2025\)](https://arxiv.org/html/2609.13980#bib.bib12)manually translates and adapts three thousand MMLU\-Redux question\-answer pairs into five dialects \(Syrian, Egyptian, Emirati, Saudi, and Moroccan\), confirming persistent dialectal gaps across nineteen open\-weight models\. Together these efforts define precisely where Mizan departs\. Provenance: AraDiCE derives items by translating existing benchmarks; Mizan authors every item natively, because translated items inherit the source benchmark’s cultural content and carry translationese artifacts \- the item asks what the translator wrote, not what a speaker would say\. Coverage: Iraqi Arabic is absent from AraDiCE’s dialectal tasks and cultural regions, and from all five dialects of DialectalArabicMMLU \- the largest Arab country by population after Egypt appears in neither effort\. Scope: no existing benchmark, dialectal or otherwise, measures official\-document field extraction, human\-rubric generation quality with inter\-annotator agreement, or safety behavior in a specific national context \- including the dialect\-triggered over\-refusal mode we document in Section 7, which is invisible by construction to translated MSA\-derived test sets\. Governance: Mizan couples its dataset to a live platform with an enforced integrity protocol \- immutable snapshots, verification certificates, a human publication gate, and public retraction \- exercised during this very study\. ### 2\.4Positioning Mizan is complementary to this landscape, not competitive with it\. Shared\-Arabic competence is measured well and at scale by OALL, HELM Arabic, and BALSAM; regional dialect breadth by AraDiCE and DialectalArabicMMLU\. Mizan contributes depth in one nation’s linguistic and civic reality which was originally authored, dually reviewed, statistically audited, sovereignly anchored and, empirically, the discriminative signal that the saturated MSA axis no longer provides\. To our knowledge, Mizan is the first comprehensive, originally\-authored evaluation framework dedicated to Iraqi Arabic and the Iraqi civic context\. ## 3The Mizan Benchmark ### 3\.1Design principles Five principles govern the benchmark, each a response to a documented failure mode\. Original authorship: every item is written natively in its target variety by the authoring team and individually reviewed, corrected, and approved by native Iraqi speakers among the authors; nothing is translated from existing benchmarks, because a translated item measures the translator’s output rather than the speaker’s language, and inherits the source benchmark’s cultural frame[Koto et al\. \(2024\)](https://arxiv.org/html/2609.13980#bib.bib2);[Mousi et al\. \(2025\)](https://arxiv.org/html/2609.13980#bib.bib8)\. Dual review with accumulated rulings: every item passed a second linguistic review, and the rulings this produced \- lexical authenticity constraints, register boundaries, orthographic conventions \- accumulated into binding authoring guidelines applied to all subsequent items\. Contamination tiering: items carry a tier field separating the public development set from a reserved private\-test tier; the entire pilot bank is deliberately public\_dev, and the private tier remains empty until the next development phase, a limitation we state rather than obscure \(Section 9\)\. Statistical balance: correct\-answer positions in multiple\-choice items are balanced by design and audited by a chi\-square goodness\-of\-fit test against the uniform distribution, which fails to reject uniformity for every \(track, axis\) group \(all chi\-square≤1\.20\\leq 1\.20, df=3=3,p\>0\.05p\>0\.05; actual audited counts reported, not planned ones\)\. Human judging where automation misleads: generative axes are scored by human judges with numbered rubrics and inter\-annotator agreement measured by Krippendorff’s alpha[Krippendorff \(2019\)](https://arxiv.org/html/2609.13980#bib.bib9); automatic judges serve only as secondary signals, given documented same\-family stylistic bias in LLM\-as\-judge settings\. ### 3\.2Two tracks of equal standing Every axis runs on two tracks\. The MSA track is the comparative baseline: it makes Mizan legible to the international literature, anchors each model’s dialect gap to its own standard\-Arabic ceiling, and \- empirically \- documents saturation\. The Iraqi track is the discriminator: originally authored Iraqi Arabic across Baghdadi, southern, and mixed registers, with a region tag on every item \(Mosuli coverage is one of the next phase targets\)\. The design prevents the framework from collapsing into either another MSA benchmark or an untethered dialect exercise: a model’s headline story on Mizan is the pair \(MSA score, Iraqi score\) and the distance between them\. ### 3\.3Six axes Dialect comprehension \(multiple choice, auto\-scored\) covers lexical meaning, morphology and syntax, idioms, contextual inference, and MSA\-Iraqi false friends \- e\.g\., asking the meaning of hassa \(”now”\) in a natural sentence\. Dialect generation \(open generation, human\-judged\) elicits Iraqi text under scenario constraints; its central rubric dimension penalizes the characteristic failure we call fusha in disguise \(fusha = MSA/Classical Arabic\): fluent MSA cosmetically sprinkled with dialect markers\. Bidirectional translation \(open generation, human\-judged with Character n\-gram F\-score \(chrF\) as a secondary signal\) tests MSA\-to\-Iraqi and Iraqi\-to\-MSA equally, since the two directions fail differently\. Iraqi knowledge \(multiple choice, auto\-scored\) spans history, geography, constitution and institutions, and popular culture \- including time\-sensitive facts tagged as such \(Iraq’s nineteen governorates following Halabja’s promotion to governorate status in 2025\)\. Official\-document extraction \(structured extraction, auto\-scored\) presents a complete simulated official letter and requires six fields: issuing authority, reference number, date, addressee, subject, and required action \- with deliberately variable fields \(letters lacking an addressee or an action\) and Iraqi administrative conventions enforced verbatim; no real document enters the bank\. Safety in Iraqi context \(open generation, human\-judged\) probes neutrality across components, governorates, and symbols, refusal of genuinely harmful requests, stereotype handling, and loaded\-premise questions; harmful content is described, never instantiated, and prompts use generic framings rather than named targets\. ### 3\.4Authoring workflow and the rejected\-item discipline Items were drafted against per\-axis writing guidelines specifying valid item shapes, positive and negative examples, and reviewer checklists, then individually adjudicated by the native\-speaker reviewers\. Rejection was a working tool, not an exception: candidate items were discarded for lexical rarity \(dead or regionally opaque vocabulary\), ambiguity of key, or contested ground truth\. One rejected item became a design lesson we preserve: a geography question on the Tigris\-Euphrates confluence, classically at Qurna but hydrologically shifted toward Karmat Ali \- two defensible keys, therefore no item\. The bank’s lexical layer follows accumulated authenticity rulings favoring living, widely\-attested usage over dictionary dialect\. ### 3\.5Orthography Iraqi Arabic has no standard orthography; the same word varies across attested spellings\. The pilot adopts documented internal conventions \(including the Persian\-derived letters for Iraqi phonemes\) applied consistently across the bank, and treats a unified, publishable Iraqi orthography guide as a distinct future deliverable rather than claiming one prematurely\. ### 3\.6Composition The pilot bank comprises 340 items: 265 Iraqi\-track and 75 MSA\-track\. By format: 140 multiple\-choice \(50 Iraqi comprehension, 50 Iraqi knowledge, 20 MSA comprehension, 20 MSA knowledge\), 50 official\-document extraction, and 150 open\-generation items pending human judging \(50 generation, 50 translation, 50 safety\)\. Auto\-scored coverage is therefore 190 items per model; the generative axes publish only after the human\-judging campaign \(in progress\), and the leaderboard displays them as pending rather than substituting automatic proxies\. ## 4Platform and Integrity Protocol ### 4\.1Architecture Mizan runs as a bilingual public platform \(Arabic\-default, right\-to\-left \(RTL\) layout; English also supported\) backed by a relational store of items, runs, and per\-axis results, with a public leaderboard, per\-axis tables, confidence intervals, and analytical charts\. Evaluation is deliberately decoupled from the platform: an offline runner executes a bank against a model and emits a self\-describing results file \(schema mizan\-results\-v1\) carrying the harness commit hash, provider, and configuration; the platform imports results and publishes them through a separate explicit action\. Item schemas are structurally compatible with lm\-evaluation\-harness conventions[Gao and others \(2023\)](https://arxiv.org/html/2609.13980#bib.bib10)a widely used open\-source framework for standardized LLM evaluation\. ### 4\.2Uniform conditions and the failure ladder To distinguish genuine model refusals from transient infrastructure failures, every model faces identical prompts, identical scoring, and an identical completion\-budget policy: a 2048\-token base with a uniform retry ladder that, on an empty completion or transient error, retries at 8192 tokens up to twice with backoff, logging every recovery so the console log constitutes a complete audit trail\. The runner additionally emits a per\-item detail sidecar \(item, raw response, score\) for every run \- the substrate of Section 7’s error analysis\. The failure policy is one sentence long and applied without exception: transient failures are retried and recovered; deterministic refusals are scored as failures on the affected items and reported separately as over\-refusal\. ### 4\.3Snapshots, certificates, and public verification A published run is a dated, immutable snapshot of one model version on one bank version under one harness commit\. New model versions produce new runs; nothing is overwritten, so the leaderboard doubles as a longitudinal record\. Every imported run receives a SHA\-256 verification certificate which is a cryptographic fingerprint that lets anyone independently confirm a published score has not been altered; a public verification page checks any certificate and displays revoked ones as revoked rather than hiding them\. ### 4\.4The human gate, retraction, and supersession \- exercised in practice No result reaches the public leaderboard without an explicit human publication decision\. The protocol’s value is easiest to show by the two occasions this study exercised it\. First, retraction: the leaderboard’s diagnostic charts exposed two published runs with near\-zero single tracks \- the signature of infrastructure failure, not model ability \(one model scored below random chance on one track while healthy on the other\)\. Both runs were retracted in public view, their certificates revoked but preserved, and both models re\-evaluated cleanly\. Second, supersession and re\-verification: after hardening the runner, we re\-executed the entire board under the hardened harness with detail sidecars; the resulting runs supersede the originals \(which remain in the record\), and the two independent rounds furnish the run\-to\-run agreement analysis of Section 6\. Retraction is reserved for measurement\-wounded runs; sound runs are superseded, never erased\. ### 4\.5Self\-hosted models and the submission model The runner speaks to any OpenAI\-compatible endpoint via a configurable base URL, so self\-hosted models evaluate under conditions identical to API models with no size limit \- the path used here for the open and sovereign Arabic models via local quantized serving, and available to any laboratory\. Access follows a three\-line model: reading is open to all; reproduction is free on the researcher’s own resources with the public code and development set; publication on the official leaderboard passes through a human\-verified submission gate, because unattended automation is how leaderboards accepting external results lose their integrity\. ## 5Experimental Setup We evaluate 27 systems on the frozen pilot bank under the protocol of Section 4\. Selection follows four stated criteria: general\-purpose instruction\-tuned systems; accessibility through a unified API layer or the self\-hosted path at evaluation time; deliberate coverage of closed frontier models, open weights across size tiers \(7B to frontier\-scale\), and Arabic\-focused systems; and, within the freeze window, the newest releases \- GPT\-6 Astra entered evaluation three days after its public release\. The Arabic group spans three worlds: a commercial API system \(Mistral Saba\), an open specialized model \(SILMA\-9B\), and a sovereign state model \(ALLaM\-7B, SDAIA\), the latter two served locally through the OpenAI\-compatible path in 4\-bit quantization \(Q4\_K\_M, a common compressed\-precision format for efficient local serving\) \- a deployment\-realistic condition we state plainly\. Two Arabic candidates could not be evaluated fairly and are reported as such: AceGPT\-v2\-8B, whose only artifact available through the standard self\-hosting channel combined destructive 2\-bit quantization with a missing chat template and degenerated on a smoke test; and the large Arabic models \(Jais\-class\), absent from the unified API layer entirely \- an infrastructure\-absence observation we return to in Section 8\. Multiple\-choice items are scored by exact answer\-letter match; extraction items by strict per\-field matching against ground truth, with Iraqi administrative conventions enforced verbatim; the per\-item score is the matched\-field fraction\. Every published proportion carries a 95% Wilson interval[Wilson \(1927\)](https://arxiv.org/html/2609.13980#bib.bib11)\(a small\-sample\-safe way of expressing how much a score could plausibly shift on repeat measurement\), chosen over the normal approximation for small samples near the boundaries; for the extraction axis, whose per\-item scores are\[0,1\]\[0,1\]fractions rather than Bernoulli outcomes, the binomial treatment is a deliberately conservative bound \(Bernoulli variance is maximal for bounded variables\), which we disclose\. Track\-level intervals pool corrects counts across the track’s auto\-scored axes \(equal n per axis makes the pooled proportion equal the unweighted mean\)\. Each model was evaluated in two independent full rounds \- an original round and a re\-verification round under the hardened harness \- and the tabled results are the verified round; the pair yields the agreement analysis of Section 6\.4\. Prompt\-sensitivity is not studied in the pilot \(Section 9\)\. ## 6Results Table 1 presents the frozen board: 27 systems, MSA and Iraqi track means over the auto\-scored axes, and the overall macro average\. Four regularities organize what follows\. Overall is the macro\-average across all five auto\-scored axes \(MSA comprehension, MSA knowledge, Iraqi comprehension, Iraqi knowledge, and document extraction\), not the mean of the two track scores; the lower\-scoring extraction axis pulls Overall below the simple MSA/Iraqi average\. Table 1:The frozen pilot board \(auto\-scored axes; generative axes pending human judging\)\. Ranks follow the declared tie\-break rule: display\-precision overall, then the Iraqi track; ties share a rank\.Table 1 \- The frozen pilot board \(auto\-scored axes; generative axes pending human judging\)\. Ranks follow the declared tie\-break rule: display\-precision overall, then the Iraqi track; ties share a rank\. ### 6\.1Saturation on one track, discrimination on the other Eight systems score a perfect 100\.0 on the MSA track and the remainder compress above 80, while the Iraqi track spreads across more than twenty\-three points\. Every model carries a large internal gap between its own two tracks \- 14 to 18 points for most of the field \(Figure 1\) \- and the two strongest systems tie at display precision \(90\.9 overall, 84\.9 Iraqi\) and swapped ranks between the two independent rounds: the top of contemporary AI is statistically inseparable on this benchmark, and the newest frontier model, evaluated 72 hours post\-release, neither broke the Iraqi ceiling nor escaped the gap\. Figure 2 gives the Iraqi track with pooled 95% intervals; overlapping intervals across the leading cluster mean exact top ranks are not settled at the pilot’s sample size \- the tiers, however, are\. Figure 1:The two\-track gap: Every model, MSA vs\. Iraqi\.Figure 2:Iraqi Track with 95% confidence intervals\. ### 6\.2The document wall Official\-document extraction confines every system tested to between 32 and 56 \(Figure 3\)\. The best score on the axis \- 56\.4 \- belongs to neither of the overall leaders; the newest frontier model reaches 53\.7; and the 7B tier collapses toward 32\. The single axis built for a state’s working need is the single axis where the entire field fails, and reading the raw responses \(Section 7\) shows the failures are substantive: well\-formed outputs that miss field precision and administrative convention, not formatting accidents\. Figure 3:The Document Wall: all 27 systems confined to 32\-56 ### 6\.3The specialization law Three data points align\. SILMA\-9B \(Iraqi 63\.2\) and the sovereign ALLaM\-7B \(64\.7\) \- both marketed and built as Arabic\-specialized \- fall roughly ten points below a size\-matched generalist \(ministral\-8b, 73\.9\) on the Iraqi track while matching it on MSA\. Mistral Saba, by contrast, exceeds its own same\-family generalist sibling by about six Iraqi points, yet remains four to five points below the frontier, with a two\-track gap as wide as everyone else’s\. Arabic specialization, as currently practiced from startups to states, is MSA specialization; targeted specialization helps at matched scale, and no specialization closes the dialect gap\. ### 6\.4Run\-to\-run agreement Across the two independent full rounds, per\-model overall scores moved by a mean absolute 0\.5 points \(maximum ~1\.2\), every difference falling well inside the Wilson intervals[Wilson \(1927\)](https://arxiv.org/html/2609.13980#bib.bib11); the tied leaders exchanged ranks, exactly as overlapping intervals predict\. The measurement is stable; the reported uncertainty is honest\. ## 7Error Analysis ### 7\.1A dialect\-triggered over\-refusal mode The pilot’s most unexpected finding began as an anomaly: the safety\-hardened tier of the newest model family \(claude\-fable\-5\) placed sixteenth overall with a perfect MSA score and an Iraqi comprehension score far below its two same\-family siblings\. Per\-item records showed the misses were not wrong answers but empty completions \- deterministic on the same seven items across three independent runs, surviving a fourfold budget escalation\. A dual\-path probe \(the aggregation gateway and the vendor’s direct API\) returned the same official envelope on every one: stop\_reason ”refusal”, category ”cyber”, zero output tokens, with an explanation citing restrictions on ”violative cyber content”\. The seven items ask the meaning of everyday Iraqi words \- hassa \(”now”\), khosh \(”good”\), ya\-m’awwad \(a friendly vocative\), the continuous\-aspect particle da, the copula chan, jahhal \(”children”\), and the traditional chaykhana \(”tea house”\)\. The model answered every Iraqi item it did not refuse \- its effective comprehension of answered items approaches 100% \- and refused nothing on the MSA track; the other 26 systems, including two same\-family models without the hardened tier’s additional measures, produced zero refusals on 190 items each\. Safety alignment calibrated on distributions that exclude a dialect can convert that dialect’s ordinary vocabulary into anomalous, blockable input: a model that goes silent on ”what does khosh mean” fails the Iraqi citizen as surely as one that answers wrongly \- and no MSA benchmark can observe this failure\. Under the declared policy \(Section 4\.2\) the refusals score as failures on the affected items and are reported here as a 12% over\-refusal rate on Iraqi comprehension \(6/50; 3\.7% of the model’s auto\-scored items overall; 0% MSA; 0% for all other systems\)\. ### 7\.2The document wall is substantive Reading extraction sidecars across models shows a consistent failure shape: syntactically valid structured output whose issuing\-authority field reproduces the letterhead’s full hierarchy instead of the conventional issuing office, whose required\-action field paraphrases rather than extracts, and whose optional fields hallucinate values on letters deliberately authored without them\. Strict field matching is applied identically to all systems, so the wall ranks models fairly; its height reflects genuine distance from Iraqi administrative convention\. ### 7\.3Residual genuine errors After separating refusals and recovered transients, genuine mistakes on the auto\-scored Iraqi axes are sparse at the frontier \(single items \- e\.g\., one knowledge miss for the hardened\-tier model\) and concentrate, as expected, in the small\-model tier, where comprehension errors show false\-friend interference and knowledge errors cluster on post\-2003 institutional facts\. ## 8Discussion Four blind spots, one lesson\. The pilot documents four phenomena that the mature MSA evaluation stack cannot observe in principle: a 14\-18\-point dialect gap carried by every contemporary system including one released 72 hours earlier; a civic\-document register on which the entire field fails; a specialization economy in which ”Arabic” means MSA \- two dedicated Arabic models, one sovereign, trailing a same\-size generalist on the dialect track; and a safety\-alignment mode that converts a dialect’s everyday vocabulary into blockable content\. Each was invisible until an instrument was built to see it, and each carries a practical address\. For model developers, the gap and the specialization law argue that dialect competence is a data problem that neither scale alone nor MSA\-centric fine\-tuning solves\. For safety teams, the over\-refusal case shows that alignment distributions need dialectal coverage: a filter that has never seen khosh in training treats it as anomaly\. For governments, the document wall converts procurement intuition into measurement \- the capability states most need is the one no vendor currently delivers, and no vendor benchmark currently reports\. And for the evaluation community, the infrastructure observation stands on its own: most models marketed as Arabic are unreachable through the standard access channels researchers actually use \- absent from unified API layers, or published in unusable quantized artifacts \- which is itself a finding about where the Arabic NLP ecosystem invests\. ## 9Limitations We state the pilot’s limits plainly\. Scale: fifty items per axis yields wide intervals \(roughly±\\pm10 points at the axis level\); every published score carries its interval, and exact rankings inside the leading cluster are explicitly unsettled\. Generative axes: the 150 open\-generation items \(generation, translation, safety\) await the human\-judging campaign \- 4,050 model responses under dual judging with Krippendorff’s alpha[Krippendorff \(2019\)](https://arxiv.org/html/2609.13980#bib.bib9)\- and publish nothing meanwhile; the benchmark’s most distinctive axes are its least complete, a sequencing we chose deliberately over delaying the framework\. Dialect coverage: Baghdadi, southern, and mixed registers dominate; Mosuli coverage is zero pending native authors, so ”Iraqi” here under samples Iraq’s own variation\. No human baseline yet anchors the score scale\. Contamination: the entire pilot bank is public development tier; a sealed private set arrives with v1\.0, and until then leaderboard results are explicitly development\-tier\. Measurement: two independent runs per model bound run\-to\-run variance but prompt\-format sensitivity is unstudied; self\-hosted Arabic models ran in 4\-bit quantization, a realistic but lossy condition; and strict field matching on extraction, while uniform, may undercount partially\-correct administrative phrasings\. ## 10Future Work Version 1\.0 scales the bank past 1,000 items through a national authoring platform with regional native authors \- closing the Mosuli gap \- and introduces the sealed private test tier\. The unified Iraqi orthography guide graduates into a standalone linguistic\-resource publication\. A human\-baseline study anchors interpretation\. The judged axes complete and upgrade this paper’s journal version\. Operationally: scheduled re\-evaluation rounds tracking model updates longitudinally; a formal external\-submission protocol behind the human verification gate; and GPU\-hosted evaluation of the large Arabic models the API layer omits\. ## 11Ethics and Availability All official documents in the bank are simulated end to end \- authorities, numbers, names, and subjects are fictional; no real record, personal datum, or classified material was used\. Safety items describe harm categories without instantiating harmful content and use generic framings rather than named targets\. The over\-refusal analysis relies exclusively on the vendor’s official API envelopes, reproduced verbatim\. On publication, the platform code releases under Apache\-2\.0 and the public development set with a datasheet; the leaderboard remains live with per\-run verification certificates\. The released set is development\-tier by design, so leaderboard integrity ultimately rests on the forthcoming sealed tier rather than on secrecy of the published items\. ## References - Abdul\-Mageedet al\.\(2020\)M\. Abdul\-Mageedet al\.The NADI shared tasks: nuanced Arabic dialect identification \(series\)\.Note:Proceedings of the WANLP / ArabicNLP Workshops, Association for Computational LinguisticsAnnual shared\-task seriesCited by:[§2\.2](https://arxiv.org/html/2609.13980#S2.SS2.p1.1)\. - Al\-Mathamet al\.\(2025\)R\. Al\-Mathamet al\.BALSAM: a platform for benchmarking Arabic large language models\.External Links:2507\.22603,[Document](https://dx.doi.org/10.48550/arXiv.2507.22603)Cited by:[§1](https://arxiv.org/html/2609.13980#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13980#S2.SS1.p1.1)\. - Almazroueiet al\.\(2023\)E\. Almazrouei, R\. Cojocaru, M\. Baldo, Q\. Malartic, H\. Alobeidli, D\. Mazzotta, G\. Penedo, G\. Campesan, M\. Farooq, M\. Alhammadi, J\. Launay, and B\. NouneAlGhafa evaluation benchmark for Arabic language models\.InProceedings of ArabicNLP 2023,pp\. 244–275\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.arabicnlp-1.21)Cited by:[§2\.1](https://arxiv.org/html/2609.13980#S2.SS1.p1.1)\. - Altakroriet al\.\(2025\)M\. H\. Altakrori, N\. Habash, A\. A\. Freihat, Y\. Samih, K\. Chirkunov, M\. AbuOdeh, R\. Florian, T\. Lynn, P\. Nakov, and A\. F\. AjiDialectalArabicMMLU: benchmarking dialectal capabilities in Arabic and multilingual language models\.Note:Accepted to LREC 2026External Links:2510\.27543,[Document](https://dx.doi.org/10.48550/arXiv.2510.27543)Cited by:[§2\.3](https://arxiv.org/html/2609.13980#S2.SS3.p1.1)\. - Bouamoret al\.\(2018\)H\. Bouamor, N\. Habash, M\. Salameh, W\. Zaghouani, O\. Rambow, D\. Abdulrahim, O\. Obeid, S\. Khalifa, F\. Eryani, A\. Erdmann, and K\. OflazerThe MADAR Arabic dialect corpus and lexicon\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\),Miyazaki, Japan\.External Links:[Link](https://aclanthology.org/L18-1535/)Cited by:[§2\.2](https://arxiv.org/html/2609.13980#S2.SS2.p1.1)\. - Elfilaliet al\.\(2024\)A\. Elfilaliet al\.Open Arabic LLM leaderboard \(OALL\)\.Note:Hugging Face / Technology Innovation InstituteExternal Links:[Link](https://huggingface.co/spaces/OALL/Open-Arabic-LLM-Leaderboard)Cited by:[§2\.1](https://arxiv.org/html/2609.13980#S2.SS1.p1.1)\. - Gaoet al\.\(2023\)L\. Gaoet al\.A framework for few\-shot language model evaluation \(lm\-evaluation\-harness\)\.EleutherAI / Zenodo\.Note:Version v0\.4\.0External Links:[Document](https://dx.doi.org/10.5281/zenodo.10256836)Cited by:[§4\.1](https://arxiv.org/html/2609.13980#S4.SS1.p1.1)\. - Kotoet al\.\(2024\)F\. Koto, H\. Li, S\. Shatnawi, J\. Doughman, A\. Sadallah, A\. Alraeesi, K\. Almubarak, Z\. Alyafeai, N\. Sengupta, S\. Shehata, N\. Habash, P\. Nakov, and T\. BaldwinArabicMMLU: assessing massive multitask language understanding in Arabic\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 5622–5640\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.334),[Link](https://aclanthology.org/2024.findings-acl.334/)Cited by:[§2\.1](https://arxiv.org/html/2609.13980#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.13980#S3.SS1.p1.1)\. - Krippendorff \(2019\)K\. KrippendorffContent analysis: an introduction to its methodology\.4th edition,SAGE Publications\.Cited by:[§3\.1](https://arxiv.org/html/2609.13980#S3.SS1.p1.1),[§9](https://arxiv.org/html/2609.13980#S9.p1.1)\. - Mousiet al\.\(2025\)B\. Mousi, N\. Durrani, F\. Ahmad, Md\. A\. Hasan, M\. Hasanain, T\. Kabbani, F\. Dalvi, S\. A\. Chowdhury, and F\. AlamAraDiCE: benchmarks for dialectal and cultural capabilities in LLMs\.InProceedings of the 31st International Conference on Computational Linguistics \(COLING 2025\),Abu Dhabi, UAE,pp\. 4186–4218\.External Links:[Link](https://aclanthology.org/2025.coling-main.283/),[Document](https://dx.doi.org/10.48550/arXiv.2409.11404)Cited by:[§2\.3](https://arxiv.org/html/2609.13980#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2609.13980#S3.SS1.p1.1)\. - Stanford CRFM and Arabic\.AI \(2025\)Stanford CRFM and Arabic\.AIHELM Arabic\.Note:Stanford Center for Research on Foundation ModelsExternal Links:[Link](https://crfm.stanford.edu/helm/arabic/latest/)Cited by:[§2\.1](https://arxiv.org/html/2609.13980#S2.SS1.p1.1)\. - Wilson \(1927\)E\. B\. WilsonProbable inference, the law of succession, and statistical inference\.Journal of the American Statistical Association22\(158\),pp\. 209–212\.External Links:[Document](https://dx.doi.org/10.1080/01621459.1927.10502953)Cited by:[§5](https://arxiv.org/html/2609.13980#S5.p1.1),[§6\.4](https://arxiv.org/html/2609.13980#S6.SS4.p1.1)\.
Similar Articles
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
MCBench is a new benchmark for assessing the safety of omnimodal large language models across vision, audio, and text modalities. It includes 1196 scenarios and finds current models struggle with cross-modal safety reasoning.
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
This paper introduces MameLoshnLM, the first open-source 8B-parameter Yiddish language model, along with the Oytser pretraining corpus and Kashes evaluation benchmark. It demonstrates that continued pretraining on high-quality Yiddish data outperforms general multilingual models, highlighting the value of dedicated low-resource language modeling.
Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy
This paper introduces AI-MASLD, a stress-audit framework for medical LLMs that reveals how benchmark accuracy can hide serious safety failures, and demonstrates that open-weight models can match or exceed proprietary ones on safety dimensions.
RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants
RobustMAD introduces a benchmark to evaluate the real-world robustness of multimodal small language models for deployable industrial anomaly detection assistants. It reveals critical failure modes and provides guidance for next-generation systems.