YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

arXiv cs.CL Papers

Summary

YallaMorph is a large-scale benchmark for evaluating controlled Arabic morphological generation in large language models, revealing significant challenges with cliticized and morphologically rare forms.

arXiv:2609.10153v1 Announce Type: new Abstract: Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:19 AM

# A Benchmark for EvaluatingArabic Morphological Generation in Large Language Models
Source: [https://arxiv.org/html/2609.10153](https://arxiv.org/html/2609.10153)
## YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

1Salam Khalifa12\] Reham Marzouk3Nizar Habash1Affiliation:Computational Approaches to Modeling Language \(CAMeL\) LabAffiliation:1New York University Abu Dhabi,Affiliation:2Stony Brook University,3Mohamed bin Zayed University of Artificial IntelligenceAffiliation:\{mahmoud\.ali,salam\.khalifa,nizar\.habash\}@nyu\.edu,Reham\.Marzouk@mbzuai\.ac\.ae

###### Abstract

Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control\. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature\-based input\. We introduceYallaMorph, a large\-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations\. We evaluate multilingual and Arabic\-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries\. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms\.

## 1Introduction

Arabic morphology presents a major challenge for language generation\. The interaction of templatic and concatenative morphology with morphosyntactic and cliticization features creates a large space of productive forms\. Our central question is not simply whether Large Language Models \(LLMs\) can generate plausible Arabic, but whether they can systematically produce appropriate inflected forms for specific morphological features\.

Most evaluations of Arabic LLMs focus on downstream tasks or broad generation qualitynagoudi\-etal\-2022\-arat5;almazrouei\-etal\-2023\-alghafa;koto\-etal\-2024\-arabicmmlu, providing limited insight into whether models can systematically generate correct Arabic word forms from explicit lexical and morphological specifications\. Existing morphological inflection benchmarks provide more controlled evaluation settingskodner\-etal\-2022\-sigmorphon;goldman\-etal\-2023\-sigmorphon, but they do not capture the scale and Arabic\-specific feature space required for evaluating modern LLMs\.

We introduceYallaMorph, a large\-scale benchmark for controlled Arabic morphological generation\.111The benchmark data is publicly available:[https://github\.com/CAMeL\-Lab/YallaMorph](https://github.com/CAMeL-Lab/YallaMorph)As illustrated in Figure[1](https://arxiv.org/html/2609.10153#S1.F1), the task maps a lemma, part\-of\-speech, gloss, and target feature bundle to the corresponding Arabic surface form, or to an indication that the requested configuration is morphologically invalid\. The benchmark covers verbs, nouns, adjectives, and their cliticized forms, and contains over 600K evaluation instances sampled across root class, stem complexity, paradigm completeness, and lemma frequency\.

Figure 1:Morphological realization of Arabic lemmaبَتَكkatab ‘write’ asاَهَنوُبُتْكَيَوwayak\.tubuwnahaA ‘and they write it’ through morphosyntactic features and clitics\. Transliterations follow HSB\(Habash:2007:arabic\-transliteration\)\.We evaluate proprietary, multilingual, and Arabic\-oriented instruction\-tuned LLMs under both diacritized and undiacritized settings\. Results show that controlled Arabic morphological generation remains challenging even for strong modern LLMs, particularly for cliticized, unseen, and morphologically rare forms\.

Our contributions are: \(a\) introducing alarge\-scale benchmarkfor controlled Arabic morphological generation designed using a linguistically motivated sampling framework; and \(b\)evaluating multilingual and Arabic\-oriented LLMsand providing detailed analysis across linguistic and distributional dimensions\.

## 2Related Work

#### Morphological Inflection

Morphological inflection is a well\-established standalone task in NLP\. The SIGMORPHON morphological \(re\)inflection shared tasks have served as a central benchmark series for this problemcotterell\-etal\-2016\-sigmorphon;cotterell\-etal\-2017\-conll;cotterell\-etal\-2018\-conll;mccarthy\-etal\-2019\-sigmorphon;vylomova\-etal\-2020\-sigmorphon;pimentel\-ryskina\-etal\-2021\-sigmorphon;kodner\-etal\-2022\-sigmorphon;goldman\-etal\-2023\-sigmorphon\. These tasks evaluate systems that map lemmas and morphosyntactic feature bundles to inflected forms, and later editions increasingly emphasized generalization across typologically diverse languages, unseen lemmas, and unseen feature combinations\. The 2022 and 2023 shared tasks in particular strengthened the evaluation setup through data splits targeting generalization to unseen lemmaskodner\-etal\-2022\-sigmorphon;goldman\-etal\-2023\-sigmorphon\. This line of work is closely tied to UniMorphbatsuren\-etal\-2022\-unimorph, a broad\-coverage universal schema for morphological annotation and inflection tables that represent forms through a lemma and a bundle of morphosyntactic features\. UniMorph 4\.0 covers 182 languages, including Arabic varieties such as Modern Standard Arabic \(MSA\), Egyptian Arabic, and Gulf Arabic in recent shared\-task data\. However, while the SIGMORPHON shared tasks yielded a diverse set of baselines and state\-of\-the\-art systems, they were not designed as LLM benchmarks\. Moreover, for our purposes, UniMorph does not provide exhaustive lexical coverage, full Arabic paradigm coverage, or cliticized inflected forms\. In this work, we instead use Arabic\-specific morphological resources that provide broader lexical coverage and richer paradigm generation\.

#### Morphological Generators for Arabic

Morphological analyzers and generators have been central to Arabic NLP since its early stages, providing explicit linguistic representations for a morphologically rich and complex language\. Early work included finite\-state and templatic approaches that modeled Arabic root\-and\-pattern morphologyBeesley:1989:two\-level;Kiraz:1994:multi\-tape;Beesley:1998:arabic;Habash:2006:magead;Smrvz:2007:elixirfm\. A parallel line of work adopted lexicon\- and compatibility\-table\-based resources, most notably the Buckwalter Arabic Morphological Analyzer and its successorsBuckwalter:2002:buckwalter;Maamouri:2010:ldc\. Aragen/ALMORGEANA further extended Buckwalter\-style lexical resources toward generation from lexeme\-and\-feature representationsHabash:2005:morphological\. More recently,khairallah\-etal\-2024\-camelintroduced CamelMorph, a comprehensive open\-source morphological analyzer and generator for MSA that builds on this tradition, reporting over 100K lemmas with rich morphological features and broad paradigm coverage\. In this work, we use the CamelMorph lexicon and database to construct our controlled benchmark set, and we use its generation engine, exposed through CAMeL Toolsobeid\-etal\-2020\-camel, to generate and validate reference inflected forms\.

#### LLM Evaluation for Morphological Inflection

Recently, benchmarking LLMs for morphological generation, particularly inflection, has gained traction as researchers increasingly probe LLMs for linguistic knowledge\. Early work in this area used variants of the Wug testBerko\-1958\-wugto evaluate morphological productivity and generalization in LLMs across typologically diverse languagesweissweiler\-etal\-2023\-counting;anh\-etal\-2024\-morphology\.ismayilzada\-etal\-2025\-evaluatingextended this line by evaluating morphological compositional generalization through both generative and discriminative tasks\. However, these studies do not focus on Arabic or on controlled generation from explicit Arabic morphosyntactic feature specifications, which is the focus of our work\.

Very few studies have focused on Arabic morphological generation in LLMs\. IMPACTsaeed2025impactinflectionalmorphologyprobesintroduced a multilingual evaluation framework for inflectional morphology across five morphologically rich languages, including Arabic, but its Arabic component targets a narrower set of agreement and inflectional phenomena\. The work byalakeel2026morphemesbordersevaluatingrootpatternfocuses solely on Arabic, specifically MSA, evaluating how LLM tokenizers align with Arabic morphological structure and how LLMs perform on productive root–pattern generation\. It constructs a controlled test set with real and nonce roots and finds that tokenizer–morpheme alignment is neither necessary nor sufficient for successful morphological generation\. In contrast, our work targets controlled Arabic morphological generation over an explicitly structured feature space, focusing on inflected\-form realization and systematic sampling rather than root–pattern productivity alone, gender\-focused generation, or broad multilingual probing\. Furthermore, our benchmark dataset is substantially larger and more comprehensive\.

## 3Linguistic Background & Terminology

Arabic morphology is characterized by bothrichnessandcomplexity\. Its richness stems from large inflectional paradigms and productive cliticization involving conjunctions, prepositions, articles, and pronouns, yielding many possible surface forms for a single lemma \(Table[2](https://arxiv.org/html/2609.10153#S4.T2)\)\. Its complexity arises from interactions between templatic \(root and pattern\) and concatenative \(affix or clitic\) morphemes, often accompanied by orthographic and morphophonological alternations\.

#### Relevant Terminology

We briefly summarize the main terminology used throughout the paper, following CamelMorphkhairallah\-etal\-2024\-camel\. Alemmarepresents an abstraction over all inflectional forms of a lexical itemhabash\-etal\-2022\-morphotactic\. Arootis an abstract consonantal sequence encoding core lexical meaning, while apatternspecifies the vocalic and templatic structure used to derive grammatical forms\. Astemis the form produced by combining roots and patterns before affixation\. Morphological representation is further organized through functional features such as gender, number, case, and state, as well aspart\-of\-speech \(POS\), which specifies the grammatical category, and thegloss, which provides an English semantic description\. Finally, the framework distinguishes betweenaffixes, which realize core morphosyntactic features within thebaseword, andclitics, which include conjunctions, prepositions, definite article, and object and possessive pronouns\.

#### Example

Consider the noun lemmaةَنْجَلlaj\.naℏ\\hbar‘committee’ which illustrates multiple types of interactions\. It is derived from the sound root l\.j\.n using the singular pattern 1a2\.3aℏ\\hbar\. Its plural is the broken pluralناَجِلlijaAn, which preserves the root while changing the pattern to 1i2aA3\. Clitic attachment further increases surface variation\. For example, attaching the enclitic pronounاَهhaA ‘her’ producesاَهُتَنْجَلlaj\.n\+at\+u\+haA ‘her committee \[nominative\]’, whereةaℏ\\hbarsurfaces asَتat\. Single proclitic attachment yields forms such asِناَجِّللاAl\+l∼\\simijaAni ‘the committees \[genitive\]’ andٍناَجِلِلli\+lijaAnĩ ‘for committees \[genitive\]’\. Combining both triggers orthographic assimilation, producingِناَجِّلِلli\+l∼\\simjaAni ‘for the committees \[genitive\]’\.

## 4YallaMorphBenchmark Design

CategoryTypeExplanationPOS GroupLemmaFrequencyHighHigh\-frequency lemmaAllMediumMedium\-frequency lemmaAllLowLow\-frequency lemmaAllRootClassSoundRoot with no weak letters, hamza, or geminationAllGeminatedRoot with identical second and third radicalsAllHamzatedRoot containing a hamza consonantAllWeak InitialRoot whose first radical is weak \(w or y\)AllHollowRoot whose second radical is weakAllDefectiveRoot whose third radical is weakAllNTWSNon\-templatic word stemNouns, AdjectivesParadigmCompletenessFullFully inflecting paradigmAllMasculineMasculine only nominal paradigmNounsFeminineFeminine only nominal paradigmNounsStemComplexityRegularNo major orthographic or morphological alternationsAllDefective FamilyStem allomorphs exhibit defective\-family alternationsAllHamza FamilyStem allomorphs exhibit hamza\-related alternationsAll\#N FamilyStem ends with letterنn\.Verbs\#T FamilyStem ends with letterتt\.Perfective VerbsL\# FamilyStem begins with letterلl\.Nouns, AdjectivesDiptote FamilyAt least one stem allomorph is a diptoteفرصلانمعونمم\.Nouns, AdjectivesTable 1:Definitions of sampling categories used for lemma selection\.In this section, we present the design details of theYallaMorphbenchmark\.

### 4\.1Design Principles

The design ofYallaMorphis centered on evaluating Arabic morphological generation at scale while maintaining linguistic balance and interpretability\. The benchmark includes over 600K sampled forms covering the major open POS classes of Arabic and capturing both inflectional and cliticization phenomena\. To ensure broad linguistic coverage, the dataset is balanced across a wide range of morphophonological dimensions, including root classes, stem complexity, and paradigm completeness\. In addition,YallaMorphincorporates frequency\-balanced sampling over lemmas and morphological categories, enabling controlled evaluation of model generalization across both frequent and long\-tail forms\.

### 4\.2Data Source

We construct our benchmark from the lemma inventory and morphological analyses provided by CamelMorph MSAkhairallah\-etal\-2024\-camel\. Specifically, we use the lemma entries, part\-of\-speech labels, glosses, and morphological features to support sampling, categorization, and paradigm generation\. Access to these analyses is handled through CAMeL Toolsobeid\-etal\-2020\-camel\.

### 4\.3POS Groups & Morphological Features

YallaMorphfocuses on the three major open POS classes in Arabic: verbs, nouns, and adjectives\. Within the verbal domain, we distinguish five core configurations: active perfective, passive perfective, active imperfective, passive imperfective, and command\. The various POS groups differ in inflectional behavior, stem complexity, and compatibility with clitic attachment\. To capture both core morphology and cliticization attachment phenomena, the benchmark is organized into two complementary subsets: a baseword subset focusing on inflectional morphology, and a cliticization subset focusing on interactions between stems and attached clitics\.

### 4\.4Sampling Motivation

Arabic morphology exhibits substantial lexical and inflectional richness\. Combining the large lemma inventory of CamelMorph MSA with feature bundles and clitic configurations yields an impractically large number of possible test instances \(Table[2](https://arxiv.org/html/2609.10153#S4.T2)\)\. We therefore construct the benchmark through guided sampling rather than exhaustive enumeration\. The following subsections describe the linguistic and distributional dimensions underlying this sampling process\. Our benchmark comprises over 600K carefully sampled entries\.

Hypothetical MaximumPOSLemmas\- Clitics\+ CliticsVerbs9,3332,183,454754,417,944Nouns19,8441,071,5762,730,375,648Adjectives6,924373,896952,687,008Total36,1013,628,9264,437,480,600Table 2:Lemma and maximum instance counts by lexical class, with and without clitics\.
### 4\.5Sampling Dimensions

We sample lemmas along linguistically motivated dimensions capturing root class, lemma frequency, paradigm completeness, and stem complexity\. Each eligible lemma is assigned categorical labels for the relevant dimensions, and combinations of these labels define the sampling strata\. Table[1](https://arxiv.org/html/2609.10153#S4.T1)summarizes all category dimensions and their possible values, while Table[3](https://arxiv.org/html/2609.10153#S4.T3)presents representative noun and active perfective verb \(PVA\) examples illustrating their combinations\.

GroupFrequencyRoot ClassParadigmStem ComplexityExample LemmaGlossVerbHighSoundFullRegularعَجَرrajaς\\varsigmareturnVerbHighHamzatedFull\#Hamza FamilyأَدَتْبِاAib\.tadaÂbeginVerbMediumGeminatedFull\#N FamilyّنَتْمِاAim\.tan∼\\simbe gratefulVerbMediumHollowFullRegularفاَنnaAfexceedVerbLowSoundFullRegularرَمْلَبَتtabal\.marbe polymerizedVerbLowDefectiveFullRegularوُخَرraxuwbe looseNounHighSoundFull\#Hamza FamilyكيِرَشšariykpartnerNounHighNTWSMasculineL\# Familyرْتيِلliyt\.rliterNounMediumHollowFeminine\#Defective FamilyةَّيِنيِصSiyniy∼\\simaℏ\\hbarporcelainNounMediumWeak InitialMasculine\#Diptote Familyروُشْوَمmaw\.šuwrprismNounLowSoundFeminineL\# FamilyةَّيِماَسِقْنِااَلlaAAin\.qisaAmiy∼\\simaℏ\\hbarindivisibilityNounLowGeminatedFullRegularظيِظَكkaĎiyĎoverfilledTable 3:Representative verb and noun lemmas illustrating sampling categories across lemma frequency, root class, paradigm completeness, and stem complexity\. Verb categories are from the perfective active subset\.#### Lemma Frequency

Lemma frequency provides a corpus\-based estimate of how often each lemma is attested in a reference corpus\. We estimate frequency using BAREC\-10M\(elmadani\-etal\-2026\-large\), a large, balanced, multi\-domain corpus of Arabic that is automatically morphologically tagged and lemmatized\.

Our lemma\-frequency analysis was conducted in parallel with the BAREC\-10M annotation effort\. The final BAREC\-10M annotations and their additional lemma\-cluster resolution procedure were not used because they were not yet available at the time of our analysis\. We therefore independently analyzed the raw corpus using the morphological disambiguation component of CAMeL Tools\.222We used CAMeL Tools v1\.5\.5\(obeid\-etal\-2020\-camel\), with BERT\-Disambig\(inoue\-etal\-2022\-morphosyntactic\)and the CamelMorph MSA v1\.1 database\(khairallah\-etal\-2024\-camel\)\.For each token, we retained the top\-ranked morphological analysis\.

Lemmas for which our analysis recovered no occurrences were assigned a count of zero\. For sampling, we group lemmas into three frequency bands: low, medium, and high\. The low band contains lemmas with frequency 0; the medium band contains lemmas with frequencies from 1 to 20; and the high band contains lemmas with frequencies above 20\.

#### Root Class

Root class captures the phonological type of a lemma’s root\. We use seven root\-class labels\. Each lemma is assigned to exactly one root\-class label\. For lemmas with identifiable roots, overlapping classifications are resolved using the following priority hierarchy: geminated\>\>hamzated\>\>weak\-initial\>\>hollow\>\>defective\>\>sound\. Lemmas with no identifiable root are treated as non\-templatic word stems \(NTWS\) and assigned to the NTWS class\. This procedure yields a single, consistent root\-class label for each lemma and supports balanced sampling across major root types\.

#### Paradigm Stem Complexity

Paradigm stem complexity captures broad stem\-related patterns that affect how forms are realized across a lemma’s paradigm\. We derive this dimension from stem\-related information provided in CamelMorph MSA\. We assign lemmas to one of seven stem\-complexity families and some lemmas may satisfy more than one stem\-complexity condition\. Overlapping classifications are resolved using the following priority hierarchy: \#N Family\>\>\#T Family\>\>L\# Family\>\>Defective Family\>\>Hamza Family\>\>Diptote Family\>\>Regular\. This procedure yields a single, consistent stem\-complexity label for each lemma and supports non\-overlapping sampling across stem\-complexity families\. For example, the lemmaكيِرَشšariyk ‘partner’ has a broken pluralءاَكَرُشšurakaA’ which ends in a Hamza \(glottal stop\)\. This places the whole lemma in the Hamza Family group\. Table[1](https://arxiv.org/html/2609.10153#S4.T1)defines the various stem\-complexity families\. Some of these apply to all POS, such as Hamza Family, while others are very specific, e\.g\., \#T Family only applies to perfective verbs and captures the reduced spelling of t\-initial suffixes with t\-final verbs due to orthographic gemination \(Shadda\):ُّتُفfut∼\\simu \(fut\+tu\) ‘I entered’\.

#### Paradigm Completeness

Paradigm completeness captures the extent to which a lemma realizes the expected range of gender–number combinations in its paradigm\. We derive this dimension from stem\-related information provided in CamelMorph MSA\. Lemmas are assigned to one of three labels: masculine\-only, feminine\-only and full\. Verbal and adjectival paradigms are always full in Arabic, whereas nouns can be of any of the three types\. The inclusion of masculine\-only and feminine\-only paradigms leads to the possibility of Null forms, i\.e\., feature combinations that cannot be accommodated such as the feminine ofبَتْكَمmak\.tab ‘office’ is a masculine\-only lemma\. Simply adding the feminine ending to produceةَبَتْكَمmak\.tabaℏ\\hbar‘library’ yields an incorrect result\. Around 13% of all the entries inYallaMorphcorrectly have a null gold reference\.

### 4\.6Lemma Sampling Process

After assigning sampling dimension labels to all lemmas, we group them into strata defined by joint combinations of the sampling dimensions across the benchmark POS groups\. Since these strata vary greatly in size, we adopt a fixed per\-stratum strategy, sampling up to 10 lemmas from each non\-empty stratum while including all lemmas for smaller strata for the baseword subset of the benchmark\. This preserves broad linguistic coverage while limiting over\-representation of highly productive categories\. The cliticization subset is constructed separately due to the large expansion introduced by proclitic and enclitic combinations\. Instead of exhaustively enumerating all cliticized forms, we derive a reduced sample from the baseword subset while preserving coverage across lemma frequency and stem complexity, and we only include full paradigm completeness cases\. For each relevant category, we sample up to 5 lemmas, yielding a manageable yet linguistically diverse clitic\-focused evaluation set\.

### 4\.7Arabic Morphological Generation Entries

We formulate the task as slot\-level Arabic morphological generation\. Given a lemma, POS, gloss, and a target feature bundle \(henceforth, morphological specification\), the model generates the corresponding inflected surface form\. Outputs may consist of a single valid form, multiple valid realizations, or an indication that the requested configuration is invalid \(null reference\)\.

Multiple\-reference cases account for 3\.2% of the benchmark entries\. In these cases, the gold reference set contains more than one valid realization for the same morphological specification\.

Morphological specifications are POS\-dependent\. Verbal forms are specified using aspect, person, gender, number, voice, and mood, while nouns and adjectives use gender, number, case, and state\. Clitic configurations are represented separately using positional proclitic \(prc0–3\) and enclitic \(enc0–1\) slots, with POS\-specific constraints\. The full list of features is provided in Appendix\. The selected clitic inventory and compatibility restrictions are provided in Appendix\.

LemmasInflected FormsPOS GroupBasewordCliticsBasewordCliticsActive PV513759,23445,738Passive PV498708,9648,820CV433607,7946,480Active IV4356039,150181,800Passive IV4356039,15037,800Nouns1,88775101,89867,950Adj5947532,07676,950Subtotal4,795475238,266425,538Total4,795663,804Table 4:Benchmark subset sizes\.
### 4\.8Benchmark Statistics

FrequencyInstance CountPercentage10M\+430\.01M–10M7210\.1100K–1M5,0400\.810K–100K14,0682\.11K–10K25,6993\.9101–1K36,0225\.411–10043,7466\.61–1055,9088\.40397,17959\.8Null85,37812\.9Total663,804100Table 5:Frequency distribution of entries\.Table[4](https://arxiv.org/html/2609.10153#S4.T4)summarizes benchmark sizes across POS groups and experimental subsets\. Although the cliticization subsets contain fewer lemmas, they are substantially larger in total size due to the multiplicative effect of generating all inflected baseword forms and their cliticized variants\. Table[5](https://arxiv.org/html/2609.10153#S4.T5)shows the frequency distribution of undiacritized target forms based on the CAMeLBERT Frequency ListKhalifa:2021:Camel\_Frequency, highlighting the strong long\-tail nature of the benchmark, with 59\.8% of unseen forms, while an additional 12\.9% correspond to null reference entries\.

## 5Evaluation

### 5\.1Metrics

Our main metric isAny Match Accuracy, computed at the instance level by comparing each prediction against its corresponding gold form or set of gold forms\. For multi\-form outputs, AMA counts a prediction as correct if at least one of the generated forms matches a gold form\. Target configurations with a Null reference, representing morphologically invalid lemma–feature combinations, are counted as correct only when the model explicitly predicts a Null output\.

We additionally report micro\-averagedPrecision,Recall, andF1to capture partial correctness in multi\-form outputs\. We evaluate predictions under three orthographic settings: diacritized, undiacritized, and Alif/Ya/Hamza/Ta\-Marbuta normalized\. We treat the diacritized setting as primary because it provides the strictest evaluation of Arabic morphological generation\. Table[6](https://arxiv.org/html/2609.10153#S5.T6)reports AMA and F1, while the corresponding Precision and Recall results are provided in Appendix\.

10\-ShotZero\-ShotAny Match AccuracyF1 ScoreAny Match AccuracyF1 ScoreModelDiacUndiacNormDiacUndiacNormDiacUndiacNormDiacUndiacNormGPTen51\.967\.067\.754\.470\.571\.346\.261\.662\.749\.266\.067\.3Geminien49\.760\.761\.452\.464\.264\.939\.153\.754\.341\.557\.057\.7Fanaren21\.536\.537\.222\.538\.339\.15\.119\.419\.95\.320\.320\.9Fanarar15\.826\.827\.216\.628\.128\.52\.715\.515\.82\.816\.116\.5Jais\-70Ben9\.120\.921\.29\.521\.722\.18\.215\.615\.98\.816\.516\.9Jais\-70Bar8\.217\.517\.88\.618\.318\.63\.810\.711\.34\.111\.412\.1ALLaMen4\.710\.711\.04\.710\.911\.24\.48\.68\.83\.57\.27\.3ALLaMar4\.510\.010\.34\.19\.39\.53\.19\.09\.22\.88\.48\.6Qwenen4\.513\.013\.64\.914\.415\.02\.08\.79\.12\.19\.29\.7Jais\-8Ben3\.410\.110\.23\.39\.910\.02\.76\.46\.72\.35\.76\.0Jais\-8Bar2\.38\.28\.52\.17\.67\.92\.67\.77\.92\.16\.46\.6Table 6:Overall benchmark results under the 10\-shot and zero\-shot prompting settings\. Subscripts ar and en indicate the prompt language\.
### 5\.2Compared Models

We evaluate seven instruction\-tuned LLMs spanning commercial multilingual, open\-weight multilingual, and Arabic\-focused model families:GPT\-5\.4,Gemini 3\.1 Flash\-Lite,Qwen2\.5\-7B\-Instruct,Fanar\-2\-27B\-Instruct,Jais\-2\-8B\-Chat,Jais\-2\-70B\-ChatandALLaM\-7B\-Instruct\-preview333Abbreviations used throughout: GPT\-5\.4 \(GPT\), Gemini 3\.1 Flash\-Lite \(Gemini\), Qwen2\.5\-7B\-Instruct \(Qwen\), Fanar\-2\-27B\-Instruct \(Fanar\), Jais\-2\-70B\-Chat \(Jais\-70B\), Jais\-2\-8B\-Chat \(Jais\-8B\), and ALLaM\-7B\-Instruct\-preview \(ALLaM\)\.\. This selection provides broad coverage of widely used general\-purpose systems and models developed specifically for Arabic\.

#### Prompt Design

We use POS\-specific prompts for verbs, nouns, and adjectives, with separate variants for clitic\-aware experiments\. Inputs include the lemma, POS, gloss, and the relevant morphological feature bundle\. Models are instructed to return outputs in a fixed JSON schema\. We evaluate bothzero\-shotandfew\-shotprompting settings, with the few\-shot setting including 10 in\-context examples\. Prompts remain lightweight, with only minimal clarifications for features that showed consistent ambiguity in pilot experiments, such as nominal state distinctions, emphatic verbal moods, and clitic ordering\.

We used English and Arabic versions of the prompts \(Appendix\)\. Arabic and English prompts are used forALLaM,Fanar,Jais\-70BandJais\-8B; all other models use English prompts only\.

#### Output Processing

Model outputs are parsed using a unified post\-processing pipeline that extracts candidate Arabic strings, handles minor formatting deviations from the requested JSON schema, and removes duplicate forms\.

### 5\.3Results

We first compare overall model performance, then analyze the best\-performing model across linguistic and distributional dimensions\.

#### Overall Performance

Table[6](https://arxiv.org/html/2609.10153#S5.T6)shows that GPT substantially outperforms all other models across both Any Match Accuracy and F1 metrics in all evaluation settings, followed by Gemini\. In contrast, the Arabic\-oriented open models perform considerably worse overall, suggesting that Arabic specialization alone does not guarantee accurate morphological generation\. English prompts generally perform better, especially in the 10\-shot setting\. However, Arabic prompts perform better for ALLaM and Jais\-8B on several zero\-shot undiacritized and normalized metrics\.

Similar Articles

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

arXiv cs.CL

This paper introduces MameLoshnLM, the first open-source 8B-parameter Yiddish language model, along with the Oytser pretraining corpus and Kashes evaluation benchmark. It demonstrates that continued pretraining on high-quality Yiddish data outperforms general multilingual models, highlighting the value of dedicated low-resource language modeling.