The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation
Summary
A two-dialect finite-state morphological analyzer for the Dungan language is presented, with a multi-genre evaluation measuring inflection, ambiguity, and lexical coverage.
View Cached Full Text
Cached at: 08/03/26, 07:33 AM
# A Two-Dialect Finite-State Model and a Multi-Genre Evaluation Source: [https://arxiv.org/html/2607.28766](https://arxiv.org/html/2607.28766) ## The Morphological Core of Dungan: A Two\-Dialect Finite\-State Model and a Multi\-Genre Evaluation Anton M\. AlekseevSt\. Petersburg Department of the Steklov Mathematical Institute, RASSt\. Petersburg State UniversityKyrgyz State Technical University named after I\. RazzakovSergey I\. NikolenkoSt\. Petersburg Department of the Steklov Mathematical Institute, RASSt\. Petersburg State University ###### Аннотация Dungan, a Sinitic language of Central Asia written in a Cyrillic\-based script, is described in detail in the grammatical literature, yet the quantitative properties of its morphology in actual usage, to the best of our knowledge, have never been measured systematically\. This paper uses a finite\-state morphological analyzer as a measuring instrument\. Implemented with HFST and covering both dialect groups — the Gansu variety \(the literary standard\) and the Shaanxi variety — the model offers no new grammatical description; it formalizes the knowledge accumulated in Dungan studies and makes it measurable on corpora of three genres\. Three results follow\. Overt inflection is rare and limited: only 9\.3% of recognized tokens in the encyclopaedic register have an overt marker, the system has just ten categories, and degree marking is almost absent\. Ambiguity is genuine but sharply localized: 78\.1% of tokens receive a single analysis, and the residue sits almost entirely on two clitics,\-ди\(genitive / progressive\) and\-ни\(locative / prospective\)\. And the grammatical core proves effectively closed, the claim the instrument is really needed for: between 78% and 95% of the tokens the analyzer fails on, depending on register, are simply absent from the lexicon, and the phenomena the model deliberately declines to implement account for at most 4\.5% of those failures\. Held\-out coverage \(80–85%\) is no lower than development coverage \(73%\), while a stem list with no morphology already reaches 67\.4%, so the morphology is worth 5\.2 points\. The open frontier of Dungan is lexical\. The analyzer, its sources and every evaluation script are released openly\. ## 1Introduction Dungan is a Sinitic language whose speakers — descendants of Hui \(Chinese\-speaking Muslim\) migrants who moved to Central Asia in the second half of the 19th century — live mainly in Kazakhstan, Kyrgyzstan and Uzbekistan\(Zavyalova,[1996](https://arxiv.org/html/2607.28766#bib.bib22); Rimsky\-Korsakoff Dyer,[1979](https://arxiv.org/html/2607.28766#bib.bib14)\)\. Unlike other Chinese languages it is written in a Cyrillic\-based orthography, adopted in 1953 in place of a Latin script that had itself replaced, in 1928–1932, the Arabic\-script*xiaoerjing*tradition\(Zavyalova,[1996](https://arxiv.org/html/2607.28766#bib.bib22)\)\. Grammatically Dungan is strongly isolating: inflection is minimal, grammatical meanings are carried mostly by function words and a small set of clitics, and the written norm does not mark lexical tone, which gives rise to extensive homography\. Although Dungan is well described in the grammatical literature, we have been unable to locate an open computational*morphological processor*for it: a check of OLAC, Glottolog, Universal Dependencies, the Apertium and GiellaLT repositories, and a GitHub code search forlexc/twolsources taggeddng, finds neither an analyzer nor a treebank, and every negative claim of this kind below is bounded by what those searches returned\.111A Dungan text corpus of roughly 600,000 tokens is reported to exist at Northwest Normal University \(Lanzhou, PRC\); we were unable to download it and verify its size or access conditions independently and note it only for completeness\.Dungan is not, however, absent from computational resources, and it would be wrong to imply otherwise: the Crúbadán web crawl\(Scannell,[2007](https://arxiv.org/html/2607.28766#bib.bib8)\)supplies a word\-form frequency list with source URLs; English Wiktionary holds some 5,400 Dungan lemma pages carrying part\-of\-speech, tone and hanzi correspondence \(a source this work itself draws on\); Salmi’s*Dungan–English Dictionary*\(Salmi,[2018](https://arxiv.org/html/2607.28766#bib.bib16)\)is an electronic lexicographic resource; and the codedngappears in recent massively multilingual corpora and language\-identification models\. What none of these provides is a grammar: a system that segments a word form, assigns it grammatical categories and generates it back\. That is the gap this work fills, and precisely for less\-resourced languages such a system is the basic building block for lemmatization, corpus annotation, spell checking, machine translation, and lexicography\. By origin and basic lexicon Dungan belongs to the Central Plains group of Mandarin \(中原官话; the Guanzhong and southern Gansu varieties\) and is close to Standard Mandarin\(Zavyalova,[1996](https://arxiv.org/html/2607.28766#bib.bib22); Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17)\)\. What matters computationally is not genealogy but the script\. Lexicon also shows Turkic, Iranian and Slavic loans, typical of contact but less important for morphological analysis\. Chinese characters suppress written homonymy, e\.g\. 飯 ‘food’ and 犯 ‘to violate’ are graphically distinct though homophonous, while the phonographic Cyrillic of Dungan exposes it, especially because it does not record tone, the distinctive feature of any Sinitic phonological system\. A large share of lexical stems are homographic due to that, and without dedicated means an analyzer cannot tell them apart\. Inflection in the traditional sense is practically absent: no declension, no conjugation, no agreement\. Grammatical meanings are expressed by word order and by particles and clitics that behave in many respects like independent words but attach phonetically to a content word\. A finite\-state rather than paradigm\-based approach is therefore natural: the task reduces to describing the attachment of a limited marker set to stems and resolving the resulting homography\(Koskenniemi,[1983](https://arxiv.org/html/2607.28766#bib.bib4)\)\. The aim of this work is not to give a new description of Dungan grammar \(the language is described in detail\) but to*measure*the properties of its morphological core in actual usage: how compact and productive the inventory of inflectional categories is, what the inflectional load is and whether its profile is stable across genres, and, crucially, how*closed*the core is, i\.e\. whether a finite set of markers exhausts the morphology of real text, leaving only the lexical frontier open\. The instrument is a finite\-state analyzer and generator covering both major dialect groups: analysis maps a word form to a lemma with grammatical tags \(жынму→\\toжын<n\><an\><t1\><pl\>, ‘‘lemmaжын‘person’, noun, animate, tone I, plural’’\), generation is the inverse mapping\. The analyzer is implemented with HFST\(Lindénet al\.,[2011](https://arxiv.org/html/2607.28766#bib.bib5)\)and accompanies this paper as supplementary material\.222The complete analyzer — lexc/twol sources, both dialect lexicons, the gold test suite, the CI configuration and every evaluation script that produces a number reported here — is available at[https://github\.com/alexeyev/dungan\-finite\-state\-morphology](https://github.com/alexeyev/dungan-finite-state-morphology)\. A principled commitment of this work is that the model does not claim a new grammatical description of Dungan but converts into computable form the knowledge already accumulated in Dungan studies\. Every formalized phenomenon is traced, wherever possible, to an existing description — plural, genitive, aspect markers, locative, classifiers all rest on published accounts\(Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19); Kalimov,[1968](https://arxiv.org/html/2607.28766#bib.bib31); Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28); Zevakhina and Imazov,[1997](https://arxiv.org/html/2607.28766#bib.bib23)\), checked case by case \(Section[5](https://arxiv.org/html/2607.28766#S5), Appendix[B](https://arxiv.org/html/2607.28766#A2)\) — and decisions with no description in the accessible literature are explicitly flagged as provisional\. The analyzer is released accordingly: as a reproducible resource open to verification, correction and extension by Dungan speakers and specialists\. Our contributions are: a two\-dialect finite\-state model of Dungan morphology grounded in the existing descriptions, released with tests and CI; what is, to our knowledge, the first*quantitative profile*of inflection in usage — frequency, productivity and load across three genres \(Section[7](https://arxiv.org/html/2607.28766#S7)\) — reported with the sensitivity of each figure to how the metric is defined, rather than as a point estimate; a quantification of*ambiguity*and of where it sits \(78\.1% of recognized forms are unambiguous; the residue is concentrated on the clitics\-диand\-ни\); and a direct test of*core completeness*\(Section[8](https://arxiv.org/html/2607.28766#S8)\) that classifies every recognition failure as lexical or morphological instead of inferring closure from a coverage figure — at most 4\.5% of failures are attributable to phenomena the model deliberately omits, against a no\-morphology baseline that the model beats by 5\.2 points\. ## 2Related Work The finite\-state approach goes back to Koskenniemi’s two\-level model\(Koskenniemi,[1983](https://arxiv.org/html/2607.28766#bib.bib4)\), in which one description of the lexical–surface relation works in both directions;Kaplan and Kay \([1994](https://arxiv.org/html/2607.28766#bib.bib2)\)showed that both ordered rewrite rules and two\-level grammars are regular relations realisable by transducers, which is what makes analysis and generation invertible in a single mechanism\. The practical toolkit is presented byBeesley and Karttunen \([2003](https://arxiv.org/html/2607.28766#bib.bib1)\)and surveyed byKarttunen and Beesley \([2005](https://arxiv.org/html/2607.28766#bib.bib3)\)\. The two largest platforms of this kind, Apertium\(Khannaet al\.,[2021](https://arxiv.org/html/2607.28766#bib.bib38)\)\(machine translation between closely related languages\) and GiellaLT\(Moshagenet al\.,[2013](https://arxiv.org/html/2607.28766#bib.bib39)\)\(Uralic and Siberian areas\), provide the infrastructure for building and testing such analyzers; the present work follows their conventions \(tag format, lexc/twol organization, coverage and precision testingWashingtonet al\.,[2012](https://arxiv.org/html/2607.28766#bib.bib40)\)\. Recent analyzers in this tradition include those for Kurmanji and Central Kurdish\(Ahmadi and Hassani,[2020](https://arxiv.org/html/2607.28766#bib.bib9); Naserzadeet al\.,[2023](https://arxiv.org/html/2607.28766#bib.bib42)\)and Maithili\(Rahiet al\.,[2020](https://arxiv.org/html/2607.28766#bib.bib10)\)\. Finite\-state lexica have also been used to recover a suprasegmental the script omits, in both standard cases for stronger reasons than ours:Meurer \([2011](https://arxiv.org/html/2607.28766#bib.bib6)\)incorporates Abkhaz stress because stress position governs the surface realization of \[\\textschwa\] \(schwa\), so the transducer cannot parse orthographic forms without it, andReynolds \([2016](https://arxiv.org/html/2607.28766#bib.bib7)\)generates Russian stress for language learners, where Constraint Grammar disambiguation over the transducer raises accuracy\(Bick and Didriksen,[2015](https://arxiv.org/html/2607.28766#bib.bib41)\)\. Our tone tag is lexical, never touching the surface; we return to the difference in Section[5](https://arxiv.org/html/2607.28766#S5)\. Dungan grammar is described in detail in a body of \(mostly Russian\-language\) scholarship\. The first scientific description is byDragunow and Dragunowa \([1936](https://arxiv.org/html/2607.28766#bib.bib12)\)andDragunov and Dragunova \([1937](https://arxiv.org/html/2607.28766#bib.bib18)\), who characterized Dungan as an*independent*language, distinct from the other known Chinese languages, with the Gansu dialect as its leading, literary, variety\.333“…das Dunganische…als eine selbständige Sprache zu betrachten, die sich von allen übrigen uns bekannten chinesischen Sprachen unterscheidet”; “In dialektischer Hinsicht ist es die Gansu\-Mundart, die bereits zu einer Literatursprache wird”\(Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 35\)\. As the authors note, at the time of writing Dungan had not yet been scientifically studied \(“…da das Dunganische bis jetzt noch nicht wissenschaftlich untersucht worden ist”\)\.The same first decade of research produced Polivanov’s contributions: alongside his part in the 1928–1932 Latinization of the script, he wrote the first school grammars of the emerging literary language \(Frunze, 1935–1936\) and two studies in the 1937 Frunze orthography volume — the phonological system of the Gansu dialect\(Polivanov,[1937a](https://arxiv.org/html/2607.28766#bib.bib44)\)and the tonal analysis this model draws on throughout\(Polivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32)\)\.Dragunov \([1940](https://arxiv.org/html/2607.28766#bib.bib19)\)then first described the categories of aspect and tense, followed by Kalimov\(Kalimov,[1958](https://arxiv.org/html/2607.28766#bib.bib30),[1968](https://arxiv.org/html/2607.28766#bib.bib31)\), Imazov\(Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28),[1987](https://arxiv.org/html/2607.28766#bib.bib29)\), Zavyalova\(Zavyalova,[1979](https://arxiv.org/html/2607.28766#bib.bib21),[1996](https://arxiv.org/html/2607.28766#bib.bib22)\)and Salmi\(Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15),[2018](https://arxiv.org/html/2607.28766#bib.bib16)\); a summarising encyclopaedic description isZevakhina and Imazov \([1997](https://arxiv.org/html/2607.28766#bib.bib23)\), and individual categories, in particular the adjective, have been studied in fieldwork\(Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\)\. Among recent work,Honkasalo \([2024](https://arxiv.org/html/2607.28766#bib.bib13)\)presented a contact\-linguistic description of spoken Kazakhstani Gansu Dungan from 2022–2023 field recordings, which documents deep Russian influence and is a reminder that written \(literary\) Dungan, on which the present model rests, and the living spoken varieties differ considerably\. We have likewise been unable to locate computational resources carrying grammatical annotation; the electronic resources that do exist \(Section[1](https://arxiv.org/html/2607.28766#S1)\) supply word lists and lexicography rather than grammar; for languages in a comparable position the standard route to such a corpus is precisely a rule\-based analyzer feeding a corpus platform such as Tsakorpus\(Arkhangelskiy,[2019](https://arxiv.org/html/2607.28766#bib.bib45)\), which is the downstream use the present analyzer is built for\. The present work adds almost nothing to this corpus of descriptions but builds on it, translating established knowledge into a reproducible computational form and*measuring*it in usage, filling what remained outside its predecessors’ scope only as explicitly flagged provisional decisions\. ## 3Language Material and Sources The main lexical source is Yanshansin’s*Concise Dungan–Russian Dictionary*\(Yanshansin,[2009](https://arxiv.org/html/2607.28766#bib.bib37)\), openly available as a PDF: about twelve thousand entries, each with a tone mark \(Roman numerals I–III\) and a Russian gloss; 6,987 of them are mined into the analyzer’s lexicon\.444The lexicon is not exclusively dictionary\-derived, and the remainder is small but not negligible\. Beside the 6,987 Yanshansin entries the stem files hold 288 lemmas harvested from English Wiktionary’s Dungan categories \(CC BY\-SA; see the Ethics Statement\), 159 forms from Salmi’s grammatical material and 47 from his dictionary, 145 proper names, 210 shared function words, 42 recovered adjective readings \(see Limitations\) and 23 dialect and core\-vocabulary entries\.Grammatical information, especially the aspect markers and counting suffixes, draws on the Soviet–Russian descriptive tradition\(Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19); Kalimov,[1968](https://arxiv.org/html/2607.28766#bib.bib31); Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28)\)and on Salmi’s work on aspect and the lexicon\(Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15),[2018](https://arxiv.org/html/2607.28766#bib.bib16)\)\. For quality evaluation and for recovering individual constructions we used a small trilingual corpus \(Dungan text with Chinese\-character and Russian versions\), from which we also built the correctness control set of Section[8](https://arxiv.org/html/2607.28766#S8); the character correspondence is key — since Cyrillic does not record tone, it is the character that identifies unambiguously which morpheme underlies a word form\. The bulk of the development corpus is 126 articles from the Dungan Wikipedia incubator, 12,337 tokens; the parallel material adds 384 \(folk narrative 156, the Shyvaza poem 47, glossed sentences 181\), for a development\-corpus total of 12,721\. One dependency bears on Section[8](https://arxiv.org/html/2607.28766#S8)\. The dictionary and both held\-out texts are published by the Institute for Bible Translation, and 38 of the 126 Wikipedia articles \(15\.9% of development tokens\) are themselves religious in subject; that classification is deliberately conservative\. Biblical toponyms, ethnonyms, Roman officials and scriptural coinage are all excluded, so the figure is a lower bound\. ‘‘Held\-out’’ below means*not used during development*; it does not mean lexically or stylistically independent\. ## 4Analyzer Architecture Dungan falls into two dialect groups, Gansu and Shaanxi, whose differences, though regular, affect the most frequent items: copula, personal pronouns, classifiers\. The literary standard is the three\-tone Gansu variety, because, asSalmi \([2023](https://arxiv.org/html/2607.28766#bib.bib17)\)remarks crediting Yanshansin, the first teachers of the Dungan school in Frunze \(now Bishkek\) were Gansu speakers\. The design is ‘‘shared core \+ thin dialect layer’’: the majority of the morphology is described once, in a shared module, and dialect differences are factored into separate files defining uniquely named hook lexicons referenced from the shared root lexicon\. Grammatical tags are multichar symbols following Apertium conventions: part of speech \(<n\>,<vblex\>,<adj\>, …\), nominal features \(<an\>/<nn\>,<sg\>/<pl\>\), verbal aspect \(<pst\>,<prog\>,<fut\>,<exp\>\), and others\. Рис\. 1:Build pipeline of the ‘‘shared core \+ thin dialect layer’’ architecture: lexical sources compile to a lexical transducer, composed with the twol rules; analyzer and generator are derived from the sharedmorph\.The build is a pipeline of finite\-state operations \(Figure[1](https://arxiv.org/html/2607.28766#S4.F1)\): the root morphotactics, the shared stem files and one dialect stem file are concatenated and compiled byhfst\-lexc, composed with the twol rules to yieldmorph, and from it the generator and the analyzer \(by inversion, after composition with a case\-folding transducer\) are derived\. Таблица 1:Regular dialect correspondences formalized in the model\.The split costs little: the two analyzers differ by 1\.2 points on the same corpus, and essentially the whole difference is one entry —вый‘to be, become’, absent from the Gansu lexicon, occurs 147 times \(1\.16% of the corpus\)\. The comparison measures a lexicon gap rather than a grammatical distance\. ## 5Modelled Morphology Inflection is modest, but the system of parts\-of\-speech and of the bound morphemes attaching to them is well developed:Salmi \([2023](https://arxiv.org/html/2607.28766#bib.bib17)\)observes that Dungan has visibly more morphological markers than other Sinitic varieties\.Imazov \([1979](https://arxiv.org/html/2607.28766#bib.bib27)\)distinguishes fourteen parts of speech; the analyzer realises eleven of them \(participle, converb and preposition are not tagged:ба把 is a particle andлян連 a conjunction\), splits Imazov’s single conjunction class into the two tags<cnjcoo\>and<cnjsub\>, and adds<np\>and<cop\>, so that its own inventory also numbers 14 tags, which is a coincidence\. Tables[4](https://arxiv.org/html/2607.28766#A1.T4)and[5](https://arxiv.org/html/2607.28766#A1.T5)\(Appendix[A](https://arxiv.org/html/2607.28766#A1)\) give the fourteen part\-of\-speech tags and the ten overtly marked inflectional categories, listing separately the five tags the analyzer emits that are lexical or zero\-marked \(<sg\>,<pers\>,<p1\>–<p3\>\), since Section[7](https://arxiv.org/html/2607.28766#S7)turns on that distinction\. Every string shown there with tags is verbatimhfst\-lookupoutput; cells giving only a lemma, a hanzi \(a logogram\) and a gloss identify the example word and are not analyzer output\. A key boundary is that between inflection and word formation\. Two\-syllable nouns are formed by compounding, suffixation \(the ‘‘substantive’’\-зы,\-р\) and reduplication, all yielding lexicalized, idiomatic units\(Tsunvazo,[1955](https://arxiv.org/html/2607.28766#bib.bib35); Zevakhina,[2018](https://arxiv.org/html/2607.28766#bib.bib25),[2019](https://arxiv.org/html/2607.28766#bib.bib26)\); adverbs are formed with multifunctional markers as well \(\-ди,\-ха,\-шон,\-му, the postposition\-ни\) rather than with a dedicated suffix as in Chinese, and the formation is lexicalized\(Tsunvazo,[1963](https://arxiv.org/html/2607.28766#bib.bib36)\)\. Such words are stored as whole dictionary stems \(<adv\>\)\. What is modelled productively is inflection in the strict sense: number, case clitics, aspect, degrees of comparison\. Таблица 2:The five\-member aspect–tense paradigm of the verb\.The aspect–tense system is obligatory: a verb without one of the five markers forms not a sentence but an incomplete phrase555Sharply formulated byDragunow and Dragunowa \([1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 48\); quoted in full in Appendix[C](https://arxiv.org/html/2607.28766#A3)\.\(Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12); Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19); Kalimov,[1968](https://arxiv.org/html/2607.28766#bib.bib31); Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15),[2023](https://arxiv.org/html/2607.28766#bib.bib17)\)\. The five markers are summarized in Table[2](https://arxiv.org/html/2607.28766#S5.T2)\. Every modelled phenomenon is traced to its sources with page\-level anchors in Appendix[B](https://arxiv.org/html/2607.28766#A2), which presents the model as a single decision table:*phenomenon*→\\to*what the sources say*→\\to*how it is formalized*\. Argumentation that does not reduce to a table row follows as notes; verbatim quotations underpinning each rule, with page and paragraph anchors, are in Appendix[C](https://arxiv.org/html/2607.28766#A3)\. Because much of this literature is hard to locate, Appendix[E](https://arxiv.org/html/2607.28766#A5)additionally surveys the Dungan resource landscape with locators\. ### Decisions needing argument\. Four choices do not reduce to a table row and are set out in full in Appendix[D](https://arxiv.org/html/2607.28766#A4)\. Derivation is treated lexically, since\-зыis largely lexicalized and segmenting it would create spurious homonymy\(Tsunvazo,[1955](https://arxiv.org/html/2607.28766#bib.bib35)\)\. The tag<pst\>is a practical label for an aspectual marker\(Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 15\), as<fut\>is for a prospective one\. Our five\-way system differs from Salmi’s\([1984](https://arxiv.org/html/2607.28766#bib.bib15)\): it distinguishes progressive\-диниfrom stative\-диand does not include Salmi’s past habitual category\. And, important for the following, Dungan lets a marker scope over a whole phrase: an adjective predicates with aspect markers \(жә\-дини‘\(it is\) hot’;Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\) and in the verb–object construction the marker attaches to the object noun,ще зы\-дини‘is writing’ \(寫字底呢\)\(Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19)\)\. The noun therefore admits the aspect clitics, so the nominal genitive and the predicative progressive*compete*for one surface\-ди, an ambiguity quantified in Section[7](https://arxiv.org/html/2607.28766#S7)\. ### Degrees of comparison\. The adjective admits the comparative\-щер\(да‘big’→\\toдащер‘bigger’\) and the intensive\-elative\-дихын\(得很 ‘very’:да→\\toдадихын‘very big’\); Imazov’s own examples areхи‘black’→\\toхищерandбый‘white’→\\toбыйдихын\(Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27),[1982](https://arxiv.org/html/2607.28766#bib.bib28)\)\. We keep the released analyzer’s tag name<sup\>, but as a label only:Imazov \([1982](https://arxiv.org/html/2607.28766#bib.bib28)\)places\-дихынin the complex comparative andZevakhina \([2001](https://arxiv.org/html/2607.28766#bib.bib24)\)calls it a suffix of full predication; the analytic superlativeзыюis not modelled\. ### Invertibility and tag order\. Analysis and generation are invertible: the tag sits on the upper side of the transducer, the surface marker on the lower, soжынмуanalyses asжын<n\><an\><t1\><pl\>and that tag string generates exactlyжынму\. The tone tag belongs to that upper string rather than annotating it, so the generator requires it:жын<n\><an\><pl\>, identical but for the missing<t1\>, is rejected\. Inversion is exact but not one\-to\-one:жынмуalso analyses asжынму<n\><an\><t1\><sg\>, from a plural the dictionary lexicalizes as an entry of its own, and both tag strings generateжынмуback\. Invertibility guarantees that every analysis returns to the form it came from, not that the analysis is unique; Section[7](https://arxiv.org/html/2607.28766#S7)quantifies how often it is not\. Further examples, the tag\-order convention, and the transducer graph of Figure[2](https://arxiv.org/html/2607.28766#A4.F2)are in Appendix[D](https://arxiv.org/html/2607.28766#A4)\. The markers discussed above attach to the stem in a fixed order; Table[6](https://arxiv.org/html/2607.28766#A1.T6)\(Appendix[A](https://arxiv.org/html/2607.28766#A1)\) summarises this order for the four main classes as a morphotactic template\. ## 6Morphophonology and Cyrillic Orthography Formalizing morphology as a finite\-state machine forces explicit decisions where a prose description can remain indeterminate; they cluster at the interface of morphophonology and the Cyrillic script\. The first, the central one is the lexical\-tone decision of Section[5](https://arxiv.org/html/2607.28766#S5)\. The second is erization \(兒化, final\-р\), which unlike tone*is*written\. Here, attested erized forms are listed as separate stems\. We do not model erization productively because this would also require an account of the accompanying tone alternations, which are not represented in the orthography\. The third decision is the affix–clitic boundary\. The aspect and case markers \(\-ли,\-ди,\-ни,\-му\) sit between affix and clitic: phonetically bound to the content word, syntactically in many ways autonomous\. The analyzer treats these elements as bound inflectional markers rather than as separate syntactic words\. This keeps the model simple and makes the analysis directly testable on corpora\. However, it does not capture uses that depend on sentence context, such as the omission of an aspect marker after preposed negation \(see Limitations\)\. These decisions leave almost nothing for the two\-level component, which is itself a result\. Each dialect’s*twol*file has exactly two active rules, deleting the affix boundary%\>and the clitic boundary%\+\. Two further constraints \(theў/уalternation after labials, the Shaanxi palatalizationдь/ть→\\toҗь/чь\) are documented but vacuous, since the lexicon stores stems in final orthography\. Dungan has essentially no morphophonemic alternation at the boundaries this model draws: all descriptive work is done by the lexc morphotactics, and composition with*twol*is as of now a formality, the expected shape for a strongly isolating language\. ## 7A Morphological Profile of Usage A finite\-state model can not only analyze word forms but also*measure*how the category inventory of Section[5](https://arxiv.org/html/2607.28766#S5)behaves in real text\. We applied the analyzer to three corpora: an encyclopaedic corpus consisting of Dungan Wikipedia articles and parallel texts \(12,543 tokens666Single\-letter tokens, which are text\-extraction artefacts, are excluded here\. The resulting total is slightly lower than the development\-corpus total reported in Section[8](https://arxiv.org/html/2607.28766#S8)\(12,721 tokens\), where no such filter is applied\. This difference does not affect the conclusions\.\); a folkloric corpus of folk proverbs \(14,016 tokens\); and an Old Testament narrative corpus consisting of the Dungan Pentateuch \(119,920 tokens\)\. For each corpus, we counted the inflectional tags assigned to every recognized token\. Two properties of the instrument shape every figure below\. 1. \(i\)What counts as*inflection*? The analyzer emits five tags with no overt exponent: the zero\-marked singular<sg\>and the pronominal<pers\>,<p1\>–<p3\>\. Thus a bareжын‘person’ analyses asжын<n\><an\><t1\><sg\>\. Counting such tokens as inflected would measure the tagset rather than the language: on the encyclopaedic register those five tags alone contribute 5\.9 of the 15\.2 points an unrestricted definition yields\. We count a token as inflected only if it carries an overtly marked category\. 2. \(ii\)*Which*analysis is counted?hfst\-lookupreturns all analyses at weight zero, so its first output line reflects transducer path order rather than any disambiguation; we report the first\-reading figure together with the any\-reading figure, which bracket the true value\. On that definition the share of recognized tokens bearing an overt inflectional marker is 9\.3% in the encyclopaedic register \(any\-reading upper bound 13\.7%\), 13\.4% in the folkloric \(14\.2%\) and 27\.1% in the narrative \(27\.6%\)\. The inflectional*load*is low and strongly genre\-dependent, densest in narrative; the bulk of every text consists of unmarked stems and function words, as the isolating character of the language predicts\. encycl\.folklorenarrativecategory%lem\.%lem\.%lem\.genitive\-ди2\.23953\.291614\.82395perfective\-ли1\.43792\.531093\.44458progressive\-ди/\-дини1\.36672\.501244\.05360prospective\-ни0\.82502\.961533\.00468locative\-ни/\-шон1\.63391\.53433\.12153plural \(anim\.\)\-му1\.46200\.2797\.1135possessive\-ди\(prn\)0\.6660\.2742\.778experiential\-гуә0\.0220\.0980\.1317elative\-дихын0\.0430\.0000\.0415comparative\-щер0\.0110\.0000\.001*overt inflect\. load*9\.3%13\.4%27\.1%Таблица 3:Profile of inflectional categories in three genres: frequency \(% of recognized tokens carrying the category in the first analysis\) and productivity \(distinct carrier lemmas\)\. Only categories with an overt exponent are counted; the zero\-marked<sg\>and the pronominal subclass tags are excluded \(see text\)\. Gansu analyzer\.Composition and*productivity*complete the picture \(Table[3](https://arxiv.org/html/2607.28766#S7.T3)\)\. The attributive\-genitive\-диleads every register on both counts, spanning 95 to 395 lemmas, though the three aspect markers taken together exceed it\. The rest are stratified: the locative and animate plural are moderately frequent — the plural spikes to 7\.1% in the Pentateuch, where collective reference to peoples is constant, but on 35 lemmas only, frequent rather than productive — while the experiential and the two degree markers are*vacant*, together under two tenths of a percent everywhere\. The rank order is reasonably stable*across genres*\. Over the ten categories of Table[3](https://arxiv.org/html/2607.28766#S7.T3), Spearman’sρ\\rhois 0\.88 between the encyclopaedic and narrative domains \(n=10n=10, one\-sided permutationp=0\.001p=0\.001\), 0\.74 between encyclopaedic and folkloric \(p=0\.009p=0\.009\) and 0\.70 between folkloric and narrative \(p=0\.015p=0\.015\)\. The poles stay fixed: the genitive and the aspect markers at the top, the degree markers and the experiential at the bottom\. The middle of the ranking reorders freely with genre, most visibly the plural\. With ten categories the numbers are illustrative; the same small inventory, with the same extremes, describes texts of widely different topic and style; the ordering itself is not genre\-invariant\. Finally,*ambiguity*is limited: 78\.1% of recognized development\-corpus tokens receive a single analysis, 15\.9% two, the remaining 6\.0% three or more\. Tone tags contribute little: ignoring them raises the unambiguous share only to 78\.9%, and at the lemma level alone 98\.4% of tokens are unambiguous\. The residue sits mostly on the two polyfunctional clitics\. Among genuinely suffixed\-диforms \(299 tokens; particleдиexcluded\), 49\.5% receive both a genitive and a progressive reading, 20\.1% genitive\-only, and 30\.4% progressive\-only; among suffixed\-ниforms \(77 tokens\), 68\.8% receive both a locative and a prospective reading and the entire unambiguous residue is prospective\. These two clitics are the natural targets for contextual disambiguation, typically a Constraint Grammar layer over the transducer\(Bick and Didriksen,[2015](https://arxiv.org/html/2607.28766#bib.bib41)\)in this tradition, which is future work\. ## 8Core Completeness and Evaluation The analyzer is evaluated on*coverage*\(the share of tokens receiving at least one analysis\) and on the correctness of the readings it returns\. Correctness has two sides and we keep them apart:*recall*is the share of control items whose contextually correct reading is among the readings returned,*precision*the share of returned readings that are correct\. The distinction matters here because the transducer deliberately leaves ambiguity unresolved, so a reading that is wrong*in context*may still be a legitimate reading of the string\. These are the conventional metrics for finite\-state analyzers in the Apertium/GiellaLT tradition\(Washingtonet al\.,[2012](https://arxiv.org/html/2607.28766#bib.bib40); Khannaet al\.,[2021](https://arxiv.org/html/2607.28766#bib.bib38)\)\. We add a third measure, core\-vocabulary completeness \(Section[8](https://arxiv.org/html/2607.28766#S8)\), specific to the question this paper targets\. We measured coverage on a single corpus of real Dungan text \(folk narrative, a poem by Ya\. Shyvaza, glossed examples and 126 Dungan Wikipedia articles; 12,721 tokens\) for both dialect analyzers, making the figures directly comparable: 72\.6% for Gansu \(9,235/12,721\), 73\.8% for Shaanxi \(9,389/12,721\)\. Type coverage is much lower: 39\.3% and 39\.4%, as a Zipfian distribution against a fixed lexicon predicts\. We report it because the token figure alone flatters any analyzer\. On the literary subcorpus \(203 tokens\) coverage is 83\.3% and 84\.7%, but that sample is too small to press: the 95% Clopper–Pearson interval on 83\.3% of 203 is \[77\.4, 88\.1\]\. For comparison, the first Apertium Kyrgyz release reported 82–87%\(Washingtonet al\.,[2012](https://arxiv.org/html/2607.28766#bib.bib40)\); the Dungan figures sit in a similar band on a much smaller lexicon\. For an isolating language a bare word list is already a strong baseline\. Matching tokens against the 7,207 surface stem forms of the Gansu lexicon, with no continuation classes and no clitics, already recognizes 67\.4% of development\-corpus tokens against the transducer’s 72\.6%\! The modelled morphology is worth 5\.2 points \(5\.2 for Shaanxi too, 68\.6%→\\to73\.8%\), which is a meaningful but modest contribution, and demonstrates how thin Dungan inflection is\. We grew the lexicon against the same corpus on which coverage is reported, so that figure may be optimistically biased\. To assess generalization, we measured coverage on two Dungan texts*never used*in development — publications of the Institute for Bible Translation:777Electronic Dungan editions \(IBT\):[ibtrussia\.org/Dungan/bible](https://arxiv.org/html/2607.28766v1/ibtrussia.org/Dungan/bible)\. The texts are used solely to count recognition statistics; they are not part of the model or of any released material\.a collection of Dungan folk proverbs \(folk\-literary register, 14,016 tokens\) and the Dungan Pentateuch \(Old Testament narrative, 119,920 tokens; the edition’s glossary excluded\)\. Token coverage is 85\.2% on the proverbs and 79\.8% on the Pentateuch, i\.e\. both are*above*the development corpus\. Coverage is thus not an artefact of fitting to the development corpus\. Two qualifications limit what ‘‘held\-out’’ means here\. Domain difficulty differs: proverbs rest on frequent vocabulary, while the encyclopaedic text is dense with terminology and proper names\. And, as Section[3](https://arxiv.org/html/2607.28766#S3)notes, they are not independent of the development material: same publisher as the dictionary, and at least 15\.9% of development tokens religious in subject\. These samples establish that the analyzer transfers to unseen text of related register, leaving transfer to Dungan at large untested\. A direct check supports the same reading, and in a stronger form than a first estimate suggested\. Rebuilding the analyzer with every corpus\-mined sublexicon emptied — 122 stems, added because they surfaced as frequent unknowns during development — costs 15\.6 points of coverage on the development corpus \(72\.6%→\\to57\.0%\) but only 2\.6 on the proverbs \(85\.2%→\\to82\.6%\) and 2\.0 on the Pentateuch \(79\.8%→\\to77\.7%\)\. Those additions are therefore overwhelmingly specific to the corpus they were mined from, and what carries the held\-out figures is the lexicon core\. For correctness we assembled a control set of 104 annotated items over 99 distinct word forms from Dungan text \(the trilingual parallel material of Section[3](https://arxiv.org/html/2607.28766#S3): the glossed sentences and the folk narrative\), fixing each reference reading by its Chinese character correspondence\. The set is our own, assembled for this evaluation because no annotated Dungan corpus exists to sample from\. Every item is a form attested in that material, five of them only inside a longer word form, and its reference reading is fixed from the hanzi of the parallel version \(for four items, whose character that version does not supply, from the dictionary entry\)\. It ships with the analyzer astests/gold\-precision\.txt, one row per item with the hanzi and a gloss, so every label can be inspected; the annotation is the authors’ own, AI\-assisted and unreviewed by a native speaker \(Ethics Statement\)\. Recall is 104/104, i\.e\. for every item the reference reading is among those returned, and precision 106/127 = 83\.5%: of the 127 readings returned, a mean of 1\.22 per item, 21 are not the reference reading\. Both are tone\-insensitive; with the tone tag required to match they become 69\.2% \(72/104\) and 56\.7% \(72/127\), and we report both pairs, since the tone tags are a central design decision and the tone\-blind figures do not test them\. The 95% Clopper–Pearson intervals are \[96\.5, 100\] and \[75\.8, 89\.5\]\. Precision here is a lower bound, and diagnostic rather than damning: the reference fixes one reading per item, so a reading legitimate for the string but wrong in context counts against it, and that is what all 21 are, six of them the\-диand\-ниclitic ambiguities quantified in Section[7](https://arxiv.org/html/2607.28766#S7), nine lexical homographs \(эр二 ‘two’ besideэр兒 ‘son’\), six part\-of\-speech splits on function words \(ниas pronoun你 beside postposition 裏 and particle 呢\)\. None is a malformed analysis: what precision measures here is the residual contextual ambiguity a Constraint Grammar layer would resolve\. Both figures rest on a single, circular source: Yanshansin’s dictionary supplies both the analyzer’s stems \(stem, part of speech from the Russian gloss, tone, character correspondence\) and the reference readings, fixed from the same character by the same operations\. Agreement therefore certifies*internal consistency*, not correctness: it catches implementation errors — a twol rule that failed to fire, a typo in a continuation class, lexicon–reference drift — but not errors shared by both sides, whether an inaccuracy of the dictionary itself or a linguistic decision applied uniformly to both\. Only an independent reference and adjudication by an independent party — a native speaker or Dungan\-studies expert — can break the loop; both remain necessary future work, and a kit for the second path ships with the analyzer: a stratified worksheet of word forms, analyses and contexts, biased towards affixed and ambiguous forms, prepared for native\-speaker annotation\. Coverage is free of this loop: measured on external text, it registers only the fact of recognition\. As a floor on basic vocabulary the analyzer recognizes a form for each of 34 concepts of the 40\-item ASJP core list\(Wichmannet al\.,[2016](https://arxiv.org/html/2607.28766#bib.bib11)\)\. The check is weak: six concepts are untested, it asks only whether some analysis exists rather than whether it matches the concept, and its concept\-to\-spelling mapping comes from the same dictionary as the lexicon\. It probes lexicon completeness and leaves the circularity above intact\. Coverage counts recognitions; it cannot by itself support the claim that the grammatical core is closed, since a high coverage figure is equally consistent with a large lexicon and an incomplete marker inventory\. The claim is about what the*failures*are made of, so we measure that\. For every unrecognized token we ask whether stripping a candidate affix leaves a form the analyzer recognizes, classifying the failure as \(A\) reachable by morphology the model implements — a morphotactic gap; \(B\) reachable by a documented but deferred phenomenon \(past habitual\-лэ, erization\-р, derivational\-зы\); \(C\) reduplicationXXXXwithXXknown; or \(D\) a purely lexical gap\. \(A\)–\(C\) are upper bounds — a string ending in\-лэneed not contain the marker — which is the direction the argument needs\. Category \(D\) — a purely lexical gap — accounts for 94\.7% of failures in the encyclopaedic register, 78\.4% in the folkloric and 80\.5% in the narrative; \(A\) for 4\.0, 17\.1 and 15\.9% respectively, \(B\) for 0\.8, 3\.2 and 2\.3%, and \(C\) for 0\.5, 1\.3 and 1\.3%\. Between 78% and 95% of all recognition failures are purely lexical, and the phenomena the model deliberately omits — past habitual, erization, reduplication, derivational suffixation — together account for at most 1\.3% of failures in the encyclopaedic register and 4\.5% in the folkloric\.*This*is the evidence for a closed core: the marker inventory of Section[5](https://arxiv.org/html/2607.28766#S5)is not visibly missing anything that running text demands, and the open frontier is lexical\. On the development corpus the residue is dominated by proper names and terminology; the single largest unrecognized type in the Gansu build isвый‘to be, become’, the Shaanxi form discussed in Section[4](https://arxiv.org/html/2607.28766#S4)\. The measurement also shows where the model should grow next\. Category \(A\) is far larger on held\-out text \(16–17%\) than on the development corpus \(4\.0%\), and inspection shows it dominated by one construction: the aspect clitics\-ли/\-ниon predicative adjectives, 69 tokens in the folkloric register alone \(дуәли,лоли,дали,лынли\)\. These are documented\(Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\), and the model declines to generate them on the grounds that they are barely attested in the development corpus — a judgement the held\-out text shows to be an artefact of that corpus rather than a property of the language\. A typology of failures is a work\-list, which is what the instrument buys beyond a percentage\. ## 9Conclusion The measurement adds up to a coherent statement:*the morphological core of Dungan is compact, productive in only a few categories, and effectively closed*\. Overt inflection touches 9–27% of recognized tokens by register, ten categories exhaust the system and two are all but unused \(Section[7](https://arxiv.org/html/2607.28766#S7)\); 78\.1% of forms are unambiguous and the residue sits on two clitics\. Closure is established not by the coverage figure but by the composition of the failures behind it \(Section[8](https://arxiv.org/html/2607.28766#S8)\): 78–95% of unrecognized tokens are lexical gaps, and everything the model omits accounts for at most 4\.5% of them; the open frontier of Dungan is lexical\. The claim is modest: it is a quantification of what the descriptive literature already asserts, the value added being that each count is tied to an explicit and inspectable*grammatical*decision rather than to a proxy such as a type–token ratio\(Bentzet al\.,[2016](https://arxiv.org/html/2607.28766#bib.bib43)\)\. Methodologically, finite\-state formalization proves a convenient*measuring instrument*for a low\-resource grammar: it forces decisions, yields testable quantitative estimates instead of qualitative ones, and converts its own gaps into a ranked work\-list through the failure typology\. Known phenomena described but not implemented are itemized in the Limitations section\. The analyzer, source code and tests accompany this paper and will be released openly, as a basis for corpus annotation, lexicography and machine translation and a departure point for native\-speaker verification\. ## Limitations The work rests on written sources and the descriptive literature, not on work with native speakers; every decision in the model therefore has the status of a hypothesis, checked against existing descriptions and the corpus but not verified by a speaker\. This is why the analyzer is released openly with tests, designed for correction by Dungan specialists and speakers; a prepared adjudication worksheet for native\-speaker annotation ships with it\. The correctness measurement is circular in the sense detailed in Section[8](https://arxiv.org/html/2607.28766#S8): reference readings derive from the same dictionary that supplies the lexicon, so the reported 100% recall and 83\.5% precision certify internal consistency, not independently validated correctness; with the tone tag required to match, the pair is 69\.2% and 56\.7%\. The ASJP core\-vocabulary check does not escape the loop either, since the ASJP Dungan wordlist is itself compiled from a Yanshansin dictionary\. Coverage figures are free of this loop, but the held\-out evaluation validates coverage and its transfer only, not correctness\. The domain \(register\) difficulty of the held\-out texts differs from the development corpus, and, as Section[3](https://arxiv.org/html/2607.28766#S3)already notes, the held\-out material shares a publisher and a subject domain with the development corpus and the source dictionary, so it is unseen but not independent\. The degree\-marker counts are the weakest cells in Table[3](https://arxiv.org/html/2607.28766#S7.T3)and their vacancy is partly an artefact of the lexicon rather than a fact about Dungan\. Three mechanisms contribute\. The source dictionary lists 29\-щер/\-дихынforms as whole entries, 26 of them tagged as nouns, soзощер‘earlier’ andщинщер‘newer’ receive only a nominal reading and never reach the<comp\>slot\. Automatic part\-of\-speech assignment from the Russian gloss misfires on exactly the relevant class, in two ways: an entry whose gloss is a bare qualitative adjective but whose illustrative phrase is nominal was filed as a noun \(хи‘black’, gloss*чёрный*, illustrated byхи кўзы‘black trousers’\), and an adjective homographic with a numeral or a verb lost the slot to the competing reading \(бый白 ‘white’ againstбый百 ‘100’;го高 ‘high’ againstго告 ‘complain’\)\. The released analyzer carries an additive recovery layer of 42 adjective readings: the original reading is kept and an<adj\>reading is added beside it, so no analysis is lost, admitted only for short, underived, gradable stems whose gloss opens with a bare Russian qualitative adjective\. Both of Imazov’s canonical degree examples consequently build:хищерanalyses asхи<adj\><t1\><comp\>andбыйдихынasбый<adj\><t2\><sup\>\. The whole\-entry storage of the 29 dictionary degree forms is not fixed by that layer, and the layer is itself a hand\-audited patch over an automatic assignment, not a re\-derivation of it\. Separately,\-ли/\-ниon predicative adjectives\(Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\)are not modelled at all, so such tokens fall outside the table entirely\. We have not quantified the lexicon’s overall part\-of\-speech error rate; doing so is necessary future work\. The profile of Section[7](https://arxiv.org/html/2607.28766#S7)is a measurement*under the model*, with two specific dependencies\. Category counts are taken from the first analysishfst\-lookupreturns, and since all analyses carry weight zero that choice is an artefact of transducer path order; we bound the effect by reporting any\-reading figures alongside, but we do not resolve it, and a Constraint Grammar layer would\. The counts inherit the lexicon’s errors: part of speech is assigned automatically from the Russian gloss, derived forms are sometimes stored as whole stems \(notably the degree forms discussed in Section[7](https://arxiv.org/html/2607.28766#S7)\), and we have not quantified the resulting error rate\. Measured ambiguity is likewise bounded from below by lexicon incompleteness: homographs the lexicon does not contain cannot show up as ambiguity, so the 78\.1% unambiguous figure is an upper bound on the true share\. The model covers written literary language\. Spoken Kazakhstani Gansu Dungan differs considerably: recent fieldwork documents deep Russian influence and, for instance, the collapse of the classifier system to a singleгә\(Honkasalo,[2024](https://arxiv.org/html/2607.28766#bib.bib13)\), so performance on transcribed speech will be worse in ways the written\-text evaluation does not show\. Within the written language, a set of documented phenomena is not modelled \(each flagged as deferred where it arises\): the aspect clitics\-ли/\-ниon predicative adjectives\(Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\), which Section[8](https://arxiv.org/html/2607.28766#S8)identifies as the largest gap on held\-out text; noun reduplication in its distributive, diminutive\-hypocoristic \(лянлянзы‘little face’\) and kinship uses\(Tsunvazo,[1949b](https://arxiv.org/html/2607.28766#bib.bib34); Zevakhina,[2018](https://arxiv.org/html/2607.28766#bib.bib25)\), which in Northern Chinese and Jin dialects comes with tone alternations and erization\(Zavyalova,[1996](https://arxiv.org/html/2607.28766#bib.bib22)\); tone alternations under reduplication and compounding, and*déplacement*of the expiratory stress generally; the past habitual\-лэ/\-дилэ; productive erization and derivational suffixation; participles, converbs, imperatives; and the prefixal ordinalsту\-/ди\-/чу\-\(Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28)\)\. Negation is partial: the prohibitiveбә\(叵/嫑\) enters the model with<neg\>, but the full paradigm in which preposedбу/мә/бәtrigger regular loss of the aspect marker \(лэни‘will come’ —бу лэ;лэли‘came’ —мә лэ\), which have been described systematically byDragunov \([1940](https://arxiv.org/html/2607.28766#bib.bib19)\)and repeatedly confirmed\(Tsunvazo,[1949a](https://arxiv.org/html/2607.28766#bib.bib33); Kalimov,[1968](https://arxiv.org/html/2607.28766#bib.bib31); Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28)\), is syntactic and deliberately not treated at the morphological level\. The tone tags resolve only tonal homography, not full homonymy of stems identical in tone\. Finally, the lexicon derives from a single dictionary of the Gansu standard; its extraction from a PDF text layer, while checked, inherits any inaccuracies of the source\. ## Ethics Statement The model and tests are built exclusively on openly available published materials\. The development corpus uses Dungan Wikipedia incubator text \(CC BY\-SA\) and short published examples; a further 288 lemmas in the lexicon are harvested from English Wiktionary’s Dungan categories, whose content is CC BY\-SA 4\.0\. The analyzer is released under GPLv3, which Creative Commons designates as a one\-way compatible licence for BY\-SA 4\.0 material, so the ShareAlike condition on that portion of the lexicon is satisfied by the release as a whole\. The held\-out texts of the Institute for Bible Translation are fetched from the publisher’s site at evaluation time, used only to compute recognition statistics, and are not redistributed with the analyzer\. One test file contains a short excerpt of a published Dungan poem used as a recognition probe; it is flagged with a rights caveat in the source documentation and can be removed without affecting the build\. The work concerns a low\-resource minority language; the analyzer is intended as an open community resource, explicitly designed for inspection and correction by Dungan speakers and specialists, and makes no claims about speakers or communities beyond the linguistic properties of published text\. AI assistants \(Anthropic’s Claude\) were used throughout the project, and the scope was not limited to polishing text: for auxiliary programming \(extraction of the lexicon from the dictionary PDF, the build and evaluation scripts\), for editing the prose, and for assembling the 104\-item correctness gold set of Section[8](https://arxiv.org/html/2607.28766#S8), whose reference readings were assigned from the hanzi correspondences by the same assisted pipeline that built the lexicon\. All linguistic examples, numbers and bibliographic entries were verified by the authors against the primary sources, no previously unpublished data were shared with these services, and all substantive decisions are the authors’\. ## Список литературы - Towards finite\-state morphology of Kurdish\.Note:arXiv:2005\.10652External Links:2005\.10652Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - T\. Arkhangelskiy \(2019\)Corpora of social media in minority Uralic languages\.InProceedings of the Fifth International Workshop on Computational Linguistics for Uralic Languages,Tartu, Estonia,pp\. 125–140\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - K\. R\. Beesley and L\. Karttunen \(2003\)Finite state morphology\.CSLI Publications,Stanford, CA\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - C\. Bentz, T\. Ruzsics, A\. Koplenig, and T\. Samardžić \(2016\)A comparison between morphological complexity measures: typological data vs\. language corpora\.InProceedings of the Workshop on Computational Linguistics for Linguistic Complexity \(CL4LC\),D\. Brunato, F\. Dell’Orletta, G\. Venturi, T\. François, and P\. Blache \(Eds\.\),Osaka, Japan,pp\. 142–153\.External Links:[Link](https://aclanthology.org/W16-4117/)Cited by:[§9](https://arxiv.org/html/2607.28766#S9.p1.1)\. - E\. Bick and T\. Didriksen \(2015\)CG\-3 — beyond classical constraint grammar\.InProceedings of the 20th Nordic Conference of Computational Linguistics \(NODALIDA 2015\),B\. Megyesi \(Ed\.\),Vilnius, Lithuania,pp\. 31–39\.External Links:[Link](https://aclanthology.org/W15-1807/)Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1),[§7](https://arxiv.org/html/2607.28766#S7.p6.1)\. - A\. A\. Dragunov and E\. N\. Dragunova \(1937\)Dunganskii yazyk \[The Dungan language\]\.Zapiski Instituta vostokovedeniya AN SSSR6,pp\. 117–131\.Note:In RussianCited by:[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - A\. A\. Dragunov \(1940\)Issledovaniya v oblasti dunganskoi grammatiki\. Ch\. 1: Kategoriya vida i vremeni v dunganskom yazyke \(dialekt Gan’su\) \[Studies in Dungan grammar\. Part 1: The category of aspect and tense in the Dungan language \(Gansu dialect\)\]\.Trudy Instituta vostokovedeniya AN SSSR,Izdatel’stvo AN SSSR,Moscow; Leningrad\.Note:In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.15.12.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.16.13.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.17.14.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.18.15.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.19.16.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.21.18.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.22.19.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.17.13.2.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px4.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px5.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px6.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px8.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px9.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.28766#S1.p6.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p3.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - A\. A\. Dragunov \(1952\)Issledovaniya po grammatike sovremennogo kitaiskogo yazyka\. Ch\. 1: Chasti rechi \[Studies in the grammar of modern Chinese\. Part 1: Parts of speech\]\.Izdatel’stvo AN SSSR,Moscow; Leningrad\.Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.10.6.2.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px14.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px3.p1.1)\. - A\. Dragunow and E\. Dragunowa \(1936\)Über die dunganische Sprache\.Archiv Orientální8,pp\. 34–48\.Note:In GermanCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.2.2.3.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.14.11.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.6.3.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.14.10.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.7.3.2.1.1),[Приложение B](https://arxiv.org/html/2607.28766#A2.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px22.p1.2),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px7.p1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§5](https://arxiv.org/html/2607.28766#S5.p3.1),[footnote 3](https://arxiv.org/html/2607.28766#footnote3),[footnote 5](https://arxiv.org/html/2607.28766#footnote5)\. - S\. Honkasalo \(2024\)Kazakhstani Gansu Dungan as a contact language: an analysis of Russian influence\.Languages9\(2\),pp\. 59\.Note:[https://doi\.org/10\.3390/languages9020059](https://doi.org/10.3390/languages9020059)Cited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.14.10.2.1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p7.1)\. - M\. Kh\. Imazov \(1979\)O chastyakh rechi v dunganskom yazyke \[On parts of speech in the Dungan language\]\.InAktual’nye voprosy dunganovedeniya,pp\. 74–83\.Note:In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.1.1.3.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.11.8.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.8.5.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.15.11.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.19.15.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.20.16.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.21.17.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.22.18.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.23.19.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.24.20.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.25.21.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.4.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.8.4.2.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px1.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px10.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px12.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px17.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px2.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px20.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px3.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px4.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px4.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px2.p1.4),[§5](https://arxiv.org/html/2607.28766#S5.p1.1)\. - M\. Kh\. Imazov \(1982\)Ocherki po morfologii dunganskogo yazyka \[Essays on the morphology of the Dungan language\]\.Ilim,Frunze\.Note:In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.20.17.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.3.3.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.15.11.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.16.12.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.9.5.2.1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2607.28766#S1.p6.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px2.p1.4),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - M\. Kh\. Imazov \(1987\)Ocherki po sintaksisu dunganskogo yazyka \[Essays on the syntax of the Dungan language\]\.Ilim,Frunze\.Note:In RussianCited by:[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - A\. Kalimov \(1958\)Schetnye suffiksy v sovremennom dunganskom yazyke \[Numeral \(counting\) suffixes in modern Dungan\]\.InVoprosy grammatiki i istorii vostochnykh yazykov,Note:In RussianCited by:[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px17.p1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - A\. Kalimov \(1968\)Dunganskii yazyk \[The Dungan language\]\.InYazyki narodov SSSR\. Vol\. 5: Mongol’skie, tunguso\-man’chzhurskie i paleoaziatskie yazyki,Note:In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.20.17.2.1.1),[§1](https://arxiv.org/html/2607.28766#S1.p6.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p3.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - R\. M\. Kaplan and M\. Kay \(1994\)Regular models of phonological rule systems\.Computational Linguistics20\(3\),pp\. 331–378\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - L\. Karttunen and K\. R\. Beesley \(2005\)Twenty\-five years of finite\-state morphology\.InInquiries into Words, Constraints and Contexts,A\. Arppe, L\. Carlson, K\. Lindén, J\. Piitulainen, M\. Suominen, M\. Vainio, H\. Westerlund, and A\. Yli\-Jyrä \(Eds\.\),pp\. 71–83\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - T\. Khanna, J\. N\. Washington, F\. M\. Tyers, S\. Bayatlı, D\. G\. Swanson, T\. A\. Pirinen, I\. Tang, and H\. A\. i Font \(2021\)Recent advances in Apertium, a free/open\-source rule\-based machine translation platform for low\-resource languages\.Machine Translation35\(4\),pp\. 475–502\.External Links:[Document](https://dx.doi.org/10.1007/s10590-021-09260-6)Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1),[§8](https://arxiv.org/html/2607.28766#S8.p1.1)\. - K\. Koskenniemi \(1983\)Two\-level morphology: a general computational model for word\-form recognition and production\.Ph\.D\. Thesis,University of Helsinki, Department of General Linguistics\.Note:Publication No\. 11Cited by:[§1](https://arxiv.org/html/2607.28766#S1.p4.1),[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - K\. Lindén, E\. Axelson, S\. Hardwick, T\. A\. Pirinen, and M\. Silfverberg \(2011\)HFST — framework for compiling and applying morphologies\.InSystems and Frameworks for Computational Morphology \(SFCM 2011\),C\. Mahlow and M\. Piotrowski \(Eds\.\),Communications in Computer and Information Science, Vol\.100,pp\. 67–85\.Cited by:[§1](https://arxiv.org/html/2607.28766#S1.p5.1)\. - P\. Meurer \(2011\)A finite state approach to Abkhaz morphology and stress\.InLogic, Language, and Computation \(TbiLLC 2009\),Lecture Notes in Computer Science, Vol\.6618,pp\. 271–282\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - S\. N\. Moshagen, T\. A\. Pirinen, and T\. Trosterud \(2013\)Building an open\-source development infrastructure for language technology projects\.InProceedings of the 19th Nordic Conference of Computational Linguistics \(NODALIDA 2013\),S\. Oepen, K\. Hagen, and J\. B\. Johannessen \(Eds\.\),Oslo, Norway,pp\. 343–352\.External Links:[Link](https://aclanthology.org/W13-5631/)Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - M\. Naserzade, A\. Mahmudi, H\. Veisi, H\. Hosseini, and M\. MohammadAmini \(2023\)CKMorph: a comprehensive morphological analyzer for Central Kurdish\.International Journal of Digital Humanities5\(2\),pp\. 187–232\.External Links:[Document](https://dx.doi.org/10.1007/s42803-022-00062-7)Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - E\. D\. Polivanov \(1937a\)Fonologicheskaya sistema gan’suiskogo narechiya dunganskogo yazyka \[The phonological system of the Gansu dialect of the Dungan language\]\.InVoprosy orfografii dunganskogo yazyka,Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.2.2.4.1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - E\. D\. Polivanov \(1937b\)Muzykal’noe slogoudarenie, ili «tony» dunganskogo yazyka \[Musical syllable stress, or the ‘tones’ of the Dungan language\]\.InVoprosy orfografii dunganskogo yazyka,pp\. 41–58\.Note:In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.12.9.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.2.2.4.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px18.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px19.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px1.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px5.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px5.p2.2),[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - R\. Rahi, S\. Pushp, A\. Khan, and S\. K\. Sinha \(2020\)A finite state transducer based morphological analyzer of Maithili language\.Note:arXiv:2003\.00234External Links:2003\.00234Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - R\. J\. Reynolds \(2016\)Russian natural language processing for computer\-assisted language learning\.Ph\.D\. Thesis,UiT The Arctic University of Norway\.Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1)\. - S\. Rimsky\-Korsakoff Dyer \(1979\)Soviet Dungan kolkhozes in the Kirghiz SSR and the Kazakh SSR\.Oriental Monograph Series,Australian National University,Canberra\.Cited by:[§1](https://arxiv.org/html/2607.28766#S1.p1.1)\. - O\. Salmi \(1984\)The aspectual system of Soviet Dungan\.Folia Fennistica & Linguistica11\.Cited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.14.11.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.20.17.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.23.20.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.24.21.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.9.6.2.1.1),[Приложение B](https://arxiv.org/html/2607.28766#A2.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p3.1)\. - O\. Salmi \(2018\)Dungan–english dictionary\.Eastbridge Books,Manchester\.Cited by:[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px21.p1.1),[§1](https://arxiv.org/html/2607.28766#S1.p2.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1)\. - O\. Salmi \(2023\)The syntax of Central Asian Dungan\.Note:Unpublished manuscript \(draft of a Licentiate’s thesis\); deposited with the Digital Archive of Dungan StudiesCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.17.14.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.3.3.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.7.4.2.1.1),[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.9.6.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.3.3.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.26.22.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.9.5.2.1.1),[Приложение B](https://arxiv.org/html/2607.28766#A2.p1.1),[§1](https://arxiv.org/html/2607.28766#S1.p3.1),[§4](https://arxiv.org/html/2607.28766#S4.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p3.1)\. - K\. P\. Scannell \(2007\)The Crúbadán project: corpus building for under\-resourced languages\.InBuilding and Exploring Web Corpora: Proceedings of the 3rd Web as Corpus Workshop,Vol\.4,Louvain\-la\-Neuve,pp\. 5–15\.Cited by:[§1](https://arxiv.org/html/2607.28766#S1.p2.1)\. - Y\. Tsunvazo \(1949a\)K voprosu o sredstvakh vyrazheniya otritsaniya v dunganskom yazyke \[On the means of expressing negation in the Dungan language\]\.Vestnik Akademii nauk Kazakhskoi SSR,pp\. 99–106\.Note:No\. 10 \(55\)\. In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.21.18.2.1.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - Y\. Tsunvazo \(1949b\)Povtorenie v dunganskom yazyke \[Reduplication in the Dungan language\]\.Vestnik Akademii nauk Kazakhskoi SSR,pp\. 67–73\.Note:No\. 7 \(52\)\. In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.12.8.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.28.24.2.1.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - Y\. Tsunvazo \(1955\)K voprosu o sposobakh slovoobrazovaniya v dunganskom yazyke \[On methods of word formation in the Dungan language\]\.Izvestiya Akademii nauk Kazakhskoi SSR\. Seriya filologii i iskusstvovedeniya,pp\. 75–84\.Note:No\. 3–4\. In RussianCited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.2.2.3.1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p2.1)\. - Y\. Tsunvazo \(1963\)O morfologicheskom sposobe slovoobrazovaniya narechiya v dunganskom yazyke \[On the morphological method of adverb formation in the Dungan language\]\.Trudy Instituta yazykoznaniya AN Kazakhskoi SSR3\.Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.16.12.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.27.23.2.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px15.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px16.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.p2.1)\. - J\. Washington, M\. Ipasov, and F\. Tyers \(2012\)A finite\-state morphological transducer for Kyrgyz\.InProceedings of the Eighth International Conference on Language Resources and Evaluation \(LREC’12\),N\. Calzolari, K\. Choukri, T\. Declerck, M\. U\. Doğan, B\. Maegaard, J\. Mariani, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Istanbul, Turkey,pp\. 934–940\.External Links:[Link](https://aclanthology.org/L12-1642/)Cited by:[§2](https://arxiv.org/html/2607.28766#S2.p1.1),[§8](https://arxiv.org/html/2607.28766#S8.p1.1),[§8](https://arxiv.org/html/2607.28766#S8.p2.1)\. - S\. Wichmann, E\. W\. Holman, and C\. H\. Brown \(2016\)The ASJP database\.Max Planck Institute for the Science of Human History,Leipzig\.Note:Available at[https://asjp\.clld\.org/](https://asjp.clld.org/)Cited by:[§8](https://arxiv.org/html/2607.28766#S8.p8.1)\. - Y\. Yanshansin \(2009\)Kratkii dungansko\-russkii slovar’ \[Concise Dungan–Russian dictionary\]\.2, revised and enlarged edition,Institute for Bible Translation,Moscow\.Note:In Russian and DunganExternal Links:ISBN 978\-5\-93943\-139\-2Cited by:[Таблица 7](https://arxiv.org/html/2607.28766#A2.T7.3.6.3.3.1.1),[§3](https://arxiv.org/html/2607.28766#S3.p1.1)\. - O\. I\. Zavyalova \(1979\)Dialekty Gan’su \[The dialects of Gansu\]\.Nauka, GRVL,Moscow\.Note:In RussianCited by:[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - O\. I\. Zavyalova \(1996\)Dialekty kitaiskogo yazyka \[The dialects of Chinese\]\.Nauchnaya kniga,Moscow\.Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.12.8.2.1.1),[§1](https://arxiv.org/html/2607.28766#S1.p1.1),[§1](https://arxiv.org/html/2607.28766#S1.p3.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - T\. S\. Zevakhina and M\. Kh\. Imazov \(1997\)Dunganskii yazyk \[The Dungan language\]\.InYazyki Rossiiskoi Federatsii i sosednikh gosudarstv\. Entsiklopediya\. Vol\. 1,pp\. 349–362\.Note:In RussianCited by:[§1](https://arxiv.org/html/2607.28766#S1.p6.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1)\. - T\. S\. Zevakhina \(2001\)Funktsional’no\-grammaticheskaya parametrizatsiya prilagatel’nogo \(po dannym polevogo issledovaniya dunganskogo yazyka\) \[Functional\-grammatical parametrization of the adjective \(based on fieldwork on the Dungan language\)\]\.Yazyk, soznanie, kommunikatsiya20,pp\. 69–86\.Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.17.13.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.7.3.2.1.1),[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.8.4.2.1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px11.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px13.p1.1),[Приложение C](https://arxiv.org/html/2607.28766#A3.SS0.SSS0.Px14.p1.1),[Приложение D](https://arxiv.org/html/2607.28766#A4.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2607.28766#S2.p2.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.28766#S5.SS0.SSS0.Px2.p1.4),[§8](https://arxiv.org/html/2607.28766#S8.p11.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p5.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - T\. S\. Zevakhina \(2018\)Reduplikativy dunganskogo yazyka v slovare i tekste \[Reduplicatives of the Dungan language in dictionary and text\]\.InMalye yazyki v bol’shoi lingvistike\. Sbornik trudov konferentsii 2017,K\. P\. Semenova \(Ed\.\),pp\. 57–62\.Note:In RussianCited by:[Таблица 8](https://arxiv.org/html/2607.28766#A2.T8.4.28.24.2.1.1),[§5](https://arxiv.org/html/2607.28766#S5.p2.1),[Limitations](https://arxiv.org/html/2607.28766#Sx1.p8.1)\. - T\. S\. Zevakhina \(2019\)Semanticheskie relyatsii slozhnogo slova v dunganskom yazyke: kolichestvennyi analiz \[Semantic relations of the compound word in the Dungan language: a quantitative analysis\]\.InSlovo\. Slovar’\. Termin\. Leksikograf,pp\. 245–250\.Note:In RussianCited by:[§5](https://arxiv.org/html/2607.28766#S5.p2.1)\. ## Приложение ATag Inventory and Morphotactic Template Tables[4](https://arxiv.org/html/2607.28766#A1.T4)and[5](https://arxiv.org/html/2607.28766#A1.T5)give the full tag inventory of the analyzer, and Table[6](https://arxiv.org/html/2607.28766#A1.T6)the morphotactic template\. Таблица 4:Parts of speech in the model: tag, class, example\. Cells containing tags are analyzer output; cells giving only a lemma and a gloss name the example word\.TagMeaningExample*overt inflection \(counted in Section[7](https://arxiv.org/html/2607.28766#S7)\)*<pl\>plural \(animate\)жынму‘people’<gen\>genitive / attributiveсамарияди‘of Samaria’<loc\>locativeдащүәни‘at the university’<px\>possessive \(pronoun\)вәмуди‘our’<pst\>perfective \(demarcational\)нянли‘read \(pfv\)’<prog\>progressive\-durativeщедини‘is writing’<fut\>prospective\-futureщени‘will write’<exp\>experientialнянгуә‘has read \(repeatedly\)’<comp\>/<sup\>comparative / elativeдащер,дадихын*lexical and zero\-marked \(not counted\)*<an\>/<nn\>animacy / inanimacyжын<an\>,вынхуа<nn\><sg\>singular \(zero\-marked\)жын<sg\><t1\>/<t2\>/<t3\>lexical toneзадё‘suck out’/‘fence off’/‘blow up’<neg\>prohibitive negationбә<pers\>,<p1\>–<p3\>pronoun personвә<pers\><p1\><incl\>inclusive \(pronoun\)заму‘we \(incl\.\)’<dem\>/<itg\>demonstr\. / interrog\.нагә‘that’,са‘what’Таблица 5:Features assigned on top of the part\-of\-speech tag, split by whether they correspond to an overt exponent\. Only the upper group is counted as inflection in Section[7](https://arxiv.org/html/2607.28766#S7)\. Cells containing tags are analyzer output; cells giving only a lemma and a gloss name the example word\.Таблица 6:Morphotactic template: linear order of marker attachment \(optional positions in parentheses, alternatives separated by\|\|\)\. The tone tag is emitted after the part\-of\-speech and animacy tags and before inflection\. Note that the aspect slot on the verb is optional in the transducer, so a bare<vblex\>form is generable even though aspect marking is obligatory in the language \(Section[5](https://arxiv.org/html/2607.28766#S5)\); and that the adjective admits only the progressive, not the full aspect set\. ## Приложение BModelling Decisions with Page\-Level Provenance Tables[7](https://arxiv.org/html/2607.28766#A2.T7)and[8](https://arxiv.org/html/2607.28766#A2.T8)gather every modelled phenomenon into a single decision table: the phenomenon, the source\(s\) with page\-level anchors, and how it is formalized in the analyzer\. Page anchors are given for editions whose scans were read directly; where no page number can be substantiated \(e\.g\. the author’s web version ofSalmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15)does not preserve the journal pagination\) the pages column shows a dash rather than a guess; pages known only through citation in another work are marked ‘‘ap\.’’ \(apud\)\. The German\-languageDragunow and Dragunowa \[[1936](https://arxiv.org/html/2607.28766#bib.bib12)\]is cited by its own pagination \(S\.\), the manuscriptSalmi \[[2023](https://arxiv.org/html/2607.28766#bib.bib17)\]by draft pages \(ms\.\)\. PhenomenonSource\(s\), pagesFormalization*General*Dungan an independent language; Gansu dialect as literary standardDragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 34–35Lexicon extracted from a Gansu dictionary\[Yanshansin,[2009](https://arxiv.org/html/2607.28766#bib.bib37)\]; Gansu is the lexical base and three\-tone standard of the model, Shaanxi a separate hook lexicon\.Gansu as the literary norm \(first school teachers from Frunze\)Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. p\. 12Same architectural decision \(Gansu as base lexicon\), not a separate tag\.Inventory of 14 parts of speechImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), pp\. 74–7511 of Imazov’s 14 realized as POS tags \(Table[4](https://arxiv.org/html/2607.28766#A1.T4)\); participle, converb and preposition are not tagged \(ба把 is<part\>,лян連 is<cnjcoo\>\);<np\>and<cop\>are added\.More morphological markers than in ChineseSalmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. p\. 27;Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15), —General observation motivating finite\-state modelling; no dedicated tag\.*Nominal morphology*Criteria for the noun; animacy \(сый/саtest\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), pp\. 78–83Tag<n\>\+ feature<an\>/<nn\>\(Table[5](https://arxiv.org/html/2607.28766#A1.T5)\); animacy gates the plural\-му\.Animate plural\-му; genitive\-диImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), pp\. 78–83<pl\>on\-му\(жын→\\toжынму\);<gen\>on\-ди\. Analytic number \(ну,хошо,хиге\) is syntactic, unmarked\.Substantive suffixes\-зы/\-р\(lexicalized, not segmented\)Tsunvazo,[1955](https://arxiv.org/html/2607.28766#bib.bib35), pp\. 75, 82–83;Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 44Stored as whole dictionary stems \(до→\\toдозы\); see §[5](https://arxiv.org/html/2607.28766#S5)and Appendix[D](https://arxiv.org/html/2607.28766#A4)\.\-зы\-marking correlates with stem tonePolivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32), pp\. 41–58Treated as an explanation of lexicalization choice, not a separate rule; see Appendix[D](https://arxiv.org/html/2607.28766#A4)\.Classifier\-гәa suffix, not a counting wordSalmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. p\. 46;Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28), p\. 66 \(ap\. Salmi 2023\)Tag<cls\>; 8 lexicon units \(гә/ба/җон/тё/җы/пи/бын/фу\), 5 productive; numeral\+classifier a single complex \(йи\+ 個→\\toйигә\)\.*Verbal morphology \(aspect–tense system\)*Obligatory aspect marking: bare verb makes an incomplete utteranceDragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 48;Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15), —The five markers of Table[2](https://arxiv.org/html/2607.28766#S5.T2)are the<vblex\>continuation classes\. NB the transducer leaves the slot optional, so a bare<vblex\>form is generable; obligatoriness is not enforced\.Perfective\-лиas ‘‘demarcation point’’ \(aspect, not tense\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), pp\. 15, 26Tag<pst\>a practical label; aspectual analysis and examples in §[5](https://arxiv.org/html/2607.28766#S5)and Appendix[D](https://arxiv.org/html/2607.28766#A4)\.Inceptive reading of\-лиwith statives \(‘came to know’, ‘became cold’\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), pp\. 15–17, 26Same<pst\>; the inceptive reading is a contextual consequence of stem semantics, not a separate tag\.Double\-лиwith numeral objects \(perfect of persistent situation\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 32;Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. §6\.3Formalized in both loci at once: both verb and object noun admit<pst\>/aspect clitics \(§[5](https://arxiv.org/html/2607.28766#S5), analytic constructions\)\.Imperfective/future\-ни: conditioned predicate; restriction with modalsDragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), pp\. 39–40, 44Tag<fut\>on\-ни; the modal restriction among the optionality cases \(next row\)\.Progressive\-диниvs stative\-диcontrast \(background vs point\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 26Both markers receive<prog\>\(not distinguished by tag\); cf\. the divergence from Salmi’s five\-way analysis \(Appendix[D](https://arxiv.org/html/2607.28766#A4)\)\.Experiential\-гуә\(過, ‘at least once’\)Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15), —;Kalimov,[1968](https://arxiv.org/html/2607.28766#bib.bib31), —;Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28), —Tag<exp\>on\-гуә\(нянгуә\)\.Aspect\-marker loss under preposed negation \(мә/\-ли,бу/\-ни\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 58;Tsunvazo,[1949a](https://arxiv.org/html/2607.28766#bib.bib33), pp\. 101–104NOT formalized at the morphological level — syntactic, requires context beyond the stem; see Section[9](https://arxiv.org/html/2607.28766#S9)\.Aspect optionality \(modals, statives, imperative, nominal predicate\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 44Not blocked by a dedicated tag; control rests with the lexical entry of each stem\.Five\-way aspect system \(different partition of categories\)Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15), —Divergence discussed in Appendix[D](https://arxiv.org/html/2607.28766#A4)\(progressive/stative as two tags instead of Salmi’s imperfective\)\.Past habitual\-лэ/\-дилэ\(not modelled\)Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15), — \(crediting Dragunov 1952\)NOT in the lexicon — nearly unattested in the development corpus; deferred \(Section[9](https://arxiv.org/html/2607.28766#S9)\)\.Таблица 7:Modelling decisions with page\-level provenance \(part 1 of 2\)\. A dash in the pages column means the edition has no citable pagination \(web version\), not a missing source\.PhenomenonSource\(s\), pagesFormalization*Adjective*Adjective as predicator; copular\-сыa copula suffix from 是Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 42, 44;Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24), pp\. 72–73, 75–76<adj\>admits aspect clitics \(predicator status\); the copula is a separate tag<cop\>\(сы/шы, Table[4](https://arxiv.org/html/2607.28766#A1.T4)\)\.Predicative\-диon the adjectiveImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), pp\. 78–83;Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24), pp\. 72–73, 75–76The same marker\-диas the nominal genitive, in predicative\-attributive role on<adj\>\(marker polyfunctionality\)\.Degrees of comparison\-щер/\-дихынImazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28), pp\. 58, 60–61 \(ap\. Salmi 2023\);Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. p\. 48<comp\>on\-щер,<sup\>on\-дихын; three competing readings of\-дихын: Appendix[D](https://arxiv.org/html/2607.28766#A4)\.Tonal criterion distinguishing adjective from verbDragunov,[1952](https://arxiv.org/html/2607.28766#bib.bib20), — \(ap\. Zevakhina 2001\)Not encoded as a tag; cited as an additional \(for the model, non\-morphological\) argument for<adj\>as an autonomous class\.*Tone and morphophonology*Lexical tone: three tones in isolation; privative analysisPolivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32), pp\. 41–58;Polivanov,[1937a](https://arxiv.org/html/2607.28766#bib.bib44)Tags<t1\>/<t2\>/<t3\>on≈\\approx7,100 stems \(≈\\approx91% of the lexicon\); the privative reading of tone I: Appendix[D](https://arxiv.org/html/2607.28766#A4)\.Three tones in isolation / four in connected speech; sandhi ‘‘first→\\tosecond’’Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. pp\. 11–12 \(crediting Zavyalova 1979\)NOT modelled \(sandhi requires connected\-speech context\); see Appendix[D](https://arxiv.org/html/2607.28766#A4)\.Tone sandhi under reduplication \(not modelled\)Tsunvazo,[1949b](https://arxiv.org/html/2607.28766#bib.bib34), pp\. 67–70;Zavyalova,[1996](https://arxiv.org/html/2607.28766#bib.bib22), pp\. 64–65NOT modelled — reduplication as a whole outside this version; see also Appendix[D](https://arxiv.org/html/2607.28766#A4)\.*Other*Full and short forms \(elision of /i/, /\\textschwa/\); relevant to tokenizationHonkasalo,[2024](https://arxiv.org/html/2607.28766#bib.bib13), §3\.2;Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 38 \(ap\. Honkasalo\)Handled at lexicon\-extraction level \(tokenization\), not by a grammatical tag\.Noun locatives\-ни\(裡 ‘in’\) and\-шон\(上 ‘on’\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), pp\. 78–83;Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28)<loc\>on\-ни\(гонзы\-ни‘in the bucket’,фули\-ни‘in the forest’\) and\-шон\(дезы\-шон‘on the plate’\)\.Genitive\-диpolyfunctional \(attributive, substantivising, adverbial;гўр\-ди‘with a sharp sound’\)Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28);Tsunvazo,[1963](https://arxiv.org/html/2607.28766#bib.bib36)The one marker appears in several roles in the model: nominal genitive, predicative\-attributive on adjectives \(motivates multi\-role\-ди\)\.Noun admits aspect clitics \(predicative marking of the V–O group\)Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19);Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)<n\>continuation classes include the aspect clitics\-дини/\-ди/\-ни\(§[5](https://arxiv.org/html/2607.28766#S5)\)\.Collective ethnonymsгрек\-жын\-му‘Greeks’ \(lit\. ‘Greek\-person\-PL’\)corpus\-attested; provisionalModelled asжын人 \+\-муattachment to ethnonym stems; explicitly flagged as a*provisional*generalization\.Pronoun paradigm: person/numberвә, ни, та, вәму, ниму, таму; inclusiveза/заму\(咱/咱們\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<prn\>with subtypes<pers\>\(person/number features\) and<incl\>; pronouns take no preposed modifiers\.Possessive pronoun forms in\-ди\(вәму\-ди‘our’,ниму\-ди‘your’,таму\-ди‘their’\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<px\>slot in the pronominal template \(Table[6](https://arxiv.org/html/2607.28766#A1.T6)\)\.Demonstrativesҗыгә/җәгә\(這個\),нагә/ныйгә\(那個\); interrogativesса\(啥\),сый\(誰\),зуа,залиImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<dem\>and<itg\>subtypes of<prn\>\.Copulaсы\(Gansu\) /шы\(Shaanxi\) marking the nominal predicateImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)Separate tag<cop\>, dialect hook lexicon\.Postpositions: a closed class \(литу‘inside’,вэту‘outside’,туни‘in front’,хуту‘behind’,диха‘below’\); may take the genitive \(либян→\\toлибян\-ди\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<post\>closed class in the lexicon; genitive continuation enabled\.Particles: a closed modal\-focus class \(ба把, interrogativeма,ла, …\)Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<part\>closed class\.Conjunctions \(coordinating/subordinating\) and interjectionsImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)<cnjcoo\>/<cnjsub\>and<ij\>closed classes\.Function\-word richness compensating minimal inflectionImazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)Motivates the closed function\-word lexica above; no dedicated tag\.Dialect differences \(copula, personal pronouns, classifiers\)Salmi,[2023](https://arxiv.org/html/2607.28766#bib.bib17), ms\. p\. 13Realized as dialect hook lexicons \(Table[1](https://arxiv.org/html/2607.28766#S4.T1):сы/шы,вә/ңә,вәму/ңәму–заму,таму/ана–тана,гә/гуә–гы\)\.Morphological formation of adverbsTsunvazo,[1963](https://arxiv.org/html/2607.28766#bib.bib36), pp\. 1–2 of scanAdverbs stored as whole dictionary stems with<adv\>; no productive adverbial suffix in the model \(unlike Chinese\)\.Reduplication as a productive mechanism \(not modelled\)Tsunvazo,[1949b](https://arxiv.org/html/2607.28766#bib.bib34), pp\. 67–70;Zevakhina,[2018](https://arxiv.org/html/2607.28766#bib.bib25), —NOT modelled — outside the current lexicon version; see Section[9](https://arxiv.org/html/2607.28766#S9)\.Таблица 8:Modelling decisions with page\-level provenance \(part 2 of 2\)\. ## Приложение CVerbatim Source Quotations for the Modelled Rules For each modelled rule this appendix gives the source, page and paragraph \(paragraphs are counted from the top of the page; the opening words of the paragraph are given in parentheses for verification\), followed by a verbatim quotation\. For rules supported by several sources, several quotations are given\. Quotations keep the orthography and transcription of the originals \(Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19)uses Latin transcription for the markers\-li/\-ni/\-dini;Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27)Cyrillic\-ли/\-ни/\-дини\)\. ### POS annotation: 14 parts of speech\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 78, para\. 6 \(«Таким образом, разбираемые…»\): «…разбираемые слова обладают рядом семантических, синтаксических и морфологических признаков \(значение предметности, функция подлежащего или дополнения…наличие категории числа и категории одушевлённости — неодушевлённости\), которые отличают их от слов других классов\. А такие слова…обычно относят к существительным\.» ### Animate plural\-му<pl\>\(only with nouns denoting persons\)\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 78, para\. 2 \(«Отдельные из разбираемых…»\): «…морфему\-му, которая выражает значение множественности, например:хәйуан\(учитель\) —хәйуанму\(учителя\)…Правда, такой показатель множественности могут иметь лишь те немногие из них, которые обозначают лица\.» ### Animacy<an\>/<nn\>\(theсый?/са?test\)\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 78, para\. 5 \(«…в дунганском…»\): «…в дунганском же различение одушевлённых и неодушевлённых имён существительных получило своё грамматическое выражение в постановке вопроса при установлении отнесённости слова к той или иной части речи\.» \(the questionsсый?‘who?’ /са?‘what?’\) ### Verbal aspect markers\-дини/\-ни/\-ли\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 79, para\. 5 \(«Разбираемые слова имеют…»\): «Разбираемые слова имеют также характерные морфологические признаки — суффиксы\-дини,\-ни,\-ли, например:фадини\(играет\),фани\(будет играть\),фали\(играл\);подини\(бежит\),пони\(будет бежать\),поли\(бежал\)…» Same rule, second source:Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 15, para\. 2 \(«Несмотря на кажущееся…»\): «…не требуют обязательного оформления на видовые предикативные суффиксы, т\. е\. на суффиксы \-dini, \-li, \-ni\.» ### \-ли<pst\>— perfective, demarcation point \(not a true past\)\. Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 15, §I, para\. \(«Перфективное значение…»\): «Перфективное значение суффикса \-li во всех этих примерах отчётливо выражено: то или иное действие или состояние имело место до такого\-то определённого момента, после чего наступило новое действие или состояние\. Переломный, демаркационный момент между концом прежнего \(предыдущего\) и началом нового \(последующего\) действия или состояния суффикс \-li как раз именно и выражает\.» ### Predicative / modal\-ли<pst\>on the nominal predicate \(aspect clitic on the object noun in the V–O construction; marks completion of the measure, not of the action, which may continue\)\. Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 32, para\. \(«Конструкция с числительным…»\): «…отнюдь не к самому действию, а к дополнению с числительным определением\. Действие же как таковое может и продолжаться» \(ibid\. the example «к данному моменту я живу во Фрунзе уже три года»; in our corpus —жын\-ли人哩, the hanzi gloss of the folk tale 把三個女子給哩人哩\)\. ### Obligatory aspect marking \(a bare verb is an unfinished word combination\)\. Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12), S\. 48: «Insofern sie keine Suffixe haben, sind sie im Dunganischen … sinnlos und keineswegs als Sätze zu betrachten, sondern als unvollendete Wortverbindungen\. Hierin liegt einer der am schärfsten ausgeprägten Unterschiede zwischen dem Dunganischen und dem Chinesischen\.» ### \-ни<fut\>— prospective / conditioned \(aspectual, not temporal\)\. Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 39, para\. \(«Случаи этого рода…»\): «Случаи этого рода отлично иллюстрируют различие в значении суффиксов \-li и \-ni\. В каждом из приведённых примеров оба действия относятся к будущему, однако первое оформлено на суффикс \-li, поскольку оно является сказуемым обусловливающим, второе же — на суффикс \-ni, поскольку оно является сказуемым обусловленным\.» ### \-ниis not an interrogative marker\. Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 39, para\. \(«Пример этот интересен…»\): «…приписывать дунганскому суффиксу \-ni «вопросительное» значение, как это иногда делается по аналогии с китайским языком, было бы неверно…» ### Degrees of comparison\-щер<comp\>,\-дихын<sup\>\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 80, para\. 4 \(«Последним присущи…»\): «…хи\(чёрный\) —хищер\(чернее\),го\(высокий\) —гощер\(выше\),бый\(белый\) —быйдихн\(очень белый\),ван\(мягкий\) —вандихн\(очень мягкий\)…Морфемы\-щери\-дихн…являются отличительным морфологическим признаком…прилагательных\.» ### \-дихынas intensifier of ‘‘full predication’’ \(refinement of<sup\>\)\. Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24), p\. 75, para\. 4 \(«Первая из этих…»\): «…адъективный суффикс полной предикации, который в конструкциях такого вида является необязательным, но весьма желательным…» \(of the marker\-дихын\) ### Predicative\-диon the adjective with copulaсы\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 80, para\. 3 \(«Слова данной категории…»\): «Слова данной категории в предложении являются частью составного именного сказуемого, если употребляются с суффиксом\-ди, а определяемый ими член предложения…употребляется со связкойсы, например:Ванваньсы мутуди\(пиала деревянная\)…Суффикс\-див данном случае определяет отнесённость…к прилагательным\.» ### Adjective<adj\>— a predicator but an autonomous part of speech\. Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24), p\. 75, para\. 6 \(«Близость китайского…»\): «Близость китайского прилагательного по своим морфологическим свойствам к глаголам позволила ряду исследователей говорить об общей категории предикативов \(ср\. \[Драгунов 1952: 12\]\)…признать за прилагательными дунганского языка статус, хотя и особой, но всё же самостоятельной части речи\.» ### Morphological criterion distinguishing adjective from verb — tone under reduplication\. Dragunov,[1952](https://arxiv.org/html/2607.28766#bib.bib20)\(apudZevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24), p\. 76, para\. 1, «…различие отражается…»\): «Он отмечает, что их различие отражается в тональной структуре сложных слов, образованных путём удвоения качественной, с одной стороны, и глагольной, с другой стороны, морфем \[Драгунов 1952: 168–170\]\.» ### Polyfunctional\-ди\(function element; with onomatopoeia\)\. Tsunvazo,[1963](https://arxiv.org/html/2607.28766#bib.bib36), p\. 52, para\. 2 \(«В дунганском языке…»\): «В дунганском языке звукоподражательные слова сами по себе, ни в суффиксальном оформлении самостоятельно не употребляются\. Почти всегда они используются в связанной речи, обычно в сочетании с суффиксом\-ди\.» ### Adverbs stored as whole stems \(no dedicated adverbial suffix\)\. Tsunvazo,[1963](https://arxiv.org/html/2607.28766#bib.bib36), p\. 53, last para\. \(«Наречия, обозначающие…»\): «Наречия, обозначающие одни и те же понятия, в дунганском и китайском языках в ряде случаев имеют различное оформление \(разные предлоги, суффиксы и знаменательные морфемы\)\.» ### Counting words \(classifiers\) with the numeral\. Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 81, para\. 1 \(«…количества в языке…»\): «…чаще всего сочетаются со счётными словами типафон\(пара\),ба\(горсть\)…например:Ве мали лён\-фон хэ\(я купил две пары обуви\)…» Same rule, second source:Kalimov,[1958](https://arxiv.org/html/2607.28766#bib.bib30), title of the work: «Счётные суффиксы в современном дунганском языке» \(a dedicated study of the counting markers\)\. ### Lexical tone<t1\>/<t2\>/<t3\>\(Gansu: three tones, notation I/II/III\)\. Polivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32), p\. 41, §1, para\. 1 \(«Характерным отличием…»\): «Характерным отличием всех «тибето\-китайских» языков является наличие так называемых «тонов», т\. е\. различных мелодий голосового тона, присущих определённым слогам\-морфемам…» ### Three tonemes, numbered I, II, III\. Polivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32), p\. 56, §9 \(heading\): «§ 9\. Примеры дунганских односложных слов под разными \(I, II, III\) ГС тонами» \(followed by the rowstan‘field’ \(I\) / ‘carpet’ \(II\) / ‘coal’ \(III\), etc\.\)\. ### twol: deletion of the morpheme boundaries%\>,%\+→\\to∅\\emptyset\. Engineering rule; grounding:Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27), p\. 78, para\. 6: markers attach to the stem without morphophonological changes \(Dungan is analytic\); the boundary is needed only for morphotactics and is erased on the surface \(«…наличие у некоторых из них определённых флексий…», Imazov 1979, p\. 78\)\. ### twol \(documented, vacuous\):ў/уafter labials\. Salmi,[2018](https://arxiv.org/html/2607.28766#bib.bib16); Russian Wikipedia, «Дунганский язык»; descriptive phonology: the rule is carried as a documented constraint and is vacuous in the model \(stems are stored in final orthography\)\. After the labialsб п м ф вthe closed /u/ is writtenу\(бу不,фу服\), notў\. No verbatim quotation from the accessible Soviet sources; the rule is not active\. ### twol \(documented, vacuous\): Shaanxiдь/ть→\\toҗь/чь\. Dragunow and Dragunowa,[1936](https://arxiv.org/html/2607.28766#bib.bib12); descriptive phonology: carried as a documented constraint, vacuous in the model \(Shaanxi stems are stored already palatalized\)\. The historical softдь\-/ть\-yields Shaanxiҗь\-/чь\-; at syllable levelдьа→\\toҗьа,тьи→\\toчьи\. The rule is not active; no specific dictionary pair is cited, to avoid inaccuracy\. ## Приложение DExtended Notes: Derivation, Tone History and the Transducer Graph ### Derivational suffixes\-зы/\-р\(full discussion\)\. A special case is derivational, not inflectional, suffixation\. Most disyllabic Dungan nouns arise by compounding, in which the tone of one component often changes, and a number of monosyllabic roots become autonomous only when dressed with the ‘‘substantive’’ suffixes\-зы\(子\) or\-р\(兒\):до‘sword’ —дозы‘knife’, and the bound rootкў\-\(free only in compounds, cf\.кўтуй‘trouser leg’, whereуйis ‘leg’\) —кўзы‘trousers’\[Tsunvazo,[1955](https://arxiv.org/html/2607.28766#bib.bib35), p\. 79\]\. The model treats such formations as ready\-made dictionary units rather than productively detachable markers:\-зыis largely lexicalized in the modern language \(cf\.фонзы‘house’,гонзы‘bucket’, which have lost their diminutive value\), and segmenting it productively would create spurious homonymy\. Interestingly, the very choice of\-зы\-marking correlates statistically with stem tone:Polivanov \[[1937b](https://arxiv.org/html/2607.28766#bib.bib32)\]observes that stems under the longer tone \(in Gansu, tone I\) more often remain monosyllabic, while stems under the short tones gravitate to\-зы\-marking\. Modelling derivation is therefore deferred\. ### Perfective\-лиand prospective\-ни\(full discussion\)\. The model’s tag<pst\>\(‘past’\) is a practical label:\-лиis in essence an aspectual, not a temporal marker\.Dragunov \[[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 15\]analyses it as a ‘‘demarcation point’’ — the turning point between a previous and a new state\. This explains why with atelic stems\-лиyields not a past but an inceptive reading: with the stative verbsзы‘know’ andю‘have’ the meanings ‘came to know’, ‘came to have’ arise \(юли‘\(I\) came to have’\), and with a predicative adjective, a change of state \(лын\-ли‘it became cold’\)\. The same accounts for the combination of\-лиwith the negatorбу, which conveys cessation \(‘quit smoking’\) rather than a plain past\. Symmetrically, the prospective\-ниis above all aspectual rather than temporal:Dragunov \[[1940](https://arxiv.org/html/2607.28766#bib.bib19), pp\. 39–40\]notes that in conditional and temporal constructions, even when both events are in the future, the conditioning \(prior\) predicate takes perfective\-лиand the conditioned \(subsequent\) one takes prospective\-ни\(‘when I arrive \[\-ли\], I will speak \[\-ни\]’\); the\-ли/\-ниopposition contrasts completion and prospection, not past and future\. Hence also the ‘not yet’ construction \(хан мә V\-ни\), where the negatorмәand prospective\-ниregularly co\-occur: the event has not yet arrived but is in prospect\. The obligatoriness of aspect marking has a known list of exceptions \(negated verb, modal verbs, statives, the imperative, nominal and temporal\-age predicates\), in which the marker is optional\[Dragunov,[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 44\]\. The systematic correspondence of positive and negative forms, in which a preposed negative particle triggers regular loss of the aspect marker, was described byDragunov \[[1940](https://arxiv.org/html/2607.28766#bib.bib19), p\. 58\]; see Section[9](https://arxiv.org/html/2607.28766#S9)\. ### Relation to Salmi’s analysis \(full discussion\)\. The five\-member system is close to Salmi’s independent analysis, which also distinguishes five aspect–tense categories for Central Asian Dungan, but with a different partition: past habitual\-лэ/\-дилэ, experiential perfect, imperfective, perfective and future\[Salmi,[1984](https://arxiv.org/html/2607.28766#bib.bib15)\]\. The divergence is twofold\. First, our model keeps progressive\-диниand stative\-диas separate markers, whereasSalmi \[[1984](https://arxiv.org/html/2607.28766#bib.bib15)\]unites them into an imperfective \(non\-final form\-ди, final\-дини\)\. Second, the past habitual he describes \(earlier alsoDragunov,[1952](https://arxiv.org/html/2607.28766#bib.bib20)\) —\-лэin stative clauses \(with the copula, with stative adjectives, withмәю‘not have’\) and\-дилэin actional ones — is*not*formalized in the present model: it is nearly unattested in our corpus of written Dungan and is deferred\. The stative/actional split itself, with whichSalmi \[[1984](https://arxiv.org/html/2607.28766#bib.bib15)\]explains the defectiveness of the aspect paradigm in stative contexts, agrees with the optionality contexts listed in the preceding discussion of\-ли/\-ни\. ### Degrees of comparison \(full discussion\)\. The adjective carries<adj\>and admits the comparative\-щер\(Imazov’s examples:хи‘black’ —хищер‘blacker’,го‘tall’ —гощер‘taller’\), the superlative \(elative\)\-дихын\(得很:бый‘white’ —быйдихын‘very white’\)\[Imazov,[1979](https://arxiv.org/html/2607.28766#bib.bib27),[1982](https://arxiv.org/html/2607.28766#bib.bib28)\], and predicative use with aspect particles; byZevakhina \[[2001](https://arxiv.org/html/2607.28766#bib.bib24)\]the predicative adjective takes the whole verbal aspect set \(\-ли,\-ни,\-дини, cf\.Ни хо\-дини ма?‘How do you do?’, lit\. ‘are you in a good state?’\), though\-ли/\-ниon adjectives are nearly unattested in our corpus and not modelled productively\.\-дихынis described both as a superlative\[Imazov,[1982](https://arxiv.org/html/2607.28766#bib.bib28)\]and — matching its etymology as the intensifier 得很 ‘very’ — as an optional ‘‘suffix of full predication’’\[Zevakhina,[2001](https://arxiv.org/html/2607.28766#bib.bib24)\], cf\.Кў\-зы та\-шон зый\-дихын‘the trousers are \(too\) tight for him’; all descriptions agree on the intensive\-predicative reading, and the model adopts the superlative tag\. ### Lexical tone: orthographic history and deplacement\. Introducing tone tags can be seen as restoring — at the level of grammatical annotation rather than orthography — the distinctive feature that the 1928–1932 Latin script could in principle have conveyed — but did not, being aligned with Latinxua Sin Wenz, which programmatically leaves tone unwritten\. Tellingly, leaving tone unwritten was a deliberate decision:Polivanov \[[1937b](https://arxiv.org/html/2607.28766#bib.bib32)\]himself argued against*obligatory*tone marking, holding that it would encumber Dungan writing as obligatory stress marks would encumber Russian \(cf\.за́мок‘castle’ /замо́к‘lock’\), and allowed only an optional mark in ambiguous cases and in pedagogical literature\. The homography our tags resolve is thus built into the very norm of the script\. Nor is the incompleteness of the orthographic record exhausted by that\.Polivanov \[[1937b](https://arxiv.org/html/2607.28766#bib.bib32)\]analyses the three\-tone system*privatively*: tone I is ‘‘null’’ \(level, no pitch movement\), opposed to the ‘‘positive’’ tones II \(falling\) and III \(rising\)\. With this null character of tone I he links the phenomenon he calls*déplacement*\(депласация\) — the tone\-conditioned transfer of the expiratory stress\[Polivanov,[1937b](https://arxiv.org/html/2607.28766#bib.bib32), p\. 53\]: in suffixal formations stress falls on the root by default, but if the root carries tone I it shifts to the suffix \(ма1‘hemp’ \+\-ди→\\toма\-ди́, butма2‘horse’ \+\-ди→\\toма́\-ди\)\. Since Cyrillic marks neither tone nor stress, pairs such asполи‘fled’ \(tone I, suffix stress\) andполи‘ran’ \(tone II, root stress\) are doubly indistinguishable in writing\. ### Invertibility examples and tag order\. The tonal disambiguation and round\-trip ofзадё: \(3\)задё→\\toзадё<vblex\><t1\>‘suck out’;задё<vblex\><t2\>‘fence off’;задё<vblex\><t3\>‘blow up’\(4\)задё<vblex\><t3\><pst\>↔\\leftrightarrowзадёли‘blew up’ The formжынмуis analysed asжын<n\><an\><t1\><pl\>, and the same tag string generates exactlyжынму: \(5\)жын\-мужын<n\><an\><t1\><pl\>‘people’\(6\)задё\-лизадё<vblex\><t3\><pst\>‘blew up’\(7\)самария\-дисамария<np\><gen\>‘of Samaria’ Tag order follows Apertium and GiellaLT practice: part of speech first, then the lexical tone feature, then inflectional features \(number, aspect\)\. This order is an engineering convention of the annotation, not a claim about the structure of the Dungan word; the lexically inherent feature \(tone\) sits closer to the stem than the inflectional ones\. ### The transducer as a graph\. Figure[2](https://arxiv.org/html/2607.28766#A4.F2)shows two fragments of the transducer\. The upper fragment illustrates nominal morphotactics: the stemжынtakes the automaton from the initial state to a state from which optional number and genitive slots lead to an accepting state; each arc is labelled with a ‘‘surface segment : upper\-level tag’’ pair, so the pathжынмуreads on the lower level as the surface form and on the upper level as the analysisжын<n\><an\><t1\><pl\>\. The empty transition \(ε\\varepsilon\) corresponds to the unmarked value \(singular; no genitive\)\. The lower fragment illustrates tonal disambiguation: one and the same surface stringзадёmaps to three analyses differing only in the tone tag — it is these parallel arcs that restore the distinction the orthography does not record\. Рис\. 2:Two transducer fragments\. Top: nominal morphotactics \(the pathжынму↔\\leftrightarrowжын<n\><an\><t1\><pl\>\); bottom: resolution of tonal homography of the stemзадё\(one surface spelling — three analyses differing only in the tone tag\)\. Arcs are labelled ‘‘lower level : upper level’’;ε\\varepsilonis the empty transition\. ## Приложение EA Survey of Dungan Language Resources Tables[9](https://arxiv.org/html/2607.28766#A5.T9)–[12](https://arxiv.org/html/2607.28766#A5.T12)give an annotated survey of 70 resources on Dungan and adjacent topics, grouped by type\. Availability:*open*= freely accessible electronic version;*part\.*= accessible with caveats \(registration, partial, cache only\);*print*= print only\. Titles are kept in their original language\. The Oslo Digital Archive of Dungan Studies \(ODADS, row 53\) hosts scans of most of the classic literature\. Таблица 9:Dungan resource survey, part 1: reference works and dictionaries\.Таблица 10:Dungan resource survey, part 2: primary grammatical literature \(first half\)\.Таблица 11:Dungan resource survey, part 3: primary grammatical literature \(second half\)\.Таблица 12:Dungan resource survey, part 4: corpora and archives, FST parallels, tools, background\.
Similar Articles
Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition
Dango is a 1.8B-parameter LLM trained strictly on Japanese (L1) then fine-tuned on English (L2) to study language transfer effects in second language acquisition. The model filters English contamination from the pretraining corpus and shows human-like L2 production patterns.
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
Researchers introduce MORPHOGEN, a multilingual benchmark testing LLMs’ ability to rewrite first-person sentences in the opposite gender while preserving meaning across French, Arabic, and Hindi.
YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models
YallaMorph is a large-scale benchmark for evaluating controlled Arabic morphological generation in large language models, revealing significant challenges with cliticized and morphologically rare forms.
5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
The paper presents 5-Dialects-BN, a manually annotated benchmark dataset for five Bangla dialects with aligned transliterations and translations, designed to improve evaluation of dialect-aware LLMs.
DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
This paper introduces DiaLLM, a framework for adapting LLMs to English dialects, revealing a gap between dialectal robustness (understanding) and generation (producing dialectal text), and showing that explicit variety-targeted alignment improves generation but not necessarily human preference.