ThaiTrees:泰语跨领域句法依存树

arXiv cs.CL 论文

摘要

ThaiTrees引入了一个3.42亿词的自动解析泰语跨领域语料库,采用Universal Dependencies下的可复现流水线,以支持句法研究和分析。

arXiv:2609.27558v1 Announce Type: new Abstract: Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.
查看原文
查看缓存全文

缓存时间: 2026/09/24 09:24

# ThaiTrees: Thai Syntactic Dependency Trees Across Domains
Source: [https://arxiv.org/html/2609.27558](https://arxiv.org/html/2609.27558)
Papatchol ThientongAffiliation:Chulalongkorn University

###### Abstract

Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale\. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research\. We present ThaiTrees, a 342M\-token corpus drawn from news, Wikipedia, spoken transcripts, and social media\. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework\. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions\. We release a frequency lexicon and CoNLL\-U parses in machine\-readable formats suitable for both AI\-assisted and conventional programmatic analysis\.

## 1Introduction

Studying syntactic patterns in naturally occurring language requires a corpus large enough to yield stable frequency estimates and support statistical inference\. A treebank or other parsed corpus makes grammatical relations queryable, enabling analyses of syntactic distributions, valency patterns, and alternations\([Lehmann and Schneider, 2013](https://arxiv.org/html/2609.27558#bib.bib17)\)\. However, manual syntactic annotation \(or treebanking\) requires rare trained annotators and is costly to create and maintain at scale\([Marcus et al\., 1993](https://arxiv.org/html/2609.27558#bib.bib16)\), so they are mainly used for training automatic parsers\. The parser can be applied automatically to a much larger corpus\. Building such automatically parsed corpora is consequently important for extending quantitative syntactic research beyond the limited size of gold\-standard annotation\.

Dependency grammar and Universal Dependency provide a convenient representation for computational analysis because it encodes syntax directly as labeled head\-dependent relations between words\([de Marneffe et al\., 2021](https://arxiv.org/html/2609.27558#bib.bib11);[Nivre et al\., 2017](https://arxiv.org/html/2609.27558#bib.bib12)\)\. In a basic dependency tree, words are the nodes and head–dependent relations are the arcs; unlike phrase\-structure trees, no intermediate constituent nodes need be introduced\. The representation remains a rooted tree, but the fixed one\-node\-per\-word structure means that parsing connects existing lexical nodes directly\([Nivre, 2010](https://arxiv.org/html/2609.27558#bib.bib13)\)\. Automatically parsed dependency corpora can therefore be searched for grammatical collocations and used to examine constructional alternations such as the active–passive, verb–prepositional\-phrase, and dative alternations\([Uhrig et al\., 2018](https://arxiv.org/html/2609.27558#bib.bib14);[Lehmann and Schneider, 2013](https://arxiv.org/html/2609.27558#bib.bib17)\)\. At web scale, such corpora have also supported syntax\-based distributional models, open information extraction, and question answering\([Panchenko et al\., 2018](https://arxiv.org/html/2609.27558#bib.bib15)\)\.

Thai already has a manually annotated dependency treebank: Thai\-TUD contains over 3,600 sentences and provides a foundation for training and evaluating Thai parsers\([Sriwirote et al\., 2025](https://arxiv.org/html/2609.27558#bib.bib7)\)\. It does not, however, provide the large automatically parsed corpus needed for corpus\-scale syntactic analysis\. We therefore create ThaiTrees, a 342M\-token corpus drawn from news, Wikipedia, spoken transcripts, and social media\. Its sentences are parsed automatically under the Universal Dependencies framework\. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text, and release the frequency lexicon and CoNLL\-U parses in machine\-readable formats suitable both for AI\-assisted and conventional programmatic analysis\.111The corpus is available at[https://github\.com/nlp\-chula/thaitrees](https://github.com/nlp-chula/thaitrees)

## 2Related Work

Thai resources provide gold\-standard annotation at a scale suited to model development\. UD Thai\-TUD contains 3,627 manually annotated dependency trees \(77,215 tokens\) drawn from the Thai National Corpus and Thai Wikipedia\([Sriwirote et al\., 2025](https://arxiv.org/html/2609.27558#bib.bib7)\); we use a dependency parser trained on it\. ORCHID provides manually checked sentence boundaries, word segmentation, and POS tags for technical prose\([Charoenporn et al\., 1997](https://arxiv.org/html/2609.27558#bib.bib9)\); LST20 provides segmentation, POS, named entities, and clause and sentence boundaries for 3\.2 M words of news\([Boonkwan et al\., 2020](https://arxiv.org/html/2609.27558#bib.bib10)\)\. The Thai National Corpus supplies a general written reference corpus and frequency interface; the 2009 progress report documented 14 M processed words and described collection constraints from copyright clearance\([Aroonmanakun, 2007](https://arxiv.org/html/2609.27558#bib.bib8);[Aroonmanakun et al\., 2009](https://arxiv.org/html/2609.27558#bib.bib23)\)\. These resources are valuable for training and evaluating models, but their scale, single\-register coverage, or lack of dependency annotation limits the corpus\-linguistic claims they can support\. ThaiTrees therefore adds a much larger, four\-domain corpus for distributional analysis, while relying on Thai\-TUD as the gold\-standard source for parser training\.

Large corpora with automatic dependency parses already support corpus\-linguistic research in English and other languages\. Sketch Engine, for example, is designed to query dependency\-parsed corpora and derive grammatical\-relation summaries\([Kilgarriff et al\., 2014](https://arxiv.org/html/2609.27558#bib.bib18)\); its preloaded collections include English pukWaC, parsed with MaltParser\([Sketch Engine, 2026](https://arxiv.org/html/2609.27558#bib.bib19)\)\. The earlier WaCky web corpora provided large, automatically linguistically processed resources for English, German, and Italian\([Baroni et al\., 2009](https://arxiv.org/html/2609.27558#bib.bib20)\), and the related TenTen family extends this web\-corpus approach across languages\([Jakubíček et al\., 2013](https://arxiv.org/html/2609.27558#bib.bib21)\)\. Thai currently lacks a comparably broad corpus with automatically produced, queryable dependency analyses\.

Because Thai lacks explicit word and sentence boundaries, most processing stages rely on machine\-learning models tuned to Thai data\. AttaCut\([Chormai et al\., 2020](https://arxiv.org/html/2609.27558#bib.bib2)\)performs word segmentation with a convolutional neural network trained on annotated Thai syllable/word boundaries\. PyThaiNLP and its CRFcut sentence segmenter\([Phatthiyaphaibun et al\., 2023](https://arxiv.org/html/2609.27558#bib.bib1);[Chumpolsathien, 2020](https://arxiv.org/html/2609.27558#bib.bib3)\)use CRF\-backed sentence\-boundary models \(with internal newmm features\) benchmarked on corpora such as ORCHID and TED transcripts\. AttaParse 1\.0\([Sriwirote et al\., 2025](https://arxiv.org/html/2609.27558#bib.bib7)\)wraps Stanza’s graph\-based UD parser\([Qi et al\., 2020](https://arxiv.org/html/2609.27558#bib.bib4)\)and parses pre\-tokenized input using models trained on UD Thai\-TUD\. A PhayaThaiBERT transformer, fine\-tuned on UD Thai\-TUD, supplies UPOS tags that overwrite the parser’s column\([Sriwirote et al\., 2024](https://arxiv.org/html/2609.27558#bib.bib6)\)\. These ML\-based tools perform decently well on in\-domain material, and together show that Thai has a wide array of automatic linguistic annotation tools available for use\.

## 3ThaiTrees: Corpus Construction

### 3\.1Data Sources

Table 1:The four sub\-corpora that make up ThaiTrees\.The four domains are journalistic prose \(news\), encyclopedic prose \(Wikipedia\), conversational spoken language \(YouTube podcasts, livestreams, and film transcripts\), and informal written conversation \(online forum and sentiment\-dataset posts\)\. The Spoken sub\-corpus has fewer documents but each is roughly an order of magnitude longer, because each document is the transcript of one long recording \(Table[1](https://arxiv.org/html/2609.27558#S3.T1)\)\.

The news sub\-corpus consists of over 100,000 articles from the ThaiPBS public\-broadcaster website\. We only include articles that are made available publicly on the website in December 2025\.

Thai Wikipedia articles were sampled via MediaWiki’s random\-article endpoint in batches of 500, namespace 0, with disambiguation pages excluded and reference sections truncated\. Sampling was done in December 2025\.

Spoken\-language transcripts were retrieved from YouTube through the youtube\-transcript\.io API\. We manually select the channels that include manual transcription so that we get the highest quality transcription without using automatic speech recognition\. The sources are the Thai PBS Podcast network, independent podcasts \(bigboung, BeSider\), film subtitles, Prachatai political\-commentary livestreams, and SaltymanTH livestreams\.

The social\-media sub\-corpus combines the Wisesight sentiment dataset and Pantip forum used in training large language models such as WangchanBERTa\([Lowphansirikul et al\., 2021](https://arxiv.org/html/2609.27558#bib.bib5)\)and PhayaThaiBERT\([Sriwirote et al\., 2024](https://arxiv.org/html/2609.27558#bib.bib6)\)\.

### 3\.2Processing Pipeline

raw ThaitextCRFcutsentencesegmentationAttaCuttokenization\(per sentence\)AttaParse 1\.0dependencyparsingPhayaThaiBERTPOSoverwriteCoNLL\-UAttaCuttokenization\(whole document\)frequencylexiconlexicon branchparsing branchFigure 1:The processing pipeline\. The raw text runs through two branches: the lexicon branch \(top\) runs AttaCut over each whole document, producing the frequency lexicon; the parsing branch \(bottom\) runs CRFcut for sentence segmentation, re\-applies AttaCut to each sentence, parses with AttaParse 1\.0 \(tokenize\_pretokenized=True\), and overwrites the POS column with PhayaThaiBERT\.The lexicon and parsed corpus are produced in separate branches that apply word segmentation to different units of text: whole documents for the lexicon and individual sentences for parsing \(Figure[1](https://arxiv.org/html/2609.27558#S3.F1)\)\. The resulting token boundaries can differ, so the two releases share document identifiers but not token identifiers\. For each processing step that requires an machine\-learning\-based tool, we select the tools that achieve state\-of\-the\-art results on some Thai benchmark data except for Thai POS tagger \(Table[3](https://arxiv.org/html/2609.27558#S3.T3)\)\.

#### 3\.2\.1Word segmentation

We use a CNN\-based state\-of\-the\-art Thai word segmenter, AttaCut\([Chormai et al\., 2020](https://arxiv.org/html/2609.27558#bib.bib2)\), to segment each whole document for the lexicon branch\. We store its output in Parquet, replacing thetextcolumn with pipe\-delimited tokens and adding atoken\_countcolumn\. Lexicon construction then removes punctuation, pure\-numeric tokens, emoji, whitespace\-only tokens, and forms whose total corpus frequency is at most five\. This filtering is applied to the document\-level AttaCut output, independently of the parsing branch\.

#### 3\.2\.2Sentence segmentation and tokenization

The parsing branch reads the raw text directly\. It collapses newlines and runs of spaces to a single space and caps social\-media documents at 500,000 characters before sentence segmentation; this truncates 149 long, concatenated documents\. PyThaiNLP’s CRFcut\([Phatthiyaphaibun et al\., 2023](https://arxiv.org/html/2609.27558#bib.bib1);[Chumpolsathien, 2020](https://arxiv.org/html/2609.27558#bib.bib3)\)then returns sentence strings\. CRFcut uses dictionary\-based word segmentation internally to extract CRF features but discards those token boundaries when returning the strings, but sentence segmentation itself does not remove text from its normalized input\. This separate branch avoids carrying CRFcut’s internal token boundaries into the downstream word\-segmentation step\.

#### 3\.2\.3Dependency parsing

We parse with a graph\-based neural dependency parser: AttaParse’s Thai\-specific Stanza model uses PhayaThaiBERT contextual representations and the no\-POS configuration identified in Thai\-TUD\([Qi et al\., 2020](https://arxiv.org/html/2609.27558#bib.bib4);[Sriwirote et al\., 2025](https://arxiv.org/html/2609.27558#bib.bib7)\)\. For each CRFcut\-delimited sentence, AttaParse receives the AttaCut tokens as pretokenized input and predicts labeled head–dependent arcs while preserving those token boundaries\. Crucially, it requires neither POS tags nor lemmas: the POS and lemma processors are disabled,UPOSandXPOSare set to\., andFEATSto\_before parsing\. The releasedXPOScolumn consequently remains\.\.

Whitespace\-only tokens are discarded at the parser’s pretokenized entry point\. Parsing output is written as 5 M\-token CoNLL\-U chunks on document boundaries; completed chunks are skipped and writes are atomic to allow resumption\. One deterministic parser failure on Wikipedia chunk 16 is excluded from the final release\.

#### 3\.2\.4POS tagging and benchmark performance

To our knowledge, no off\-the\-shelf Thai Universal POS \(UPOS\) tagger has been benchmarked on Thai\-TUD with a directly comparable held\-out evaluation\. We therefore train our own tagger on Thai\-TUD\([Sriwirote et al\., 2025](https://arxiv.org/html/2609.27558#bib.bib7)\), fine\-tuning a state\-of\-the\-art encoder\-only model for Thai, called PhayaThaiBERT\([Sriwirote et al\., 2024](https://arxiv.org/html/2609.27558#bib.bib6)\), for four epochs with AdamW and a learning rate of5×10−55\\times 10^\{\-5\}\. The tagger achieves 90\.64% held\-out accuracy and 81\.34% macro F1 across 15 UPOS classes \(Table[2](https://arxiv.org/html/2609.27558#S3.T2)\)\. Only the first sub\-word prediction for each word is retained for the prediction head\.

Table 2:Fine\-tuning configuration for the PhayaThaiBERT POS tagger \(phayathaibert\-thai\-pos\-tagger\)\. The best checkpoint was selected by held\-out accuracy; macro F1 is over the 15 UPOS classes present in the treebank\.The predictions overwrite column 4 \(UPOS\) of the AttaParse CoNLL\-U output\. Since the dependency parser does not use POS features, this changes no edges: tags and edges remain independent annotation layers over the same tokens\.

Table 3:In\-domain benchmarks for the four pipeline stages\.†\\daggerAttaParse 1\.0 = the no\-POS graph\-based PhayaThaiBERT configuration \(row GP\) of[Sriwirote et al\. \(2025\)](https://arxiv.org/html/2609.27558#bib.bib7), Table 3\.

### 3\.3Released Artifacts

Computing syntactic frequency at this scale produced three artifacts, which we release together so the analyses in §[4](https://arxiv.org/html/2609.27558#S4.T4)and §[5](https://arxiv.org/html/2609.27558#S5)can be reproduced\. Document identifiers are shared across all three, and sentence identifiers across the parsed corpus, so results can be joined at the document level\. Token identifiers do not align across the lexicon and the parsed corpus, which index different tokenizations of the same raw text \(§[3\.2](https://arxiv.org/html/2609.27558#S3.SS2)\)\.

The raw text corpus contains 366,120 documents and 341,967,133 AttaCut\-segmented tokens in Parquet format\. Its schema is\{doc\_id, domain, text, token\_count\}, wheretextuses pipe delimiters to mark token boundaries\. Per\-domain files and a sentence\-segmented variant \(4,519,775 sentences\) are included\.

The frequency lexicon contains 451,252 unique word forms and 255,434,917 filtered tokens, released as Parquet and SQLite\. Each row carries 34 columns\. A five\-column core holds the form, its total count, total rank, frequency per million, and cross\-corpus document count\. Six columns per domain give the count, rank, freq/M, document count, document frequency, and IDF\. Five further columns record how many domains the form appears in, a cross\-domain flag, character length, IDF, and the dominant domain\. This artifact covers the full raw corpus through the document\-level lexicon branch\. 39,465 word forms \(8\.7%\) appear in all four domains\.

The dependency\-parsed corpus contains 199,836,464 dependency edges in 10\-column CoNLL\-U, chunked into 5 M\-token archives\. A per\-pattern edge frequency table is released alongside; it lists the 5,944 distinct patterns in the corpus \(head POS, relation, dependent POS\) with per\-domain counts\. The parsed corpus contains 203,892,200 tokens\.

## 4Word Frequency

NewsWikipediaSpokenSocial Media1\.2\.3\.4\.5\.6\.7\.8\.9\.10\.11\.12\.13\.14\.15\.16\.17\.18\.19\.20\.232221222921

Figure 2:Top 20 words by freq/M in each domain \(union of all four top\-20 lists\)\. Lines trace each word’s rank across domains and end where the word falls outside a domain’s top 20; when a word re\-enters later, a dashed line bridges the gap, with its true rank in the skipped domain given in gray\.Table 4:Top 20 words across the corpus by overall frequency per million\.### 4\.1Methodology

Word frequency is measured in occurrences per million tokens \(freq/M\) against the raw AttaCut token total of each domain \(Table[1](https://arxiv.org/html/2609.27558#S3.T1)\)\. The released lexicon uses the same denominator, so any freq/M in Table[4](https://arxiv.org/html/2609.27558#S4.T4)can be recomputed from the lexicon’s counts\.

To identify the distinctive vocabulary of each domain we use the odds ratio, the effect\-size keyness statistic\([Pojanapunya and Watson Todd, 2018](https://arxiv.org/html/2609.27558#bib.bib22)\)\. Log\-likelihood ranks words that are frequent in the target domain even when they are frequent everywhere; the odds ratio ranks words whose frequency differs most between the target and the reference\. We want the second, so we take the odds ratio and report its natural logarithm for symmetry around zero\.

For each target domainTTand a reference corpusRRformed by the union of the other three domains, the log odds ratio of wordwwis

log⁡OR⁡\(w\)=ln⁡\(a⋅db⋅c\),\\log\\mathrm\{OR\}\(w\)\\;=\\;\\ln\\\!\\left\(\\frac\{a\\cdot d\}\{b\\cdot c\}\\right\),\(1\)where, writingNTN\_\{T\}andNRN\_\{R\}for the total token counts ofTTandRR,

a\\displaystyle a=count of​w​in​T,\\displaystyle=\\text\{count of \}w\\text\{ in \}T,b\\displaystyle b=count of​w​in​R,\\displaystyle=\\text\{count of \}w\\text\{ in \}R,c\\displaystyle c=NT−a\(all other tokens inT\),\\displaystyle=N\_\{T\}\-a\\quad\(\\text\{all other tokens in \}T\),d\\displaystyle d=NR−b\(all other tokens inR\)\.\\displaystyle=N\_\{R\}\-b\\quad\(\\text\{all other tokens in \}R\)\.
Table 5:Top 20 words of each domain by log odds ratio against the union of the other three domains, among words occurring at least 10,000 times in both the target and reference and in more than ten documents of the target domain\.NOUNNOUNteamnationnmod

NOUN\-nmod\-NOUN “national team”

VERBNOUNreadnewsobj

VERB\-obj\-NOUN “read the news”

VERBVERBgetreceivecompound

VERB\-compound\-VERB “to receive”

NOUNVERB\-erperformacl

NOUN\-acl\-VERB “actor”

ADPNOUNinyearcase

NOUN\-case\-ADP “in the year”

Figure 3:The five most frequent dependency\-edge patterns, each shown with a real example from the corpus\. The edge points from head to dependent and is labeled with the Universal Dependencies relation; the pattern is writtenhead\-relation\-dependent\. Sources, left to right:news\_507:506,news\_5052:302,news\_5028:1,news\_5052:289,news\_5034:94\.NewsWikipediaSpokenSocial Media1\.2\.3\.4\.5\.6\.7\.8\.9\.10\.11\.12\.13\.14\.15\.16\.17\.18\.19\.20\.VERB\-obj\-NOUNVERB\-obj\-NOUNVERB\-obj\-NOUNVERB\-obj\-NOUNNOUN\-nmod\-NOUNNOUN\-nmod\-NOUNNOUN\-nmod\-NOUNNOUN\-nmod\-NOUNNOUN\-acl\-VERBNOUN\-acl\-VERBNOUN\-acl\-VERBNOUN\-acl\-VERBVERB\-compound\-VERBVERB\-compound\-VERBVERB\-compound\-VERBVERB\-compound\-VERBNOUN\-case\-ADPNOUN\-case\-ADPNOUN\-case\-ADPNOUN\-case\-ADPVERB\-obl\-NOUNVERB\-obl\-NOUNVERB\-obl\-NOUNVERB\-obl\-NOUNVERB\-nsubj\-NOUNVERB\-nsubj\-NOUNVERB\-nsubj\-NOUNVERB\-nsubj\-NOUNVERB\-mark\-SCONJVERB\-mark\-SCONJVERB\-mark\-SCONJVERB\-mark\-SCONJVERB\-aux\-AUXVERB\-aux\-AUXVERB\-aux\-AUXVERB\-aux\-AUXVERB\-advmod\-ADVVERB\-advmod\-ADVVERB\-advmod\-ADVVERB\-advmod\-ADVNOUN\-nummod\-NUMNOUN\-nummod\-NUMVERB\-advcl\-VERBVERB\-advcl\-VERBVERB\-advcl\-VERBVERB\-advcl\-VERBVERB\-ccomp\-VERBVERB\-ccomp\-VERBVERB\-ccomp\-VERBVERB\-conj\-VERBVERB\-conj\-VERBVERB\-conj\-VERBVERB\-conj\-VERBNOUN\-compound\-VERBNOUN\-nmod\-PROPNNOUN\-nmod\-PROPNVERB\-cc\-CCONJVERB\-cc\-CCONJVERB\-cc\-CCONJNOUN\-compound\-NOUNNOUN\-compound\-NOUNNOUN\-compound\-NOUNNOUN\-amod\-ADJNOUN\-amod\-ADJNOUN\-amod\-ADJVERB\-nsubj\-PRONVERB\-nsubj\-PRONVERB\-nsubj\-PRONVERB\-nsubj\-PRONNOUN\-det\-DETNOUN\-det\-DETNOUN\-conj\-NOUNNOUN\-punct\-PUNCTVERB\-compound\-ADVVERB\-compound\-ADVPROPN\-punct\-PUNCTVERB\-punct\-PUNCTVERB\-advmod\-PARTVERB\-advmod\-PARTPART\-fixed\-PART312623

Figure 4:Top 20 dependency\-edge patterns by freq/M in each domain \(union of all four top\-20 lists\)\. Lines trace each pattern’s rank across domains and end where it falls outside a domain’s top 20; when a pattern re\-enters later, a dashed line bridges the gap, with its true rank in the skipped domain in gray\.Table 6:Top 20 dependency\-edge patterns across the corpus by overall frequency per million\. Patterns are writtenhead\-relation\-dependent\.Table 7:Top 20 dependency\-edge patterns of each domain by log odds ratio against the union of the other three domains\. Patterns are writtenhead\-relation\-dependent\(POS of head, dependency relation, POS of dependent\); only patterns with at least 10,000 edges in both the target and reference domains are considered\.Words withlog⁡OR\>0\\log\\mathrm\{OR\}\>0are concentrated in the target relative to the reference; words withlog⁡OR<0\\log\\mathrm\{OR\}<0are under\-represented in the target\. Because the formula contains “everything else” terms on both sides \(ccanddd\), corpus\-size differences cancel and no further normalization is required\.

Two cut\-offs constrain which words are scored\. We keep only words with at least 10,000 occurrences in both the target and the reference: the odds ratio inflates for rare words, and without a floor they would dominate the top of the ranking\. Each word must also appear in more than ten documents of the target domain, so that no keyword comes from a single prolific source\.[Pojanapunya and Watson Todd \(2018\)](https://arxiv.org/html/2609.27558#bib.bib22)recommend a minimum\-frequency threshold and offer a minimum number of texts as an alternative; we apply both\. We additionally normalize whitespace within word forms and exclude single\-character tokens, tokens containing ASCII punctuation \(abbreviations such as \.\. and \.\. are not treated as lexical words here\), URL fragments, tokens consisting entirely of non\-Thai, non\-Latin script, and pure\-Latin tokens of fewer than three characters\. We report the top 20 words per domain, ranked bylog⁡OR\\log\\mathrm\{OR\}descending\.

These exclusions leave 354,430 scored word forms over 247,324,621 tokens \(news 39,068,831; Wikipedia 66,405,646; spoken 25,912,233; social media 115,937,911\)\. These are the totalsNTN\_\{T\}andNRN\_\{R\}of Eq\.[1](https://arxiv.org/html/2609.27558#S4.E1)\. The keyness rankings and cross\-domain rank comparisons use these filtered forms\. The released 451,252\-form lexicon uses a less restrictive filter: the repetition mark , for example, is among its twenty most frequent forms \(Table[4](https://arxiv.org/html/2609.27558#S4.T4)\) but is excluded from keyness scoring\.

### 4\.2Results and Discussion

The corpus\-wide word distribution follows the familiar Zipfian pattern: a small number of forms account for a large share of all tokens\. Accordingly, the top 20 words are predominantly function words, as is typical of frequency analysis in a large corpus \(Table[4](https://arxiv.org/html/2609.27558#S4.T4)\)\. Five high\-frequency verbal forms also have function\-like uses: can mark ability, and can express deictic or directional meanings, is a copula, and can function like a preposition\. The formally nominal forms and are likewise highly productive elements that combine with a wide range of verbal material\. Thus, the upper end of the frequency list is dominated not simply by lexical categories, but by forms with broad grammatical and combinatory roles\. The same high\-frequency forms recur in all four domains, but their rankings differ just slightly \(Figure[2](https://arxiv.org/html/2609.27558#S4.F2)\)\.

The keyness results make the domain\-specific vocabulary explicit \(Table[5](https://arxiv.org/html/2609.27558#S4.T5)\)\. News is characterized by law, government, and incident\-reporting vocabulary, including “police”, “case”, and “legal section”, alongside attribution verbs such as “announce” and “confirm”\. Wikipedia favors encyclopedic dates and names, with month names and sport or entertainment terms such as “football” and “album”\. Spoken language is marked by conversational particles and informal pronouns \(, , \), while social media is associated with finance and commerce \( “stocks”, “bank”, “buy”\) as well as buyer–seller politeness \(, “thank you”\)\.

## 5Syntactic Frequency

### 5\.1Methodology

We treat each dependency edge as an instance of a pattern: the triplet⟨head​\_​POS,relation,dep​\_​POS⟩\\langle\\mathrm\{head\\\_POS\},\\mathrm\{relation\},\\mathrm\{dep\\\_POS\}\\rangleof the head’s part\-of\-speech tag, the Universal Dependencies relation, and the dependent’s part\-of\-speech tag\. That means we do not count the subtrees\. We count the edges along with the POS tags from the heads and the dependents of the edges\.

We apply the same keyness calculation method to dependency edges\. The counted unit is now a syntactic pattern or a subtree instead of a word form:aaandbbare the pattern’s counts in the target domainTTand the referenceRR, and the totalsET,ERE\_\{T\},E\_\{R\}are edge counts in place of the token countsNT,NRN\_\{T\},N\_\{R\}\. As in the lexical analysis we keep only patterns with at least 10,000 edges in bothTTandRR\. Unlike word keyness calculation, we set no minimum number of documents here\.

### 5\.2Results and Discussion

The five most frequent patterns include noun modification, verb–object relations, and compounding \(Figure[3](https://arxiv.org/html/2609.27558#S4.F3)\)\. The illustrated word pairs are corpus instances selected from the most frequent pairs for each pattern\.

The 20 most frequent patterns are all well\-formed and familiar Thai constructions \(Table[6](https://arxiv.org/html/2609.27558#S4.T6)\)\. This does not make every individual automatic parse correct—the parser’s attachment accuracy is below 90%—but it provides a useful check that its most common output reflects ordinary Thai grammar, supporting aggregate downstream analyses\. Twelve of the patterns are headed by verbs and eight by nouns\. The leading noun pattern,NOUN\-nmod\-NOUN, reflects the pervasive modification and compounding of nominal expressions; the leading verbal pattern,VERB\-obj\-NOUN, is the expected structure of a verb with a nominal object\.

Two other high\-ranking patterns are particularly revealing\. The third\-rankedVERB\-compound\-VERBpattern represents serial\-verb constructions, whose 9\.7M instances show that they are one of the more common and central grammtical constructions in the Thai language\. Fourth\-rankedNOUN\-acl\-VERBcovers a verb modifying a noun, including relative clauses and nominalized clauses\. Given the high corpus frequency of the nominalizers and , many instances likely involve nominalization rather than relative clauses

The remaining frequent patterns describe similarly expected verbal and nominal structure: adverbial modification \(VERB\-advmod\-ADV\), oblique nominal dependents \(VERB\-obl\-NOUN\), auxiliaries \(VERB\-aux\-AUX\), and subordinate clauses \(VERB\-advcl\-VERB,VERB\-mark\-SCONJ\)\. Taken together, the top\-20 list presents common grammatical constructions of Thai despite imperfect automatic parsing\.

Looking at the top 20 syntactic patterns in each of the four domains, we see that news and Wikipedia form one cluster, while spoken transcription and social media form another \(Figure[4](https://arxiv.org/html/2609.27558#S4.F4)\)\. News and Wikipedia show similar rankings of dependency edges\. In spoken transcription,VERB\-advmod\-ADV,VERB\-nsubj\-PRON, andVERB\-aux\-AUXrank higher\. Adverbial modification covers many syntactic phenomena\. Further inspection of the data shows thatVERB\-advmod\-ADVoften involves , a general\-purpose discourse connective that can signal several discourse relations\([Prasertsom et al\., 2024](https://arxiv.org/html/2609.27558#bib.bib24)\)\. Speakers may use it more often to make connections between sentences clear and maintain coherence despite the natural disfluencies of speech\. Pronouns and auxiliary verbs are also more common in spoken language than in news and Wikipedia\. This may reflect the personal and interactive nature of speech, as well as the tendency of news and Wikipedia to present information directly and with certainty\.

## 6Conclusion

We built ThaiTrees, a 342 M\-token corpus of Thai across four domains, to measure how word and syntactic frequency vary between them\. By word frequency and syntactic frequency, the domains split into a formal pair \(news, Wikipedia\) and a conversational pair \(spoken, social media\)\. The corpus is available in machine\-readable formats, with a frequency lexicon\.

## References

- Aroonmanakunet al\.\(2009\)W\. Aroonmanakun, K\. Tansiri, and P\. NittayanuparpThai National Corpus: a progress report\.InProceedings of the 7th Workshop on Asian Language Resources,Suntec, Singapore,pp\. 153–160\.External Links:[Link](https://aclanthology.org/W09-3422/)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p1.1)\.
- Aroonmanakun \(2007\)W\. AroonmanakunCreating the Thai National Corpus\.Manusya: Journal of Humanities10\(3\),pp\. 4–17\.Note:Special Issue 13\.Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p1.1)\.
- Baroniet al\.\(2009\)M\. Baroni, S\. Bernardini, A\. Ferraresi, and E\. ZanchettaThe WaCky wide web: a collection of very large linguistically processed web\-crawled corpora\.Language Resources and Evaluation43\(3\),pp\. 209–226\.External Links:[Document](https://dx.doi.org/10.1007/s10579-009-9081-4),[Link](https://doi.org/10.1007/s10579-009-9081-4)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p2.1)\.
- Boonkwanet al\.\(2020\)P\. Boonkwan, V\. Luantangsrisuk, S\. Phaholphinyo, K\. Kriengket, D\. Leenoi, C\. Phrombut, M\. Boriboon, K\. Kosawat, and T\. SupnithiThe annotation guideline of LST20 corpus\.Note:arXiv:2008\.05055NECTEC Technical Report\.External Links:[Link](https://arxiv.org/abs/2008.05055)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p1.1)\.
- Charoenpornet al\.\(1997\)T\. Charoenporn, V\. Sornlertlamvanich, and H\. IsaharaBuilding a large Thai text corpus — Part\-of\-speech tagged corpus: ORCHID\.InProceedings of the Natural Language Processing Pacific Rim Symposium 1997 \(NLPRS ’97\),pp\. 509–512\.Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p1.1)\.
- Chormaiet al\.\(2020\)P\. Chormai, P\. Prasertsom, J\. Cheevaprawatdomrong, and A\. RutherfordSyllable\-based neural Thai word segmentation\.InProceedings of the 28th International Conference on Computational Linguistics,D\. Scott, N\. Bel, and C\. Zong \(Eds\.\),Barcelona, Spain \(Online\),pp\. 4619–4637\.Note:Peer\-reviewed publication of what was earlier preprinted as “AttaCut” \(arXiv:1911\.07056\); we refer to the tool by the AttaCut name throughout the paper\.External Links:[Link](https://aclanthology.org/2020.coling-main.407/),[Document](https://dx.doi.org/10.18653/v1/2020.coling-main.407)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.2\.1](https://arxiv.org/html/2609.27558#S3.SS2.SSS1.p1.1),[Table 3](https://arxiv.org/html/2609.27558#S3.T3.2.2.4)\.
- Chumpolsathien \(2020\)N\. ChumpolsathienCRFcut: Thai sentence segmentation with conditional random fields\.Note:Software,[https://github\.com/vistec\-AI/crfcut](https://github.com/vistec-AI/crfcut)Default sentence\-segmentation engine in PyThaiNLP\. No separate paper; cite as software\.Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.2\.2](https://arxiv.org/html/2609.27558#S3.SS2.SSS2.p1.1),[Table 3](https://arxiv.org/html/2609.27558#S3.T3.2.3.4)\.
- de Marneffeet al\.\(2021\)M\. de Marneffe, C\. D\. Manning, J\. Nivre, and D\. ZemanUniversal dependencies\.Computational Linguistics47\(2\),pp\. 255–308\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00402),[Link](https://aclanthology.org/2021.cl-2.11/)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.
- Jakubíčeket al\.\(2013\)M\. Jakubíček, A\. Kilgarriff, V\. Kovář, P\. Rychlý, and V\. SuchomelThe TenTen corpus family\.InProceedings of the 7th International Corpus Linguistics Conference,Lancaster, UK,pp\. 125–127\.External Links:[Link](https://ucrel.lancs.ac.uk/cl2013/doc/CL2013-ABSTRACT-BOOK.pdf)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p2.1)\.
- Kilgarriffet al\.\(2014\)A\. Kilgarriff, V\. Baisa, J\. Bušta, M\. Jakubíček, V\. Kovář, J\. Michelfeit, P\. Rychlý, and V\. SuchomelThe Sketch Engine: ten years on\.Lexicography1,pp\. 7–36\.External Links:[Document](https://dx.doi.org/10.1007/s40607-014-0009-9),[Link](https://www.sketchengine.eu/wp-content/uploads/The_Sketch_Engine_2014.pdf)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p2.1)\.
- Lehmann and Schneider \(2013\)H\. M\. Lehmann and G\. SchneiderBNC dependency bank 1\.0\.InAspects of Corpus Linguistics: Compilation, Annotation, Analysis,S\. Hoffmann, P\. Rayson, and G\. Leech \(Eds\.\),External Links:[Link](https://varieng.helsinki.fi/series/volumes/12/lehmann_schneider/)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p1.1),[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.
- Lowphansirikulet al\.\(2021\)L\. Lowphansirikul, C\. Polpanumas, N\. Jantrakulchai, and S\. NutanongWangchanBERTa: pretraining transformer\-based Thai language models\.External Links:2101\.09635,[Link](https://arxiv.org/abs/2101.09635)Cited by:[§3\.1](https://arxiv.org/html/2609.27558#S3.SS1.p5.1)\.
- Marcuset al\.\(1993\)M\. P\. Marcus, B\. Santorini, and M\. A\. MarcinkiewiczBuilding a large annotated corpus of English: the Penn treebank\.Computational Linguistics19\(2\),pp\. 313–330\.External Links:[Link](https://catalog.ldc.upenn.edu/docs/LDC95T7/cl93.html)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p1.1)\.
- Nivreet al\.\(2017\)J\. Nivre, D\. Zeman, F\. Ginter, and F\. TyersUniversal dependencies\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts,Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.
- Nivre \(2010\)J\. NivreDependency parsing\.Language and Linguistics Compass4\(3\),pp\. 177–191\.External Links:[Document](https://dx.doi.org/10.1111/j.1749-818X.2010.00187.x),[Link](https://doi.org/10.1111/j.1749-818X.2010.00187.x)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.
- Panchenkoet al\.\(2018\)A\. Panchenko, E\. Ruppert, S\. Faralli, S\. P\. Ponzetto, and C\. BiemannBuilding a web\-scale dependency\-parsed corpus from CommonCrawl\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\),Miyazaki, Japan\.External Links:[Link](https://aclanthology.org/L18-1286/)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.
- Phatthiyaphaibunet al\.\(2023\)W\. Phatthiyaphaibun, K\. Chaovavanich, C\. Polpanumas, A\. Suriyawongkul, L\. Lowphansirikul, P\. Chormai, P\. Limkonchotiwat, T\. Suntorntip, and C\. UdomcharoenchaikitPyThaiNLP: Thai natural language processing in Python\.InProceedings of the 3rd Workshop for Natural Language Processing Open Source Software \(NLP\-OSS 2023\),L\. Tan, D\. Milajevs, G\. Chauhan, J\. Gwinnup, and E\. Rippeth \(Eds\.\),Singapore,pp\. 25–36\.External Links:[Link](https://aclanthology.org/2023.nlposs-1.4/),[Document](https://dx.doi.org/10.18653/v1/2023.nlposs-1.4)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.2\.2](https://arxiv.org/html/2609.27558#S3.SS2.SSS2.p1.1)\.
- Pojanapunya and Watson Todd \(2018\)P\. Pojanapunya and R\. Watson ToddLog\-likelihood and odds ratio: Keyness statistics for different purposes of keyword analysis\.Corpus Linguistics and Linguistic Theory\.External Links:[Document](https://dx.doi.org/10.1515/cllt-2015-0030)Cited by:[§4\.1](https://arxiv.org/html/2609.27558#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.27558#S4.SS1.p5.1)\.
- Prasertsomet al\.\(2024\)P\. Prasertsom, A\. Jaroonpol, and A\. T\. RutherfordThe Thai discourse treebank: annotating and classifying Thai discourse connectives\.Transactions of the Association for Computational Linguistics12,pp\. 613–629\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00650),[Link](https://aclanthology.org/2024.tacl-1.34/)Cited by:[§5\.2](https://arxiv.org/html/2609.27558#S5.SS2.p5.1)\.
- Qiet al\.\(2020\)P\. Qi, Y\. Zhang, Y\. Zhang, J\. Bolton, and C\. D\. ManningStanza: A Python natural language processing toolkit for many human languages\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations,Online,pp\. 101–108\.External Links:[Link](https://aclanthology.org/2020.acl-demos.14/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-demos.14)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.2\.3](https://arxiv.org/html/2609.27558#S3.SS2.SSS3.p1.1)\.
- Sketch Engine \(2026\)Sketch EngineList of corpora\.Note:Online corpus catalogueAccessed 22 September 2026\.External Links:[Link](https://www.sketchengine.eu/corpora-and-languages/corpus-list/)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p2.1)\.
- Sriwiroteet al\.\(2025\)P\. Sriwirote, W\. Q\. Leong, C\. Polpanumas, S\. Thanyawong, W\. C\. Tjhi, W\. Aroonmanakun, and A\. T\. RutherfordThe Thai Universal Dependency treebank\.Transactions of the Association for Computational Linguistics13,pp\. 376–391\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00745),[Link](https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00745/128939/)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p3.1),[§2](https://arxiv.org/html/2609.27558#S2.p1.1),[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.2\.3](https://arxiv.org/html/2609.27558#S3.SS2.SSS3.p1.1),[§3\.2\.4](https://arxiv.org/html/2609.27558#S3.SS2.SSS4.p1.1),[Table 3](https://arxiv.org/html/2609.27558#S3.T3),[Table 3](https://arxiv.org/html/2609.27558#S3.T3.2.4.4)\.
- Sriwiroteet al\.\(2024\)P\. Sriwirote, J\. Thapiang, V\. Timtong, and A\. T\. RutherfordPhayaThaiBERT: enhancing a pretrained Thai language model with unassimilated loanwords\.ACM Transactions on Asian and Low\-Resource Language Information Processing\.Note:Preprint: arXiv:2311\.12475\.External Links:[Link](https://arxiv.org/abs/2311.12475)Cited by:[§2](https://arxiv.org/html/2609.27558#S2.p3.1),[§3\.1](https://arxiv.org/html/2609.27558#S3.SS1.p5.1),[§3\.2\.4](https://arxiv.org/html/2609.27558#S3.SS2.SSS4.p1.1)\.
- Uhriget al\.\(2018\)P\. Uhrig, S\. Evert, and T\. ProislCollocation candidate extraction from dependency\-annotated corpora: exploring differences across parsers and dependency annotation schemes\.InLexical Collocation Analysis: Advances and Applications,P\. Cantos\-Gómez and M\. Almela\-Sánchez \(Eds\.\),pp\. 111–140\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-92582-0%5F6),[Link](https://doi.org/10.1007/978-3-319-92582-0_6)Cited by:[§1](https://arxiv.org/html/2609.27558#S1.p2.1)\.

相似文章