Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies
Summary
This paper introduces UD_Czech-PDTC, a large and genre-diverse treebank for Czech in the Universal Dependencies framework, derived from the Prague Dependency Treebank-Consolidated. It describes the conversion process and differences between annotation schemes.
View Cached Full Text
Cached at: 06/24/26, 07:46 AM
# Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies Source: [https://arxiv.org/abs/2606.24337](https://arxiv.org/abs/2606.24337) [View PDF](https://arxiv.org/pdf/2606.24337) > Abstract:Czech has been part of Universal Dependencies since its first release in 2015\. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks\. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the "Prague Dependency Treebank\-Consolidated" \(PDT\-C\)\. In comparison to the original PDT, PDT\-C is more than twice as large, but it is also much more diverse in terms of genres and domains\. In this paper, we describe the conversion of the new resource to Universal Dependencies\. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories\. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion\. We argue that while PDT is less "universal" and more tightly bound to one language, its multi\-layer annotation is rich and provides all information needed for basic UD trees, and much more\. ## Submission history From: Milan Straka \[[view email](https://arxiv.org/show-email/897090c7/2606.24337)\] **\[v1\]**Tue, 23 Jun 2026 09:22:42 UTC \(988 KB\)
Similar Articles
Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme
We present the second consolidated version of the Prague Dependency Treebank, a 4-million-token manual multilingual annotation resource covering morphology, syntax, semantics, coreference, and discourse, along with compatible lexicons.
ThaiTrees: Thai Syntactic Dependency Trees Across Domains
ThaiTrees introduces a 342M-token automatically parsed corpus of Thai text across domains, using a reproducible pipeline under Universal Dependencies to enable syntactic research and analysis.
AthDGC: An Open Diachronic Greek Treebank with Indo-European Parallels
This paper introduces AthDGC, the first openly licensed dependency-parsed treebank of Greek spanning eight diachronic periods, with verse-level cross-alignment to four ancient Indo-European languages using NLP tools like Stanza, LaBSE, and multilingual-BERT.
A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text
This paper presents a reproducible pipeline for building Universal Dependencies-style parsing resources for Katharevousa Greek parliamentary text, including OCR reconstruction, LLM-assisted annotation, and evaluation of multiple parsers. The best model (XLM-R) achieves 0.8893 UPOS accuracy and 0.5162 LAS, significantly outperforming off-the-shelf baselines.
Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations
This paper introduces two new Czech corpora, Hlava Cor and Hlava AD, designed to study human label variation in coreference and discourse relations. The corpora feature multiple annotations and annotator explanations, achieving 60-65% inter-annotator agreement and revealing systematic differences in interpretation.