Tag
ThaiTrees introduces a 342M-token automatically parsed corpus of Thai text across domains, using a reproducible pipeline under Universal Dependencies to enable syntactic research and analysis.
This paper investigates why cross-direction pairing in bidirectional LSTMs underperforms same-direction pairing for dependency relation-type classification, using frozen-trunk diagnostics to analyze representational redundancy and directional information decay.
This paper introduces SiPE, a lightweight method that injects syntactic priors from dependency parses into transformer positional embeddings, improving syntactic generalization (up to 10.3% on SyntaxGym) and language understanding (up to 8.2% on GLUE) without increasing inference cost.
This paper investigates whether dependency parsing of non-human primate vocalizations or gestures can be evaluated without a gold standard. Using network science, the authors show that the proportion of correct edges retrieved by a parser is necessarily high due to the fast decay of sequence length distributions in non-human primates, making evaluation feasible, unlike for human language.
We present the second consolidated version of the Prague Dependency Treebank, a 4-million-token manual multilingual annotation resource covering morphology, syntax, semantics, coreference, and discourse, along with compatible lexicons.
This paper introduces AthDGC, the first openly licensed dependency-parsed treebank of Greek spanning eight diachronic periods, with verse-level cross-alignment to four ancient Indo-European languages using NLP tools like Stanza, LaBSE, and multilingual-BERT.
AfriSUD is a new dependency treebank collection for African languages, following the Surface-Syntactic Universal Dependencies (SUD) framework, designed to evaluate NLP models on languages like Naija, Wolof, and Yorùbá.
This paper presents a reproducible pipeline for building Universal Dependencies-style parsing resources for Katharevousa Greek parliamentary text, including OCR reconstruction, LLM-assisted annotation, and evaluation of multiple parsers. The best model (XLM-R) achieves 0.8893 UPOS accuracy and 0.5162 LAS, significantly outperforming off-the-shelf baselines.