Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks
Summary
This paper presents the development of Cantonese and Irish treebanks within the ParGram Project and investigates the potential and limitations of using multilingual LLMs (OpenAI's gpt-oss-120b) for grammar engineering tasks such as translation and syntactic structure generation.
View Cached Full Text
Cached at: 08/10/26, 08:05 AM
# Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks Source: [https://arxiv.org/abs/2608.07283](https://arxiv.org/abs/2608.07283) [View PDF](https://arxiv.org/pdf/2608.07283) > Abstract:Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross\-linguistic consistency with language\-specific properties\. This paper presents the development of Cantonese and Irish treebanks within the Parallel Grammar \(ParGram\) Project, where linguistic parallelism is maintained at an abstract functional level\. We also investigate the methodological potential and limitations of using multilingual LLMs to support grammar engineering, focusing on Cantonese\-Irish translation and the generation of formal syntactic structures using OpenAI's gpt\-oss\-120b model\. The results show that translation performance was generally unsatisfactory and unaffected by prompt language\. For syntactic structure generation, the model produced some structurally meaningful outputs, but performed poorly on tasks requiring cross\-linguistic abstraction\. Nonetheless, LLM\-generated outputs may still offer some reference value by suggesting alternative analyses and \(partially\) capturing predicate\-argument relations\. Overall, our findings highlight both the potential and limitations of using LLMs in collaborative grammar engineering, while underscoring the continued importance of expert\-driven analysis and verification\. ## Submission history From: Elaine Uí Dhonnchadha \[[view email](https://arxiv.org/show-email/c3e83368/2608.07283)\] **\[v1\]**Fri, 7 Aug 2026 14:44:19 UTC \(729 KB\)
Similar Articles
LLMs for automatic annotation of Mandarin narrative transcripts
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
A comprehensive dual-aspect evaluation framework for large language models on Vietnamese legal text simplification, combining quantitative benchmarking (Accuracy, Readability, Consistency) with qualitative error analysis across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1.
Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts
This paper proposes a hybrid framework combining LLM summarization, a linguistically grounded Expert Index, and STAC classification to generate causal graphs from narrative texts, outperforming GPT-4o and Claude 3.5 on benchmark stories.
@stanfordnlp: There are still opportunities for using detailed linguistics in the age of LLMs!
Stanford NLP announces that a linguistics-informed NLP paper titled 'The Imperfective Paradox in LLMs' has won a Best Paper Award at ACL 2026, encouraging more detailed linguistic evaluation of large language models.
Exploring the Capability Boundaries of LLMs in Mastering Chinese Chouxiang Language
This paper introduces Mouse, a specialized benchmark for evaluating LLMs on Chinese Chouxiang Language tasks across six NLP domains, revealing that current state-of-the-art models have significant limitations with this subcultural internet language despite performing well on contextual understanding tasks.