Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
Summary
A deep learning framework is developed to analyze grammatical gender evolution from Latin to Romance languages, focusing on low-resource historical settings using lexical and contextual analysis.
View Cached Full Text
Cached at: 06/02/26, 03:35 PM
Paper page - Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
Source: https://huggingface.co/papers/2605.09156
Abstract
A deep learning framework is developed to analyze the grammatical gender system evolution from Latin to Romance languages, examining both lexical and contextual factors in a low-resource historical setting.
The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration (masculine, feminine, neuter) to a bipartite one (masculine, feminine) in most Romance languages. In this work, we introduce an interpretabledeep learning frameworkto investigate this phenomenon at both lexical andcontextual levels. First, we show that conventionaltokenizationstrategies are insufficiently robust for this low-resource historical setting, and that our proposed tokenizer improves performance over these baselines. At thelexical level, we evaluate the contribution ofmorphological featurestogender prediction. At thecontextual level, we quantify the contributions of differentpart-of-speech categoriesto grammaticalgender prediction. Together, these analyses characterize the distribution of gender information between the lemma and its sentential context. We make our codebase, datasets, and results publicly available at https://github.com/ahan-2000/Lost-in-Translation-{https://github.com/ahan-2000/Lost-in-Translation-}.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2605\.09156
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.09156 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.09156 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.09156 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
Researchers introduce MORPHOGEN, a multilingual benchmark testing LLMs’ ability to rewrite first-person sentences in the opposite gender while preserving meaning across French, Arabic, and Hindi.
Probing Character-level Transformers for the Spanish L-shaped Morphome
This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.
LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language
The article introduces LLiMba, a 3B parameter model adapted from Qwen2.5 for Sardinian using continued pretraining and supervised fine-tuning on a single consumer GPU. It evaluates various LoRA configurations, finding that adapter capacity significantly impacts performance and factual accuracy in low-resource language adaptation.
Mitigating Gender Bias in English to Romanian Machine Translation
This paper proposes a hybrid pipeline combining fine-tuned LLaMA-based gender classification with tag-aware neural machine translation to mitigate gender bias in English-to-Romanian MT, introducing new datasets and improving gender accuracy by over 40 points on benchmarks.
Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages
This research investigates how multilingual large language models internally handle subject-verb agreement across languages, finding that models reuse shared computational structure for languages with overt inflection, indicating cross-lingual overlap in morphosyntactic processing.