Tag
This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.
Introduces MORFES, a benchmark of 500 expert-verified items for testing productive inflectional competence in Modern Greek, and evaluates open language models including their own Sophea-Genesis-1, which leads on inflectional morphology.
This paper presents Morpheus, a neural tokenizer and word embedder for Turkish that learns morpheme boundaries without string normalization, achieving lossless tokenization and competitive embeddings for lexical retrieval, while using less GPU memory than subword tokenizers.
Introduces PACUTE, a diagnostic benchmark of 4,600 tasks evaluating morphological understanding in Filipino, revealing that even frontier models struggle with morpheme decomposition and productive morphological composition.
A personal account of finding a fossil seashell in the Saudi desert and using computational methods to analyze its shape compared to thousands of shell species.
Presents a novel pattern-and-root model for describing Arabic noun inflection, focusing on broken plurals, with a taxonomy of 160 classes and an encoding scheme applied to 3,200 entries, aiming to improve computational language resources.
This paper investigates how character-level transformer models generalize to irregular verb subtypes in Japanese past-tense inflection. Controlled experiments show that including irregular examples can improve generalization, challenging the assumption that regularity simplifies learning.