Tag
This paper presents the first neural benchmark for Sinhala Sandhi splitting using character-level seq2seq models on the SandhiLex dataset, with a bidirectional LSTM encoder reaching 68.40% exact-match accuracy on hard lexicalized/derivational Sandhi versus 94% on regular affixational cases.
The microGPT-C project provides a minimal, dependency-free implementation of a character-level transformer in pure C, achieving over 10 million tokens per second inference speed on Apple M5 hardware.
The paper introduces Stoicheia, a 405M-parameter character-level masked diffusion encoder for Ancient Greek that unifies textual restoration, parsing, and metrical scansion in a single model, outperforming prior systems like Ithaca on benchmark tasks.
This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.
Presents CHiPS, a lightweight character-level authorship attribution method for Romanian texts that uses character histograms and positional signals without tokenization or pretrained language models, achieving 0.9310 accuracy on a closed-set benchmark.