标签
本文提出了首个面向僧伽罗语连声(Sandhi)切分的神经网络基准任务,采用字符级 seq2seq 模型在 SandhiLex 数据集上进行实验,其中双向 LSTM 编码器在困难的词汇化/派生性连声上达到 68.40% 的完全匹配准确率,而在规则性的词缀切分案例上则达到 94%。
microGPT-C项目是一个基于纯C语言的最小化、无依赖的字符级Transformer实现,在Apple M5硬件上可达到超过10M tps的推理速度。
The paper introduces Stoicheia, a 405M-parameter character-level masked diffusion encoder for Ancient Greek that unifies textual restoration, parsing, and metrical scansion in a single model, outperforming prior systems like Ithaca on benchmark tasks.
This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.
提出CHiPS,一种轻量级的字符级作者归属方法,适用于罗马尼亚语文本,采用字符直方图和位置信号,无需分词或预训练语言模型,在封闭集基准测试中达到0.9310准确率。