subword-tokenization

Tag

Cards List
#subword-tokenization

TokEval: A Tokenizer Evaluation Suite

arXiv cs.CL · 2026-08-19 Cached

This paper introduces TokEval, a framework for evaluating language model tokenizers using intrinsic metrics that correlate with downstream task performance.

0 favorites 0 likes
#subword-tokenization

QuechuaTok: Morphological Boundary Accuracy as a Necessary Metric for Tokenizer Evaluation in Agglutinative Low-Resource Languages

arXiv cs.CL · 2026-06-24 Cached

This paper presents QuechuaTok, a benchmark for evaluating tokenization strategies for Southern Quechua, and introduces Morphological Boundary Accuracy (MorphAcc) as a necessary metric. It shows that BPE achieves low fertility but poor morphological accuracy, while a morphology-aware PRPE tokenizer achieves 83% MorphAcc, demonstrating that fertility rate alone is insufficient for agglutinative languages.

0 favorites 0 likes
#subword-tokenization

Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation

Hugging Face Daily Papers · 2026-05-14 Cached

This paper investigates the impact of subword tokenization on LLM training efficiency and performance by conducting controlled byte-level pretraining experiments. It reveals key factors such as training throughput and the integration of subword boundaries as linguistic priors.

0 favorites 0 likes
← Back to home

Submit Feedback