Tag
A comprehensive survey on tokenization in modern NLP, compiled by 32 tokenizer researchers, covering algorithms, evaluations, multilinguality, encodings, theory, and alternatives like latent or visual tokenization, plus adjacent topics such as constrained generation and tokenizer security.
This paper quantifies the compression cost of pre-tokenisation boundary rules by bounding minimum token counts from both sides, certifying that regex boundaries increase optimal token counts by 28.3–36.8% on English Wikipedia, and showing that compression-optimal token dictionaries do not necessarily improve language-model prediction quality.
This paper examines how using reserved tokens versus subword tokens in chat templates affects the success of prompt injection attacks against LLM agents, finding that reserved tokens significantly increase attack authority and success rates.
The paper shows that unequal single-token support for names in large language models leads to biased concept accessibility across demographics, and introduces NameTrace to measure lexical comparability.
This paper investigates how American English becomes the default in large language models, examining structural bias across pretraining data, tokenization, and generation stages through a controlled study with American and British English variants.
This tool compiles fonts where each LLM token is of equal width, with presets for various tokenizers and integration options for Discord and Slack.
The paper addresses tokenization errors in large language models when applied to Brahmic scripts by formalizing orthographic constraints in Agda and developing a provably correct fix for tokenization, with practical implementations in SentencePiece and a Rust library.
Hugging Face has released version 1 of their tokenizers library, featuring multiple language support, multi-thread scaling, and minimal package size.
This paper introduces a linguistically motivated phonemic tokenizer for Vietnamese and Chinese that factorizes syllables into onset, rime, and tone components, reducing vocabulary size and improving efficiency. The proposed PhonemicBERT models demonstrate competitive or superior performance in language understanding tasks compared to existing tokenizers and pretrained models.
Hugging Face releases tokenizers v1, a major performance update for the tokenization library, with benchmarks showing significant speed improvements over previous versions.
This article details building a world model for Super Mario Land using a transformer to predict next game frames from tokenized inputs, and demonstrates autoregressive 'dreaming' to generate future frames.
This research paper investigates why pretraining in large language models fails to transfer knowledge across languages, identifies disjoint token spaces as a fundamental barrier, and proposes mapping languages to a shared token space to improve cross-lingual generalization.
This paper measures the tokenization cost for Sanskrit compared to English and Hindi, finding that Sanskrit requires more tokens per proposition due to its information density, but the penalty decreases with larger vocabulary sizes.
This paper presents a large-scale study comparing distilled byte and token models, finding that byte models achieve higher performance ceilings with more compute and are more data-efficient than token models.
This article explores using Byte-Pair Encoding (BPE) token efficiency as a more effective alternative to entropy for detecting secrets in code, focusing on statistical rarity over randomness.
BNB Chain led all blockchain networks in the growth of tokenized U.S. Treasury bill market capitalization over the past year, adding $3.9 billion.
This paper introduces MEL, a coordinate-preserving EEG tokenization framework for translating EEG to fMRI, addressing representation-interface mismatch and improving prediction over baselines through explicit modeling of hemodynamic latency and spectral-spatial dynamics.
The study finds that neural language models degrade similarly under word-level noise but differently under character-level noise, with tokenization identified as the key hidden variable. It provides a method to predict model robustness without noisy evaluation and suggests noise-augmented training for install robustness.
This paper audits extractive prompt compressors across ten languages, revealing that English-trained models exhibit significant performance gaps on non-English text at high compression rates and proposes a translate-then-compress pipeline as a more effective alternative.
This paper investigates the fairness of crosslingual evaluation methods for language models, showing that common normalized metrics can be biased due to tokenization and orthographic differences, and proposes using sentence-level negative log likelihood on semantically equivalent sequences for more consistent crosslingual comparisons.