tokenization

Tag

Cards List
#tokenization

Tokenization: A Survey for Modern NLP [R]

Reddit r/MachineLearning ↗ · 4h ago

A comprehensive survey on tokenization in modern NLP, compiled by 32 tokenizer researchers, covering algorithms, evaluations, multilinguality, encodings, theory, and alternatives like latent or visual tokenization, plus adjacent topics such as constrained generation and tokenizer security.

0 favorites 0 likes
#tokenization

The Price of Token Boundaries: Compression Certificates and Prediction

arXiv cs.AI ↗ · 18h ago Cached

This paper quantifies the compression cost of pre-tokenisation boundary rules by bounding minimum token counts from both sides, certifying that regex boundaries increase optimal token counts by 28.3–36.8% on English Wikipedia, and showing that compression-optimal token dictionaries do not necessarily improve language-model prediction quality.

0 favorites 0 likes
#tokenization

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Hugging Face Daily Papers ↗ · 2d ago Cached

This paper examines how using reserved tokens versus subword tokens in chat templates affects the success of prompt injection attacks against LLM agents, finding that reserved tokens significantly increase attack authority and success rates.

0 favorites 0 likes
#tokenization

Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

Hugging Face Daily Papers ↗ · 2d ago Cached

The paper shows that unequal single-token support for names in large language models leads to biased concept accessibility across demographics, and introduces NameTrace to measure lexical comparability.

0 favorites 0 likes
#tokenization

How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline

Hugging Face Daily Papers ↗ · 2d ago Cached

This paper investigates how American English becomes the default in large language models, examining structural bias across pretraining data, tokenization, and generation stages through a controlled study with American and British English variants.

0 favorites 0 likes
#tokenization

Generate fonts where every LLM token is the same width

Hacker News Top ↗ · 4d ago Cached

This tool compiles fonts where each LLM token is of equal width, with presets for various tokenizers and integration options for Discord and Slack.

0 favorites 0 likes
#tokenization

Type-Driven Tokenization for Brahmic Scripts

arXiv cs.CL ↗ · 2026-09-22 Cached

The paper addresses tokenization errors in large language models when applied to Brahmic scripts by formalizing orthographic constraints in Agda and developing a provably correct fix for tokenization, with practical implementations in SentencePiece and a Rust library.

0 favorites 0 likes
#tokenization

tokenizers v1 (rust)

Reddit r/LocalLLaMA ↗ · 2026-09-21

Hugging Face has released version 1 of their tokenizers library, featuring multiple language support, multi-thread scaling, and minimal package size.

0 favorites 0 likes
#tokenization

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper introduces a linguistically motivated phonemic tokenizer for Vietnamese and Chinese that factorizes syllables into onset, rime, and tone components, reducing vocabulary size and improving efficiency. The proposed PhonemicBERT models demonstrate competitive or superior performance in language understanding tasks compared to existing tokenizers and pretrained models.

0 favorites 0 likes
#tokenization

tokenizers v1: encode, decode and scaling, measured

Hugging Face Blog ↗ · 2026-09-21 Cached

Hugging Face releases tokenizers v1, a major performance update for the tokenization library, with benchmarks showing significant speed improvements over previous versions.

0 favorites 0 likes
#tokenization

World Models From Scratch 2: Model Training and Dreaming [P]

Reddit r/MachineLearning ↗ · 2026-09-19 Cached

This article details building a world model for Super Mario Land using a transformer to predict next game frames from tokenized inputs, and demonstrates autoregressive 'dreaming' to generate future frames.

0 favorites 0 likes
#tokenization

Why Pretraining Fails to Share Cross-Lingual Knowledge

arXiv cs.CL ↗ · 2026-09-18 Cached

This research paper investigates why pretraining in large language models fails to transfer knowledge across languages, identifies disjoint token spaces as a fundamental barrier, and proposes mapping languages to a shared token space to improve cross-lingual generalization.

0 favorites 0 likes
#tokenization

Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition

arXiv cs.CL ↗ · 2026-09-14 Cached

This paper measures the tokenization cost for Sanskrit compared to English and Hindi, finding that Sanskrit requires more tokens per proposition due to its information density, but the penalty decreases with larger vocabulary sizes.

0 favorites 0 likes
#tokenization

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

arXiv cs.CL ↗ · 2026-09-14 Cached

This paper presents a large-scale study comparing distilled byte and token models, finding that byte models achieve higher performance ceilings with more compute and are more data-efficient than token models.

0 favorites 0 likes
#tokenization

Rare Not Random Using Token Efficiency for Secrets Scanning

Lobsters Hottest ↗ · 2026-09-12 Cached

This article explores using Byte-Pair Encoding (BPE) token efficiency as a more effective alternative to entropy for detecting secrets in code, focusing on statistical rarity over randomness.

0 favorites 0 likes
#tokenization

@TCryptochicks: BNB Chain led all chains in tokenized U.S. Treasury bill market cap growth over the past year, adding an impressive $3.…

X AI KOLs Following ↗ · 2026-09-10 Cached

BNB Chain led all blockchain networks in the growth of tokenized U.S. Treasury bill market capitalization over the past year, adding $3.9 billion.

0 favorites 0 likes
#tokenization

MEL: Coordinate-Preserving EEG Tokenization for fMRI Translation

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper introduces MEL, a coordinate-preserving EEG tokenization framework for translating EEG to fMRI, addressing representation-interface mismatch and improving prediction over baselines through explicit modeling of hemodynamic latency and spectral-spatial dynamics.

0 favorites 0 likes
#tokenization

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

arXiv cs.CL ↗ · 2026-08-28 Cached

The study finds that neural language models degrade similarly under word-level noise but differently under character-level noise, with tokenization identified as the key hidden variable. It provides a method to predict model robustness without noisy evaluation and suggests noise-augmented training for install robustness.

0 favorites 0 likes
#tokenization

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper audits extractive prompt compressors across ten languages, revealing that English-trained models exhibit significant performance gaps on non-English text at high compression rates and proposes a translate-then-compress pipeline as a more effective alternative.

0 favorites 0 likes
#tokenization

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper investigates the fairness of crosslingual evaluation methods for language models, showing that common normalized metrics can be biased due to tokenization and orthographic differences, and proposes using sentence-level negative log likelihood on semantically equivalent sequences for more consistent crosslingual comparisons.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback